How Google Crawls and Indexes Websites: From URL Discovery to Search Results

UC_From Crawl Issues to Organic Growth

Publishing a website does not automatically mean that it will appear on Google.

Before a page can rank for keywords, generate organic traffic, or bring in inquiries, Google must first discover the page, access its content, understand what it is about, and decide whether it should be added to the search index.

This process is generally divided into three stages:

  1. URL discovery
  2. Crawling and rendering
  3. Indexing

Understanding these stages is especially important for businesses dealing with Google indexing issues, declining organic traffic, or newly published pages that do not appear in search results.

A professional technical SEO audit does more than check keywords. It determines whether Google can actually find, process, and index the pages that matter to the business.

What is the Difference Between Crawling and Indexing?

Crawling and indexing are related, but they are not the same thing. 

CrawlingIndexing
It happens when Googlebot visits a URL and downloads the page’s content. It happens after crawling, when Google analyses the page and considers storing it in its search index. 

A page can therefore be:

  • Discovered but not crawled
  • Crawled but not indexed
  • Indexed but not ranking well
  • Indexed under a different canonical URL
  • Blocked before Google can properly process it

This distinction explains why submitting a URL through Google Search Console does not guarantee that it will appear in search results. Google states that requesting a crawl does not guarantee immediate inclusion or inclusion at all, because its systems prioritize useful, high-quality content.

For business owners searching for why a website is not appearing on Google, the problem is not always the keyword strategy. It may begin much earlier in the crawling and indexing process.

Stage One: How Does Google Discover Website URLs?

UC_From Crawl Issues to Organic Growth

Google does not maintain a central directory containing every page published on the internet. It continuously discovers new and updated URLs through several methods. 

Internal Links

One of the main ways Google discovers pages is by following links from pages it already knows.

For example, imagine that a Malaysian digital marketing agency publishes a new page about its Google Ads management service. If that page is linked from:

  • The main navigation
  • The digital marketing services page
  • Relevant blog articles
  • The website footer
  • A service-category page

Googlebot has several paths through which it can discover the URL.

However, if the service page is published without being linked from anywhere, it becomes an orphan page. The URL may exist, but Google may take much longer to discover it.

Google specifically identifies links from known pages and submitted sitemaps as important methods of URL discovery.

This is why internal linking should not be treated as an afterthought. It helps users navigate the website while also showing Google how pages are connected.

XML Sitemaps

An XML sitemap provides Google with a list of URLs that the website owner considers important. 

It is particularly useful for new websites, large websites, and sites with pages that are not strongly connected through internal links. However, Google makes it clear that including a URL in a sitemap does not guarantee that the URL will be crawled or indexed. 

A sitemap is a discovery signal, not an indexing command.

External Links

Google may also discover a page when another website links to it.

For example, a newly launched corporate website may be discovered faster when it is linked from:

  • A recognized business directory
  • An industry association
  • A partner’s website
  • A media publication
  • A company’s verified social or business profile

External links can therefore support both page discovery and authority building. However, backlinks cannot repair a page that is blocked by technical SEO problems.

Stage Two: How Does Google Crawl and Render a Page?

After discovering a URL, Google may place it into a crawl queue. Googlebot then decides when to request the page based on factors such as crawl demand, website quality, server capacity, and how frequently the content changes.

Google does not crawl every discovered URL immediately. It also adjusts its crawling behavior to avoid overwhelming a website’s server. Repeated server errors, slow responses, or availability problems can therefore reduce crawling efficiency.

Google Checks the Server Response

When Googlebot requests a page, the server returns an HTTP status code.

Common responses include:

  • 200: The page loaded successfully
  • 301 or 308: The URL permanently redirects elsewhere
  • 302 or 307: The redirect is temporary
  • 404: The page cannot be found
  • 410: The page has been intentionally removed
  • 500: The server encountered an error
  • 503: The service is temporarily unavailable

A service page that consistently returns a 500 error cannot be crawled reliably. Similarly, a page that redirects through several unnecessary URLs wastes crawling resources and creates a slower path to the final destination.

A website SEO audit should therefore check status codes, redirect chains and server errors, not just page titles and keyword density.

Google Reads the Robots.txt File

Before accessing a URL, Googlebot checks whether crawling is permitted by the website’s robots.txt rules.

A common mistake is assuming that robots.txt can reliably remove a page from Google.

It cannot.

Google explains that robots.txt is mainly used to manage crawler access and prevent unnecessary crawling. A blocked URL may still appear in search results if Google discovers it through links elsewhere. To prevent indexing, website owners should use methods such as a noindex directive, password protection, or complete removal.

Incorrectly blocking CSS or JavaScript files can also make it harder for Google to understand a page’s layout and content. 

Google Renders JavaScript

Modern websites frequently use JavaScript to load product listings, service details, navigation links, calculators, and interactive content.

Google can process JavaScript using a modern rendering system based on Chromium. Its JavaScript processing sequence generally includes crawling, rendering, and indexing. However, Google cannot render scripts or pages that it was prevented from crawling.

For example, suppose a landing page contains its main service description inside a JavaScript widget that only loads after the visitor clicks “Learn More.” Users may be able to access the content, but Google may not process it as part of the initially available page.

Critical SEO content should therefore be included in crawlable HTML whenever possible.

Googlebot Has a Fetching Limit

A less commonly discussed technical factor is the amount of page data Googlebot processes.

In a March 2026 technical explanation, Google stated that Googlebot currently fetches up to 2 MB for an individual URL, excluding PDFs. Content beyond that point is not fetched, rendered, or indexed. Google also recommends keeping HTML lean and placing critical elements, such as titles, canonicals, and structured data, higher in the document.

Most normal business websites will never approach this limit.

Stage Three: How Does Google Decide What to Index?

After crawling and rendering the page, Google analyses its primary content, title, headings, images, videos, links, and other signals.

Google then determines whether the page should be indexed and which URL should represent the content.

Google Evaluates the Main Content

A technically accessible page is not automatically valuable enough to index. 

If every page contains almost identical text with only the location name changed, Google may treat some of the pages as duplicate or low-value content.

A stronger location page would include genuinely useful local information such as:

  • Services relevant to businesses in the area
  • Local project experience
  • Industry examples
  • Local market challenges
  • Customer testimonials
  • Location-specific FAQs
  • Relevant case studies

Indexability is not only a technical matter. The page must also justify why it deserves to exist separately.

Google Selects a Canonical URL

When Google finds several similar versions of a page, it may select one URL as the canonical version. 

A canonical tag helps indicate the preferred URL, but it should be supported by consistent internal links, redirects, and sitemap entries. 

When these signals conflict, Google may select a different canonical URL from the one requested by the website owner. 

Google Uses the Mobile Version for Indexing

Google uses the mobile version of a website’s content for indexing and ranking.

This means the mobile page should contain the same essential content, metadata, structured data, and images as the desktop version. Hiding important content from mobile users may also prevent Google from using that information during indexing.

Content can be placed inside accordions or tabs to save space, but it should remain accessible in the page’s rendered HTML.

Case Study: Saramin Increased Organic Traffic by Fixing Crawling Issues

Saramin Increased organic traffic by fixing crawling issues

Saramin, a large Korean employment platform, provides a documented example of how technical crawling and indexing improvements can affect business performance.

After registering its website with Google Search Console, Saramin initially focused on identifying crawling errors and ensuring that Googlebot could properly access pages for indexing.

According to Google’s published case study, these initial fixes contributed to a 15% increase in organic traffic.

Saramin then expanded its SEO work by:

  • Removing unhelpful keyword-heavy meta tags
  • Implementing canonical URLs
  • Removing duplicate content
  • Correcting structured data issues
  • Improving mobile usability
  • Monitoring technical errors in Search Console

As the number of valid pages increased, the company saw stronger business results. During its peak hiring season in September 2019, Saramin recorded a 102% year-over-year increase in organic Google traffic, a 93% increase in new sign-ups from organic search, and a 9% increase in the organic conversion rate.

The important lesson is that the result did not come from repeatedly requesting indexing.

Saramin improved Google’s ability to crawl, interpret, and trust its website. It also connected technical SEO work with conversion outcomes such as registrations, not traffic alone.

Crawling and Indexing Are the Foundation of SEO

Google cannot rank content that it cannot properly discover, crawl, render or index.

Keywords, backlinks and content creation remain important, but they work most effectively when the technical foundation is stable.

For a Malaysian SME, the most valuable SEO question is not simply:

“Have we published enough articles?”

A better set of questions is:

  • Can Google find our most important pages?
  • Are we making those pages easy to crawl?
  • Does Google understand which URL is canonical?
  • Is the mobile content complete?
  • Are our service pages useful enough to index?
  • Do our indexed pages satisfy commercial search intent?

Solving these issues can produce more than an increase in indexed URLs. As the Saramin case demonstrates, better crawling, clearer indexing signals, and stronger page quality can contribute to increased organic traffic, registrations, and conversions.

Businesses facing persistent Google indexing problems should begin with a structured technical SEO investigation rather than repeatedly submitting URLs or publishing more generic content.

Related Posts