indexar en Google
02 / 10 / 2026

Indexing on Google in 2026: Key Takeaways from Search Central Live

Bruno Díaz Marketing Manager
Bruno Díaz
—
Marketing Manager

Indexing is no longer a formality: Google only stores a small portion of what it finds, and JavaScript, AI and changes in the SERP are making it harder.

Google invited us to Search Central Live Deep Dive Europe 2026, the congress it organizes for SEO professionals and which this year was held in Barcelona. The first day was dedicated to how Google crawls the web. The second, to how it interprets it and decides what to store in the index. This article covers that second day, organized with a common thread: understanding how Google processes the content of our website to know what to do to get it indexed.

It is the most technical format of Search Central Live: long sessions with the Google Search team and space to ask questions. For an agency that works with SEO and development, it is an opportunity to hear first-hand how their systems work and contrast it with what we see every day in client projects.

Escenario de la Search Central Live Deep Dive Europe 2026 de Google con dos ponentes sentados

Indexing is no longer a formality. Google only saves a small part of what it finds, and the context has changed: there are more and more websites and applications built with JavaScript, many already generated with AI; the results page incorporates summaries and modules that make not everything a blue link (we deal with it in SEO in 2026), and AI assistants read the website with different capabilities than Google. If the page does not reach the index well, the rest of the SEO work is useless.

The article follows the journey a page takes from when Google finds it to when it decides whether to save it. At each step, we explain what Google does, where websites tend to fail, and what should be checked. These are our lecture notes, not official Google documentation: for the fine details, Search Central documentation is still the reference. If any concepts don't ring a bell, our SEO Dictionary lists them.

The journey from a page to the table of contents

e to the table of contents

From when Google finds a URL until it decides to save it, there are six steps:

  1. Google decides which URLs you visit and at what pace, and downloads them.
  2. Parsing of HTML. It reads the code, separates the main content from the rest, looks at the head, and extracts the links and robot meta tags.
  3. If the page relies on JavaScript and CSS, it goes through a second queue where it runs and looks like a user would see it.
  4. Group pages with the same or very similar content and choose the most representative one.
  5. Feature and signal extraction. Structured data, images, video, language, country, freshness and quality.
  6. Selection for the table of contents. With all of the above, decide whether the page enters.

Google Process Diagram: Crawl Queue, Crawler, Render, Render, and Index

A problem in one step conditions the next. A page that Google cannot render will not arrive well interpreted in deduplication, or in signals, or in the index. That is why the most useful thing is to know in which step each page is lost.

Tracking: Access and Budget

Google doesn't crawl the entire web or always at the same rate. What you visit depends on what it finds and how the server responds, and that's the crawl budget. On the first day of the congress we talked about this most, and the server's responses are what affect it the most:

Reply

What's up

What to do

5xx (500, 502, 503) and 429

Googlebot slows down the pace. If they persist, it may stop indexing URLs

In temporary maintenance, serve 503

Slow response times or timeouts

The longer the server takes, the fewer URLs it crawls

Improve performance

5xx in the robots.txt

Google can stop crawling the entire web until it recovers

Monitor file availability

3xx

Each jump is an extra request. Chains are especially bad

Link to the final URL

Soft 404 (200 with "not found" content)

Google continues to track them

Return a real 404 or 410

404 / 410

They don't stop crawling, but Google keeps revisiting them. The 410 is removed a little faster

Use 410 if they won't come back

200 with duplicate or poor content (parameters, filters, infinite paginations)

Consume budget without adding value

Control it with robots.txt, canonical and linked

304 Not Modified (with ETag or Last-Modified)

Google doesn't have to redownload the page

Implement these headers

To see these answers in practice, we have the Google Search Console guide and a post on how to fix 5xx and 4xx errors.

Other decisions that were clarified:

  • HTML Sitemaps. They don't need to be done, and poorly implemented they can be harmful.
  • Sitemap in the robots.txt. It's preferable, but it's open to everyone, even to those who aren't interested. If we only want Google to find it, just send it from Search Console.
  • The robots.txt is not a system for deindexing. If we block a URL by robots, Google can still index it, but without content. To remove a page from the search you need noindex, and for Google to read it the URL has to be crawlable. We explain it in the post how to deindex a Google URL.
  • Out of stock products. It depends on the case. If the product is important, exclusive, difficult to find and the user can wait, the safest thing to do is to keep the 200 with preorder or an action that does not generate frustration. If it will return soon, 200 with the stock indicated in the structured data; no 404 or redirection.

HTML: what Google reads and in what order

When Google lands on a page, it parses the HTML following the elements of the DOM. First, it separates what it considers to be the main content from the rest. Then it reads the head (canonical, hreflang, title), extracts the links, and queries the robot meta tags.

The links

Links are Google's way of discovering new pages, and they have to be . Any other solution is a risk.

Table of links that Google can extract, with a href, versus those that can't, such as routerLink or span

On websites made with JavaScript frameworks it is one of the first things to look at: if the menu or products are opened with an onclick or a routerLink, Google does not see them as links.

Robots meta tags

In the head, you can indicate rules for all robots or for a specific one. The ones you need to know:

Directive

What it does

noindex

Does not index the page

Nofollow

It doesn't follow links. At the page level it makes little sense; it has it more applied link by link

none

Equals noindex plus nofollow

nosnippet

Indexes the page, but without snippets or summaries in AI Overviews

data-nosnippet

HTML attribute (not meta tag) that excludes only a portion of the content from the results

max-snippet

Limits the length of the fragment, usually with a number of characters; -1 is unlimited

max-image-preview

Set the size of the image: none, standard, or large. Large  can help you gain visibility in Discover

max-video-preview

Sets the maximum preview duration

Notranslate

Avoid translation in results and Chrome. Unusual use

noimageindex

It indexes the content, but not the images. Rare case

 Throttling directives are a business decision: if we give the whole answer on the results page, the user has no reason to click. Still, for most websites the goal is maximum visibility, and the combination recommended by the presentation is this:

max-image-preview:large, max-snippet:-1

Google's recommendation for maximum visibility: max-image-preview:large and max-snippet:-1

Two more clarifications. The robots.txt rules: if it blocks access to resources that we later want to enhance with robots meta tags, they are useless. And robot meta tags can be included using JavaScript, but it is unstable, especially in the face of changes. Better to avoid it.

JavaScript, the long way

Google has greatly improved the crawling and indexing of content in JavaScript, but it's still not perfect. What it wants is to understand what a person sees to know if the page responds, and it doesn't always succeed. The first time, almost never.

A page with JS goes through two queues: the HTML queue and then the rendering queue. The path is a bit longer, but if the content is well served, Google ends up picking it up. What you need to be clear about is that CSS and JS change how the page looks, and the result can be very different from bare HTML. If Google can't render well, most links and menus don't pick them up.

Whether your website is made with React, Vue, or Next.js, our guide to JavaScript-based SEO goes  into more detail.

What to watch out for

  • Submit basic SEO attributes such as the title and meta description by JS. In the presentation, a real case was shown where they were different in the raw HTML and in the rendering.
  • Update information with JS (data, pricing), which can alter what Google indexes.
  • Inject text from database fields.
  • Rendering timeouts.
  • Content Security Policy (CSP) violations.
  • Resources blocked by JS or by robots.txt.
  • An excess of external JS that we do not control and that we will hardly be able to offer in HTML.

Comparison between raw HTML and rendering on a Lidl page where the title and meta description change with JavaScript

Case shown in the presentation: the title and meta description of the raw HTML do not match those of the rendering.

The mistakes that are most seen

  1. Content that is not in the rendered HTML. If it is not in the final DOM, Google does not see it. It happens when the JS is not accessible to Googlebot or when the APIs that need to be loaded are blocked or unresponsive. It is common in infinite scrolls, where Google only picks up the first part.
  2. Snippets (#) in the URL when the content changes. Google doesn't process them. The solution is the History API, with URLs with path or parameter: /path/products and /?path=products work, /#path=products don't.
  3. Soft 404. The application shows a "not found" but answers 200, or the page arrives broken. Before they were empty or poor pages; now they are often real errors served with a 200 via JS. You have to review the server responses of the pages that Google crawls and adjust them to reality.
  4. Blocked resources. There are still many robots.txt that block the JS or API endpoints that the page needs to render.

Summary of blind spots with JavaScript (content changes, invisible links, timeouts, CSPs, and blocked resources)

How to debug it

Three steps to find the JavaScript that prevents content from being displayed: Network tab, call that supplies it, and search in sources

Method to find the request behind the missing content.

    1. In Chrome, open the Network tab, filter by Fetch/XHR, and reload the page.
    2. Locate the call that supplies the missing content: Has it been executed? Has it returned data? Has it been blocked?
    3. If it is still missing, search for the text in all scripts loaded with Cmd/Ctrl + Shift + F.
    4. In Search Console, do a URL inspection, test the published URL, and view the tested page. Look at the screenshot, rendered HTML, and JavaScript console messages. The Rich Results Test works just as well.
    5. Compare the source code with the DOM and review the blocked resources.

A detail to remember: the server can respond 200 and the JS content is broken, and then Google treats it as a soft 404. We tested it during the session with a page on our website: the URL was available to Google and only showed two console warnings for preloaded resources that were not used.

JavaScript, AI, and server-side rendering

Google's systems collect content in JavaScript quite well, but other AI assistants still don't. That's why it's safest to serve it with SSR, i.e. server-side rendering. It's also worth asking whether AI bots can access all the content on the web, if you want them to.

The headless CMS is clean and secure, but the result is often not very readable by Google, even if the user sees it well. This can be partly compensated for by good structured data markup.

For AI-generated applications that then have to be delivered to be indexed, the presentation presented reframing with web fragments, which serve the app as isolated fragments within the page. It is still difficult to find standards and there is still a long way to go, so it is advisable to follow it from afar. WebMCP also came out, which is gaining importance in the age of management.

How Google understands a page

Google separates three areas when it reaches a page: header, navigation and main content. And it gives more and more weight to the third. That's why, many times, the title and description it shows in the results are not the ones we have given it in the header.

The consequence is practical: everything that is important has to go in the main content, but with criteria, without loading it up with everything. In the presentation they illustrated it with a post from a personal blog: the title and the text of the body are considered important, while the categories or the footer weigh less.

Main content versus secondary content, with a post from John Mueller's blog as an example.

Example of a blog post where the title and text are marked as important and the categories or footer as less important

What Google saves in the index is not the page as text either. It tokenizes the content: it separates each word, records its position and the area of the page where it is (header, centerpiece), and assigns each word to the URLs where it appears. This is how it tries to understand what the content is about and what format it is in. It has serious problems with languages that do not separate words with spaces.

Example of tokenization of English, German, and Thai phrases in Google's index

How words are separated when saving a page in the table of contents, and why languages without spaces complicate it.

This is not the chunking that AI models do. Google looks for context and does not need the content cut into fragments: it does not make sense to optimize it in blocks with Google in mind. For other AI models it can, for the time being.

Blog text converted to numerical tokens for an AI model

The tokenization that AI models do, which is not what Google stores in its index.

Duplicates and canonicals

Google believes that the user doesn't want to see the same pages, and deduplicating also saves crawling time to read new pages. The process has three steps:

  1. Understands content clusters, relates them and differentiates them within the cluster.
  2. Group the pages, choose the most representative and index it.
  3. Check additional signals (links, etc.) for more context.

If the pages are the same, it canonicalizes. If they are in different languages, it consults the hreflang and respects it. It also saves the brands: in a rebranding, try to relate the old brand with the new one, so that whoever searches for the old one sees the new one. And the same criterion applies to content migrations.

Where duplicates usually appear

  • Poorly done redirects. They have to go to equivalent content. If it isn't, the accumulated SEO is not carried over.
  • The same website served with and without www, or with http and https. It seems unbelievable, but it is still found often.
  • Unclear URL structure. In the presentation they illustrated it with /buy/fax, /buy/typewriter and /office-equipment/ as possible duplicates, and with /barcelona/services, /madrid/services and /valencia/services as almost identical pages. If only the name of the city changes, you have to ask yourself if it is worth tracking them all.
  • The same content for different countries, or automatic geolocation redirects. Here you have to use hreflang.
  • Bot blockers. Some also affect Google's crawling.

Canonicals: best practices

You should always put the canonical rel, although Google checks it and doesn't always follow what we tell it. Discrepancies are checked in Search Console. Specifically:

  • Group the canonized URLs to understand the cases (pagination, parameters, language).
  • Avoid strings of canonical and canonicals that point to a URL that does not respond 200.
  • Check that the canonical of the HTML matches the one that is served later with JavaScript. When they don't match, Google can get messy, and it's a new phenomenon.
  • Always do the internal linking to the canonical URL. If we link non-canonical URLs, Google ends up getting to the good one, but we waste time and budget.
  • In pagination and filtering, canonicalizing everything to the first page usually doesn't make sense.

When the canonical one that Google chooses is not ours, the cause is usually one of these:

  1. An attempt at hijacking: duplicate domains, open staging environments that shouldn't be, or content theft.
  2. Serious issues on the page: loading speed, security, HTTPS.
  3. Incorrect signals in the header, sitemap, or robots.txt.

The first thing to do is check that the page we want is indexable and that the HTML and rendering are correct. If none of this explains it, we can try redirecting the indexed URL to the good one.

With hreflang there is a nuance: on pages with the same language for many countries, Google ends up indexing one or a few URLs, no matter how much we mark them all. It is advisable to decide which is the main URL and work to make it the one that Google chooses.

The presentation closed with seven recommendations: use redirects in migrations, use HTTP response codes and do not block agents, review canonical rels, use hreflang to help locate, report strange canonicals in forums, make secure pages that work and keep canonical signals clear.

Structured data, images, and video

When Google already has the HTML, it does a feature extraction: it takes the general HTML, saves the images and videos separately and extracts the structured data. The current results page is full of modules that are not the ten blue links, and almost all of them come out of here.

Structured Data

It's easier, and cheaper, for Google to read structured data than to interpret visible content, as AI does. In addition, they allow you to give information that is not on the page and help bots focus on relevant content. They make the page eligible for rich results, for site and page functions (site name, breadcrumb), and for visibility in Discover. Several studies suggest that a good implementation can increase organic traffic.

How to work them:

  • Follow schema.org standards.
  • Recommended format: JSON-LD. There are other options, but it is easier to operate and gives fewer errors.
  • An excess does not penalize, but it is advisable not to add secondary elements and focus on what is important.
  • Google doesn't use all of them to show results. The search gallery indicates which ones.
  • Validate with the Rich Results Test.

For e-commerce there are new features: now you can send the loyalty program, the shipping policy and the product category.

Images

Images are increasingly present in Search, Discover, and other surfaces. Google extracts them from HTML with the elements and . Those that are only referenced as CSS background-images are not reliably indexed. Better ; also picks them up, but it's best to avoid it.

Within the image, Google uses:

  1. SRC: It has to be available and not blocked.
  2. alt: you must describe the image to someone who does not see it, without filling it with keywords. More details in the ALT text guide.
  3. title: can provide complementary context to the alt.

The rest of the attributes are ignored. It does read the words surrounding the image to get context. As good practices, it is advisable to include them in sitemaps, offer various formats and sizes, use AVIF and WebP whenever possible and seek the balance between quality and compression. We have more details in the post on SEO for images.

Video

The weight of the video depends on the search intent. It appears in Search, YouTube, and Discover, now with previews and key moments. Video markup in HTML is very important. Best practices include:

  • A URL for each video.
  • Elaborate titles and descriptions and attractive thumbnails.
  • Structured data and videos in sitemaps.
  • Preferred mp4 format.
  • The video, within the main content and at the top of the page, with the text that surrounds it worked on.
  • Keep an eye on page load time.
  • Platforms like YouTube or Vimeo, which already distribute them, and move them through other channels such as social networks, because it generates signals. For YouTube, we have an SEO post on YouTube.

Images and video have their own robots in Google, so they can be blocked separately with the robots.txt or managed with robots meta tags. For Discover, max-image-preview:large. And for AI-generated content, the recommendation is to review it before publishing it.

Language and market: International SEO

Google identifies the language by the content of the page, not by the URL or the hreflang. That's why you should avoid more than one language in the same URL. To find out which country a website is targeting, look at several factors:

  1. Geographic top-level domain (.es, .sg...). This is the strongest signal. Having your own domain per country is the best, but it is not always viable; if you don't have it, you have to strengthen the rest.
  2. Hreflang, in tags, headers, or sitemaps.
  3. Server location (IP address).
  4. Other signs: language, currency, links, and business profile.

Google doesn't vary the location of the crawler to detect different versions of the same page, and ignores geo.region location meta tags.

Multi-language SEO is gaining importance in the context of AI. You can translate with AI, but you need to check the quality of the translation. And often translating is not enough: there are legal, contextual, and cultural issues. Currency, offers and promotions, opinions and reviews, brand reputation, or product durability all weigh differently depending on the country. The presentation illustrated this with data from Think with Google on what buyers in the United States and Europe value, and with differences between countries such as Germany and France. Translating with AI may fall short for these types of factors, and will depend on the resources and risks of each project.

International SEO is not solved with translations, or even with a good hreflang:

  • The same catalog, with the same prices and photos, can sell a lot in one country and nothing in another.
  • If there are several countries with the same language and dialects, it must be adapted: Spanish, English or French in different countries.
  • We must value the authority we have in each country. Do they know us? Have we won awards? Do we have physical stores? Do we do promotions? A good hreflang does not make up for it.
  • Sometimes photos, prices or promotions will have to be changed to adapt to the market.

To go deeper, we recommend Alizée Baudez's article on international SEO.

On how to structure the website by markets, we have a post dedicated to international SEO: domain, subdomain or folder.

Signals and index selection: what goes in and what doesn't

Google only indexes a very small part of the pages it finds. It tries to keep the ones that give users the most satisfaction, and the selection is made at the end, when it has already analyzed content and signals. To decide this, you need to know if the content is reliable and relevant.

The essence hasn't changed: useful, original, well-structured, well-written content, with good metadata and schema. PageRank still exists, but it's been refined. The signs that stood out are four:

  • Country and language. It detects the language and the target country or region. Indexing can vary by domain, by presence or because there are fewer resources for minority languages. It is normal for a website in Spanish and Catalan to index better in Spanish.
  • It is always taken into account, and more so in searches that need recent results. A page can be indexed but without visibility, and in the long run it can be deindexed.
  • Safe Search. Be very careful with explicit content.
  • Google finds spam pages in droves and takes it as a priority, especially with AI: artificial, massive, dangerous content or impersonations. Even so, it has difficulty detecting it.

There are also negative signals that take a page out of the race: noindex, expired content (unavailable_after), soft 404, duplicates, and spam signals. User signals also count.

Search Console statuses

Two statuses of the page indexing report have a clear read:

  • Discovered, currently not indexed. Google knows the URL exists but hasn't crawled it at all. It's in the queue, and crawl budget has been spent.
  • Crawled, currently not indexed. Google has already analyzed it and does not consider it eligible. Resubmitting it is useless if there is no substantial change. It is usually poor, duplicate or outdated content.

What happens in the index

Google tokenizes the content and tags each token with the area of the page where it is (for example, header or centerpiece). Then it builds a posting list: each word is associated with the URLs where it appears, tagged by format. This way it tries to understand what the content is about and what format it is in.

Quick responses to common cases

Location

Recommendation

Quickly remove pages from search

Request deletion in Search Console

Parameters we don't want to be crawled (variants, search filters)

Block them with robots.txt. Canonicals help index well, but they don't control crawling

Google rewrites titles and we don't like it

Write better ones. Usually take the meta title or H1; if they are good, it will not take others (or yes)

Out of stock product that will be back soon

200 with schema out of stock, no 404 or redirect. Indicate on the page that it is not now and, if possible, offer preorder

Self-referential canonical

Yes, just in case

Disavow tool

No need

Website deindexed by mistake that we want to recover quickly

Submit your top URLs through Search Console

If you want to understand why Google rewrites titles, our post on how Google chooses page titles details it.

Content is still king, according to the paper, and it makes sense: the whole system described so far exists to decide what content is worth saving.

To decide what content is worth creating, a good source is Google Trends, which was also presented in the presentation. It has data from 2004 to a few minutes ago, combines text search, images and shopping, and serves both long-term trends and current topics. It is a good tool for marketing studies, for making SEO or SEM decisions and for presenting data to the client. It does not include Discover, because the user does not search there.

Where to start

All of the above can be turned into an orderly review. This is the order we propose, from faster and with more impact to more work.

Quick Reviews

  1. txt. That it doesn't block JS, CSS, or APIs required for rendering, and that the sitemap is submitted to Search Console.
  2. Search Console. Look at soft 404s and the "Discovered" and "Crawled, currently not indexed" statuses. Do not forward URLs without changes. This review is part of any SEO audit.
  3. Navigation links. Let them all be , also in the menus and filters.
  4. Self-referential, without strings, towards URLs that respond to 200, and with internal linking always to the canonical version. Check www and without www, http and https.
  5. Robots meta tags. max-image-preview:large and max-snippet:-1, except where there is a reason to limit it.

Revisions with more work

  1. Compare raw HTML and rendered DOM. Title, meta description, canonical, and links must be in the initial HTML. On JS-heavy websites, value SSR. For development teams, we have some essential SEO guidelines.
  2. Structured data. JSON-LD validated with the Rich Results Test, focused on what's important.
  3. Images and video. with descriptive alt, WebP or AVIF formats, sitemaps, and a URL and markup for each video.
  4. City landings with the same content, versions by country and hreflang.

In new projects

  1. One-to-one redirects to equivalent content. More details on web migration and SEO.
  2. Multi-country websites. Adapt prices, currency, promotions and reviews, and assess the local authority in each market.
  3. AI-generated applications. Think about how the content will be delivered by design so that Google indexes it.

Google Documentation

The official sources of each topic, to contrast what we have explained. All the documentation is in English.

And two technical references mentioned in the presentation: the History API (MDN) for URLs in single-page applications and Web Fragments for the delivery of AI-generated apps.

How we work at La Teva Web

Indexing is a topic in which SEO and development have to be discussed. Many of the decisions that determine whether Google sees a website well (how it is rendered, how links are generated, what the server responds to) are development and are made before anyone thinks about SEO. At La Teva Web we work with our own team, without subcontracting, and that allows us to review these points with the person who builds the website.

If you want to know at which step of the journey the pages of your website are lost, write to us and we will do an indexability review. You can see how we work on SEO in our SEO agency in Barcelona and review the basics in the SEO Guide.

Bruno Díaz Marketing Manager
About the author
Bruno Díaz — Marketing Manager
Professional with a long career as a communication and digital marketing consultant, specializing in SEO, SEM and web projects. As Marketing Manager of the agency, I coordinate a great team of digital marketing technicians of which I am very proud.

Related news

Hello! drop us a line