How-to and workflows

The AI SEO checklist for 2026

The AI SEO checklist for 2026

In short. An AI SEO checklist has four layers, and they are worth doing in order. Make sure AI crawlers can reach and render your pages. Structure each page so a self-contained passage can be lifted as an answer. Make your entity consistent everywhere it appears, on your site and off it. Then measure visibility rather than only rankings. Skipping straight to layer two is the most common mistake, because a page no crawler can read cannot be quoted however well it is written.

Most AI SEO checklists are the same on-page checklist from 2019 with the word AI added. This one is ordered by what actually blocks results, starting with the layer that silently disqualifies everything above it. It assumes you have the fundamentals working, because Google's own guidance on helpful content has not changed and nothing here replaces it. Forty checks, four layers, and a note on which ones do not matter.

Layer 1: can AI systems reach your content?

This is the layer people skip, and it is the only one that can make everything else irrelevant. A page that a retrieval crawler cannot fetch or cannot render is not a candidate for citation at any quality level.

  • robots.txt allows the retrieval crawlers you want to be cited by: OAI-SearchBot, PerplexityBot, ClaudeBot and the user-triggered fetchers.
  • No stray noindex left over from staging. Check with Search Console's URL Inspection, not by reading the template.
  • Core content is present in the server-rendered HTML, not injected client-side after load. Many AI crawlers do not execute JavaScript, and content they cannot see does not exist for them.
  • No content locked behind a cookie banner, an interstitial or a soft paywall that a crawler receives instead of the article.
  • Server returns 200 for crawler user agents, not 403. Check your access logs, not your browser.
  • Sitemap contains only canonical, indexable, 200-status URLs, with honest lastmod values.
  • No redirect chains on important pages; each hop is a chance to be dropped.
  • Pages load in a reasonable time under load. Crawlers time out.
  • hreflang is reciprocal if you run multiple locales, or your markets will cannibalise each other.
  • Server logs actually show AI crawlers arriving. If they do not, something above is wrong.

Our robots.txt generator handles the first item with all fifteen agents explained, and the crawlers in our technical crawler comparison will find the rest in one pass.

Layer 2: is each page structured to be quoted?

Retrieval systems extract passages, not pages. The unit of optimization is therefore the section, and the question for each one is whether it would still make sense pasted into an answer with no surrounding context.

  • A direct answer to the page's main question appears in the first 80 words, before any preamble.
  • Headings are phrased as the questions readers actually type, where the section answers one.
  • Each section answers its own heading in the first two to four sentences, then develops.
  • No section depends on 'as we saw above'. Every one stands alone.
  • Average sentence length under 25 words. Long, clause-heavy sentences resist extraction.
  • Key terms are defined explicitly, in the form 'X is …', near the top.
  • At least one table or list per substantial page. Structured blocks survive extraction; prose walls do not.
  • Concrete numbers with named, linked sources rather than qualitative claims.
  • Something only you could write: your own test, your own data, your own screenshots.
  • Published and modified dates visible on the page and in Article markup, and honest.
  • Author or publisher stated, with a page a reader can check.
  • Internal links to the related pages that complete the topic.

The GEO content scorer checks most of this list automatically and names the fix for each item you fail.

Layer 3: is your entity consistent everywhere?

Assistants assemble recommendations from many sources, not from your homepage. If your brand is described three different ways across the web, the system has three weakly supported claims instead of one strong one.

  • Organization schema site-wide with sameAs links to every profile that corroborates you.
  • Company name, description and category identical across your site, LinkedIn, Crunchbase, G2, review sites and directories.
  • Your product is described in the same words on your site and on the third-party pages that list it.
  • Comparison and alternatives pages exist on your own domain and are accurate about competitors.
  • You appear on the third-party roundups your buyers read. Assistants quote those lists heavily for 'best X' prompts.
  • Reviews exist on the platforms your category uses, because those platforms are cited sources.
  • Community presence where your buyers discuss the category, since community threads are heavily quoted.
  • Documentation and glossary pages are public and indexable. They are quoted more than marketing pages.
  • Pricing is public. Assistants asked 'how much does X cost' will name a competitor whose price they can read.
  • Founders or experts are identifiable people with a footprint, not anonymous bylines.

This layer is slow and it compounds. Our guide to why AI recommends competitors is essentially a diagnostic for the gaps here.

Layer 4: are you measuring the right things?

Rank tracking answers a narrower question every year. It is still worth doing, and it is no longer sufficient on its own.

  • A fixed prompt panel, 20 to 100 questions your buyers actually ask, run on a schedule and never edited casually.
  • Mention, citation and recommendation logged separately. They are different achievements.
  • Competitor share of voice on the same panel.
  • Which sources the engines cite instead of you. That list is your digital-PR roadmap.
  • AI referral traffic isolated in GA4 with a regex on session source.
  • Search Console CTR on informational queries, tracked over time as your exposure signal.
  • Impressions as a first-class metric, not a footnote, because they now move independently of clicks.
  • Server-log monitoring of AI crawler visits.
  • Conversion rate of AI-referred sessions compared with ordinary organic.
  • Reporting that shows all of it together, so nobody concludes the channel is worthless from the volume line alone.

The setup is covered step by step in how to track AI traffic in GA4 and Search Console, and the paid options start around $20 a month in our tracker comparison.

What is on this checklist that should not be?

Three items appear on most AI SEO checklists and do not deserve their placement.

llms.txt. Crawler-log studies show negligible pickup, and no engine has confirmed using it for retrieval. It costs ten minutes and it is not a priority. Our honest verdict on llms.txt goes into why.

Publishing more. Volume was the answer when the constraint was production capacity. The constraint is now differentiation, and adding undifferentiated pages makes your average worse.

Chasing a content score to 100. Optimization scores measure similarity to what already ranks. Maximising similarity produces keyword-stuffed prose that reads badly to humans and to the systems quoting passages. Use the score as a checklist of missing subtopics and stop there.

And one thing that is not on the checklist because it is not optional: being right. A hallucinated statistic or a wrong price will keep you out of answers far more reliably than any technical misconfiguration, and it costs a great deal more to repair.

Conclusion

Work the layers in order. Crawlability first, because it silently disqualifies everything else. Structure second, because it decides whether a passage can be lifted. Entity consistency third, because it decides whether you are the one named. Measurement last, because it tells you which of the three to work on next. Most teams have layers two and four partly done and have never checked layer one. Start there: it takes an afternoon, and it is the only layer where a single misconfiguration undoes a year of content work.

Score a page against the structure checks Free SEO tools and calculators

Frequently asked questions

What is the first thing to fix for AI SEO?

Crawlability. Confirm in your server logs that the AI crawlers you care about are actually reaching your pages and receiving a 200 with the content in the HTML. A page that cannot be fetched or rendered cannot be cited, however good it is.

How is an AI SEO checklist different from a normal SEO checklist?

The fundamentals overlap almost completely. What is genuinely new is passage-level structure, entity consistency across third-party sources, AI crawler access as a separate concern from Googlebot, and measuring visibility inside answers rather than only positions in a list.

How often should I run through this?

Layer one quarterly and after any template or hosting change, because it breaks silently. Layer two on every new page, as part of the brief. Layer three quarterly, since it moves slowly. Layer four continuously, because it is a dashboard rather than an audit.

Sources

Every figure in this article traces back to one of these. We link them so you can check the original rather than take our summary of it.

Free tools for this

Everything below runs in your browser, with no signup and nothing uploaded.

Definitions: Content brief · Internal linking · Content cluster · Structured data

← All articles · Glossary · Statistics

From reading to choosing

Thirty AI SEO tools compared on verified pricing, capabilities and AI search support.

See the 2026 ranking Find your tool in 60 seconds