Book an audit

Four layers between you and the answer.

What ChatGPT, Perplexity, Claude, Gemini and Google AI Overviews need from a website before they will use it in an answer: the technical layer, the content layer, the access layer and the citation layer, in the order they fail. Each layer filters out the sites that failed the one below it, so the order of the work matters more than the amount of it.

Layers
4
Research network
50 sites
AI bot logs
200+ days
Outside studies cited
11

Four layers, one order.

An AI answer engine does four things with a website, always in the same order. It tries to reach the page. It cuts what it finds into passages. It reads the structured signals that say what the page is, who wrote it and when. Then it picks the few sources it trusts for the answer. Each step filters out the sites that failed the step before, which is why the order of the work matters more than the amount of it.

Every layer below answers three questions: what the machine actually does at that layer, what technically ideal looks like, and how to test it yourself without buying a tool. Every number carries its source. Numbers from my own research network (50 live sites across 9 industries, 200+ days of CDN-level AI bot logs) are labelled as mine. Everything else names the study, the sample and the date, because this field changes every quarter and a figure without a date is a guess.

The four layers, the question each one asks, and where sites fail
LayerThe question the machine asksWhere sites fail
01 TechnicalCan I reach this page quickly and read it without running any code?Slow server response, blocked bots, JavaScript-only content, redirect chains.
02 ContentDoes one passage from this page answer the question on its own?Answers buried under intros, pronouns instead of names, claims with no source or date.
03 AccessCan I find, date and identify this content without guessing?No feed or an excerpt-only feed, sitemaps with fake dates, broken or contradictory schema.
04 CitationsWhich sources do I already trust for this question, and is this brand in them?Never measured. Absent from the roundups, directories and communities the engines actually cite.

The ratio that surprises people: 70/30.

Across my research network, AI visibility behaves as roughly 70% infrastructure and 30% content. That is the reverse of the usual SEO ratio. AI fetchers spend most of their requests probing the shape of a site before they decide what to read, so content investment on top of a broken technical layer is money spent on pages the bots never get to read.

Layer 01.
Technical

If it can't reach you, it can't cite you.

A live AI answer is a race against a clock. The engine does not read one website: it rewrites the question into many searches, shortlists pages from an index and opens them live while a person waits. A page that is slow, blocked or empty on first load is not ranked lower. It is dropped from the answer.

When someone asks ChatGPT, Perplexity, Claude or Google AI Mode a question that needs current information, the engine breaks it into sub-questions and runs a search for each one. Google documents this as query fan-out, and its Deep Search mode can issue hundreds of searches for a single request. ChatGPT typically runs a handful and grounds each sentence of its answer in one URL. Every candidate on the shortlist is a live request against a time budget.

Between the question and the answer: six steps.

Every AI answer that uses the live web runs the same six steps. Step 4 is where technical debt turns into invisibility: no engine publishes its timeout, so design for the most impatient one.

  1. Question. A person asks. Nothing has been fetched yet.
  2. Fan-out. The question becomes many searches: a handful in ChatGPT, dozens to hundreds in Google's Deep Search.
  3. Shortlist. Candidates come from an index: mostly Bing for ChatGPT, Google for Gemini and AI Overviews, Brave for Claude.
  4. Live fetch, against the clock. Shortlisted pages are requested in parallel under a time budget. Slow, blocked or empty responses are abandoned.
  5. Passages. Whatever loaded is cut into passages and matched against each sub-question.
  6. Answer. A few passages shape the answer. Their URLs become the citations.

The bots behind each engine.

Since 2026 the major labs run separate bots for training and for retrieval. Blocking the training column is a legitimate policy choice with a trade-off: content kept out of training is missing from future models' built-in knowledge. Blocking the other two columns is invisibility. Google-Extended is the exception in the table below: it is a robots.txt token, not a crawler, and it opts you out of Gemini training, not out of AI Overviews.

AI bots by vendor and job, September 2026
VendorTrains the modelBuilds the search indexFetches live for a user
OpenAIGPTBotOAI-SearchBotChatGPT-User
AnthropicClaudeBotClaude-SearchBotClaude-User
PerplexityNonePerplexityBotPerplexity-User
GoogleGoogle-Extended (robots.txt token)GooglebotGooglebot

Slow is not ranked lower. Slow is absent.

Three independent datasets point the same way: an unreliable, slow or JavaScript-dependent page loses its citations, it does not merely rank lower.

Evidence that fetch reliability decides citations
FigureWhat was measuredSource
18xFewer citation events for unreliable pages. Across 700,000 pages, pages that failed on more than 75% of AI fetch attempts earned roughly 18 times fewer citations than stable pages, and many earned none at all.Profound (J. Singh, J. Blyskal), April 2026, reported by iPullRank, May 2026
99%Of HTTP 499 timeouts came from ChatGPT-User. A 499 means the client gave up before the server answered, and almost all of them came from the fetcher that runs while a real person waits. Some sites lost 5% of all ChatGPT visits this way.Oncrawl data, J. Salomon, 2025
0Major AI crawlers that run JavaScript. GPTBot and Claude's crawler download JavaScript files but never execute them, so anything that appears only after client-side rendering (prices, specs, reviews, FAQ answers) does not exist for them.Vercel and MERJ, "The rise of the AI crawler", December 2024

Six technical requirements.

Six requirements define a technical layer that AI engines can use:

  • The server answers fast. Aim for a time to first byte under 200 ms on the pages that matter and treat 500 ms as the ceiling. Serve cached HTML from a CDN edge near your buyers instead of a cold render on every request.
  • The answer is in the first HTML. Server-render or pre-render every page that carries a fact you want quoted. If a number only appears after JavaScript runs, a browser shows it and an AI fetcher does not. Test with curl, not with your browser.
  • AI bots are allowed at every gate. robots.txt is only the first gate. CDN "block AI bots" toggles, WAF rules matching any user agent that contains "bot", and security plugins block silently. In my network roughly 30% of sites blocked at least one major AI bot without knowing it, and none of it showed in Search Console. In one client case, a firewall nobody had configured for AI blocked 2 in 5 of Claude's crawler visits.
  • Rate limits spare the request that matters. A limit tuned for a crawler's burst will also block the single live visit ChatGPT-User makes while someone is asking about you, which is the exact request where a citation was about to happen. Limit by IP, allow bursts, return Retry-After.
  • One host, clean status codes. Pick one host (www or apex) and one protocol, and redirect everything else in a single hop. AI crawlers already waste requests: in Vercel's data about a third of GPTBot and Claude fetches hit 404s, and ChatGPT spent another 14% following redirects.
  • You are in the index the engine searches. An engine can only fetch what its search step returned. Seer Interactive matched 87% of ChatGPT citations to Bing's top results. Verify Bing Webmaster Tools, push changes with IndexNow, and check that your key pages rank outside Google too.

Two files and two commands.

The robots.txt below allows the bots that put you in answers and treats the training bots as the policy decision they are. The two curl commands show what ChatGPT's live fetcher gets from your server, and whether the sentence you want quoted exists before any JavaScript runs. If your CDN has its own AI bot controls, check them too: my Cloudflare piece covers the settings that block silently.

A robots.txt that keeps you in the answers
# Search and live-fetch bots: allow. These are the ones that put you in answers.
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: PerplexityBot
User-agent: Perplexity-User
Allow: /

# Training bots: a policy decision. Allowing is shown; blocking is legitimate.
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: CCBot
Allow: /

User-agent: *
Allow: /
Disallow: /admin/

Sitemap: https://yoursite.com/sitemap.xml
Test it yourself: two commands
# 1. Status code and time to first byte, as ChatGPT's live fetcher sees them
curl -s -o /dev/null -w "%{http_code} %{time_starttransfer}s\n" \
  -A "Mozilla/5.0 (compatible; ChatGPT-User/1.0; +https://openai.com/bot)" \
  https://yoursite.com/pricing
200 0.184s      # good: 200, well under half a second

# 2. Is the sentence you want quoted in the raw HTML, before any JavaScript?
curl -s https://yoursite.com/pricing | grep -c "Plans start at"
0               # bad: the text only exists after rendering
Layer 02.
Content

Written to survive the cut.

An AI engine never reads your page the way a visitor does. It cuts the page into passages, retrieves the few that match the question and writes its answer from those passages alone. The headline, the intro and the paragraph above are gone. A passage that opens with "This is where we're different" has lost whatever "this" was, so the engine drops it or credits it to "one provider".

Four things can happen to your content in an AI answer, from least to most valuable:

  • Absorbed. Your passage supplied the facts. The answer neither names nor links you. This is the default fate of clear but generic copy.
  • Mentioned. Your brand is named in the answer, with no link.
  • Cited. Your URL appears as a source. This is where the click comes from.
  • Grounded and cited. Your passage shaped a specific sentence and is linked right next to it.

Two levers move copy up that ladder.

Extractability (a named subject, the answer first, a passage that stands alone) gets a passage retrieved and used. Distinctiveness (a proprietary number, a dated fact, a named source, a specific mechanism) gets it named. Structure alone earns absorption. Structure plus distinctiveness earns citations.

The full writing guide, with before-and-after rewrites from a real client engagement, is in How AI reads your pages. The short version follows.

The scissors test: five checks for every sentence.

Cut any sentence out of the page, hand it to a stranger and ask what it says. If they have to ask who, which, says who or compared to what, the sentence isn't finished. The examples use Ledgerly, a fictional bookkeeping app.

The five checks, with a failing and a passing sentence for each (Ledgerly is fictional)
CheckFailsPasses
WhoWe're the #1 choice for freelancers.Ledgerly is a bookkeeping app for freelancers and sole traders.
WhatFast, easy setup.Ledgerly connects to a bank account and imports 12 months of transactions in under 10 minutes.
Where and whenAvailable in most countries.Ledgerly supports bank connections in the US, Canada and the UK as of September 2026.
Says whoUsers save hours every month.In Ledgerly's 2026 user survey, freelancers reported saving a median of 4 hours a month on bookkeeping.
Compared to whatCheaper than the big players.Ledgerly costs $9 a month, against $30 for the entry plans of the two largest accounting suites.

One passage, rewritten.

Before, absorbed at best. Why choose us? We've been helping businesses like yours for years. Our platform is fast, easy and built for the way you work. That's why thousands of customers trust us every single day.

After, citable. Who is Ledgerly for? Ledgerly is a bookkeeping app for freelancers and sole traders who file their own taxes. It imports bank transactions automatically, sorts them into tax categories and prepares a year-end summary for a self-assessment return. As of September 2026 Ledgerly supports banks in the US, Canada and the UK and costs $9 a month.

Why the rewrite works. The heading is the question a buyer types. The first sentence names the product and answers it. Every claim is specific enough to check, scoped by country and dated. Cut out alone, the passage still says exactly who, what, where and how much.

Ten passage rules.

Google's guidance for its AI features says there is no need to chunk content or rewrite it "for AI". Nothing below contradicts that. These rules are clear writing: a person skimming the page benefits exactly as much as a retrieval system does, and that is the test I use for every rewrite. Ten rules make a passage survive the cut:

  • Answer first. The first sentence under a heading states the full answer: subject, claim, scope. Explanation comes after it.
  • Headings are questions. Written in the words a buyer actually types, not clever labels like "The Acme Way".
  • One idea per passage, finished. Two questions deserve two passages. Fan-out gives a page that answers three sub-questions three chances; blended, it gets none.
  • Say the name again. Never open a passage with we, our, it, this or they. Name the company or product once, in the first sentence.
  • Nothing points off the page. "As mentioned above" and "see the table below" die in the cut. Restate or remove.
  • Use the exact words. Retrieval matches meaning and exact terms. Model numbers, product names, places and prices belong in the text, not in an image or a script.
  • One distinctive element. Each key passage carries a sourced number, a dated fact, a named mechanism or an honest disagreement with the consensus.
  • Date what goes stale. "Prices confirmed September 2026." Prices, coverage and counts without a date read as unverified.
  • Lists carry their subject. A lead-in sentence names what the list is about, because the items alone are fragments.
  • Don't stuff. Exact terms once per passage. Keyword stuffing measured below baseline in the only controlled study.

What the passage studies found.

Four studies point the same way: put the answer high, phrase headings as questions, add sourced numbers and date what goes stale.

Passage-level citation studies
FigureFindingSource
55%Of AI Overview citations came from the top 30% of the page, in a 100-page study.CXL, 2026
2xQuestion-style headings were cited about twice as often as other headings: 18% against 8.9%.Ahrefs, 2026
+37%Visibility lift from adding statistics. Keyword stuffing scored about 10% below baseline.Aggarwal et al., "GEO: Generative Engine Optimization" (Princeton), KDD 2024
80%Of cited passages carried a visible date, against 53% of uncited ones. Small sample: 112 passages.Advanced Web Ranking, August 2026
Layer 03.
Access

Hand the machine a map.

Before an AI system reads a single article, it reads the shape of the site: the feed, the sitemap, robots.txt and the structured data. In my research network these discovery endpoints took more AI fetcher traffic than the content pages themselves. A broken discovery layer asks the bots to guess, and most of them don't bother.

Where AI fetchers spend their requests, as a share of every logged AI fetcher request by endpoint type across my network (CDN-level logs, 9 industries; the full breakdown is in What AI bots actually read):

RSS / Atom feeds
40%
HTML pages
25%
Sitemap XML
14%
Schema / JSON-LD
9%
Images and media
6%
PDFs and documents
4%
robots.txt and llms.txt
2%

The feed is the most requested endpoint. Treat it like one.

RSS is structured, chronological and cheap to parse, so when an AI system wants to know what is new on a site, the feed answers in a fraction of the bandwidth of crawling pages. Most CMS defaults break it in quiet ways. A feed is only real when it passes the first five checks below; the sixth is the next level.

  • Valid XML. Passes the W3C Feed Validator. No broken CDATA blocks, unencoded characters or duplicate items.
  • Linked in the head. A <link rel="alternate" type="application/rss+xml"> tag on every page, so no bot has to guess the URL.
  • Full content. The whole article in content:encoded. Most CMS defaults ship a 50-word excerpt, which starves the bots.
  • Honest dates. RFC 822 dates that match the page and the schema. Inconsistent dates are worse than missing ones.
  • Not blocked. No Disallow: /feed/ in robots.txt. For AI visibility the feed is your most important public endpoint.
  • Next level: per-section feeds. /research/feed, /news/feed. In my network, sites with per-section feeds drew 1.5 to 2x the AI bot engagement of single-feed sites.

Freshness: visibility decays fast.

AI bot traffic falls off quickly once a site stops publishing. The bars show approximate AI bot traffic against the site's own peak, by time since the last publish, from my research network:

Under 7 days
100%
7 to 14 days
~80%
14 to 30 days
~50%
30 to 60 days
~25%
60+ days
~10%

Real change, not date stuffing.

Sites publishing weekly drew roughly 3x the AI bot traffic of slower publishers in my network. New posts, research, product and service pages are the strongest freshness signal. Substantive updates to evergreen pages count too, as long as dateModified moves because the content moved.

Adding "Updated 2026" to every footer, or bumping dates with no change underneath, is ignored. If weekly publishing isn't realistic, a steady rhythm of genuine updates to the pages that sell is the next best thing.

The sitemap: a complete map with honest dates.

Four properties separate a sitemap that helps from one that teaches bots to ignore it:

  • Every public URL, nothing else. No noindexed, redirected or 404 URLs. A sitemap missing 30% of pages silently removes 30% of the surface an AI bot would otherwise know about.
  • lastmod tells the truth. Many CMSs stamp every URL with "now" on each regeneration. Bots that triangulate freshness learn the signal is worthless and stop trusting it. In one client case, all 956 pages of a site carried the same date.
  • Declared in robots.txt. One line: Sitemap: https://yoursite.com/sitemap.xml. Large sites split it into a sitemap index by section.
  • No orphans. Every URL in the sitemap is linked from at least one other page, so bots understand where it sits in the site.

Schema: structured truth, stated once, stated right.

Schema is only about 9% of AI fetcher traffic in my network, but it is the densest source of facts on a page. When an engine needs a company name, an author, a price or a publication date, it reaches for the JSON-LD first, and it repeats whatever it finds there, mistakes included. A misspelled Organization name in schema turns up misspelled in answers.

Six rules make schema something an engine can trust:

  • JSON-LD only. About 95% of the structured data AI fetchers extracted in my network was JSON-LD. Microdata and RDFa are legacy.
  • Five types first. Organization, Person, WebSite, BreadcrumbList, Article. Then Product and Offer, a LocalBusiness subtype or SoftwareApplication where they fit.
  • One @graph, stable @id. Nodes reference each other by @id, so engines merge the same entity across every page instead of meeting a stranger each time.
  • Server-rendered. Schema injected by JavaScript validates in your browser and does not exist for fetchers that never run it.
  • Mirrors the visible page. Schema saying $99 on a page saying $79 forces the engine to pick one, or to trust neither. One client's product schema carried a placeholder price that matched no real price on the site.
  • One identity everywhere. The same name, URL and logo, plus sameAs links to real profiles, so engines connect the brand to the right domain.
The pattern: one graph, every node linked
{ "@context": "https://schema.org", "@graph": [
  { "@type": "Organization", "@id": "https://yoursite.com/#organization",
    "name": "Your Company", "url": "https://yoursite.com",
    "sameAs": ["https://www.linkedin.com/company/your-company"] },
  { "@type": "Person", "@id": "https://yoursite.com/#author", "name": "Author Name" },
  { "@type": "WebSite", "@id": "https://yoursite.com/#website",
    "publisher": { "@id": "https://yoursite.com/#organization" } },
  { "@type": "Article", "headline": "Who is Ledgerly for?", "datePublished": "2026-09-14",
    "author":    { "@id": "https://yoursite.com/#author" },
    "publisher": { "@id": "https://yoursite.com/#organization" } }
] }

Six type families do almost all the work.

Six type families (Organization and Person, Article, Product and Offer, FAQPage, BreadcrumbList, Recipe) account for about 97% of the AI schema engagement across my network. The other 700+ schema.org types share the remaining 3%. The breakdown by type is in Schema markup that AI models actually use.

Three things do nothing measurable. llms.txt drew zero requests from any AI fetcher across my network in 200+ days of logs, and Google confirmed in May 2026 that it has no effect on its AI features: harmless to keep, pointless to pay for (the full data). Speakable, HowTo and ClaimReview (unless you are a verified fact-checker) showed no measurable AI engagement. "Schema everything", fifteen types per page, is noise: fewer types deployed correctly beat more types deployed carelessly.

Layer 04.
Citations

Go where the answers come from.

~30 domains

About 30 domains capture roughly two thirds of ChatGPT's citations within a topic, according to Kevin Indig's analysis of about 98,000 citation rows (Gauge, 2026). The list you need to be on is short. You just have to find it.

Layers 01 to 03 make a site citable. They don't make it cited. For most buying questions ("best X for Y", "X vs Y", "alternatives to X") AI engines lean on a short list of third-party pages they already trust: roundups, review platforms, directories, forums and videos. Across my client work the pattern repeats: a brand's own site wins questions about the brand and head-to-head comparisons, while broad category answers come almost entirely from third-party sources. The only way to know which sources is to measure them.

Five steps from prompts to a target list.

The method I use to find that list runs in five steps, in this order:

  1. Choose the prompts. 20 to 100 real questions, weighted toward the ones that decide a purchase: "best X for Y", "X vs Y", "is X worth it", "alternatives to X", "X in [city or industry]". Seed them from three places: your Search Console queries, the questions sales and support hear every week, and a model asked what someone would type to find each of your key pages. Keep only prompts that surface real competitors, and group them under the handful of topics you must be visible for.
  2. Run them on a schedule. Every day, or twice a day, for two to four weeks, in ChatGPT, Perplexity, Google AI Mode and AI Overviews, Gemini and Claude, always with web search switched on. The same prompt returns different sources from run to run. One run is an anecdote; a month of runs is a sample.
  3. Log every answer. One row per citation: date, engine, prompt, cited URL, cited domain, the brands named and their order. Keep the full answer text too. A spreadsheet works for the first week. I run mine as a scheduled script that writes every answer to cloud storage, because a month of daily runs across five engines is thousands of rows.
  4. Build the source map. Group the citations by domain and count them. The list collapses fast: for most topics 20 to 50 domains carry the bulk of the citations. Mark where your own site appears, where competitors appear, and which sources cite competitors but not you. That gap list is the work plan.
  5. Get into the sources, then re-measure. Classify each domain by source type and work it with the method that fits that type. Re-run the same prompt set every month and compare against the baseline, so every change in visibility has a before and after.
Step 3 in practice: one row per citation
date        engine      prompt                           cited_domain      brands_named
2026-09-14  chatgpt     best bookkeeping app freelancer  review-site.com   A, B, you
2026-09-14  perplexity  best bookkeeping app freelancer  reddit.com        B, A
2026-09-14  ai-mode     ledgerly vs competitor-a         yoursite.com      you, A

Source types and how to earn a place.

Seven source types show up in AI citations, and each has its own way in:

Where AI engines find their sources, and how a brand gets into each
Source typeWhat it looks likeHow to get in
Roundups and listicles"Best X in 2026" pages on publishers and niche blogsContact the author with a real reason to include you: data, a trial account, a clear differentiator. Make your facts easy to verify on your own pages.
Review platformsG2, Capterra, Trustpilot, Tripadvisor, vertical review sitesA complete profile, accurate categories and a steady flow of genuine reviews.
Directories and registriesAssociations, marketplaces, local directories, regulator listsClaim every listing and keep the name, URL and core facts identical everywhere.
CommunitiesReddit, niche forums, Q&A sitesGenuine participation by real people from the company. Seeded fake posts are spam: Google names them explicitly and they get filtered.
VideoYouTube reviews, tutorials, explainersYour own channel, plus the creators who already appear in the citations.
Press and referenceNews, trade media, Wikipedia where you are genuinely notableDigital PR built on original data that journalists can quote.
Your own siteComparison, alternatives, pricing and FAQ pagesThe pages engines cite for brand and head-to-head questions, built to layers 01 to 03. Publishing real prices took one client from a tenfold spread in AI price answers to 3 of 3 correct.

What to track.

Five metrics show whether the work is landing. Mention rate: the share of answers that name you. Citation rate: the share that link to your site. Linked mention: named and linked in the same answer, the best outcome there is. Share of voice: your mentions and citations against named competitors, per topic. Position: where you sit in list-style answers.

Weight the deal-deciding questions. Mention rate on vague prompts is cheap to inflate; what matters is what the models say when someone asks the question that decides the purchase. Why single runs and most dashboards mislead is covered in AI citation tracking tools are mostly noise.

Each engine needs its own map.

In Writesonic's 161,286-prompt sample only about 17% of cited sources were shared across engines, and under 4% appeared in all of them. A source that dominates Perplexity can be invisible in ChatGPT, which is why step 2 runs every engine and step 4 builds a map per engine, not one blended list.

Fix it in order.

Skipping ahead is the most expensive mistake in AI visibility. Outreach for a site that times out earns citations to a page no engine can load. Rewriting content on a site that blocks ChatGPT-User changes nothing. Work the layers from the bottom up and re-measure after each one. The checklist, one layer at a time:

Layer 01: technical

  • Time to first byte under 200 ms on key pages, never above 500 ms
  • Key facts present in the raw HTML (the curl test)
  • AI search and user bots allowed in robots.txt, CDN and WAF
  • Rate limits by IP, bursts allowed, Retry-After returned
  • One canonical host, internal links resolve 200 on the first hop
  • Indexed and ranking in Bing, Google and Brave

Layer 02: content

  • Every heading is a question buyers actually type
  • The first sentence under each heading names the subject and answers
  • No passage opens with we, it, this or they
  • One distinctive, sourced element per key passage
  • Facts that go stale carry a visible date
  • Exact product, place and model names in the text

Layer 03: access

  • Valid, full-content RSS, linked in the head, not blocked
  • New or substantively updated content every week
  • Sitemap complete, honest lastmod, declared in robots.txt
  • JSON-LD in one @graph with stable @id, server-rendered
  • Organization, Person, WebSite, BreadcrumbList, Article, plus vertical types
  • Schema matches the visible page exactly

Layer 04: citations

  • 20 to 100 deal-deciding prompts, grouped by topic
  • Daily runs across five engines for two to four weeks
  • Every answer and every cited URL logged
  • Source map built: the top 20 to 50 domains per topic, per engine
  • Gap list: sources that cite competitors but not you
  • Monthly re-run against the baseline

No one can promise a citation.

I don't guarantee AI citation outcomes, and anyone who does is either uninformed or lying. The models are opaque, non-deterministic and change every quarter. What the four layers do is remove every reason for an engine to skip you, and then measure whether it stopped.

Where the numbers come from.

Third-party figures are reported as published by their authors and have not been replicated on my own network unless stated. AI answer engines change quarterly, and every figure here is dated for that reason. The sources, in the order they appear:

  • Misha Manko, AI visibility research network: 50 sites, 9 industries, 200+ days of CDN-level AI bot logs, 2026.
  • Google Search Central, guide to generative AI features in Search, May 2026; query fan-out, Google I/O 2025.
  • Profound dataset (J. Singh, J. Blyskal), 700,000 pages, April 2026, reported by iPullRank, May 2026.
  • Oncrawl data and HTTP 499 analysis, J. Salomon, 2025.
  • Vercel and MERJ, "The rise of the AI crawler", December 2024.
  • Seer Interactive, ChatGPT citations vs Bing and Google rankings, 500+ citations, 2025.
  • CXL, 100-page AI Overview citation position study, 2026.
  • Ahrefs, question-style headings and citation rates, 2026.
  • Aggarwal et al., "GEO: Generative Engine Optimization", KDD 2024.
  • Advanced Web Ranking, passage-level citation coding, 112 passages, August 2026.
  • Kevin Indig / Gauge, ChatGPT citation concentration, ~98,000 citation rows, 2026.
  • Writesonic, cross-engine citation overlap, 161,286 prompts, 2026.
Misha Manko, independent AI visibility researcher
Written by

Misha Manko

Independent researcher, AI visibility and technical SEO

I run a 50-site instrumented research network measuring how ChatGPT, Claude, Perplexity, and Google AI actually read websites - 200+ days of bot logs, real client engagements, numbers over claims. Everything I publish comes from measurement, not opinion.

Want the four layers run on your site?

Four layers, measured on your site.

The AI Visibility Audit works through all four layers with your real bot logs, your pages and 30+ live citation queries against your competitors, then ranks every fix by impact, effort and urgency.