A MACHINE WITH SCISSORS.
When ChatGPT, Google's AI Overview, Gemini, or Perplexity answer a question in your category, they do not open your website and read it. They search an index of passages, pull the ones that best match the question, and build an answer from those passages only. The headline, the paragraph above, the hero image, the sidebar: gone. Whatever survives the cut is all the engine knows.
This piece is an internal writing guide I delivered to a consulting client in September 2026 - a virtual-first pediatric therapy provider operating across 47 US states. Their name is replaced with Acme Health throughout and the clinical details are lightly disguised; the audit numbers, the before-and-after structure, and every sentence pattern are real. Before publishing I ran every claim in the guide through an evidence review against the engines' own documentation and the 2026 passage-level citation studies; where the evidence refined a claim, this version carries the correction.
One guardrail before the mechanics, because Google says it in plain words in its own generative-AI guidance: you do not need to chunk your content for AI, and you should not rewrite your site for robots. Google retrieves and ranks at the passage level internally; it explicitly tells authors not to fragment content to match. Nothing below is a chunking hack. It is a standard of clear writing - name your subject, state checkable facts, finish each thought where it starts - that humans have always rewarded and that passage retrieval now rewards twice.
The whole guide compresses into one test. Cut any sentence out of the page with scissors, hand it to a stranger, and ask what it says. If the stranger has to ask "who?", "which?", "says who?" or "compared to what?", the sentence is not finished.
THE PAGE IS CUT FIRST, MATCHED SECOND.
When an AI engine's systems fetch a page, the text is treated as passages, not as a whole. This is vendor-documented, not speculation: Google's generative-AI guide describes retrieval-augmented generation over its core search index, and Gemini's developer API exposes grounding metadata that maps specific segments of a generated answer back to specific source chunks by character position. The useful unit is short - practitioner measurements cluster somewhere between roughly 50 and 300 words, most often one heading plus the paragraphs under it - though the exact figure varies by study and should be read as a range, not a rule. What matters is what the passage does not contain: the rest of the page. A chunk that says "This is where Acme Health is fundamentally different" has lost whatever "this" was.
Retrieval then scores each passage on its own, with a mix of two methods. One is meaning-based, so a passage about "online therapy" can match a question about "virtual therapy". The other is word-based - Anthropic's own retrieval engineering combines semantic embeddings with BM25 keyword matching precisely because exact terms catch what embeddings miss. Use the exact words a user would type, at least once per passage. But do not stuff: the Princeton GEO study measured keyword stuffing at roughly 10% below baseline.
Google adds query fan-out: one user question becomes several sub-questions, each retrieving its own passages - confirmed on stage at Google I/O and in the documentation. A page that answers three sub-questions in three separate, self-contained passages gets three chances. A page that blends them into one long paragraph gets none.
MOSTLY REWRITTEN, SOMETIMES QUOTED.
The engines are less extractive than the industry assumes. The largest 2026 passage-level study reverse-mapped 2,422 AI-cited sentences back to their sources and found that 76% could not be traced to any single source passage - the engine had synthesized across sources and rewritten. Where engines do quote, the median quoted chunk is about 25 tokens, roughly 19 words. Google's AI Overview is the one surface that quotes long - median around 50 tokens, occasionally up to 500 - which is why it feels extractive. Everyone else takes your facts and rewrites your sentences.
This changes what "quotable" means. You are not writing sentences the engine will reprint; you are writing facts the engine can carry - and attribute. A vague passage does not produce a vague quote. It produces nothing, because there is no fact in it to carry.
There are four outcomes, and they are worth different amounts. Grounded means your passage shaped the answer. Cited means your URL appears as a source link - where the click comes from. Mentioned means your brand appears in the text with no link. And absorbed - the failure mode - means your page supplied the information and received no credit at all. The next finding is about what separates cited from absorbed.
Where the client stood: in the 34-query sweep, Acme Health was cited on 6 of 13 category questions - and in every single one, the cited URL was the homepage, not the relevant service page. The homepage is the only page whose opening block states what the company is, where, and for whom, in plain sentences. The service pages describe; they do not state. That gap is what the rest of this guide closes.
EXTRACTION GETS YOU USED. DISTINCTIVENESS GETS YOU NAMED.
An August 2026 study hand-coded 112 passages behind 265 citations in Google AI Overviews and Bing Copilot, comparing passages that earned a citation against passages that fed answers without one. The result cuts against most of the advice sold as GEO: extractive structure did not separate cited from uncited (76% vs 71% - noise), and definition-style formatting did not either. Clean structure gets your information absorbed into the answer. It does not get your name attached.
What did separate them: a visible fresh date (80% of cited passages had one, against 53% of uncited), a named entity at first mention (96% vs 82%), and escaping consensus - passages that merely restated what every other source said were absorbed anonymously, while passages that added something the others did not were named. Sharpest of all: zero of the seventeen uncited passages contained a hard number or a novel claim. Not one.
The practical rule this adds to everything below: every passage you care about should carry one distinctive element - a proprietary number, a named case, a dated measurement, or a documented disagreement with the consensus. The Princeton GEO experiments point the same direction from the other side: adding citations to sources lifted visibility by about 40%, statistics by about 37%, quotations by about 22%. Structure is the entry ticket. Distinctiveness is what the engine attributes.
FOUR QUESTIONS EVERY SENTENCE MUST SURVIVE.
Take the sentence a lot of companies write: "We are #1 in this industry." Read alone, it fails four ways. Who is "we"? Number one at what? Which industry? Who decided? The engine cannot answer any of those from the chunk, so it cannot use the sentence as a fact. It will either drop it or rewrite it as "one provider claims to be the largest". The finished version of the same claim reads: "Acme Corp is the #1 marketing agency in the United States by client revenue, according to the 2026 Agency Index." Same claim - now with a named subject, a scope, a measure, and a source. That sentence can be carried, attributed, and checked.
- WHO. The subject is named. "Acme Health", "your assigned clinician", "Texas Medicaid". Never "we", "it", "this", "they", or "the program" as the first word of a passage.
- WHAT. The claim is specific enough to be false. "Starts most families within 90 days" can be checked. "Fast starts" cannot.
- WHERE / WHEN. The scope is stated. In which states, for which service line, in which year, for which age group.
- SAYS WHO. The source is named in the same sentence or the next. A study, a report, a regulator, the company's own data with the year.
SAY THE NAME AGAIN.
A pronoun at the start of a passage is the single most common reason a good paragraph is unusable once it is cut out of the page. This is the best-scienced claim in the whole guide: the retrieval literature calls it the anaphoric reference problem - a chunk that opens with "it" or "we" produces a muddy embedding that fails to match queries about the named entity - and a family of 2025-2026 papers builds coreference-resolution machinery specifically to repair it, reporting 10-20% retrieval gains on entity queries. The live-answer data agrees: in the passage study above, 96% of cited passages named their entity at first mention. The fix costs nothing and is unglamorous: say the name again.
The natural worry is that repeating the company and service name will read as stiff. The rule is once per passage, in the first sentence - then pronouns are fine for the rest of that passage. The human reading the page sees a paragraph that starts with the name; the machine sees a chunk that knows what it is about. Both are served. Twice in one paragraph is already too much - this is not keyword density, it is subject-labelling.
BY A MACHINE, WITH NO MEMORY OF THE SENTENCE BEFORE IT.
EIGHT REWRITES FROM LIVE COPY.
Every "before" below was live on the client's site. Every "after" follows their claims register: real figures only, superlatives with their mechanism, third-party findings labelled as such. Names disguised, patterns intact.
The superlative with no subject.
Before: "#1 Largest Virtual Network in the U.S."
After: "Acme Health operates the largest virtual-first pediatric therapy network in the United States, with licensed clinicians in 47 states. The network was built for virtual delivery from the start rather than added to a clinic chain, which is why capacity is not limited by clinic seats."
The original is a label, not a sentence. Cut out of the page it has no subject, no scope, no cause.
"We" at the top of the page.
Before: "Whatever fits your family, care starts without a waitlist."
After: "Acme Health starts most families in virtual therapy within 90 days of the first call, with no waitlist. Acme's intake team runs the assessment and the insurance authorization remotely, so a family is not waiting on a clinic seat to open."
"Care starts without a waitlist" has no subject and no cause. The rewrite attaches the figure and the mechanism - and keeps the comparison about other providers in a separate sentence, so each can be carried on its own.
Bare numbers in stats tiles.
Before: "2x Increase in utilization · 87% Overall goal success rate · 127% Improvement in the first 20 weeks"
After: "In Acme Health's published outcomes research, children whose parents took part in every session showed a 127% improvement against their therapy goals within the first 20 weeks, and 87% of goals were met overall. Acme attributes both results to parent involvement in each session rather than to the delivery format."
A number in a stats tile has no sentence around it, so a chunk containing it says "127% improvement" and nothing else. The tiles can stay as design; the sentence under them is what the engine reads - and a passage with a sourced number is precisely what the citation studies say gets named rather than absorbed.
The comparison with no other side.
Before: "Children receiving our online therapy showed stronger gains than conventional delivery, across the skill areas that matter most to families."
After: "In a peer-reviewed study published in [year], children receiving Acme's virtual therapy with a parent present in every session made larger gains than a comparison group in conventional in-clinic care, measured on [instrument] across communication, daily living and social skills."
"Stronger gains" and "skill areas that matter most" are not measurable. Name the study, the year, the comparison group, the instrument.
The "yes" that lost its question.
Before: "Yes, we require that a legal parent or guardian is present during all sessions."
After: "A legal parent or guardian must be present for every in-home session with Acme Health. The requirement is part of the clinical model: the parent watches the technician's approach and repeats it between sessions, which is where most of the progress happens."
An FAQ answer is retrieved without its question. "Yes, we require" tells the engine nothing about what is required, by whom. Every FAQ answer restates the subject in its first sentence.
The benefit stack.
Before: "Backed by the Nation's Largest Network: faster starts, uninterrupted sessions, and care that's always within reach."
After: "Acme Health's clinical network covers 47 states, so when a family's clinician is unavailable, another licensed clinician in the same state can cover the session. That is what keeps sessions from being cancelled and what lets Acme start families without a waitlist."
Three unsupported claims in one chunk became one claim with its cause. A benefit stack reads as marketing; a mechanism reads as a fact.
The team with no titles.
Before: "Our team handles verification and authorization so you can focus on your family, not paperwork."
After: "Acme Health's intake team verifies a family's insurance coverage and obtains the therapy authorization from the plan before the first session. The family does not submit paperwork to the insurer."
"Our team" fails the who test twice: which company, which team. Name both, say what is done and in what order.
The title tag that leads with the brand.
Before: "Acme Health | Virtual Therapy"
After: "Virtual Pediatric Therapy | No Waitlist, Start in as Few as 90 Days | Acme Health"
The title is the one line that is almost always inside the retrieved unit. A title that starts with the brand spends its first words on the one thing the engine already knows. Lead with what the page answers; the brand goes last. (This one is practitioner heuristic rather than vendor-documented - flagged honestly - but the logic holds.)
RULES FOR THE PASSAGE.
The sentence test is the foundation. These rules are about the paragraph and the section - the unit the engine actually stores. None of them is a formatting trick; each one is ordinary editorial discipline that also happens to survive retrieval.
- One idea per passage. A passage is a heading plus the paragraph or two under it. If a paragraph answers "does insurance cover it" and then drifts into "how long does authorization take", split it - those are two different questions, each retrieved by a different query. Practitioner ranges for a well-sized passage cluster around 40 to 150 words; treat that as a range, not a rule.
- Answer first, then explain. The first sentence under a heading states the answer in full: subject, claim, scope. The position data backs this hard: in the citation studies, the top third of a page earns over half of all citations, and the model-level research ("lost in the middle") shows retrieval quality degrades for content buried mid-page.
- The heading is the question. A heading in the user's own words ("Does Texas Medicaid cover virtual therapy?") matches the query on both meaning and exact words. Clever headings ("The Acme Way", "Outcomes that hold up") match nothing.
- Nothing points off the page. "As mentioned above", "see the table below", "the same applies here" all point at content that is not in the chunk. A list needs a lead-in sentence that carries the subject; a table needs a caption sentence.
- Use the words the user typed. Retrieval is part exact-word matching. Spell out state names and plan names. Introduce a technical term with its plain meaning in the same sentence, so the passage matches both vocabularies - and stop there; stuffing measurably backfires.
- One distinctive element per key passage. A proprietary number, a named case, a dated measurement, or a documented disagreement. Consensus restatement gets absorbed without credit; distinctiveness is what gets named.
- FAQ answers are complete on their own. Two to four sentences, the first of which restates the subject and gives the answer. Never "Yes", "Absolutely" or "It depends" as the whole first sentence. On one of the client's pages, seven of thirteen FAQ answers opened with "Yes", "Absolutely", "This" or "That". (And skip the FAQ format entirely where the content is not genuinely question-shaped - forced FAQ blocks are cargo cult.)
THREE FINDINGS FROM THE SWEEP.
The homepage is the only page that gets quoted. Across the 13 category questions where Acme was cited, the cited URL was the homepage every time. The homepage opens with a block of plain facts with a subject; the service pages open with a headline and a form. Rewriting the top of each service page as an answer block is the single highest-value copy change on the site.
A competitor won on a named study. On one question, ChatGPT chose a competitor over Acme and cited that competitor's June 2026 peer-reviewed study by journal name. Acme has published research too - but the site describes it as "stronger gains across the skill areas that matter most", which gives the engine nothing to attribute. A study with a journal, a year and a number is a citation. A study described in adjectives is not. I have seen the same pattern from the other side: the page that states facts plainly becomes ChatGPT's favourite source even when Google ignores it.
AI Overviews do not repeat bare promises. In the captured Google AI Overviews for access questions, no provider's "no waitlist" claim was repeated as a fact; the overviews talk about waitlists in general. A promise becomes usable by an engine only when it is written as a fact: subject, figure, and mechanism in the same sentence.
WHAT THE ENGINE REJECTS, WHAT IT USES.
Seven patterns, both sides:
- Subjectless label ("#1 Largest Network") - dropped, or attributed to "one provider". Used: a named subject with scope and cause.
- Bare number in a stats tile - quoted without meaning, or skipped. Used: the number inside a sentence with source and mechanism.
- FAQ answer starting "Yes, we..." - the engine cannot tell what was answered. Used: an answer that restates the subject in sentence one.
- "This", "it", "our model" opening a section - the chunk has no topic. Used: company and service named in sentence one.
- Adjective comparison ("stronger gains") - not a fact, not carried. Used: a measured comparison with the other side named.
- Promise without cause ("no waitlist") - left out of the answer. Used: the promise with its figure and mechanism in one sentence.
- Consensus restatement (what every source already says) - absorbed without credit. Used: a passage with a dated, distinctive fact the other sources do not have.
THE PRE-PUBLISH CHECKLIST.
Run this on any page before it ships. Each item is a yes or no.
- 1. Read the first sentence under every heading on its own. Does it name the company or service and state the answer?
- 2. Does any section, FAQ answer or list start with "we", "our", "it", "this", "they", "yes" or "absolutely"? Replace the pronoun with the name.
- 3. Does every number sit inside a sentence that names its source and mechanism? Tiles and stat rows need a sentence underneath carrying the same figure.
- 4. Does every superlative carry its cause in the same sentence or the next? "Largest" needs "because".
- 5. Does every comparison name the other side and the measure?
- 6. Is every scope stated - which states, which service line, which year, which age group?
- 7. Does each passage make one point, roughly a heading plus a paragraph or two?
- 8. Is each heading a question a user would type, or a plain statement of the section's answer?
- 9. Does anything point off the page ("as mentioned above", "see below")? Restate it or cut it.
- 10. Do the exact words a user would search for appear in the passage, with each technical term glossed once?
- 11. Does every page state the same number for the same fact? The engine reads all pages and picks one. Two numbers for one fact means the engine chooses.
- 12. Does each key passage carry one distinctive element - a proprietary number, a named case, a dated claim - or is it restating the consensus?
- 13. Is a promise stated anywhere without its mechanism? Say what the company controls and what it does not.
WHERE THIS COMES FROM.
The mechanism is vendor-documented: Google Search Central's generative-AI guidance (retrieval from the standard search index, query fan-out, and - respected throughout this guide - the explicit instruction that chunking content for AI is not needed); Gemini's grounding API, which maps answer segments to source chunks by character position; and Anthropic's Contextual Retrieval engineering, which documents the hybrid semantic-plus-keyword retrieval that makes exact terminology matter.
The citation behaviour comes from the 2026 passage-level studies: the VisibilityStack reverse-mapping of 2,422 cited sentences (76% untraceable to any single passage; median quote about 25 tokens; AI Overviews the longest quoter), and the Advanced Web Ranking hand-coding of cited versus uncited passages (extractive format separating nothing; visible dates, named entities and hard numbers separating everything - including zero uncited passages containing a hard number). The Princeton GEO experiments (KDD 2024, 10,000 queries) supply the lift figures for sources, statistics and quotations, and the coreference literature supplies the science behind the pronoun penalty. Where studies disagree - word counts, format effects, per-engine citation counts - this guide states ranges and treats the numbers as dated, not permanent.
The client-specific findings come from a 34-query, 49-run AI sweep and 16 live SERP captures from the September 2026 engagement, cross-checked against my own 50-site research network's field notes. Two honest limits: AI answers are non-deterministic and vary by geography and personalization, and citation sits behind retrieval as a second, opaque filter - no on-page technique guarantees a citation. What the techniques above change is whether the engine has anything of yours worth carrying, and whether your name travels with it.
The door has to be open before any of this matters - a perfectly written page that AI readers cannot fetch is still invisible. That layer is covered in why the edge decides whether AI can read you.