Table of Contents
TL;DR
- the chats extract passages, not pages. They look for a self-contained answer block in the first part of the page, with definitive language. Indig’s analysis of 1.2M ChatGPT responses found 44.2 percent of citations come from the first 30 percent of page content.
- The answer-capsule format — a 40 to 60 word self-contained answer block under each H2, with no links, followed by a supporting paragraph — appears in 72.4 percent of cited blog posts. The format is the difference more often than the topic is.
- Question-based H2s that mirror the words customers actually type into chat surfaces are one of the strongest content signals. High heading-query match correlates with a 41 percent citation rate versus 30 percent for low-match pages.
- Definitive language outperforms hedging. Cited pages use definitive statements 36.2 percent of the time versus 20.2 percent for non-cited pages. the chats are risk-averse. They cite the source that commits.
- E-E-A-T has reshuffled. Experience leads. A small site with a named author and first-hand experience can outcite a bigger site with neither, because the engines prefer safer sources to bigger ones.
You ask Claude about your topic. Three websites come back as the cited sources. None of them is yours.
The frustrating part is that your competitors are not always bigger. They are not always better. They sometimes write less than you do, with thinner expertise behind it.
They get cited because they have stumbled into a format the chats can extract. You do not, because nobody told you the format matters more than the depth.
This piece is the short version of how the chats actually decide which sites to cite. It also covers what you can change on a page tomorrow morning to land on the cited side. The answer-capsule format and why it works. The question-based H2 rule and how to discover the right questions.
The definitive-language signal that most writers fight against. The E-E-A-T reshuffling that gives a small site a real chance against a big one. And the eight predictable reasons most pages never get cited, each one fixable in an afternoon.
Why does ChatGPT cite some sites and not others?
The chats extract passages, not pages.
They look for a self-contained block of writing that answers the question directly, in the first part of the page, with definitive language. The page as a whole matters less than people assume. The block matters more.
A site that buries the answer inside paragraph three of section four loses every time. The winner puts the answer in a clean block immediately under the heading. The winner can know less about the topic and still cite first.
Indig’s February 2026 analysis of 1.2 million ChatGPT responses put a number on this. Around 44.2 percent of citations come from the first 30 percent of page content. The format decides the citation more often than the topic does. The same article rewritten with the answer in front cites at multiples of the same article with the answer in the middle.
The implication for a site owner is small but specific. The competitive advantage is not always more knowledge or more words. The competitive advantage is putting the knowledge you already have into a format the engines can read.
What is the answer-capsule format and why does it work?
The answer-capsule format is a 40 to 60 word self-contained answer block, placed immediately under each H2 heading.
No links inside the block. No introductions. No "let me explain" preamble.
Just the answer, written so it stands on its own, in 40 to 60 words. The supporting paragraph follows. Total section length around 120 to 180 words.
The answer-capsule pattern appears in 72.4 percent of cited blog posts in Indig’s analysis. The honest caveat is that the 40 to 60 word range comes from practitioner consensus across multiple sources, not a single rigorous study. The structural rule holds. The exact numbers should be marked as guidance rather than law.
The mechanism is straightforward. The engines extract the block whole because it stands on its own. A reader who saw only that block would still get the answer, but a reader who saw a paragraph mid-section would not.
The engines treat each block as an independent ready-to-cite unit. A page with eight ready-to-cite blocks under eight H2s has eight chances to be cited — a page of unbroken prose has none. The same answer-block discipline underlies structuring content for featured snippets, the older sibling pattern that still drives Google traffic.
The format change is small. The downstream effect is large.
Take any page you wrote in the last six months. Put the first sentence of the most useful paragraph at the top of its section, in a tight self-contained block. The page is now ready to cite in a way it was not before.
How do question-based H2s actually change your citation rate?
Question-based H2s match the words customers type into chat surfaces.
The match between heading text and customer query is one of the strongest content signals across the chats. Pages with high heading-query match cite at 41 percent versus 30 percent for pages with low match in Indig’s data. The bigger the gap between heading text and the actual question, the lower the citation rate. The good answer underneath cannot save a poorly-matched H2.
The practical method takes ten minutes per page. Check the People Also Ask box on Google for the topic. Open ChatGPT and ask the question yourself, and note the suggested follow-ups it offers.
Read the actual customer questions in your inbox or chat history. The phrasings that recur are the ones to mirror as your H2 text.
Not paraphrases. Not the conceptually-equivalent corporate version. The exact phrasings, in the exact words.
The trap is the temptation to write H2s that describe the topic, like "Pricing Strategy". The mirror version is "How much should you charge for a small project?". The first feels professional.
The second cites. The engines have no patience for branded section labels. They reward heading text that looks like the question the user just typed in.
Why does definitive language matter for chat-surface citation?
Cited pages use definitive statements 36.2 percent of the time versus 20.2 percent for non-cited pages in Indig’s analysis.
Hedging phrases signal low confidence. "It appears that." "One might argue." "It could be the case that." These read as careful and academic to a human reader. They read as low-trust to a chat surface.
The engines are risk-averse. They prefer a source that says X is true to a source that says X might possibly be true under some conditions. The preference holds even when both sources are equally correct.
Writers trained on academic hedging or corporate caution have to unlearn it for AEO. The retraining is uncomfortable at first. The right discipline is to ask, of every hedging phrase, whether the hedge is honest or reflexive.
Honest hedges — where you genuinely do not know, or where the answer depends on a specific context — stay. Reflexive hedges — where you actually know the answer but the writing has been softened to feel polite — come out. The cleaned page reads more confident, and the engines reward the confidence with citations.
Certainty, where you can honestly claim it, is a structural signal. A page that commits beats a page that equivocates, even when the page that equivocates is technically more careful. The careful version belongs in the supporting paragraph. The answer block belongs in the language of certainty.
How do the chats decide which sources are trustworthy enough to cite?
The chats combine three signals — structural readability, authority signals, and freshness.
Around 96 percent of AI Overviews citations come from sources with strong E-E-A-T signals. The engines have learned to be cautious about who they quote. The cost of citing a wrong source is higher for them than the cost of returning a less-detailed answer.
They cite what is safest to repeat. That makes E-E-A-T the gating signal once the structural readability is in place.
The reshuffling matters here. Experience now leads. Expertise is second.
Trustworthiness is third. Authoritativeness — the link-equity signal traditional SEO rewarded for two decades — has dropped to last. The change is real and measurable.
A small site with a real named author who has first-hand experience and a custom-built author page can outcite a big site with neither. The engines prefer safer sources to bigger ones. The site article on the topical authority a small website can build covers the cluster-level work. The cluster work compounds the per-page citation signal once the format and the named author are in place.
The named-author point matters disproportionately. Most sites publish anonymously. The competitive bar is low.
A site that adds a real author byline, a real custom author page, Person schema, and consistent cross-platform mentions clears that bar quickly. The citations follow within weeks rather than months.
What are the most common reasons a page does not get cited?
Eight predictable reasons account for most missed citations.
The answer is buried below paragraph three. The page is one wall of text with no scannable sections. The H2s describe topics rather than questions.
The writing hedges where it could be definitive. There is no clear named author with a real bio. The page is more than a year old with no update or refresh.
The site renders the body content in JavaScript that the chat crawlers cannot read. The structured data is missing or shallow.
Each one is fixable in an afternoon. The combined effect of fixing all eight on a single page is usually visible within a month. The four-surface check starts returning the page as a cited source on a query that previously cited a competitor.
The fix sequence is straightforward. Format first — answer capsule and question-based H2s. Then language — cut the hedging.
Then attribution — add the named author. Then freshness — update the publication date once the changes are real. Then technical — confirm the body is server-rendered and the markup is present.
The mistake is treating the eight as separate small projects. The page is a single artifact — the engines read the whole picture. Fix one or two and the citation rate barely moves.
Fix all eight and the rate moves enough to be measurable on a monthly check. The investment is one focused afternoon, not eight scattered ones.
How do you check whether a single page is being cited and by which engines?
Run the four-question prompt rotation against the four major chat surfaces.
Use the actual question your customer asks, in their wording, not a paraphrased version. Ask the same question on each surface. Note whether your domain comes back in the cited sources for each one.
Repeat the check monthly with a small notebook of questions and surface results. The four-surface variation matters because the four indexes produce different source pools. A page cited on one surface is often missed by the other three.
Optimizing for one surface without checking the other three leaves visibility on the table. A page tuned only for one chat surface may miss another entirely. The optimization patterns overlap but they do not match. The four-surface check sets the realistic visibility picture for any one page.
Your customers do not all look in one place. Some ask a chat surface, some search a map, some scroll a social feed. The idea behind search everywhere optimization explained is to be findable wherever they happen to look.
Other questions worth answering
How does AI Mode pick targets differently from AI Overviews?
AI Mode and AI Overviews draw from different pools. The two Google products share only about 13.7 percent of cited URLs on the same prompt. The Overviews product lifts passages from the standing Google index. The Mode product runs a query fan-out that splits one prompt into several smaller searches.
Does cross-platform consistency in an author bio shift AI trust scores?
Yes, and the effect is bigger than most owners notice. The engines use external profiles — LinkedIn, industry directories, external bylines — to build entity confidence in an author. If the name, the credentials, and the bio do not match across the surfaces, the entity score drops. Consistent presentation is a quiet trust amplifier.
How does freshness of an update affect AI mentions?
A 30-day freshness window measurably lifts mentions. Content updated inside that window tends to be quoted more often across the engines. Indig’s 2026 dataset points to recency as a top-three signal alongside structural clarity and named authorship. Update the date when the rewrite is real, never as a stale-stamp trick.
What role does a custom About area play in lifting visibility?
The About area builds business-entity trust, separately from any single author. In 2026, Google and other engines use it to understand what the business is, why it exists, and what it can be trusted to know. A real founding story, real certifications, and a real team beat a generic stub.
How do JavaScript-heavy websites lose visibility with AI crawlers?
OpenAI’s crawler does not run JavaScript. A modern site builder that renders the article in the browser leaves the prose unreadable to the AI surfaces. As of 2026, the fix is server-side rendering or a static export, depending on the platform. Without it, the article may as well be missing from the indexable web.
Which one page should you fix first to start being cited?
Pick one page with the highest existing Google traffic in the topic area you most want to be cited for.
The page already has the topical signal Google trusts. The work is converting that signal into chat-surface readability. Rewrite the H2s to match real customer questions, in their actual wording, harvested from PAA and customer chat history.
Add a 40 to 60 word answer block under each H2, with no links and no introductions. Cut the hedging language and commit to the claims you can defend honestly. Add a real author byline with a real bio. Update the publication date once you have made the changes.
Run the four-question check after two weeks. If the page now appears on one of the four chat surfaces, the format works for your topic. You can apply it to the next ten pages.
If the page does not appear after a month, the issue is upstream of the format. Usually it is authority signals — no named author, no consistent bio across surfaces. Or freshness — the topic has moved past the page’s framing.
The format alone is necessary but not always sufficient. The second pass adds whichever upstream signal the first pass left out.
Have one page in mind and want a second opinion on which of the eight reasons is the dominant blocker for it? You can contact me here. Send me the URL.
I will read the page once and name the one or two changes I would make first, with the reasoning. No pitch.
