First paragraph answers primary question
What it checks
AI search engines score the first paragraph highest for extractive QA. Preamble text like “In this article” or “Welcome” wastes this prime position, causing agents to extract low-value content as the page’s representative answer.
Why it matters
On a content page, the first substantive paragraph may be written as a direct declarative answer, rather than an “In this article…” preamble. The claim is that such a paragraph is more likely to be the passage an AI answer engine extracts and cites than a paragraph further down the page.
Evidence
- Position matters, and it has been measured at scale. Kevin Indig analysed 1.2 million AI answers and 18,012 verified citations. The finding: “44.2% of citations come from the first 30% of content”, “31.1% come from the middle (30–70%)” and “24.7% come from the final third, with a sharp drop near the footer”. The figures are drawn from a wider corpus of 3 million ChatGPT responses and 30 million citations, matched with sentence-transformer embeddings — searchengineland.com/chatgpt-citations-content-stud… (verified 2026-08-21)
- Sentence-level positioning is a measurable structural factor: GEO-SFE attributes 15.4% of an overall 17.3% citation-rate improvement (p<0.001, Cohen’s d = 0.64) to micro-structure, which includes keyword positioning — arxiv.org/…/2603.29979v1 (verified 2026-08-21)
Limits
The same Indig dataset cuts against the audit’s framing at sentence granularity. “53% of citations come from the middle of paragraphs”, against only “24.5% come from first sentences” and 22.5% from last sentences. The lead sentence is the least cited position within a paragraph (searchengineland.com/chatgpt-citations-content-stud…). Google states plainly “You don’t need to write in a specific way just for generative AI search” and “There’s no requirement to break your content into tiny pieces for AI to better understand it” (developers.google.com/…/ai-optimization-guide).
C-SEO Bench found “Most current C-SEO methods are not only largely ineffective but also frequently have a negative impact on document ranking” (arxiv.org/…/2506.11097). Neither GEO-SFE nor the GEO benchmark ever tested opener phrasing. The weak-opener regex list is a copywriting convention, not a documented consumer input. Nothing establishes that a page whose lead reads “In this article…” is cited less than the same page with the preamble removed. All URLs verified 2026-08-21.
How it scores
The positional half of the claim has real observational support. But the specific thing this audit measures — the absence of five English preamble phrases — has no documented consumer and no measured effect. The same dataset that supplies the positional evidence argues against first-sentence extraction.
Sources
- 44% of ChatGPT citations come from the first third of content: Study — Search Engine Land (reporting Kevin Indig / Growth Memo research), study (verified 2026-08-21)
- Structural Feature Engineering for Generative Engine Optimization: How Content Structure Shapes Citation Behavior — Yu, Yang, Ding, Sato (arXiv, March 2026), study (verified 2026-08-21)
- AI features and your website — AI optimization guide (mythbusting section) — Google Search Central, vendor-doc (verified 2026-08-21)
- C-SEO Bench: Does Conversational SEO Work? — Puerto, Gubri, Green, Oh, Yun (arXiv; NeurIPS 2025 Datasets & Benchmarks), study (verified 2026-08-20)