Opens in a new tabSkip to content
Agent LighthouseAgent Lighthouse

    Searches the text of every published page. The evidence sources themselves are not in this index — search all of them on the trusted sources page.

    GitHub ↗
    Browse checks and page contents
    answer-readiness/unique-data

    Unique data or statistics

    What it checks

    AI generative engines prioritize content with original, citable data points over vague claims. Include specific statistics and metrics.

    Why it matters

    Adding quantitative statistics to a page’s body raises that page’s visibility in generative-engine answers — position-adjusted word count and subjective impression — relative to the same page carrying only qualitative claims.

    Evidence

    • GEO study, “Statistics Addition” — GEO-BENCH: Position-Adjusted Word Count 25.9 vs a 19.5 baseline (+32.8%), Subjective Impression 23.7 vs 19.3 (+22.8%). On live Perplexity.ai: Subjective Impression 33.9 vs 24.7 (+37.2%) — the largest Subjective Impression gain of any method tested there — and PAWC 26.2 vs 24.1 (+8.7%). The paper concludes that “adding relevant statistics wherever possible ensures increased source visibility” and finds the method strongest in domains such as Law & Government and Opinion — arxiv.org/…/2311.09735 (verified 2026-08-21), results at arxiv.org/…/2311.09735v3 (verified 2026-08-21)
    • The 2026 critical survey independently places quantitative content in its supported tier: “statistics, definitions, comparisons, prices, dates, and references have a plausible advantage” — arxiv.org/…/2607.14035v1 (verified 2026-08-21)
    • Google’s content guidance asks directly for the originality half of the audit’s title: “Does the content provide original information, reporting, research, or analysis?” — developers.google.com/…/creating-helpful-content (verified 2026-08-21)

    Limits

    The survey’s integrity warning applies squarely to a presence-counting check: “adding a fabricated statistic may increase reuse while degrading epistemic quality”, and it finds the gains conditional on “time-sensitive or commercial queries” while lacking universal applicability (arxiv.org/…/2607.14035v1, verified 2026-08-21). The GEO measurements are taken inside a fixed retrieval context and share an add-content confound across all three winning methods, so length is not separated from the statistic itself. Critically, nothing in the evidence supports the unique/original framing the audit’s title and guidance promise — no source distinguishes primary research from a re-quoted third-party figure, and none supports counting product prices or comma-grouped numbers as statistics.

    How it scores

    A controlled study measures a double-digit gain for exactly this edit on both a synthetic benchmark and a live engine, but no engine documents the behaviour and the effect is query-conditional.

    Sources