Source authority: why AI cites some sites
Source authority is the small set of extractable signals a model uses to decide which of the sources it retrieves are safe to name in an answer — not overall page quality, and not domain authority or backlink count. That’s why two organisations can publish equally accurate pages on the same question and only one gets named when ChatGPT or Perplexity answers it: specificity beats vagueness, a corroborated claim beats a lone one, and a clearly structured answer beats a well-argued one buried in the middle of a page. Sites with strong traditional SEO still watch competitors get cited in their place, because citation runs on a different scorecard entirely.
Why “we have the better page” isn’t the whole answer
It’s tempting to assume citation follows quality, the way rankings roughly follow quality in classic search. It doesn’t, or not directly. A generative engine works in two stages: it retrieves a set of candidate pages for a question, then an underlying model decides which of those candidates to actually name in the answer it writes. Only a fraction of what’s retrieved gets cited, and that second stage — selection, not retrieval — is where most of the gap between “we should be cited” and “we aren’t” actually opens up.
Google’s own guidance on this is notably understated: its documentation on AI features states there are no special requirements to appear in AI Overviews beyond standard indexing eligibility and content quality. That’s true as far as it goes, and worth taking at face value — this isn’t a hidden ranking factor to game. But it describes eligibility, not selection. Being retrievable is the floor. What happens next, when a model chooses which two or three of the retrieved candidates to actually name, runs on a different and more specific set of signals.
What actually decides which candidate gets named
Set aside vague notions of “authority” and the pattern that shows up across the research is consistent, and mechanical enough to design for.
Specific beats vague. A passage that states a figure, a definition, or a finding plainly is far more likely to be lifted into an answer than one that discusses the same territory in general terms. This is sometimes called information gain — does this passage tell the model something concrete it can use — and it is the single most consistently cited driver of citation across the literature on generative engines. A page that hedges (“this can vary considerably”) loses to a page that commits to a stated figure with its source.
Corroboration beats repetition. A claim made once, on one site, is treated by a model with appropriate caution — it might be wrong, biased, or simply unconfirmed. The same claim appearing on the source’s own site, in independent press, and in a recognised reference source stops being one party’s assertion and becomes something the model can safely repeat. Restating your own claim in five different places on your own site does not corroborate it. A second, independent source saying the same thing does — the mechanics of how language models decide what to say about a brand go deeper on this specific point.
Structure beats argument. A short, self-contained answer sitting under a heading that mirrors the actual question — the kind of thing a person could quote directly — is easier for a model to lift than the same information stated correctly but distributed across several paragraphs of connected argument. This is precisely why the FAQ format works: a direct question, answered plainly in two to four sentences, is close to the ideal shape for a citable passage.
Recency and named authorship help, unevenly. Freshness matters more to some engines than others, and content tied to a real, checkable author or organisation carries a trust signal a model can extract, where anonymous or unattributed content does not. Neither is decisive on its own, but both tip a close call.
There is no single “authority score”
The instinct to look for one dashboard number — an authority score, a citation rating — misreads how this actually works. Independent observation of ChatGPT, Perplexity and Google’s AI Overviews finds they retrieve and weight sources differently: some favour recency more heavily, some cite a narrower set of sources per answer, some lean harder on structural cues like heading-level answers. A source cited reliably on one engine is not automatically cited on another, because the two are not scoring the same thing.
That has a practical consequence worth being blunt about: a single “we’re cited by AI now” claim is nearly meaningless without saying which engine, on which question. The honest version of this measurement is per-engine and per-question, tracked over time — not a single score that flattens three different selection processes into one number that sounds tidier than it is. For more on doing that measurement properly, see how to measure your AI search visibility; source authority is one input into the broader discipline of generative engine optimisation.
There’s also a genuine equalising effect worth naming, because it cuts against how most organisations assume this works. The Princeton GEO study — the first controlled academic research on optimising content for generative engines, presented at ACM SIGKDD 2024 — found that citation-focused changes to a page can measurably increase how often it’s included in AI-generated answers, and that the gain is often largest for sources that are not already the dominant, page-one incumbent. Being smaller or newer is not disqualifying here in the way it can be for classic domain-authority-driven rankings. Owning a specific, well-corroborated answer to a specific question can out-cite a much larger source that only covers the question in passing. That is a genuinely different competitive dynamic from traditional SEO, not a rebrand of it.
How Morris McLane executes this digitally
Understanding the mechanism is the easy part. Applying it to a real body of content is where most organisations get stuck, because it requires knowing, concretely, which sources currently get cited for your category and where the specific gaps are — not a general sense that “AI visibility matters.”
That starts with a source and citation analysis: running the actual questions your buyers or stakeholders ask through ChatGPT, Perplexity, Gemini and Google’s AI Overviews, and recording exactly which sources get named for each one, and which don’t. That produces a concrete map — the specific questions where a competitor is cited and you aren’t, and the specific questions still open to whoever answers them best. From there, the work is source-layer, not campaign-layer: rewriting the passages that currently hedge instead of state, building the structured, answer-first pages that give a model something extractable, and earning the independent corroboration — credible press, reference sources, a consistent knowledge-graph entry — that moves a claim from “asserted” to “confirmed.” None of it is exotic. It is closer to editorial and structural discipline, applied specifically to the passages a model would need to quote, and measured per engine rather than assumed.
The short version
Citation by an AI engine isn’t a reward for having the better page overall — it’s a mechanical decision made about a specific passage: is it specific rather than vague, corroborated rather than lone, structured rather than buried, and current and attributable where that matters to the engine in question. Traditional domain authority doesn’t disqualify you, but it doesn’t decide the outcome either, which is genuinely good news for an organisation that isn’t the biggest name in its category. There’s no single authority score to chase — only a per-engine, per-question picture of who’s currently being named, and a specific, workable list of what to fix where you’re not. Morris McLane’s AI search visibility work is built around exactly that picture: what’s actually being cited, where the gaps sit, and what closes them.
For the version of this written specifically for membership organisations, see how a trade association becomes the cited authority on its issue.
Frequently asked questions
What does 'source authority' mean to an AI model?
Not domain authority in the SEO sense, and not a claim a site makes about itself. To a model deciding what to cite, source authority is a combination of signals it can extract mechanically: specific, corroborated information rather than vague claims; a clear, self-contained answer near the point a reader would look; and independent confirmation elsewhere that the claim is true. A page can rank well and still not be cited, and a lower-ranked page can be cited ahead of it, because the two systems are scoring different things.
Why does one site get cited by AI and an equally good competitor doesn't?
Usually because of a handful of extractable signals, not overall quality: the cited page states a specific fact or figure plainly, near the top, under a heading that matches the question; the losing page makes the same point but buries it in a longer paragraph, hedges it, or states it without a source. Corroboration matters too — if only one of the two pages is confirmed elsewhere (press, a reference source, a second independent site), the model treats that one as lower-risk to repeat.
Is source authority the same as E-E-A-T?
They overlap but aren't identical. E-E-A-t (experience, expertise, authoritativeness, trust) is Google's framework for judging content quality for ranking. Source authority, in the AI-citation sense, is narrower and more mechanical: does this specific passage answer this specific question clearly enough, and corroborated enough, for a model to safely repeat it. A page can satisfy E-E-A-T broadly and still lose a specific citation because a rival's passage is simply more extractable on that one question.
Does domain authority or backlink count decide whether AI cites a page?
Less than it used to for traditional ranking, and less than most site owners assume. Research on generative engines (including the Princeton GEO study) finds that citation-optimised content can improve a page's inclusion rate regardless of where it currently sits in classic search rankings, with the effect often strongest for pages that are not already page-one incumbents. A well-established domain is not disqualifying, but it is not sufficient either — the specific passage still has to earn the citation on its own terms.
Do different AI engines use the same authority signals?
No, and this is one of the more counterintuitive findings in the space. ChatGPT, Perplexity and Google's AI Overviews retrieve and select sources differently — different emphasis on recency, different willingness to cite fewer, longer-established sources versus more, narrower ones. A page cited reliably on one engine is not guaranteed a citation on another. Treating 'AI citation' as one universal score misses this; the honest approach measures presence per engine, not as a single figure.
Can a smaller or newer site out-cite a larger, more established one?
Yes, and this is the useful part for an organisation without a decade of domain authority behind it. Because citation selection rewards a specific, well-corroborated answer to a specific question rather than overall site size, a smaller source that owns a narrow question clearly and is corroborated on it can be cited ahead of a larger source that only addresses the question in passing. The Princeton research describes this as an equalizer effect: the gain from doing this well is largest for sources that are not already dominant.