Source influence: measuring which pages shape AI answers

A search ranking shows where a page appears in a search result. It does not show whether a generative system used that page, cited it or carried one of its claims into an answer. Source-influence analysis starts from those observable events.
The word influence needs care. A citation establishes association, not causation. To say that a page changed an answer requires a stronger design: a controlled content change, a stable prompt set, repeated measurements and, where possible, a comparison group.
Four events that should not be blended
| Event | What it establishes | What it leaves open |
|---|---|---|
| Page available | The source can be reached under the tested conditions. | Whether a platform retrieved or indexed it. |
| Page cited | The response contains a reference to the page. | Whether the page supplied the nearby claim. |
| Claim aligned | The answer repeats or closely matches information on the page. | Whether that page, rather than another source, caused the wording. |
| Answer changed after an edit | The timing is consistent with an effect. | Whether the edit caused it without controls for other changes. |
Keeping those stages separate prevents a citation count from being presented as proof that content “shaped” an answer.
Begin with a versioned measurement set
Source analysis is only interpretable when the demand sample is documented. Define the prompts, markets, languages, platforms and run schedule before examining which sources perform best. Preserve prompt versions so a shift in customer questions does not masquerade as a source effect.
For each response, retain the answer, cited URL, resolved URL, timestamp and the claim the citation appears to support. Domain-level totals are useful for orientation, but page-level evidence is needed to decide what to change.
Diagnose the gap before changing content
A missing or inaccurate answer can have several causes:
- the prompt asks for information the brand does not publish;
- the official page exists but cannot be retrieved or parsed reliably;
- official sources contradict one another;
- an independent source contains clearer or more current evidence;
- the platform cites a page but interprets the claim incorrectly; or
- the product is not a legitimate fit for the prompt.
The last possibility matters. Optimization should not turn every absence into a publishing task. Sometimes the measured answer is reasonable and no content change is warranted.
Test one hypothesis at a time
Meikai's workflow records a gap, proposes a change and reruns the same measurement. The sequence is straightforward:
- Observe: capture the answer, citations and factual error or omission.
- Diagnose: identify the smallest source-level explanation supported by the evidence.
- Change: update an owned page or correct an earned source where there is an editorial basis to do so.
- Measure again: rerun the same prompts and compare response-level outcomes.
A comparison page, market or prompt cohort makes the result more credible. Without one, report the outcome as a change observed after the edit, not as lift caused by the edit.
Owned and earned sources play different roles
Owned pages are the canonical place for specifications, policies, availability and approved claims. They should be current, internally consistent and technically accessible.
Independent sources can add comparison, testing or editorial context that a brand cannot credibly provide about itself. The objective is not to place the same talking point across as many domains as possible. It is to make accurate evidence available in the sources appropriate to the question.
Choosing editorial partners from citation data
The same evidence that diagnoses a gap can be used to decide where earned coverage is worth pursuing. Run the prompt set, record every cited domain and page, and rank candidates by how often they appear for the questions you care about, not by their general audience or domain authority. A publication that leads your category in traffic can be absent from the answers your buyers actually receive, and a narrow specialist title can be present in most of them.
Rank candidates on three observations, in this order.
- Does the domain appear at all in your prompt set? A domain never cited for these questions is a hypothesis, not a target.
- Is it cited next to the claims you care about, or only in passing? The four-stage table above applies here: a citation is association, and the claim-aligned stage tells you the page is doing work.
- Does the coverage carry checkable specifics? Comparative pieces with verifiable detail tend to be reused in answers more readily than announcements, though we have measured this only within our own cohort and would not present it as a general law.
Individual authors are worth tracking separately from the domains that publish them. Where a named analyst writes the comparative pieces in a category, their work can be cited across several outlets, and attributing citations to the author as well as the domain shows that pattern where a domain-level count hides it.
One caution. Choosing partners this way selects for sources a platform already reaches for, so it tells you where the answers are being formed today. It does not establish that placing coverage there will change an answer. That still needs the before-and-after design described above.
Metrics worth keeping
- Prompt-level citation rate: the share of reviewed responses citing a page or domain.
- Claim accuracy: the share of checked factual claims that match approved evidence.
- Source diversity: whether an answer depends on one source or draws from several relevant sources.
- Change persistence: whether an observed movement survives repeated runs and later measurement windows.
- Coverage: the prompts, platforms and markets for which data was actually collected.
Repeated brand mentions inside one answer are not a reliable authority metric. They can be produced by prompt wording and answer structure without indicating preference or trust.
What a communications team can measure today
Public relations reporting has traditionally counted placements, reach and equivalent advertising value. None of those describe whether an AI assistant uses the coverage. The measurable substitutes are narrower and more useful: whether the placed article is cited for your prompt set, whether the claim it carries appears in the answer, and whether either survives repeated runs over the following weeks.
A workable protocol looks like this. Fix a prompt set covering the commercial questions the campaign is meant to influence, and record its version. Measure it for a period before the placement so a baseline exists. Keep measuring after publication for longer than the campaign, because platforms differ enormously in how quickly they take content up and let it go. Report each platform separately.
Two figures then carry most of the meaning: the share of responses citing the placement, and the share where the answer repeats the specific claim the article makes. The first tells you the source was retrieved. The second is closer to influence, though still short of proof unless something was held back for comparison. Report the first as association and reserve causal language for the cases where the design supports it.
Limits
Generative systems and their retrieval layers change. A measured association can weaken without any change to the page, and two platforms can react differently to the same source. Results should retain their date, platform and prompt scope.
The practical aim is narrower than “engineering the answer”: publish accurate evidence, make it retrievable, measure how it is represented and state clearly what the measurement can and cannot prove.
Research
Frequently asked questions
What is the difference between source association and source influence?
Association means a generative system retrieved and cited a page. Influence means the content of that page changed the answer, which requires a stable prompt set, repeated measurement and ideally a comparison group. Most reporting available today establishes association.
How do you find out which sites shape AI answers about a brand?
Run a versioned prompt set, record every cited domain, page and resolved URL, and rank domains by how often they are cited for those specific questions. Ranking by general audience size or domain authority is a poor proxy, because the sources models reach for in a category are often narrower.
How can a PR team measure earned media impact in ChatGPT?
Fix a prompt set before the placement runs, measure a baseline, then track two figures afterwards: how often responses cite the article, and how often the answer repeats the claim it makes. Continue past the end of the campaign and report each platform separately instead of averaging them.
Can individual creators and analysts affect how AI describes a brand?
They can be cited, and where a named author writes the comparative pieces in a category their work may be referenced across several outlets. Tracking citations by author as well as by domain reveals that pattern, which a domain-level count hides.