AI discovery in finance: measuring shortlist inclusion responsibly

AI-assisted research can turn a broad investment question into a summary, a shortlist and a set of supporting sources. For an asset manager, the immediate measurement problem is not whether a model “likes” a strategy. It is whether the strategy appears for relevant questions and whether the answer describes its objective, process, liquidity and risks accurately.
Those are observable outcomes. They can be reviewed without treating a generated answer as investment advice or assuming that inclusion caused an allocation.
Separate visibility from suitability
A product can be visible and still be unsuitable for the person or mandate described in the prompt. It can also be absent for a defensible reason. Measurement should therefore keep four questions separate:
| Question | What to inspect |
|---|---|
| Was the product found? | Whether it appeared in the answer or supporting citations for a defined prompt. |
| Was it shortlisted? | Whether the system included it among the options that met the stated constraints. |
| Was it described accurately? | Whether objective, process, fees, liquidity, risks and availability matched approved sources. |
| Was the evidence appropriate? | Which filings, fact sheets, reports or third-party sources supported the description. |
A single share-of-voice figure collapses these outcomes. It can reward a frequent but inaccurate description and conceal the difference between a citation and a recommendation.
Build prompts from real selection criteria
Generic questions such as “What is the best fund?” are a weak measurement base. Institutional research starts with constraints. A prompt set can reflect them without asking the model to make a regulated decision.
Useful dimensions include vehicle and jurisdiction, liquidity, benchmark, fee structure, risk limit, investment horizon and the market scenario being investigated. The same underlying need should be tested with consistent constraints across platforms so that differences in the answers remain interpretable.
Versioning matters. If a prompt changes from “daily liquidity” to “weekly liquidity,” it is a different selection problem and should not be blended into the same historical trend.
Make approved sources easy to check
Accuracy depends on the source record. Current fact sheets, regulatory filings, methodology documents and risk disclosures should use consistent names, dates and definitions. Superseded documents need a clear status, and regional variants should not rely on a reader inferring which market they cover.
The objective is not to simplify away material nuance. It is to make the same nuance available wherever the product is described. If liquidity terms or benchmark names differ between a fact sheet and a marketing page, the measurement should record the conflict rather than guess which one an assistant ought to use.
Review errors at claim level
A claim-level review is more useful than a positive-or-negative sentiment label. For each monitored answer, a team can check:
- whether the named product exists in the stated market;
- whether the answer attaches the right objective and process to it;
- whether fees, liquidity and risk statements are current;
- whether a source supports the claim made from it; and
- whether an old or unofficial document is displacing the approved one.
That produces a repair list: correct a conflicting page, expose a missing disclosure, retire a stale document or narrow a claim that the evidence does not support.
Keep governance outside the model
For internal or client-facing use, the control layer cannot depend on the model policing itself. Approved-source lists, response logs, review thresholds and human sign-off belong in the surrounding process. Generated text should not replace suitability checks, compliance review or accountable investment judgement.
AI visibility measurement in finance is therefore best treated as an information-quality audit. It shows how a defined set of systems represents a product for a defined set of questions. It does not show that the product is suitable, that a user acted on the answer or that improved visibility produced investment performance.