The IAB AI Visibility Framework Is Here. This Is the Milestone GEO Needed.
In August 2026, the IAB published Measuring Visibility in the AI Era, the first cross-industry framework for how brand and publisher visibility in AI answers should be measured, disclosed and maintained over time. It comes out of Project Eidos, the IAB’s measurement standards initiative, built by a working group spanning brands (Walmart), agencies (WPP Media, PMG, Tinuiti), publishers and measurement companies.
The diagnosis in the document is blunt, and it is correct. More than 20 companies now sell AI visibility tools, each with its own methodology, its own prompt library and its own scoring, producing different answers for the same brand. Only 16% of brands systematically track AI visibility today, not because budgets are missing but because buyers have had no basis for evaluating what they are buying. Optimisation without measurement is guesswork, and until this publication there was no shared definition of what a mention, a citation or a share of voice figure even means.
That is why this publication matters more than its 36 pages suggest. A market can only mature when rigour becomes comparable. The framework gives buyers a vocabulary, a quality bar and a disclosure checklist, and it gives serious providers something we have wanted since day one: a way to compete on methodology instead of claims.
Disclosure and method. Meikai is an AI visibility measurement provider, so this article is a vendor reading a framework that applies to us. The IAB does not rate or certify providers today, and nothing here is an IAB endorsement. The Meikai capability statements below reflect product and MCP documentation available at publication; roadmap statements are identified as such. This remains Meikai’s own assessment, not an IAB validation. We encourage you to run the framework’s own suggested exercise against us and against every vendor on your shortlist: ask what a provider must disclose before you trust their data.
What the IAB framework actually says
The framework does four things, and deliberately does not do a fifth. It defines a shared metrics vocabulary, sets quality standards, specifies what providers must disclose, and describes how to run a measurement programme that stays comparable over time. It does not rate tools or prescribe vendors. Its scope is organic visibility only; paid placement measurement is flagged as an urgent adjacent gap.
The metrics vocabulary is organised as the 4 Ps of AI visibility, a causal hierarchy from appearing to acting:
| IAB layer | Core question | IAB brand metrics |
|---|---|---|
| Presence | Does the brand appear? | Mention Rate, Citation Rate, Share of Voice, Visibility Momentum |
| Prominence | Where and how prominently? | Position |
| Portrayal | In what context, with what accuracy? | Sentiment, Framing, Hallucination Rate, Factual Inaccuracy Rate |
| Persuasion | Does visibility drive action? | Recommendation Strength, Post-Citation CTR |
On top of the vocabulary sit two ideas we consider the heart of the document. First, the split between directional and decision-grade measurement: trend signals are legitimate, but budget allocation, provider selection and executive strategy demand a higher bar across sample size, query volume, intent coverage, cadence, reproducibility, validation, documentation and platform coverage. The failure mode is treating directional data as decision-grade without noticing the gap.
Second, the disclosure principle: quality tiers only work if buyers can verify them, so providers must disclose their platform coverage, prompt library construction, query sourcing, collection architecture, validation approach and baseline management. Where a provider will not disclose, the absence is itself the signal.
Why this is a milestone for GEO
Three reasons, beyond the obvious one that standards precede maturity in every measurement market.
It ends the definitional free-for-all. When two tools report different share of voice figures for the same brand, the conversation until now has been about whose marketing is more convincing. From now on it can be about whose denominator, whose competitive set definition and whose sampling method, which are answerable questions.
It makes opacity expensive. The framework instructs buyers to treat non-disclosure as a material gap in any quality claim. That single sentence changes procurement. Providers built on black-box scoring now carry the burden of proof.
It names the hard problems honestly. Non-determinism cannot be engineered away, single-response measurement is not measurement, and a share of voice figure reported as a bare point estimate implies false precision. The framework asks the industry to report distributions, ranges and per-platform figures. That is uncomfortable for anyone selling certainty, and exactly right.
We say this as a company whose name means clarity. Transparency of method has been Meikai’s founding position, and it is genuinely encouraging to see the industry’s standards body arrive at the same conclusion: no black box.
How Meikai maps to the 4 Ps
The short answer: Meikai is substantially aligned with the framework’s measurement principles and already implements much of its core vocabulary. Where formulas align directly, we say so; where Meikai uses a different taxonomy or proxy, we identify the difference. The table below is checkable against our in-platform KPI definitions, which are documented and queryable through our MCP server.
| IAB metric | Meikai today |
|---|---|
| Mention Rate | Visibility Score: responses mentioning the brand divided by total responses, per platform, market, topic and funnel stage. Same calculation the framework specifies. |
| Citation Rate | Citation tracking at domain and URL level, including share of citation, media-type classification, recency and, where scraped-page content is available, whether the cited page mentions the brand. |
| Share of Voice | Target-brand mentions divided by mentions across Meikai’s tracked competitor universe for the same filtered response set. That universe combines the target brand, active client-defined competitors and selected discovered competitors, excluding ignored entities. |
| Visibility Momentum | Core visibility KPIs are trended from daily observations across longer views. Prompt cycles are versioned so changes in the measurement set can be separated from changes in performance. |
| Position | Brand Position is the average response-level rank based on each brand’s first mention position; lower is better. It is distinct from competitive Visibility Rank and shopping-carousel position. |
| Sentiment and Framing | Brand Perception: structured 1 to 5 attribute ratings by topic, product, platform and market, versus competitors, trended. The IAB itself notes that sentiment classification is unreliable for nuanced language; this is why we replaced polarity sentiment with attribute scoring. |
| Hallucination and Factual Inaccuracy | Full raw answers are stored and surfaced. Perception signals, entity matching and cited-page evidence help teams investigate potential inaccuracies. Meikai does not yet report the IAB-defined Hallucination Rate and Factual Inaccuracy Rate as formal per-platform KPIs; that formalisation is on our roadmap. |
| Recommendation Strength | Perception and position signals cover part of this today; the IAB flags the boundary as a judgment call requiring a disclosed rubric, and we agree. |
| Post-Citation CTR | Site Scanner measures AI-referred visits through a first-party tracking pixel; CDN logs separately measure AI crawler activity. Because platforms do not expose complete click-opportunity denominators, this remains a directional bridge signal rather than a complete CTR measure. |
Against the decision-grade bar
The criteria matrix is where the framework has teeth. Here is the checkable mapping.
| Decision-grade criterion | Meikai today |
|---|---|
| Sample size | Prompts are scheduled daily on each enabled platform, with provider fill rates monitored. Meikai exposes response evidence and observation counts in relevant endpoints. For EIVA and citation-lift estimates, we publish high, medium and low sample-quality bands with documented thresholds. |
| Query volume | The Enterprise plan supports up to 500 tracked prompts per brand and market, organised into configurable topics and funnel stages and scheduled for daily measurement on enabled platforms. The framework treats fewer than 50 queries per programme as exploratory. |
| Prompt type coverage | Full-funnel by design: awareness, consideration, conversion, with the distribution disclosed and results segmentable by stage. |
| Testing cadence | Scheduled daily on enabled platforms, subject to provider availability and monitored fill rates. |
| Reproducibility | Brand-mention extraction is deterministic for a fixed answer and dated brand dictionary. Daily dictionary snapshots and multilingual fixtures support reproducible historical reprocessing. |
| Data validation | Deterministic data tests gate changes in CI, while integrity, recency, deduplication and provider fill-rate checks monitor pipeline quality. Prompt Studio lets client teams review topics and prompts before publication. |
| Methodology documentation | KPI formulas and scoring definitions are documented and accessible in-platform and via MCP. Topics and prompts are shared and validated with client teams before launch. |
| Platform coverage | Meikai’s platform recognises 13 AI answer surfaces, including Qwen, Ernie, Doubao, DeepSeek and Clova. Actual availability and daily coverage depend on market, provider and customer configuration. |
| Multi-platform aggregation | Aggregate views are accompanied by platform and model breakdowns, allowing teams to inspect divergence rather than rely on a single blended number. |
On stability, the framework’s third pillar, our March to July 2026 citation-decay cohort is the clearest evidence of how seriously we take platform-driven versus market-driven shifts: the same source-page change cut ChatGPT citations to 12.2% of baseline while Perplexity held at 109.5%, with no media running. A blended score would have booked that as brand performance. Per-platform reporting is not a feature preference; it is what makes attribution honest. The full analysis is in our citation decay study.
Where we do not fully match yet
The framework asks providers to differentiate on rigour rather than claims, so here is the honest gap list.
Intent taxonomy. The IAB segments queries into informational, comparison, recommendation and transactional. Our taxonomy is funnel-staged (awareness, consideration, conversion), which covers the same ground with different seams. We will publish an explicit mapping between the two so buyers can compare our disclosures like for like.
Named misrepresentation rates. We surface raw-answer, perception and cited-page evidence that teams can use to investigate potential inaccuracies rather than hiding the underlying responses. What we do not yet report are Hallucination Rate and Factual Inaccuracy Rate as two separately named, per-platform rates against a brand-supplied source of truth, in the IAB’s exact definitions. That formalisation is now on our roadmap, and we think the IAB’s separation of the two (platform fabrication versus faithful reflection of a wrong source) is the right cut, because the remedies differ.
Variability thresholds. The working group has deliberately left quantitative thresholds for acceptable variability open. For EIVA and citation-lift estimates, we publish sample-quality bands and confidence controls today. We will align broader published ranges with the working group’s numbers when they land.
None of this changes our reading of the document. A framework that lets a vendor claim a perfect score on day one would not be worth much.
Where we will push the industry to go further
We intend to align with these standards wherever they make measurement more comparable. We also intend to argue, inside the industry conversation, for three places the blueprint should evolve.
Paid measurement cannot stay an appendix. The framework scopes itself to organic visibility and flags standardised paid measurement as an urgent adjacent priority, because organic citations and paid placements now render on the same response surface. We agree, and we have been operating on that surface since ChatGPT Ads launched: our position is that paid and organic must share one planning layer and never share one score. The industry should standardise that separation before blended paid-organic dashboards proliferate. Our view is laid out in our GEA platform comparison.
Real consumer prompts should move from disclosure item to quality dimension. The framework already requires providers to disclose whether query sets are synthetic or grounded in real behaviour. We think the next revision should go further: a prompt library that has never been validated against what people actually ask AI can be reproducible and still precisely answer the wrong questions, which the framework itself acknowledges. Meikai uses anonymised real-prompt panel data in research and selected Prompt Studio workflows to test whether synthetic monitoring reflects genuine consumer demand; our prompt modelling methodology explains the broader approach.
Portrayal deserves better than polarity. The framework candidly notes that sentiment classification accuracy varies and that neutral descriptive language dominates commercial contexts. That is our experience too, and it is why we score structured attributes on a 1 to 5 scale instead. We will make the case for attribute-level perception entering the core vocabulary as the Portrayal layer matures.
And when the disclosure framework becomes a certification programme, as the document anticipates, we will be in the first group submitting.
What this means for buyers this quarter
Do not wait for certification. The disclosure checklist is usable in procurement today. Ask every provider on your shortlist, including us, which platforms and model versions they cover, how their prompt library is built and refreshed, whether queries are synthetic or grounded in real consumer behaviour, how mentions are detected across scripts and languages, what their data is validated against, and how baselines are managed when models update. Then match the answers to the decision you are making: directional evidence for a briefing, decision-grade evidence for a budget.
If you want to see how Meikai answers those questions on your own brand and markets, talk with our team. We will show you the raw answers behind every number, because that is the whole point.
Frequently asked questions
What is the IAB’s Measuring Visibility in the AI Era framework?
A cross-industry framework published by the IAB in August 2026 under Project Eidos. It defines a shared vocabulary for AI visibility metrics (the 4 Ps: Presence, Prominence, Portrayal, Persuasion), a two-tier quality standard separating directional from decision-grade measurement, a provider disclosure checklist for procurement, and practices for keeping measurement comparable over time. It does not rate providers or cover paid placement measurement, which it flags for future work.
What are the 4 Ps of AI visibility?
Presence (does the brand appear: Mention Rate, Citation Rate, Share of Voice, Visibility Momentum), Prominence (where and how prominently: Position), Portrayal (in what context and with what accuracy: Sentiment, Framing, Hallucination Rate, Factual Inaccuracy Rate) and Persuasion (does visibility drive action: Recommendation Strength, Post-Citation CTR). The hierarchy is causal: each layer only matters if the one above it holds.
What is the difference between directional and decision-grade measurement?
Directional measurement identifies trends and early signals and is suitable for internal briefings and competitive awareness. Decision-grade measurement meets a higher bar across sample size, query volume, intent coverage, cadence, reproducibility, validation, documentation and platform coverage, and can support budget allocation, provider selection and executive strategy. Both are legitimate; the failure the IAB warns against is treating directional data as decision-grade.
Does Meikai meet the IAB framework?
Meikai is substantially aligned with the framework across metric definitions, daily scheduled measurement on enabled platforms, prompt coverage, reproducible brand matching, methodology documentation and platform-level analysis. Enterprise plans support up to 500 tracked prompts per brand and market. Two areas remain in active alignment: mapping Meikai’s funnel taxonomy to the IAB intent taxonomy, and formalising Hallucination Rate and Factual Inaccuracy Rate as separately reported per-platform KPIs. Sample-quality bands currently apply to specific citation-lift estimates rather than every metric. This is Meikai’s own assessment; no IAB certification exists today.
What should I ask an AI visibility provider before trusting their data?
The framework’s disclosure list is the right starting point: platform coverage with model versions, prompt library construction and refresh cadence, query sourcing (synthetic versus grounded in real consumer behaviour), data collection architecture and method, panel validity where a panel is used, validation approach, mention detection and disambiguation logic, factual accuracy handling, and historical baseline management. Treat a refusal to disclose as the answer.
Method and sources
The framework description summarises the IAB’s Measuring Visibility in the AI Era, published August 2026 as part of Project Eidos, read in full at publication. Meikai capability statements reflect product and MCP documentation and the cited analyses available at publication; roadmap statements are identified as such. The citation figures cover a fixed cohort of 126 monitored brands from March to July 2026, measured as citations per 1,000 platform responses; full method and limits sit in the linked analysis. This is a vendor’s self-assessment against a published framework, not an independent audit.