How we build and evaluate Meikai's optimisation agents

Meikai connects AI visibility measurement to optimisation work. Its agents use response evidence, citations and brand guidelines to prepare recommendations that a team can review and implement. The useful question is how that evidence becomes a specific plan, and how the quality of the plan is checked.
Agents connect evidence to action
On-site workflows examine page content and the prompts it serves, then draft changes tied to that context. Off-site workflows prepare recommendations for relevant publishers and sources. Brand guidelines shape the tone, permitted claims and editorial choices.
An on-site plan can include an insertion point, a draft block, FAQ items and metadata. The contents depend on the evidence and the task. Recommendations remain proposals: the customer reviews and publishes changes through their own editorial workflow.
How we evaluate: LLM as a judge
LLM-as-a-judge evaluation uses a model to assess output against a written rubric. Meikai uses model-based evaluation during development, and its on-site generation workflow includes a critic step that reviews a draft and can request revisions. These are distinct controls; they do not establish a universal scoring gate for every recommendation across every workflow.
A useful review asks whether claims are supported by the supplied evidence, whether the change is specific enough to implement, and whether it respects the brand’s guidelines. A critic can still miss an error, so its verdict does not replace editorial review or demonstrate an improvement in live AI answers.
Improving agents through evaluation
Evaluation cases let the team compare prompt and model configurations on the same tasks. Failures help identify changes to instructions, evidence handling and output validation. The on-site critic checks factual support, brand style and other editorial requirements.
Judge scores measure performance against the chosen rubric. Actual visibility impact needs separate measurement after a change, with comparable prompts and repeated observations. A higher evaluation score is not a promise of more citations.
Grounded in published research
Published GEO research provides methods and hypotheses to test. It does not certify Meikai’s implementation or guarantee results for a particular brand. Three useful references are:
- Content optimisation methods and their domain-dependent effects: GEO: Generative Engine Optimization (Aggarwal et al., KDD 2024).
- A framework for measuring source influence: Beyond Keywords: Driving Generative Search Engine Optimization with Content-Centric Agents (Chen et al., 2025, preprint), which introduces CC-GSEO-Bench.
- Measurement variability: Don’t Measure Once: Measuring Visibility in AI Search, a topic we also explored in our 250,000-response repetition study.
What to ask when comparing platforms
Ask to see the evidence available to an agent, a completed plan, and the review process for that workflow. Clarify which checks run during generation, which are used in development evaluations, and who approves publication.
Explore the resulting work on the Optimisation page, build the measurement set in Prompt Studio, or talk with us about your category.
Frequently asked questions
What does it mean that Meikai is an agentic platform?
Agents connect measurement evidence to optimisation tasks. They use page content, citations and brand guidelines to prepare plans for customer review and implementation.
What is LLM-as-a-judge evaluation?
A model assesses output against a written rubric. Meikai uses model-based development evaluations and an on-site critic step that reviews drafts and can request revisions. Evaluation coverage depends on the workflow; a score does not guarantee factual accuracy or visibility gains.
How are the agents improved?
The team compares configurations on evaluation cases and uses failures to refine instructions, evidence handling and validation. Live visibility impact is measured separately from evaluation scores.
What should a buyer look for in an optimisation agent?
Ask for the underlying evidence, a concrete plan, the checks applied to that workflow, and the editorial approval process. A useful plan identifies what to change and gives the team enough context to review it.