A note from Jean-Baptiste, our CTO
People now ask ChatGPT, Gemini, Perplexity and other conversational apps what to buy. Brands want to show up in those answers, either by earning it or, now that ChatGPT sells ads, by paying for it. Measuring what actually moves those answers is still a young field.
Our optimisations already deliver results for our customers. There's also a solid body of research to build on, with a lot of fascinating work still ahead. We think the company that measures this best will define the category. That's what this role is for. When we tell a brand to rewrite a page, get covered by a publisher or put budget into ChatGPT ads, I want us to show what it changed in numbers, with an honest margin of error. Then we can do more of what works.
At Teads I had the chance to work with great data scientists for ten years, some of them in my team. I learned a lot from them about what good measurement takes. Until now our engineers have handled it alongside the product. This role makes it someone's full-time job.
Jean-Baptiste Pringuey, Co-Founder & CTO
About Meikai
Meikai tracks how 14 conversational apps talk about brands: ChatGPT, Gemini, Google AI Mode and AI Overviews, Perplexity, Claude, Copilot, Grok and others. We show brands which sources those apps rely on and where they're missing, then help them fix it with content, earned media and ads in ChatGPT.
Chanel, IKEA, Volkswagen and Pernod Ricard are among our customers. Our partner panel covers more than 1.5 billion real question-and-answer messages. We run 80,000+ prompts a day and have pulled 17 million+ citations out of AI answers.
We're about a dozen people: three engineers and our CTO, plus Customer Success and Sales. We're growing fast and close to profitable. Our investors include Bertrand Quesada and Jeremy Arditi, who built Teads into a global advertising platform.
AI answers are noisy by nature. The same prompt can get a different answer an hour later. Providers ship changes without telling anyone. A source can show up next to good visibility without causing it.
Questions we need answered
- A brand rewrites its product pages and its visibility in ChatGPT goes up six points a month later. Was it the rewrite?
- A brand spends on ChatGPT ads. How much of what followed would have happened anyway through organic answers?
- A brand's visibility score moves by four points. Is that real, or noise?
- A publisher shows up more often when a brand is visible. Is it worth pitching, or is it a sampling artefact?
- How many runs do we need before we show a customer a trend?
- Can we compare Gemini with Perplexity, or March with June, when the sampling differs?
- How do we measure how a brand is perceived in French, Korean or Chinese without losing the meaning?
- We change a model or a prompt inside one of our agents. Did it get better? What got worse?
About the role
You'll be our first data scientist, reporting to the CTO. Above all, we need you to show which optimisations work with numbers we can defend, then make them work better. That goes for organic work (content, pages, earned media) and paid work (ChatGPT ads today), across all 14 apps we track.
You'll also own the methods behind the rest of the product, such as how our visibility score handles noise and how we test changes to our agents. You'll work every day with our three engineers and the CTO. You'll also spend a lot of time with Customer Success, Sales and the CEO. In a given week you might write dbt models, then sit with a customer's marketing team to work out what their numbers do and don't show.
You can publish your work and give talks about it.
Your first year
Start with lift. Our customers already see results from the optimisations we recommend, organic and paid. Your first project is to make them perform even better. Find which actions drive the most lift in each app, then improve them.
We expect you to challenge our current metrics and raise their standard where it matters.
Within six months, customers should see your work in the product, for example in recommendations ranked by measured impact.
By the end of the year, every recommendation we make should come with a measured outcome.
Responsibilities
- Measure the impact of organic and paid optimisation for our customers, with experiments where we can run them and quasi-experimental designs where we can't
- Use those results to rank our recommendations, so the ones we push first have evidence behind them
- Review the metrics in our product and raise their statistical standard
- Set sampling and confidence rules for answers that change from run to run
- Own the evals for our agents: test sets, checks for known failures, comparisons between versions
- Build and evaluate models for entity resolution, citations, brand perception and claims
- Ship in Python and SQL on dbt and BigQuery, with tests and monitoring
- Explain results to engineers, customers and the founders, including what the data can't tell us
You might be a good fit if you
- Have measured the impact of something real (a campaign or a product change, say) and seen people act on your result
- Know experiment design and causal inference well enough to work with observational data without fooling yourself
- Write Python and SQL that other people can run and maintain
- Have worked with LLMs or agents and know a model's confidence score isn't calibration
- Can explain a result to a marketing director without hiding the caveats
- Want to own a problem end to end in a company that is growing fast
It helps if you've worked on
- Marketing measurement: incrementality, attribution, media mix models, geo experiments
- Ads, ranking, recommender systems or search
- Multilingual NLP, entity resolution or knowledge graphs
- dbt, BigQuery or Google Cloud
- An early-stage company, or a job with direct customer contact
How we work
We're a small team with very little between a question and a decision. You can disagree with the founders. You can also tell us a result isn't ready to show.
A recent example: we tested a change to the agent that writes our customers' ChatGPT ads. Its overall score went up 6.5 points in an offline A/B test, but the ads made claims about the brands that were less well supported. We didn't ship it.
We use Claude, Cursor and other AI tools every day. Use whatever helps you. What matters is that the analysis holds up and that someone else can follow what you did.
What's hard about this role
- You'll be the only data scientist. The CTO has worked with data science teams before, but you won't have a peer to check your work day to day. You'll need to spell out your own standards.
- Customers want clear answers. Sometimes the rigorous answer is a range or “not conclusive yet”. You'll be the one to say it.
- The apps change without notice. A metric that was stable last month can move for reasons that have nothing to do with the brand.
- Most data in this field is observational. Clean experiments are rare, so you'll lean on quasi-experimental designs.
- Ads in ChatGPT are new. There's little history and few benchmarks to lean on.
Compensation and logistics
- Pay: €75k to €110k base depending on experience, plus a meaningful grant of Series A BSPCE (French employee stock options).
- Location: Paris or Montpellier, hybrid. We spend at least four days a month together in person.
- Language: we work in English.
- Visa: we can't sponsor one for this role, so you'll need the right to work in France.
- Experience: around five years is typical, but what you've owned and how you think matter more to us than the number.
- If you're unsure because you don't tick every box, apply anyway.
- Equal opportunity: we welcome applications from everyone, whatever their background, gender, age or disability.
How to apply
Email careers@meikai.ai with your CV and a few lines on why this problem interests you. We read every application ourselves and reply within a week.
If you want to prepare, read one of our published studies (they're in French) and bring the things you'd question.
01First conversation
An hour with the CEO and the CTO. We'll talk about your work and what you want from this role.
02Working session
Two hours with a founding engineer and the CTO on a realistic measurement problem. There's no puzzle and no take-home. Use AI tools if that's how you normally work. We mostly want to see what you'd distrust in the data and what you'd do first.