Many believe that AI-simulated market research is the next most promising application for LLMs. But what does that future actually look like, and how should we think about this tool and the world it will bring?
The market it points at is not small. ESOMAR’s Global Market Research 2025, the annual report expects the insights industry to pass US$160 billion by the end of 2025. About 70% of global research spend goes to quantitative methods and 14% to qualitative. Simulation has a claim on both.
What’s Wrong With Traditional Methods
Surveys and interviews have not changed much in decades. Humans still interview humans. And even with an unlimited budget, the biggest problem does not go away. Traditional market research takes far more time than the decisions it serves can afford.
A typical custom project runs 6 to 12 weeks, and the shape is always the same. One to two weeks to brief and design, two to three to recruit, one to two in field, and about two more for analysis and reporting. Focus groups take four to six weeks end to end. Online quantitative studies take two to four. Multi-market international studies take ten to twelve.
The bottleneck is not fieldwork. It is recruitment. No-show rates run between 15% and 30%, so a study that needs twenty participants may require screening two hundred candidates and confirming twenty-five. Multiple rounds of questionnaire revision routinely add another two to four weeks. None of this is a budget problem. You cannot pay a panelist to stop being unavailable. The delay is structural. Planning, recruitment, fieldwork, analysis, and reporting run in sequence, and every stage waits on the one before it.
So what is the future? Use AI to interview AI. Let the models behave like people, and skip the part where we have to interview real ones.
Why the LLMs from Frontier Labs Will Almost Never Evolve to Impersonate a Persona
Today the valuations of the greatest AI companies live on the dream of AGI, an intelligence able to perform on most tasks as well as the best humans can, or better. So the models get trained toward a score. Every answer is graded against a “right” answer, and the ones that drift get corrected. Humans do not come with a right answer. We are irrational, we are non-normative, we can be outliers, and we can be some of the most random creatures on this planet.
Building models that behave like humans was never part of that AGI plan, because in some ways it is the opposite of it. Simulating us at the scale of a population is not something the large AI labs are doing, and not something they are going to do.
How Should We View the Inaccuracy of Simulated Research?
Simulating people is not new. Scientists have been building agent-based models for over fifty years. You fill a population with crude rule-following agents, every one of them a caricature, and run it forward to see what the crowd does. The Dirichlet model treats shopping as a coin flip with fixed odds. Nobody believes that is how people buy. But double jeopardy came out of it, and P&G and Nielsen have been benchmarking real market shares against it for forty years.
These models are wrong about every individual and right about the population. That is the standard a simulation should be held to. We would like a perfect replication of reality, but it is worth asking what we would be wishing for. A world that can be predicted exactly is a world where nobody’s choices matter and nobody has free wills. What we want out of a simulated persona is the insight, the mechanism behind why a segment moves the way it does, and that is not the same thing as a flawless forecast.
Current Landscape
This is a frontier field. And we are lucky to have Simile and Aaru working on it too. The three of us are coming at the same problem from different ends. Aaru has shown how far you can get from behavior alone; Simile has shown that a persona is worth little if it cannot tell you why. Neither of those is a small t
There are two ways to construct a persona, bottom-up and top-down. Bottom-up builds the persona out of individual interviews and surveys. Top-down generates it from aggregate information about the population distribution. Aaru and Simile sit on opposite sides of this philosophy of simulation. Aaru’s position is that words from real humans are unreliable, so personas have to be constructed from quantified data about reality rather than from what people say in interviews and surveys. What matters is the distribution and the accuracy of the aggregate. Simile takes the other side. Seeing what happens is not enough if you do not know why it happens, and the way to understand why is to learn it from real people, including the parts of them that are not normative.
We believe both approaches have their own advantages, and we want to combine them. We want the detail of why a person wants to do something, and we want collective behavior preserved at the same time. Statistically, top-down usually gives a good estimate of the mean. Bottom-up is what keeps the variance and the diversity. Combined, we will have a better picture of the target audience.