The progress of OpenAI and Anthropic models has been phenomenal. They are already taking over the space of many AI startups, especially those building AI agents. So it is fair to ask whether the same thing will happen to us.
In our earlier post, What Is the Future, we argued that human persona simulation is not the goal of frontier labs like OpenAI and Anthropic . Their goal is AGI. To back that up, we want to share an experiment we ran.
In 2025, a team from Columbia University released Twin-2K-500, a benchmark for how well LLMs simulate human responses. They used GPT-4.1 mini as the base model. We treat their result as the baseline and ask one question: what happens if we change nothing except the base model?
GPT-4.1 mini was released on April 14, 2025, almost a year and a half ago. If better simulation comes for free with better base models, we should see a large jump in performance.
We reran the benchmark with every major OpenAI release since then: GPT-5.4 mini (with no reasoning and with high reasoning), GPT-5.5, GPT-5.6 Sol, and GPT-6 Astra. The table below is ordered by accuracy, highest first.
Even GPT-6 Astra, the newest flagship and a far more capable model on almost every other benchmark, beats the baseline by less than 3 points. GPT-5.6 Sol, the flagship before it, scores 71.2% with medium reasoning, the same as a small non-reasoning model from April 2025. GPT-5.4 mini, the direct successor in the same size class, scores about 2 points lower than GPT-4.1 mini. Turning on high reasoning, one of OpenAI’s headline features, makes almost no difference: 69.6% versus 69.4%.
After a year and a half of frontier progress, the best available model is less than 3 points ahead of a small model from April 2025, the newer small model is behind it, and more reasoning does not help.
This tells us that progress in general model capability does not automatically turn into progress in behavioral simulation. Frontier labs are not optimizing for it, and they usually do not have the data for it either. The data that matters here is not on the public internet. It is past advertising campaigns, sales data, A/B test results, survey and focus group responses, CRM and customer feedback records, all of which sit inside companies. That is the data we are gathering, and it is why we believe this problem will not be solved by the next model release.