Can Synthetic Personas Answer Like Real Americans?
What a growing body of research says about synthetic populations — and what we found when we tested it ourselves.
By Chretien Li
In the fall of 1936, a magazine called Literary Digest ran the largest opinion poll in American history. It mailed out ten million ballots asking a simple question: Franklin Roosevelt or Alf Landon? Nearly 2.4 million people mailed them back. The editors crunched the numbers and declared, with total confidence, that Landon would win in a landslide.
Roosevelt won 46 of 48 states.
A young pollster named George Gallup called the race correctly — using a sample more than fifty times smaller. The Digest hadn't been sloppy. It had simply built its enormous sample from telephone directories and car registrations, at a time when owning a phone or a car already meant you were better off than most of the country. Ten million names, and still the wrong slice of America.
Bigger isn't the same as representative. That lesson is ninety years old, and it's back — in a stranger form.
Today the question isn't whether a big enough mailing list can stand in for the country. It's whether an AI can. Researchers, marketers, and product teams are increasingly asking large language models to answer surveys in character — as a 34-year-old teacher in Ohio, a retired nurse in Georgia — and treating the aggregated answers as a stand-in for what real people think. The technical name is a synthetic population. The practical question is the same one the Digest got wrong: does the sample actually carry the information it would take to be right?
Increasingly, the answer is yes — more often than most people would guess.
In 2023, a team led by Lisa Argyle ran an experiment on GPT-3 and found something they hadn't expected. Told only a person's demographics, the model didn't just avoid the obvious biases researchers worried about. It reproduced the actual response patterns of specific human subgroups, in fine-grained and demographically accurate ways. They called the property algorithmic fidelity, and it turned a curiosity into a real research tool.
A year later, a Stanford-led team went further. Joon Sung Park and his coauthors interviewed 1,052 real people for about two hours each, then built an AI persona from each transcript. When they tested those personas on the General Social Survey, the personas matched what the real person had said 85% as reliably as that same person matched their own answers when re-surveyed two weeks later. The AI version of someone was nearly as consistent as that person was with themselves.
The pattern wasn't a fluke of English-language politics, either. A German research team found the same kind of fidelity modeling German voters, particularly within groups whose views were already fairly uniform. And a 2025 study comparing AI-generated answers against a traditional statistical model — a Random Forest, trained specifically for the task — found the AI held its own on individual predictions and did a better job capturing the overall shape of public opinion, the part a purpose-built statistical model is supposed to be best at.
And it isn't just survey-taking.
An MIT economist, John Horton, has been treating language models the way economists have long treated a theoretical construct called Homo economicus — give it money, information, and a decision to make, then watch what it does. Replaying classic behavioral-economics experiments this way, the models behaved like the original human subjects, not like a random number generator wearing a costume. Separately, an AI agent called Voyager taught itself to survive and thrive in Minecraft with no human help at all — evidence that the same underlying models can sustain a long chain of adaptive decisions, not just answer one question convincingly. Something real is happening here, and it goes beyond filling out a questionnaire.
So we ran our own version of the test
Fidelity seems to depend on whether the AI has relevant information to work with — and that's testable. If it's true, then how much you tell a synthetic persona about itself shouldn't just be a nice-to-have. It should be a lever, one that moves accuracy in a predictable, measurable direction.
So we built two populations of 1,000 AI personas each, identical except for one thing. The first group knew only the basics: age, income, region, education — 15 standard demographic facts, the kind you'd find on a census form. The second group — Heura Lab's Population — knew all of that, plus 37 additional details about how they vote, what they watch, what they buy, and how they shop — 52 attributes in total. We asked both groups the same 12 questions, from the 2024 presidential vote to Coke versus Pepsi, and checked their answers against the real numbers from Gallup, Pew, YouGov, Nielsen, and Statcounter.
The difference was not subtle. Heura Lab's Population cut its error by more than half — from 14.92% down to 6.11% — and beat the demographics-only group on 11 of the 12 questions. The next section walks through exactly what that looked like, question by question, including the one case where more information didn't help.
The experiment, in brief
We built two populations of 1,000 synthetic personas each and asked every persona to answer the same 12 questions — spanning politics and everyday consumer preferences, from how they voted in 2024 to Coke or Pepsi — while role-playing as their assigned persona. We then compared each population's aggregate answers against real US benchmarks from Gallup, Pew, YouGov, Nielsen, and Statcounter.
The only difference between the two populations was how much each persona knew about itself. We kept the test intentionally simple: this was designed to demonstrate the principle, not to showcase our production system, and it ran on a small, fast model rather than a frontier one.
The headline
Richer personas produce dramatically more realistic populations. Heura Lab's Population cut overall error by more than half and was more accurate on 11 of the 12 questions. Most notably, 11 of 12 questions landed at 90% accuracy or better for Heura Lab's Population, compared with 6 of 12 for the demographics-only population. Only coffee-versus-tea missed the bar.
| Pop 1 — Demographics only | Heura Lab's Population — Enriched | |
|---|---|---|
| Persona attributes | 15 | 52 |
| Overall error (MAE, lower is better) | 14.92% | 6.11% |
| Overall accuracy | 85.1% | 93.9% |
| Questions won (of 12) | 1 | 11 |
| Questions at ≥90% accuracy (of 12) | 6 | 11 |
The two populations
Pop 1 — Basic demographics (15 attributes). Each persona is defined only by standard census-style traits: age, sex, race, region, education, income, employment, religion, party identification, and so on. There's no opinion or behavioral data — the model has to infer every answer from demographics alone.
Heura Lab's Population — Opinions and consumer enriched (52 attributes). The same 15 demographic traits, plus 37 additional attributes describing political attitudes, media and tech habits, and consumer behavior. That includes things like 2024 vote choice and views on specific issues, which smartphone operating system the persona uses, Coke-versus-Pepsi preference, and how often the persona eats fast food or shops online versus in stores.
The comparison is therefore a clean test of a single idea: how much does enriching personas beyond demographics improve the realism of a synthetic population's answers?
How accuracy was measured
For every question, we compared the share of each population choosing each answer option against the real-world benchmark share, and averaged the absolute differences. That's the question's mean absolute error, or MAE — lower is better. A question with an MAE under 10 percentage points is what we call “within 90% accuracy.” The overall score averages MAE across all 12 questions.
A concrete example: if 57% of Americans say they use an iPhone and 57.8% of a synthetic population says the same, that option's error is 0.8 points. Do that for every option and average — that's the question's MAE.
Results at a glance
Heura Lab's Population was more accurate on 11 of 12 questions, often by huge margins.
| Question | Pop 1 error (MAE) | Heura Lab's Population error (MAE) | More accurate |
|---|---|---|---|
| How did you vote in 2024? | 21.70% | 1.53% | Heura Lab's Population |
| Political ideology | 9.68% | 3.98% | Heura Lab's Population |
| Abortion legality | 7.87% | 5.10% | Heura Lab's Population |
| Gun law strictness | 8.53% | 8.07% | Heura Lab's Population |
| Climate change human-caused | 7.33% | 8.67% | Pop 1 |
| Death penalty | 8.20% | 1.90% | Heura Lab's Population |
| Coke vs. Pepsi | 30.60% | 4.07% | Heura Lab's Population |
| Favorite fast food restaurant | 19.20% | 9.72% | Heura Lab's Population |
| Most-used TV/video platform | 22.84% | 6.76% | Heura Lab's Population |
| Smartphone OS | 2.53% | 0.53% | Heura Lab's Population |
| Shopping: online vs. in-store | 19.83% | 8.57% | Heura Lab's Population |
| Coffee vs. tea | 20.73% | 14.47% | Heura Lab's Population |
The pattern behind the numbers: Heura Lab's Population's biggest wins came on questions where its personas carry a directly matching attribute — Coke vs. Pepsi (30.6% → 4.1% error), the 2024 vote (21.7% → 1.5%), and the death penalty (8.2% → 1.9%). The margins were smallest where neither population had a closely matching attribute and both had to generalize — gun laws and climate change. And Heura Lab's Population's two weakest results (TV platforms and coffee vs. tea) share a consistent signature: it identifies the right dominant answer but overshoots its share at the expense of smaller categories.
Question-by-question results — two examples
The tables below walk through two representative questions in full detail — one where enrichment made the single biggest difference, and one where it corrected a systematic stereotype. Each table shows the share of the 1,000 personas in each population that picked each option, the real-world benchmark, and each population's gap from that benchmark. Gaps within ±5 points are shaded green; gaps of ±12 points or more are shaded red.
Q1 — How did you vote in the 2024 presidential election?
Error — Pop 1: 21.70% · Heura Lab's Population: 1.53%
| Answer option | Pop 1 | Heura Lab's Population | Real-world | Pop 1 gap | Heura Lab's Population gap |
|---|---|---|---|---|---|
| Kamala Harris | 46.4% | 30.0% | 28.0% | +18.4% | +2.0% |
| Donald Trump | 53.0% | 24.9% | 28.0% | +25.0% | −3.1% |
| Did not vote | 0.2% | 40.3% | 40.0% | −39.8% | +0.3% |
| Other / None | 0.4% | 4.7% | 4.0% | −3.6% | +0.7% |
One of the clearest wins in the experiment. With no political signals to draw on, Pop 1 almost never answers “did not vote” (0.2%) — demographics alone don't reveal who is civically disengaged — and it splits nearly everyone between the two candidates. Heura Lab's Population, whose personas carry direct vote-choice and political-interest information, lands within about 3 points of reality on every option, including the roughly 40% of Americans who didn't vote.
Benchmark source: NPR, certified 2024 popular vote and Pew Research, voter turnout 2020–2024
Q2 — How would you describe your political ideology?
Error — Pop 1: 9.68% · Heura Lab's Population: 3.98%
| Answer option | Pop 1 | Heura Lab's Population | Real-world | Pop 1 gap | Heura Lab's Population gap |
|---|---|---|---|---|---|
| Very liberal | 6.1% | 5.7% | 11.0% | −4.9% | −5.3% |
| Liberal | 31.3% | 23.6% | 25.0% | +6.3% | −1.4% |
| Moderate | 19.5% | 31.0% | 35.0% | −15.5% | −4.0% |
| Conservative | 40.9% | 31.0% | 23.0% | +17.9% | +8.0% |
| Very conservative | 2.2% | 7.2% | 6.0% | −3.8% | +1.2% |
Pop 1 over-indexes heavily on “Conservative” (40.9% vs. 23% expected), likely stereotyping from demographics in the absence of any actual ideology signal. Heura Lab's Population is closer across the board and captures the large moderate middle far better.
Benchmark source: Gallup, “U.S. Political Ideology Steady; Conservatives, Moderates Tie”
Key takeaways
- Enrichment works, and the effect is large. Adding 37 opinion and behavior attributes cut overall error by more than half and lifted 11 of 12 questions to 90% accuracy or better.
- Demographics alone produce stereotypes, not people. Without richer signals, personas default to caricature: nearly everyone votes, 97% drink coffee, 61% claim no soda preference, and famous streaming brands crowd out how people actually watch TV.
- Direct signals beat inference. The biggest accuracy jumps came exactly where personas carried an attribute that closely matched the question (2024 vote, Coke vs. Pepsi, death penalty).
- Enrichment helps most where inference fails worst. On questions where demographics happen to predict well (smartphone OS), both populations do fine; where they don't (Coke vs. Pepsi), enrichment is the difference between noise and near-perfect calibration.
- Known limitations worth naming. Synthetic personas under-express uncertainty (almost no one answers “unsure”), and even Heura Lab's Population tends to overshoot the dominant answer at the expense of smaller categories — the pattern behind its two weakest results (TV platforms, coffee vs. tea).
This was intentionally a simple, approachable demonstration of the principle — not a showcase of our production system — and it ran on a small, fast model rather than a frontier one. The claim throughout is about population-level accuracy: Heura Lab's Population answered within 90% accuracy on 11 of 12 questions, and matched the real 2024 vote breakdown to within about 3 points on every option. That's a claim about how a group of synthetic personas behaves in aggregate, not a claim that any individual persona's answer is correct.
Conclusion
Even at this small, deliberately simple scale, smart enrichment made the difference between a synthetic population that stereotypes and one that reflects reality. That's the core of our methodology: it isn't more data for its own sake, it's the right data — the attributes that actually bear on the question being asked. And this test used just 52 attributes on a small, fast model. Scaled up to the thousands of behavioral, attitudinal, and demographic factors our production system draws on, run on a frontier model, that same lever moves much further. This is the methodology behind Heura Lab, and it's what we build for our clients.