Research2026

How Synthetic Population Error Is Measured

A framework for evaluating synthetic populations against real-world evidence — and for reporting where they diverge.

In 2021, Zillow shut down its home-buying algorithm after it lost more than half a billion dollars. The model wasn't imprecise — it had been trained and backtested extensively, and its estimates were tight and consistent. It was miscalibrated: systematically overpaying for homes in a shifting market, in a direction its own confidence intervals never flagged. Scale didn't fix the problem. Zillow bought thousands of houses on the model's word, and each additional purchase just meant more capital committed to the same wrong number.

That's the distinction this paper turns on. A model can be precise — stable, repeatable, confident — and still be wrong, if nothing is checking it against reality. Survey research learned this the hard way decades ago and rebuilt itself around continuous benchmark validation. Synthetic populations, built on language models instead of probability samples, now face the identical test: not whether they're confident, but whether their confidence is earned.

This paper lays out what that measurement should look like. Synthetic research is only as useful as its error is measurable. Evaluating synthetic populations requires more than asking whether their answers look plausible: it requires comparing them with real-world evidence, measuring where they diverge, and documenting the conditions under which those divergences matter.

Key takeaways

  • In Pew Research Center's benchmarking study, probability-based online panels averaged 2.6 percentage points of absolute error across 28 benchmarks, while opt-in samples averaged 5.8 — empirical reference points that synthetic populations can be scored against.
  • Generating more synthetic respondents reduces simulation noise but not systematic model error: a large synthetic sample can be highly precise and still miscalibrated.
  • A rigorous evaluation examines four dimensions of fidelity: benchmark accuracy, distributional similarity, subgroup performance, and behavioral validity.
  • The most informative result is not a single accuracy score but a bias profile showing where a synthetic population is reliable, where it is not, and how performance changes across questions and groups.

01 — Synthetic research needs a measurement standard

Synthetic populations — collections of AI-generated respondents intended to approximate a real population — are increasingly being explored for survey pretesting, audience simulation, segmentation, and other forms of behavioral research. Their usefulness depends on a basic question: how closely do their responses correspond to the people they are intended to represent?

There is no single established metric that answers that question. But evaluation does not have to begin from scratch. Survey methodology already provides a mature vocabulary for benchmarking estimates against external reference data, examining subgroup error, and distinguishing precision from accuracy. Those principles provide a strong starting point for evaluating synthetic populations, while synthetic systems also require additional tests for model-specific forms of error.

02 — Start with how surveys measure themselves

A useful reference point comes from Pew Research Center's 2023 comparison of six U.S. online survey samples, fielded in 2021. The study administered a common questionnaire to nearly 30,000 adults and compared estimates with 28 benchmarks drawn from high-quality federal surveys and administrative records.

Across those specific benchmarks, the three probability-based panels averaged 2.6 percentage points of absolute error for estimates among U.S. adults. The three commercial opt-in samples averaged 5.8 points. These figures are empirical reference points from that study, not universal error rates for every probability or opt-in panel: performance depends on the questions, weighting, population, benchmark quality, and other design choices.

The study also shows why aggregate accuracy is not enough. Among adults ages 18–29, the opt-in samples averaged 11.2 points of error, and among Hispanic adults they averaged 10.8 points. The probability-based panels averaged 3.6 points for each of those groups. A method that looks acceptable at the population level can therefore perform very differently within the segments researchers care about most.

The broader lesson is not that human surveys are error-free. It is that their error can be measured against external evidence. Synthetic populations should be evaluated with the same expectation of explicit, inspectable error.

03 — Precision is not calibration

Synthetic populations differ from conventional survey samples in an important way. Once a synthetic system has been built, additional simulated responses can often be generated at very low marginal cost. Repeated generation can reduce Monte Carlo noise — the variation caused by stochastic model outputs — but it does not automatically reduce systematic error in the underlying system.

Nor should a very large synthetic sample automatically be treated as equivalent to an equally large probability sample of independent people. Responses generated by the same model, prompt architecture, source data, or calibration process may share correlated errors. Conventional confidence intervals based only on the nominal number of synthetic respondents can therefore create a misleading impression of certainty.

The central question is calibration: whether the synthetic population reproduces the relevant properties of the real population. A miscalibrated system can produce increasingly stable estimates around the wrong answer. In that setting, more simulated respondents improve precision without improving validity.

04 — Four dimensions of fidelity

No single statistic captures every way a synthetic population can succeed or fail. A practical evaluation can separate fidelity into four complementary dimensions.

Benchmark fidelity

Real questions with credible external distributions are administered to the synthetic population, and the resulting estimates are compared with held-out benchmarks. Average absolute error is one interpretable metric for categorical survey estimates and allows results to be compared with established survey benchmarking work. The holdout requirement is essential: questions or target distributions used to construct or calibrate the population should not also be used as independent evidence of fidelity.

Distributional fidelity

Matching a mean or a headline percentage is not enough. Two populations can have the same average while producing very different response shapes. For ordinal or continuous outcomes, distribution-sensitive measures — such as total variation distance, Jensen-Shannon divergence, Wasserstein distance, or another metric appropriate to the outcome — can reveal errors that averages conceal.

Subgroup fidelity

Aggregate alignment can coexist with substantial errors within demographic or behavioral segments. Error should therefore be reported across relevant subgroups, with particular attention to groups that are smaller in the reference data or underrepresented in model training. Recent evaluations of persona-conditioned LLM respondents have found that demographic conditioning can redistribute error unevenly across questions and subgroups.

Behavioral fidelity

Static benchmark distributions test whether a synthetic population reproduces observed outcomes. Behavioral tests ask whether it responds to changes in context the way people do. Where credible human experimental evidence exists, known framing, ordering, treatment, or information effects can be replicated in the synthetic population. Negative controls are equally important: irrelevant changes should not create large effects.

05 — Interpreting error without inventing universal thresholds

Error is easiest to interpret when it is anchored to a relevant comparison rather than a universal pass/fail threshold. The human-panel figures from Pew's benchmarking study provide practical reference points:

Average absolute errorReference reading
≤ 3 pointsThe accuracy regime of premium probability panels
3–6 pointsThe regime of commercial opt-in research
6–10 pointsDirectionally informative — shapes and orderings largely correct
> 10 pointsBelow the range typically useful for quantitative work

These are reference points drawn from published human-panel benchmarking, not universal thresholds. What counts as acceptable error depends on the decision being made, the stability and quality of the benchmark, the population being modeled, and the cost of being wrong. A two-point error may be immaterial for one application and decision-changing for another. Evaluation should therefore define tolerances in advance and report performance relative to both benchmark uncertainty and intended use.

Synthetic performance should also be compared with meaningful baselines: human survey panels, simple demographic models, historical estimates, or non-LLM statistical methods. A synthetic system is most informative when it can show not only that its error is small, but that it adds value relative to simpler alternatives.

06 — Beyond the score: build a bias profile

A single accuracy number can hide the most important information about a synthetic population. A stronger evaluation produces a bias profile: a documented account of which question types, response formats, populations, and experimental conditions produce reliable alignment and which produce systematic deviations.

This matters because LLM-based respondents can exhibit structured biases rather than random mistakes. Published research has documented social-desirability effects in LLM personality survey responses, and recent synthetic-respondent studies have found heterogeneous performance across items and demographic groups. These findings do not imply that every synthetic population will fail in the same way. They show why biases must be measured rather than assumed away.

A bias profile should include error by topic and subgroup, response-distribution distortions, sensitivity to prompt wording, stability across repeated runs, calibration on held-out questions, and performance against non-LLM baselines. Version-to-version regression testing can then show whether changes to a synthetic population improve one area while degrading another.

07 — What rigorous evaluation looks like

A credible evaluation separates construction data from evaluation data, uses external benchmarks where available, reports both aggregate and subgroup results, tests full response distributions rather than only averages, and includes behavioral validation when suitable human evidence exists. It also reports uncertainty and limitations instead of treating a synthetic sample's nominal size as proof of accuracy.

The goal is not to demonstrate that a synthetic population is universally accurate. It is to establish the boundaries of its validity: what it reproduces well, where it diverges, and which decisions the evidence can reasonably support.

Survey research earned credibility by making error measurable. Synthetic research can follow the same principle. The strongest standard is not the absence of error, but transparent evidence about its size, structure, and consequences.

Sources

  • Mercer, A., & Lau, A. “Comparing Two Types of Online Survey Samples.” Pew Research Center, September 7, 2023.
  • Salecha, A., Ireland, M. E., Subrahmanya, S., Sedoc, J., Ungar, L. H., & Eichstaedt, J. C. “Large language models display human-like social desirability biases in Big Five personality surveys.” PNAS Nexus 3(12), pgae533, December 2024.
  • Taday Morocho, E. E., Cima, L., Fagni, T., Avvenuti, M., & Cresci, S. “Assessing the Reliability of Persona-Conditioned LLMs as Synthetic Survey Respondents.” Companion Proceedings of the ACM Web Conference 2026.