Benchmark
HumanPanel vs. real survey data: five coffee-market questions
Five questions from a published Verasight study with 1,519 real U.S. adults, rerun on 100 simulated respondents. How close does a cheap, fast synthetic panel get, and where does it break? Every miss is published alongside every hit.
Published September 2026 · HumanPanel run September 2026 · Real fieldwork September 2025
9.6 pts
HumanPanel mean absolute error
Across five questions, n=100. Derived: (9 + 13 + 6 + 16 + 4) ÷ 5.
23.8 pts
GPT-5 baseline mean absolute error
Same five questions, from Verasight's published synthetic sample. Derived: (9 + 13 + 45 + 35 + 17) ÷ 5.
2 of 5
Results indistinguishable from the real survey
3 of 5 are real biases, and all three run upward.
01
The question
Synthetic survey panels are fast and cheap. The open question is whether the numbers they produce are any good. Most vendors do not publish a benchmark against real fieldwork, so buyers are left to take accuracy on trust.
This page is our attempt to answer that question in a way you can check. We took five questions from a published study with a real, nationally representative sample, ran them on HumanPanel, and here report every hit and every miss. Because the original study also included a synthetic sample built with OpenAI’s GPT-5, the result is a three-way comparison: real people, a well-documented large-language-model baseline, and HumanPanel.
Every figure on this page is tagged with where it came from. Published means it is stated in the Verasight report. Measured means it came out of the HumanPanel run. Derived means we computed it from those two, and the formula is shown. If a number is not one of those three, we did not use it.
02
Method
The benchmark is Verasight’s Synthetic Sampling Report III, LLMs Misread Real Consumer Behavior, published on 6 October 2025. Verasight, a survey panel firm, fielded a coffee-market survey to 1,519 U.S. adults in September 2025, a nationally representative sample. In the same report they generated a matched synthetic sample using GPT-5, so their write-up already contains a real-versus-AI comparison.
We took the five yes/no questions from that study verbatim and ran them on HumanPanel with 100 simulated respondents, plus one open-text question asking respondents to rate Starbucks from 0 to 10. The population description was general U.S. adults, nationally representative, with no coffee-related screening. HumanPanel designed four worlds for that population: Northeast Corridor (17 respondents), Midwest Heartland (21), Southern Communities (38) and Western Markets (24).
One timing caveat to state plainly: the real fieldwork was September 2025 and the HumanPanel run was September 2026, a year later. Any real drift in coffee habits over that year is folded into what we call error here.
The run’s own summary flagged two limitations that we repeat. World bases ranging from 17 to 38 make world-level comparisons directional only. And the 0-to-10 rating was collected as free text without asking respondents to explain their score, which limits how much can be read into it.
03
Results
The honest headline: across the five questions, HumanPanel was off by an average of 9.6 percentage points. Verasight’s GPT-5 sample was off by an average of 23.8 points on the identical five questions. Mean absolute error, or MAE, means exactly that: on average, how many percentage points away from the real figure was the synthetic one? Lower is better, and zero would be a perfect match.
Figure 1
Percent answering “yes”: real respondents, HumanPanel and Verasight's GPT-5 sample
- Real respondentsn=1,519
- HumanPaneln=100
- GPT-5 baselinen not published
View as table
| Question | Real % yes | HumanPanel % yes | GPT-5 % yes |
|---|---|---|---|
| Q1. Have you heard of Starbucks? | 91 | 100 | 100 |
| Q2. Have you heard of Dunkin'? | 87 | 100 | 100 |
| Q3. Have you purchased coffee from Dunkin' in the past year? | 47 | 41 | 92 |
| Q4. Do you drink coffee at least once per day? | 56 | 72 | 91 |
| Q5. Do you ever drink coffee? | 83 | 87 | 100 |
| Question | Real % yesPublished by Verasight | HumanPanel % yesMeasured on HumanPanel | GPT-5 % yesPublished by Verasight | HumanPanel errorDerived | GPT-5 errorDerived |
|---|---|---|---|---|---|
| Q1 Have you heard of Starbucks? | 91 | 100 | 100 | +9 | +9 |
| Q2 Have you heard of Dunkin'? | 87 | 100 | 100 | +13 | +13 |
| Q3 Have you purchased coffee from Dunkin' in the past year? | 47 | 41 | 92 | −6 | +45 |
| Q4 Do you drink coffee at least once per day? | 56 | 72 | 91 | +16 | +35 |
| Q5 Do you ever drink coffee? | 83 | 87 | 100 | +4 | +17 |
Note on Q5: Verasight published this as “17% of adults reported never drinking coffee”, so the 83% is the complement, 100 − 17. Their GPT-5 run returned the “never” answer exactly zero times across the whole synthetic sample, hence 100%.
Three more derived figures frame the rest of the page. HumanPanel’s mean signed error, which keeps the direction of each miss, is +7.2 points ((+9 + 13 − 6 + 16 + 4) ÷ 5), a consistent upward lean. HumanPanel beat the GPT-5 baseline on 3 of 5 questions (Q3, Q4, Q5) and tied on the other two. And for context, Verasight reported an overall MAE of 19.8 points for GPT-5 across all questions in the coffee study, not only these five.
Figure 2
Absolute error against the real survey, in percentage points
- HumanPaneln=100
- GPT-5 baselinen not published
View as table
| Question | HumanPanel error | GPT-5 error |
|---|---|---|
| Q1. Starbucks awareness | 9 pts | 9 pts |
| Q2. Dunkin' awareness | 13 pts | 13 pts |
| Q3. Dunkin' purchase, past year | 6 pts | 45 pts |
| Q4. Daily coffee | 16 pts | 35 pts |
| Q5. Ever drinks coffee | 4 pts | 17 pts |
| Mean absolute error | 9.6 pts | 23.8 pts |
04
Where it held up
Two of the five results are statistically indistinguishable from the real survey. To decide that, we computed a 95% confidence interval for each HumanPanel figure. A confidence interval is the range the true value would plausibly sit in given that only 100 people answered. If the real figure falls inside that range, the gap is no larger than sampling noise can explain. If it falls outside, the gap is bigger than bad luck with 100 respondents can account for. It is a real bias in the simulation.
Q3, Dunkin’ purchase in the past year, is the standout. HumanPanel returned 41% against a real 47%, 6 points off, and the real value sits inside the 31.4–50.6 interval. Verasight’s GPT-5 sample returned 92% on the same question, 45 points off. This is the one behaviour question in the set, as opposed to awareness or habit, and it is where the gap between the two synthetic approaches is widest.
Q5, ever drinks coffee, also held. HumanPanel said 87% against a real 83%, 4 points off and inside the 80.4–93.6 interval. The GPT-5 baseline said 100%: not one synthetic respondent in Verasight’s sample reported never drinking coffee, where 17% of real adults do.
05
Where it didn't
The ceiling effect comes first, because it is the clearest miss. Both awareness questions came back at 100%. Real people never reach that: 9% of real U.S. adults had not heard of Starbucks and 13% had not heard of Dunkin’. With zero “no” answers in 100, the rule of three puts HumanPanel’s interval at 97–100, and the real figures of 91% and 87% fall well outside it. Verasight’s GPT-5 sample made the identical error on both questions, so on awareness the two synthetic approaches are tied, and both are wrong.
Daily consumption was overstated by 16 points. HumanPanel said 72% of adults drink coffee at least once a day; the real figure is 56%. The 63.2–80.8 interval does not contain 56, so this too is bias rather than noise. The GPT-5 baseline said 91%, 35 points off.
Figure 3
HumanPanel signed error: which direction each miss runs
View as table
| Question | HumanPanel | Real | Signed error |
|---|---|---|---|
| Q1. Starbucks awareness | 100% | 91% | +9 pts |
| Q2. Dunkin' awareness | 100% | 87% | +13 pts |
| Q3. Dunkin' purchase, past year | 41% | 47% | −6 pts |
| Q4. Daily coffee | 72% | 56% | +16 pts |
| Q5. Ever drinks coffee | 87% | 83% | +4 pts |
All three misses run in the same direction: upward, on awareness and habitual consumption. That matches an effect Verasight documented in their cross-topic omnibus. Regressing the model’s imputed proportions on the real ones gave a slope of 0.82 rather than 1.0, meaning the model systematically fails to predict extreme proportions as extreme as they really are. Simulated respondents under-produce the minority answer. The person who has never heard of Starbucks, or who owns a coffee maker but only uses it at weekends, is exactly the respondent a simulation is most likely to leave out.
Figure 4
Is each miss sampling noise or bias? HumanPanel estimate with its 95% confidence interval
- HumanPanel estimate and 95% intervaln=100
- Real valuen=1,519
View as table
| Question | HumanPanel | 95% CI | Real | Verdict |
|---|---|---|---|---|
| Q1. Starbucks awareness | 100% | 97.0 – 100.0 | 91% | Outside: Real bias: ceiling effect |
| Q2. Dunkin' awareness | 100% | 97.0 – 100.0 | 87% | Outside: Real bias: ceiling effect |
| Q3. Dunkin' purchase, past year | 41% | 31.4 – 50.6 | 47% | Inside: Indistinguishable from real |
| Q4. Daily coffee | 72% | 63.2 – 80.8 | 56% | Outside: Real bias: overstates habit |
| Q5. Ever drinks coffee | 87% | 80.4 – 93.6 | 83% | Inside: Indistinguishable from real |
06
What we can't tell you
Two parts of the HumanPanel run have no published real-world counterpart, so we cannot score them. Publishing an unverifiable comparison would defeat the purpose of this exercise, so we report what came out and stop there.
The Starbucks recommendation score. Verasight published only GPT-5’s error on the recommendation question: a mean miss of 18 points on the Net Promoter Score, with a consistent negative bias across all four brands they tested. They never published the real raw score. HumanPanel returned a mean of 4.7 out of 10, with 73 of 100 ratings between 4 and 6, an observed range of 2 to 7, and no rating above 7. World means ran from 4.5 (Southern Communities) to 5.2 (Western Markets). That clustering is consistent with the compression-toward-the-middle effect documented in the literature, but without the real score we cannot call it accurate or inaccurate.
The regional split. Verasight published no breakdown by region, so HumanPanel’s world-level figures have nothing to be checked against. We show them below because they are part of what the run produced, not because they have been validated.
Figure 6
HumanPanel only: Dunkin' past-year purchase by world
View as table
| World | Yes | Base | % yes |
|---|---|---|---|
| Northeast Corridor | 14 | 17 | 82.4% |
| Southern Communities | 17 | 38 | 44.7% |
| Midwest Heartland | 9 | 21 | 42.9% |
| Western Markets | 1 | 24 | 4.2% |
07
Limits of this exercise
Five questions is a small test. A sample of 100 carries roughly ±10 points of sampling error on its own, which is why the confidence intervals above are as wide as they are. It is one product category, one country, and there is a year between the two fieldwork dates. Two of the five questions were awareness questions where the benchmark answer is close to the ceiling, which is the hardest place for a simulation to get right and also the least informative once it gets it wrong.
The GPT-5 column is Verasight’s, not ours. We did not rerun it, and Verasight did not publish that sample’s size or prompt design, so it is a published baseline rather than a controlled comparison.
This is a data point, not a validation. It tells you how close one cheap, fast run got on one occasion, and where it broke. It does not tell you what the next run on a different topic will do.
08
What it means in practice
Where synthetic panels earn their keep, on this evidence, is in relative comparisons and direction-of-effect questions: which of two wordings confuses people, which concept a segment leans toward, what to ask before spending on fieldwork. The Q3 result is the kind of number that supports that use. A 6-point miss on purchase behaviour would not change the ranking of two products whose real gap is larger than that.
Where they do not earn their keep is absolute market sizing, penetration and awareness figures, and anything multi-select. A 100% awareness reading where the real figure is 91% or 87% is not a rounding error; a business that priced a launch on the synthetic awareness number would overstate its addressable base by 9 to 13 points of the adult population. On daily consumption, the 16-point overstatement is the difference between a majority and a strong majority. On multi-select questions Verasight found that of 77 response options, 27 (35%) were chosen by more than 5% of humans but under 1% of the model’s respondents, 24 (31%) were never picked at all, and one option chosen by 40% of real respondents got zero selections across 1,000 eligible cases. Their recommendation, which we repeat, is not to use language models for multi-response questions.
Subgroups deserve their own warning. In Verasight’s political study the topline questions were only about 4 points off, but subgroup error averaged 10 points and reached 30 for the smallest groups. Our own world-level figures, on bases of 17 to 38, should be read with that range in mind.
Figure 5
Mean absolute error across published real-versus-synthetic comparisons
View as table
| Study | MAE (pts) | Source |
|---|---|---|
| Verasight political topline (Report I) | 3.0 | Published · Key political questions |
| HumanPanel coffee, 5 questions | 9.6 | Derived · This case study, n=100 |
| Verasight cross-topic omnibus (Report IV) | 14.5 | Published · GPT-5.2, n=2,000, 305 response options from 52 questions, Jan 2026 |
| Verasight coffee, GPT-5, all questions | 19.8 | Published · Report III, Oct 2025 |
| Verasight health care category (Report IV) | 23.4 | Published · Worst category in the omnibus |
For orientation, the published record puts synthetic error anywhere from 3.0 points on polarised political toplines to 23.4 on health care, with a cross-topic average of 14.5 in Verasight’s most recent omnibus (best individual questions under 3.8, worst over 28.5). Elsewhere in the coffee study, GPT-5 overestimated awareness of the regional chain Peet’s Coffee by 25.5 points, and on a six-concept drink ranking where real respondents split almost evenly (four options within 18-21% of each other), 66% of synthetic respondents picked a single drink. Bisbee and colleagues found that ChatGPT’s averages tracked real survey averages closely but with less variation than real respondents, with regression coefficients that often differed from real estimates, and with results that shifted over a three-month period. Verasight also found that grounding the model in real respondent-level answers on correlated questions cut error from 20 to 9 points on one item and from 18 to 11 on another, with large errors remaining. A 9.6-point result on five consumer questions sits toward the better end of that record, and inside the range where the published literature says synthetic panels are useful for direction and not for absolutes.
09
Reproduce it yourself
Everything needed to check or repeat this comparison is public. The Verasight report contains the real and GPT-5 figures. The CSV below contains the five questions, all three columns, both error columns and the confidence intervals used on this page. The two Harvard Dataverse links are the open replication datasets from the two most-cited academic studies of synthetic respondents, if you want to run the same exercise on political questions.
- Download benchmark-data.csvbenchmark-data.csv (the five questions, all three columns, errors and confidence intervals)
- Verasight, "LLMs Misread Real Consumer Behavior" (Synthetic Sampling Report III), 6 Oct 2025
- Bisbee et al. 2024, Political Analysis: replication data on Harvard Dataverse
- Argyle et al. 2023, Political Analysis ("Out of One, Many"): replication data on Harvard Dataverse
To rerun the HumanPanel side, describe the population as general U.S. adults, nationally representative, with no coffee-related screening, paste the five questions verbatim, and run 100 respondents. Your numbers will not match ours exactly. Simulated panels vary between runs, as real samples do, and that variation is part of what the confidence intervals above are for.
Run your own comparison
Take a question you already have real data for, run it on HumanPanel, and see how close the simulation gets before you rely on it for anything else.