Skip to content

Research methods

Why Reproducible Audiences Matter in Synthetic Research

Synthetic answers drift between runs. What reusing the same simulated people fixes, what a seed fixes, and what neither fixes.

By the HumanPanel team · · 3 min read

Synthetic answers drift

Run the same prompt on a language model today and again next quarter, and the answers change. Bisbee et al. (2024) saw this with ChatGPT: opinion scores shifted when the researchers reran the same prompts three months later. The same study found answers moved with small wording changes and varied less than real people’s answers.

Drift hurts most when you compare two things. Say you test headline A this week and headline B next week, each on a new panel. If B scores higher, you don’t know whether the headline won or the panel changed.

Three sources of variation

A difference between two synthetic runs comes from one of three places:

  1. The people. Different respondents, with different ages, jobs and outlooks, give different answers.
  2. The questions. A change in wording, order or answer options changes the responses. Most of the time, this is the change you want to measure.
  3. The model. The AI writes each answer fresh, so one person asked the same question twice gives two slightly different answers. Model updates add more variation over time.

A fair comparison holds the first source still and changes the second on purpose. HumanPanel gives you two ways to hold the people still, and they differ in strength.

Option 1: ask the same people again

When you start a run, choose “Same people” and pick a completed run. HumanPanel sends your new survey to exactly those respondents. Their worlds, backgrounds and profiles stay as they were. Differences between the two runs now come from your questions and from answer-level variation, and a new panel is no longer part of the gap.

Saved audiences work the same way. Build an audience once, then send every version of a survey to the audience. Asking an existing population a new survey skips the step of creating people, so the run costs less than building a fresh panel.

Use this option for A/B tests of wording, price points, names and concepts.

Option 2: reuse a seed

The seed decides which traits each respondent draws from each world, such as age, job and outlook. Reuse a seed with the same number of respondents and HumanPanel repeats the selection. The worlds, the people’s life stories and their answers are still written fresh by the AI, so results differ more than with the same people.

A seed suits cases where you want a fresh panel with a matched mix of traits, for example after a small edit to the audience description. For strict A/B comparisons, choose “Same people”.

What reuse doesn’t fix

Holding the people still removes one source of drift. Two remain.

  • Answer-level variation. With 100 respondents, sampling error alone is around 10 percentage points either way. Treat a 4-point gap between version A and version B as a tie. Look for large, consistent gaps, and read the reasons behind them.
  • Model-level bias. Reusing people does nothing about errors every run shares. In our benchmark against a real survey, HumanPanel overstated awareness and daily habits on three of five questions, in the same direction each time. A saved audience keeps a bias like this in every run.

A simple protocol

  1. Build or pick one population, and name the population after the study.
  2. Run your baseline survey on the population.
  3. Change one thing per run: a word, a price, an answer option.
  4. Compare runs on the same population only.
  5. Treat gaps under 10 points on 100 respondents as ties, and read the reasons respondents give.
  6. Confirm the winning version with real people before you act on the result.

New accounts get $1 of free credit, and writing surveys is free. Sign up with your email to build your first saved audience.

Keep reading

Run your first simulation

Write your survey for free. New accounts get $1 of credit, with no credit card.