Synthetic data accuracy is one of the most debated questions in market research. Can data generated by models, instead of people, be trusted? And how does it compare to real respondent data?
There’s no shortage of opinions. But at Quest Mindshare, we’re a people business, and we wanted to see for ourselves. Think of it like shopping for a car: you don’t just read the reviews, you take it for a test drive.
So in September 2024, we presented the results of our own test drive at the ESOMAR Congress in Athens, Greece.
How We Tested Synthetic Data Accuracy
We started with a dataset we knew inside out: a U.S. study of about 1,200 respondents that we had designed, fielded, and thoroughly cleaned for an earlier conference presentation. It was an attitudinal survey, built mostly on scale questions, about how different ethnic groups in the U.S. view one another.
We then worked with a statistical modeling provider to generate synthetic data from it. Their approach isn’t a black box, and it doesn’t invent anything. Instead, it works in clear steps:
- Start with real data. The model needs real survey responses and profiling data from actual people. The more real data, the better.
- Clean the inputs. Because it’s a mathematical model, “garbage in, garbage out” applies. Dubious respondents should be removed before modeling.
- Find the best model. The system runs machine learning models, such as linear regression, logistic regression, and random forests, many times each, then selects the best performer.
- Predict sequentially. It predicts answers in the order of the questionnaire, so each prediction respects every answer that came before it.
The Two Tests
We ran two real-world tests:
- Test 1: Drop data. What if respondents quit partway through a survey? We modeled the rest of their answers to see whether partial responses could be saved instead of thrown away.
- Test 2: Fully synthetic respondents. We randomly removed about a third of the respondents (385 people) and generated their answers synthetically, then compared the results against the original data.
In both tests, we compared the average response to each question against the original data, testing at a 95% confidence level.
The Results: Surprisingly Close
We didn’t know what we’d find. The lines on our charts could have gone anywhere. Instead, they lined up almost perfectly.
- Synthetic segment on its own: looking only at the 385 synthetic respondents, results varied from the originals by about 5%.
- Full dataset with synthetic respondents: once those synthetic respondents were combined back into the total sample, the variance dropped to under 2%.
- Drop data: where we modeled the missing answers of partial completes, the variance was under 1%.
Frankly, we were surprised. Other researchers have reported bigger gaps between synthetic and real responses in their own tests, so these results were better than we expected for this type of data.
A Closer Look at the Differences
Even within that small variance, we dug deeper for patterns.
- Synthetic answers leaned to the extremes. The synthetic data tended to agree or disagree, and rarely took a neutral position. Once combined with the full dataset, though, the results fell back in line.
- Some segments differed slightly. We found no significant differences among Hispanic respondents. We did see some among respondents aged 55 and up and among Caucasian respondents. However, those differences disappeared at the total sample level, where shifting about 30 respondents out of 1,200 can change a result.
Best Practices for Synthetic Data Accuracy
This was one test, but it already points to a few best practices for researchers considering synthetic data.
Understand the Technology
Not all synthetic data works the same way. Statistical modeling builds on weighting, which the industry has accepted for years. Other providers use virtual audiences built on large language models (LLMs), which can carry built-in biases or stereotypes. In our test, the model drew directly from real data, and we saw no sign of that kind of bias.
Decide on the Right Blend
What share of synthetic to real data works for your study: 10%, 20%, 30%? Understanding those thresholds, and which use cases they suit, is key.
Build Better Profiling Into Your Surveys
Synthetic data accuracy depends on the data behind it. So consider which demographic and behavioural profiling questions to add to your surveys to make any modeling more accurate and robust.
What’s Next for Synthetic Data
This was one test of one model, with one type of survey, in one country. There’s a lot more ground to cover, and we plan to keep testing. Some of the questions we want to answer:
- Smaller budgets: if a client can only afford 400 completes instead of 1,000, can synthetic data help build a more robust dataset?
- Drop data: can partial completes, like respondents who quit 70% or 80% of the way through, be reliably saved instead of discarded?
- Refielding: if fraud removal leaves you with 700 of the 1,000 completes you need, could modeling fill the gap in days instead of weeks?
- Other markets and audiences: how does synthetic data perform outside the U.S., with B2B audiences, or with other question types?
Fraud will always be part of online research. But a tech-driven approach like this could change the dynamics of getting quantitative research done.
Have more questions about synthetic data? Download our Real Potential of Synthetic Data in Market Research FAQ for answers to the questions researchers ask most.


