Synthetic research: claims must match evidence
Over three weeks, I built a synthetic-research prototype and tested it until I found the boundary between useful machine evidence and claims about people.
Summary
Jump to a section
- Thinking machines are now research subjects
- Synthetic systems and synthetic respondents
- AI changes what market research must study
- Synthetic results are promising but bounded
- Testing stopped me from selling a convincing result
- Claims must match methods
- Hypotheses must leave the machine
- A plausible ranking failed a simple reversal test
Reader’s noteAbout 3,400 words · 17-minute read (references excluded)
1Thinking machines are now research subjects
Some people have The Lord of the Rings. I have Dune. I remember staying awake until four in the morning to finish the original novel. I loved the world, the films, and the miserable beauty of Arrakis.
In Dune's backstory, humanity created “Thinking machines”, suffered the consequences, and outlawed them after the Butlerian Jihad. Mentats, human beings trained as computational specialists, were developed to perform knowledge work that could no longer be entrusted to machines.
Thinking machines are real now. There will be negative consequences. Nothing is gained by pretending AI will not influence almost everything. I am interested in the opportunity to make something positive with these tools without pretending the potential for harm is not there. My practical question is: how can we benefit from them, and in what contexts can we depend on them?
2Synthetic systems and synthetic respondents
I was thinking of how to best describe what I do in plain language: I survey thinking machines. It reframes the work for me. Begin with the system, ask it structured questions, vary the conditions and learn whether its patterns hold.
Synthetic systems research is a term I use because it distinguishes this work from the synthetic-respondent research now attracting attention. Synthetic systems research treats the model or system itself as the thing being studied. Its basic claim is modest: this system produced this pattern under these conditions. Synthetic-respondent research makes the stronger claim that generated answers can stand in for what people, customers or a population would say.
The distinction matters. A model can generate a plausible answer for a fictional Gen Z consumer in Melbourne. That answer may contain useful regularities from human language and behaviour. It does not contain that person's experience, obligations, private knowledge, or consequences. A generated backstory can become a persona, a synthetic respondent, an audience, and eventually a “digital twin” before anyone has established what is being twinned.
3AI changes what market research must study
Part of the attraction of synthetic research is that the survey methods market researchers have relied on are under increasing strain. There are more barriers to high-quality survey work than there used to be: falling participation (Kennedy & Hartig, 2019), lengthy questionnaires that reduce both participation and later response quality (Galesic & Bosnjak, 2009), and incentives that do not always align with the researcher's need for careful, valid answers.
Commercial panel members weigh the incentive, time and ease of completion when deciding whether to participate (Brosnan et al., 2021). Payment does not inherently produce bad data: research in a probability-based panel found little reason to worry that moderate incentives or incentive-motivated participation reduced response quality (Schwarz et al., 2022). The design problem is that payment is usually conditional on completion, while the quality we need is thoughtful participation.
In one web-survey experiment, the fastest lower-education respondents were most vulnerable to response-order effects on unipolar rating scales. The same study warned that speed is only a proxy for satisficing among particular respondents and items, not a blunt exclusion rule (Malhotra, 2008).
Researchers have built attention checks, timing thresholds and consistency tests around these problems. The defences are useful, but they are not perfect. A new problem now looms: AI agents create an attack surface because one person can use them to complete instruments repeatedly and at scale. Westwood's autonomous synthetic respondent passed 99.8% of 6,000 attention-check trials while producing coherent, persona-consistent answers and simulating ordinary interaction traces (Westwood, 2025). Verifying that a panel member is a person does not establish that the answers were produced carefully by that person.
The quality of the experience matters beyond one dataset. A confusing, repetitive or disrespectful survey teaches participants that surveys are not worth their attention. Poor instruments consume the trust on which better instruments also depend. Synthetic respondents look like a solution to this burden, but substituting generated answers does not repair the human claim. It changes what was measured.
Preference is often built, not retrieved
The familiar commercial story is that people hold stable preferences which careful questioning can retrieve. Consumer-research evidence has challenged that story for decades. People often construct a preference in response to the task, the available alternatives and the information in front of them (Bettman, Luce, & Payne, 1998). A brand tracker, conjoint task or attitude scale can therefore produce a consistent answer without recovering a durable object that existed before the interview.
This matters most when the real decision is information-heavy. A buyer comparing a car, mortgage or business platform encounters explanations, alternatives, constraints and recommendations before choosing. The research task is not only to ask what is hidden inside the customer. It is also to understand the environment in which a preference takes shape.
AI is becoming part of that environment
AI is entering purchasing journeys, especially where decisions require more comparison and guidance. It is too early to call it a replacement for search: an analysis of 973 e-commerce sites found that ChatGPT referrals accounted for less than 0.2% of traffic and underperformed most established channels. The same study found stronger results in complex product categories, where people need more information, comparison and guidance (Kaiser & Schulze, 2026). These information environments are becoming research objects in their own right.
When a model is part of the decision context, its category account becomes observable. We can study what it makes prominent, what changes with the prompt, what survives repetition and where systems disagree. That is a real market-research object. It is evidence about the information environment, not a shortcut to declaring what customers think.
4Synthetic results are promising but bounded
The cleanest empirical questions concern the system itself. What does this model produce when asked this question under these conditions? Which patterns survive a change in wording, order, model or time? Those questions are real and useful. They do not require us to pretend the generated speakers are people.
The model's fluency can obscure the distinction. The material looks credible because producing credible-looking material is the task. Unless we actively search for vulnerabilities, a confident answer can pass through the workflow without anyone seeing how much it depends on the context we supplied.
There are credible positive results. Argyle and colleagues found that a language model conditioned on real survey backstories could reproduce nuanced patterns across several US human samples (Argyle et al., 2023). Park and colleagues built agents for 1,052 people using extensive self-report data. Interview-only agents reached 83% of the participants' own two-week consistency benchmark. Agents combining interviews and surveys reached 86% (Park et al., 2026).
These were not strangers invented from a handful of demographic and behavioural descriptors. The stronger simulations began with extensive data from the actual individuals being modelled. Data from real people was the foundation of the method.
Models can also predict some experimental results. Cui and colleagues attempted 156 replications from psychology and management. Models recovered 73% to 81% of main effects, although interactions were less reliable and effect sizes were consistently larger than in human studies (Cui et al., 2025). Ashokkumar and colleagues reported strong correlations between model predictions and 469 effects from 70 preregistered survey experiments. Their predictions still systematically overestimated effect sizes and performed less strongly in a second archive (Ashokkumar et al., 2026).
This supports pilot testing, hypothesis generation and deciding which interventions are worth testing with people. It does not establish that an arbitrary synthetic sample represents a population.
Average resemblance is not enough
Bisbee and colleagues generated more than 3.6 million model responses matched to 7,530 participants in the American National Election Study. Some overall averages were close. The synthetic responses had less variation and frequently produced different relationships between variables. Among coefficients that differed significantly from their human counterparts, nearly a third changed sign (Bisbee et al., 2024).
For me, a third of the materially different relationships changing direction is not an acceptable discrepancy. A reader can decide what level of error they will tolerate. I would not use that result to describe a market structure.
A category average is rarely the whole commercial decision. We care about who differs, what predicts choice, where an effect reverses and how uncertain the estimate is. A plausible mean can conceal a false market structure. Claims about demographic or geographic representativeness remain hypotheses until they are tested against the people they claim to represent. If we do not look, we do not know.
The prompt helps make the person
The prompt does not merely describe the synthetic person. It helps create the answer. That makes persona construction part of the instrument, not an administrative preface. Wang, Morgenstern and Dickerson found that models used as human replacements could misportray identity groups, flatten differences within them and turn identity prompts into essentialised representations (Wang et al., 2025). Lutz and colleagues found that portrayals varied with role format, demographic priming and model (Lutz et al., 2025).
Wording, position, model, date and persona construction belong in the method, not behind the curtain. Pezeshkpour and Hruschka found large performance gaps when the same answer options were reordered (Pezeshkpour & Hruschka, 2024). Zhuo and colleagues found prompt sensitivity across models and tasks (Zhuo et al., 2024). Bisbee and colleagues found changes after small wording variations and across a three-month collection period (Bisbee et al., 2024). If these conditions are not tested and recorded, we do not know which part of the result belongs to the question and which part belongs to its presentation.
5Testing stopped me from selling a convincing result
A model can answer thousands of questions in a day without recruitment, incentives or fieldwork. The headline result, supporting evidence and visualisations can exist before you have even engaged with a client. Human research is expensive, slow and burdensome to both the researcher and the person participating.
Synthetic research is already useful for producing fast category accounts, generating hypotheses and deciding what deserves further investigation. The mistake is not using it. The mistake is presenting that starting point as a verified account of people without collecting the evidence needed for the stronger claim.
Recently, over three weeks, I built a synthetic-research prototype that could run structured collections, repeat tasks, compare models and turn the output into an accessible plain-language report. Building it changed how I understood the benefits and limits of synthetic systems research. I intended it to generate hypotheses, not declare what customers thought.
This prototype was intended to make claims such as: this ranking or pattern is stable enough to be worth checking. Several collection and reporting components worked as intended. The implied claims did not. Its output looked convincing, but the weakness was not apparent until I tested it. I could show that the system had produced an answer. I could not rely on what that answer implied strongly enough to support the full claimed proposition.
Finding this limit was the consequence of doing the work. My background is in academic research, measurement, and applied data work. The danger would have been never finding it because the report looked finished. Once the central implication failed, I stopped preparing this new prototype as a product rather than make the evidence say more than it did.
6Claims must match methods
Synthetic research will become cheap enough for almost anyone to feed in a brief and produce a polished result. The important question is what quality control sits between the ingredients and the thing handed to a client. Familiar research disciplines already provide much of what this new instrument needs.
My own testing was exploratory: a fast attempt to find failure, not a formal validation study. It showed how much a polished final dataset can hide. Testing has to be designed into the work before the output starts looking authoritative. I advocate the following behaviours as a practical standard for synthetic research:
- Define the claim, or label the exploration. “The model ranks these brands” and “customers rank these brands” are different studies.
- Preserve the instrument and metadata. Keep prompts, examples, model identifiers, dates, parameters, transformations, exclusions and costs. Cost is not evidence, but it matters when deciding whether a method is usable.
- Test the influence of wording and position. Use semantically equivalent prompts and rotate answer choices, labels and evidence.
- Repeat across relevant conditions. One answer is not a measurement. Retest across fresh runs, model families, versions and time where stability matters, then report the degree of agreement. A more expensive model may provide a useful second opinion, but price is not a quality guarantee.
- Use controls and leakage checks. Include invented alternatives, blinded labels and comparisons where no difference should exist. Check whether the prompt already contains the conclusion.
- Track failure. Record failed responses, retries and the tasks on which they cluster. Retries can make a final table look complete while hiding how often the instrument failed along the way.
- Inspect the structure of response values, not only the mean. Compare variance, distributions, subgroup differences, interactions and sign. Tidy data can still support the wrong substantive ranking.
- Validate externally at the level of use. If the claim concerns people, purchases or markets, test it against independent human-research or behavioural data.
Synthetic methods can expose assumptions, pretest instruments, search for failure cases and help decide where scarce human research should be spent. They need extensive testing, particularly when a method is extended beyond the people, domain or conditions in which it was evaluated.
The minimum standards also have to be operationalised. Saying that a method is careful, respectful, transparent, or evidence-led does not maintain those qualities. The system has to turn each value into an observable behaviour, a threshold, an owner, evidence, a repair path, and a test at the point where the claim reaches a reader. For synthetic research, that means enforced claim boundaries, recorded provenance, visible caveats, refusal states, and checks that can stop an unsupported conclusion rather than merely warn about it in a methods note.
With human grounding and statistical correction, synthetic methods may also extend limited datasets. Krsteski and colleagues found that synthetic generation alone produced substantial bias, while allocating human observations to rectification reduced that bias below 5% in their settings. Performance varied substantially across questions and datasets: a method that helped on one item could perform poorly elsewhere. That reinforces the need to test the method in the domain where it will be used (Krsteski et al., 2026).
Use these systems for what they can demonstrate. Test the boundaries that matter. If the bridge fails, narrow the claim before the work reaches a decision.
A claim ladder for synthetic evidence
- System claim: this model produced this pattern under these recorded conditions.
- Robust system claim: the pattern survived declared wording, order, repeat, model and time contrasts.
- Externally grounded claim: the hypothesis was tested against independent human-research or behavioural data from the population, construct and setting it claims to describe.
- Decision claim: the validated relationship is current, precise and relevant enough to guide this decision.
Repeated generation can strengthen the first claim. Cross-model agreement can help with the second. Neither automatically reaches the third. A thousand samples from one underlying system do not become a thousand independently observed lives.
A blended approach requires an explicit boundary. Before collecting anything, decide which claims belong to the system, which concern people, what evidence each method can provide, and what result would force the claim to narrow. Otherwise, a polished synthesis can quietly give one source authority it never earned.
7Hypotheses must leave the machine
Synthetic systems research can help us craft better hypotheses. But a hypothesis has to leave the machine. Putting it into another model may produce a better hypothesis, or simply hypotheses about hypotheses. If the claim concerns people, we need to ask real people real questions.
No single research signal contains the whole truth. Surveys see declared answers. Conversation exposes language, context and contradiction. Behaviour shows what happened under actual constraints. A model shows what that system tends to produce from the information and framing available to it.
These signals do not become interchangeable because they point in the same direction. In psychometrics, reliability concerns the consistency and precision of scores across relevant replications of a measurement procedure. Test-retest repetition is one form, not the whole concept (American Educational Research Association et al., 2014). Validity asks whether the evidence supports the interpretation and use we intend. A model can be highly reliable and still systematically wrong about the world.
Hypotheses generated by synthetic results need a plan for what comes next. Triangulation is one way to sharpen and test them by taking more than one bearing. Use people, behaviour or a genuinely different method. If the signals disagree, investigate the disagreement instead of averaging it away. The synthesis becomes the actionable result, not whichever source produced the cleanest chart.
Organisations will become accustomed to generating possibilities in hours. They will not want to wait weeks for a long conventional survey to validate every one. We need faster and more respectful ways to collect human research data: short quantitative decisions, better qualitative elicitation, voice and contextual methods, and hybrid designs that ask people only for the information the machine cannot credibly supply.
The next challenge is to make human research faster, more respectful, and worth the time people give it. That is the subject of the companion article.
Worked example
8A plausible ranking failed a simple reversal test
This was an exploratory battery, not a formal validation study. I have withheld the subject, context and model names because the point is the failure mechanism, not the ranking the instrument produced.
The proposed claim
The new project was intended to report rankings associated with different outcomes and contexts. The output was complete and plausible. In the strongest run, a frontier model completed all 9,072 scheduled comparisons, and it used “insufficient basis” in only 2.58% of responses. Nothing about the finished ranking announced that it was unsafe.
The control
Each comparison was asked twice with the presentation positions reversed. That created 4,536 forward-and-back pairs. If the ranking reflected the responses rather than their placement, reversing the presentation should preserve the selected response.
It did not. Only 68.35% of resolved pairs selected the same response after reversal. In the remaining 31.65%, the model changed its selection while retaining the same side of the screen. It was often following position while appearing to express a stable judgement.
This problem was not specific to one model. A smaller model showed the same mechanism: among 4,025 resolved reversal pairs, 600 changed selection, and every one of those 600 switches preserved screen side. A more capable model did not remove the problem.
What I changed
I balanced presentation positions, repeated the comparisons in both directions, retained non-answers, moved from the smaller model to a stronger one and kept every raw response for rescoring. The ranking still failed the reversal threshold.
I then designed a replacement that removed side-by-side choice entirely: one response target per prompt, exact repeats and neutral paraphrases, with the contrast calculated only after collection. I chose not to pursue it. At that point it would have been a new instrument measuring something different, not evidence that the failed ranking had become valid.
Why I stopped
This is the boundary that ended the three-week prototype experiment. The new protocol could collect a large volume of structured data and turn it into a convincing story. It could also reveal that the story depended on where a label appeared on the screen.
The product I was trying to create during those three weeks needed sufficient stability to guide decisions or support hypotheses. This ranking could not survive its simplest hostile control. Continuing to tune prompts until it passed would have taught me how to manufacture a result, not whether the result deserved to be trusted. I stopped preparing this new prototype as a product and kept the more useful conclusion: a synthetic claim is only as strong as the testing it survives.
References19 sources
- American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (2014). Standards for educational and psychological testing. American Educational Research Association. Open-access PDF
- Argyle, L. P., Busby, E. C., Fulda, N., Gubler, J. R., Rytting, C., & Wingate, D. (2023). Out of one, many: Using language models to simulate human samples. Political Analysis, 31(3), 337–351. https://doi.org/10.1017/pan.2023.2
- Ashokkumar, A., Hewitt, L., Ghezae, I., & Willer, R. (2026). Large language models can predict the results of social science experiments. Nature. Advance online publication. https://doi.org/10.1038/s41586-026-10742-x
- Bettman, J. R., Luce, M. F., & Payne, J. W. (1998). Constructive consumer choice processes. Journal of Consumer Research, 25(3), 187–217. https://doi.org/10.1086/209535
- Bisbee, J., Clinton, J. D., Dorff, C., Kenkel, B., & Larson, J. M. (2024). Synthetic replacements for human survey data? The perils of large language models. Political Analysis, 32(4), 401–416. https://doi.org/10.1017/pan.2024.5
- Brosnan, K., Kemperman, A., & Dolnicar, S. (2021). Maximizing participation from online survey panel members. International Journal of Market Research, 63(4), 416–435. https://doi.org/10.1177/1470785319880704
- Cui, Z., Li, N., & Zhou, H. (2025). A large-scale replication of scenario-based experiments in psychology and management using large language models. Nature Computational Science, 5(8), 627–634. https://doi.org/10.1038/s43588-025-00840-7
- Galesic, M., & Bosnjak, M. (2009). Effects of questionnaire length on participation and indicators of response quality in a web survey. Public Opinion Quarterly, 73(2), 349–360. https://doi.org/10.1093/poq/nfp031
- Kaiser, M., & Schulze, C. (2026). Frontiers: ChatGPT referrals to e-commerce websites: How do LLMs compare against traditional channels? Marketing Science. Advance online publication. https://doi.org/10.1287/mksc.2025.0489
- Kennedy, C., & Hartig, H. (2019). Response rates in telephone surveys have resumed their decline. Pew Research Center. pewresearch.org
- Malhotra, N. (2008). Completion time and response order effects in web surveys. Public Opinion Quarterly, 72(5), 914–934. https://doi.org/10.1093/poq/nfn050
- Krsteski, S., Russo, G., Chang, S., West, R., & Gligorić, K. (2026). Valid survey simulations with limited human data: The roles of prompting, fine-tuning, and rectification. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 10887–10906). Association for Computational Linguistics. https://doi.org/10.18653/v1/2026.acl-long.498
- Lutz, M., Sen, I., Ahnert, G., Rogers, E., & Strohmaier, M. (2025). The prompt makes the person(a): A systematic evaluation of sociodemographic persona prompting for large language models. In Findings of the Association for Computational Linguistics: EMNLP 2025 (pp. 23212–23237). Association for Computational Linguistics. https://doi.org/10.18653/v1/2025.findings-emnlp.1261
- Park, J. S., Zou, C. Q., Kamphorst, J., Egan, N., Shaw, A., Hill, B. M., Cai, C., Morris, M. R., Liang, P., Willer, R., & Bernstein, M. S. (2026). LLM agents grounded in self-reports enable general-purpose simulation of individuals (Version 3) [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2411.10109
- Pezeshkpour, P., & Hruschka, E. (2024). Large language models sensitivity to the order of options in multiple-choice questions. In Findings of the Association for Computational Linguistics: NAACL 2024 (pp. 2006–2017). Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.findings-naacl.130
- Schwarz, H., Revilla, M., & Struminskaya, B. (2022). Do previous survey experience and participating due to an incentive affect response quality? Evidence from the CRONOS panel. Journal of the Royal Statistical Society: Series A (Statistics in Society), 185(3), 981–1003. https://doi.org/10.1111/rssa.12857
- Westwood, S. J. (2025). The potential existential threat of large language models to online survey research. Proceedings of the National Academy of Sciences, 122(47), e2518075122. https://doi.org/10.1073/pnas.2518075122
- Wang, A., Morgenstern, J., & Dickerson, J. P. (2025). Large language models that replace human participants can harmfully misportray and flatten identity groups. Nature Machine Intelligence, 7(3), 400–411. https://doi.org/10.1038/s42256-025-00986-z
- Zhuo, J., Zhang, S., Fang, X., Duan, H., Lin, D., & Chen, K. (2024). ProSA: Assessing and understanding the prompt sensitivity of LLMs. In Findings of the Association for Computational Linguistics: EMNLP 2024 (pp. 1950–1976). Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.findings-emnlp.108
