<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>Jonathan O'Hara</title>
    <link>https://jaohar.com.au/articles.html</link>
    <description>Articles about research, evidence, service design, and the systems around them.</description>
    <language>en-au</language>
    <lastBuildDate>Fri, 14 Aug 2026 16:36:47 +1000</lastBuildDate>
    <atom:link href="https://jaohar.com.au/feed.xml" rel="self" type="application/rss+xml"/>
    <item>
      <title>How can you operationalise rapport?</title>
      <link>https://jaohar.com.au/operationalise-rapport.html</link>
      <guid isPermaLink="true">https://jaohar.com.au/operationalise-rapport.html</guid>
      <pubDate>Fri, 14 Aug 2026 00:00:00 +1000</pubDate>
      <category>Service design</category>
      <category>Community</category>
      <category>Working note</category>
      <description>A working event experiment that gives peers a real problem to explore and a fair opportunity to demonstrate how they think and work.</description>
      <content:encoded><![CDATA[
<p class="standfirst">Networking can put strangers in the same room and make them aware of one another. The harder task is helping them discover who might be worth speaking with again.</p>
<p>Rapport sits somewhere between recognition and collaboration. An event cannot manufacture it, but it can arrange the conditions in which rapport has a better chance to emerge.</p>
<p><a href="https://jaohar.com.au/operationalise-rapport.html">Read the complete article.</a></p>
      ]]></content:encoded>
    </item>
    <item>
      <title>Where do humans fit in the future of research?</title>
      <link>https://jaohar.com.au/human-data-collection.html</link>
      <guid isPermaLink="true">https://jaohar.com.au/human-data-collection.html</guid>
      <pubDate>Wed, 29 Jul 2026 00:00:00 +1000</pubDate>
      <category>Human research</category>
      <category>Service design</category>
      <description>Better human research data begins with a better experience of collecting it.</description>
      <content:encoded><![CDATA[
<p class="standfirst">Humans belong wherever research asks about priorities, circumstances, trade-offs, and lived experience. Synthetic methods can help decide what to ask; they cannot supply those answers. Better human data begins with better-designed participation.</p>
<p class="article-byline">First published <time datetime="2026-07-29">29 July 2026</time>. By Jonathan O’Hara.</p>
<nav class="article-summary-contents" aria-label="Article summary">
    <h2>Summary</h2>
    <ol>
      <li><a href="https://jaohar.com.au/human-data-collection.html#experience"><strong>1. Participant experience is part of data quality</strong></a><p>Friction, repetition, forced answers, interruption, and unclear journeys affect what people provide. Spend human attention where it can change the decision.</p></li>
      <li><a href="https://jaohar.com.au/human-data-collection.html#mobile-first"><strong>2. Match the method to the question, person, and device</strong></a><p>Ratings, choices, voice, photographs, diaries, and artefacts capture different forms of evidence.</p></li>
      <li><a href="https://jaohar.com.au/human-data-collection.html#panel"><strong>3. Build panels with relationships</strong></a><p>Known participants, consented profiles, fair payment, predictable contact, and continuity support better repeated research.</p></li>
      <li><a href="https://jaohar.com.au/human-data-collection.html#practice"><strong>4. Set a higher standard for human research</strong></a><p>Own and test the complete participant journey, measure its failures, and improve it after every wave.</p></li>
    </ol>
  </nav>
<div class="body">
    <h2 id="experience">1. Participant experience is part of data quality</h2>

    <p>I have conducted social research for years, and I participate whenever I have the opportunity. Seeing research from both sides keeps one fact visible: the instrument is also an experience, and the quality of that experience shapes what people provide.</p>

    <h3>What participation reveals</h3>

    <p>Participating in other people's research makes recurring problems difficult to miss. Instruments can be too long, repetitive, short on context, and weakened by careless screening and branching. I can understand a weak survey produced by someone doing so for the first time. What is harder to accept is the same experience from organisations fielding surveys at large scale. They may have strict requirements for what the instrument must collect, yet no equivalent requirement for the experience to be good for the participant.</p>

    <p>One reason these surveys remain weak is that questions are crammed into one of many survey platforms without reconsidering the experience. We have largely replicated the paper questionnaire online rather than asking what an online-first experience should be. Online-first means device-first, and today that also means mobile-first.</p>

    <p>The best questionnaires move naturally, make the context of each question clear, and prepare the participant for what comes next. They make it easy to provide the information the researcher actually needs. Their quality makes the prevailing standard harder to excuse.</p>

    <h3>What the data reveals</h3>

    <p>The participant experience directly affects the data. I have spent much of my time analysing large survey datasets, so I began to see where people checked out, where a task stopped working, and where an instrument began producing answers that were technically complete but difficult to believe. A finished report rarely shows how much unusable data was removed, or how much doubtful data remained because the study still had to produce a result.</p>

    <p>I learned the same lesson earlier, while conducting computer-assisted telephone interviews. I was reprimanded for straying from the script when a question prompted someone to describe a difficult experience. The correct operational response was to acknowledge as little as possible and return them to option A or B. I was not especially good at that job because people do not naturally answer as instruments require. Older participants in particular often wanted to tell me what happened. The questionnaire demanded a fixed response.</p>

    <p>An online form hides that encounter. If none of the answers fit, a participant who wants the incentive still has to give the system what it demands: an answer. Paper surveys occasionally reveal what the form excluded. While entering data from a regional public-health study, I found comments written in the margins: profane, profound, and sometimes more immediately relevant than the questions we had asked. People wanted their stories to be heard even when the instrument had nowhere to put them.</p>

    <p>The interview and the handwritten margins revealed the same conflict. People were trying to tell us what had happened to them. The instrument was trying to turn that account into a permitted response. Good human research has to preserve enough structure to answer the research question without treating the part that does not fit as noise.</p>

    <h3>Design for interruption</h3>

    <p>Research is rarely the highest priority in a participant's life. People answer while travelling, working, caring for someone, waiting for an appointment, losing reception, or simply running out of attention. Interruption should be treated as expected behaviour, not participant failure.</p>

    <p>A well-designed service saves each meaningful contribution as it is made, shows the participant what has been retained, and lets them resume from their last completed thought rather than merely reopening the last screen. Once somebody has given us an answer, our process should not make them give it again because we failed to hold or retrieve it. Re-asking can still be legitimate when a participant is revising an answer, reconfirming something that may have changed, resolving a contradiction, or contributing to an intentional time series. The service should say which of those it is doing and why.</p>

    <p>The useful invitation is not “please start again”. It is “what you gave us has been retained; continue when you can”. That promise has to include correction, withdrawal, deletion, and recontact choices, with any limits explained before participation rather than discovered afterwards.</p>

    <p>US telephone poll response rates have fallen for decades <a class="cite" href="https://jaohar.com.au/human-data-collection.html#ref-kennedy-2019">(Kennedy &amp; Hartig, 2019)</a>. Respondents under cognitive load may satisfice rather than answer carefully <a class="cite" href="https://jaohar.com.au/human-data-collection.html#ref-krosnick-1991">(Krosnick, 1991)</a>. In a web-survey experiment, longer stated questionnaires reduced starts and completions, while questions placed later produced faster, shorter, and more uniform answers <a class="cite" href="https://jaohar.com.au/human-data-collection.html#ref-galesic-2009">(Galesic &amp; Bosnjak, 2009)</a>. Westwood showed that an autonomous AI respondent could pass standard attention checks while producing coherent, persona-consistent answers <a class="cite" href="https://jaohar.com.au/human-data-collection.html#ref-westwood-2025">(Westwood, 2025)</a>. We cannot keep offering people weak incentives and poor experiences, then treat declining participation as something participants have done to us.</p>

    <h3 id="signals">Spend human attention carefully</h3>

    <p>Where possible, do not push a raw question battery onto participants merely because the research team has not narrowed it. Synthetic methods can map a topic, compare structures, screen possibilities, and help sharpen hypotheses before human attention is spent. Client knowledge and existing evidence should do the same work.</p>

    <p>A team that begins with a hundred possible concepts might use those signals to identify twelve worth investigating, then ask people to evaluate the twelve. The synthetic stage reduces burden; it does not become evidence of what people prefer. Questions about priorities, trade-offs, circumstances, and lived experience still need answers from people.</p>

    <p>I develop that boundary in <a href="https://jaohar.com.au/synthetic-research.html"><em>Synthetic research: claims must match evidence</em></a>, including what synthetic work can reveal, how it should be tested, and what it cannot claim.</p>

    <h2 id="mobile-first">2. Match the method to the question, person, and device</h2>

    <p>A long web form is only one way to collect human research data. Some questions are suited to a rating or a short choice. Others need a person's own words, a photograph, a short recording, a diary entry, a receipt, or another artefact from the experience. Asking someone to compress a complicated experience into a number can discard the very information we hoped to understand.</p>

    <p>Photo elicitation, for example, can broaden and deepen a qualitative interview <a class="cite" href="https://jaohar.com.au/human-data-collection.html#ref-gill-2024">(Gill, 2024)</a>. The principle extends to videos, screenshots, proof of purchase, and material captured at the moment something happens. If entry into a category matters, verified evidence of that entry can be more valuable than another unsupported screening answer. That higher standard of evidence should also earn the participant a higher reward.</p>

    <h3>Use the device</h3>

    <p>These experiences should be mobile-first, not merely desktop questionnaires squeezed onto a smaller screen. A randomised crossover experiment found that smartphone respondents could provide careful answers, but small sliders and date pickers created input errors; the task has to be easy to perform on a touchscreen <a class="cite" href="https://jaohar.com.au/human-data-collection.html#ref-antoun-2017">(Antoun, Couper, &amp; Conrad, 2017)</a>. We do not have to inherit standard form controls. We can use the whole surface of the device to make one task unmistakable, show exactly what was recorded, and make correction easy.</p>

    <p>That does not mean every participant should be expected to swipe, drag, or manipulate the same control. Dexterity, sensation, vision, and familiarity with a device differ. Offer a brief practice step, clear feedback, and an alternative way to answer. The purpose is to help people respond accurately through an interaction they can use.</p>

    <h3>Interface quality is data quality</h3>

    <p>I have seen what happens when an instrument is not fit for purpose. In an academic health-data role, I worked with records entered by numerous nurses through an interface that made year of birth easy to enter incorrectly. Implausible dates appeared systematically. For studies of children's growth, date of birth was not a minor field; it was essential to the analysis. The problem was well known and remained unresolved. Interface quality had become data quality.</p>

    <p>Making a task easier or more satisfying does not require gamification. In one experience-sampling experiment, virtual rewards increased responses among the people who used them most, but made their momentary reports slightly less reliable <a class="cite" href="https://jaohar.com.au/human-data-collection.html#ref-dejonckheere-2024">(Dejonckheere et al., 2024)</a>. Points and badges can become another demand placed on the participant. The better aim is a satisfying interaction: clear progress, appropriate feedback, useful variety, and, where possible, something of value returned to the person.</p>

    <h3 id="prototypes">Two examples of designed participation</h3>

    <p>These examples put the method into visible form. The first is an interactive prototype; the second is a visual exploration. Neither is a finished product or validated instrument. They are design hypotheses about how a familiar research task might change when participation is designed for the device and the person using it.</p>

    <div class="prototype-grid">
      <a class="prototype-card" href="https://jaohar.com.au/prototypes/human-data/quantitative-human-data-collection.html">
        <img src="https://jaohar.com.au/images/og-human-data-quantitative.png" alt="Preview of a touchscreen preference task with large swipe targets">
        <span class="prototype-copy">
          <strong>Designed for the device</strong>
          <span>A mobile-first choice task exploring how the device itself can support new response modes: full-screen swipes, a “too close to call” response, and interaction data retained for later validation.</span>
          <b>Try the prototype →</b>
        </span>
      </a>
      <a class="prototype-card" href="https://jaohar.com.au/prototypes/human-data/qualitative-human-data-collection.html">
        <img src="https://jaohar.com.au/images/og-human-data-qualitative.png" alt="Preview of a mobile voice check-in with recording and payment confirmation screens">
        <span class="prototype-copy">
          <strong>A voice check-in</strong>
          <span>A visual exploration of short, paid voice reflections collected over time, with the participant's experience at the centre.</span>
          <b>View the example →</b>
        </span>
      </a>
    </div>

    <p>The choice task demonstrates one question at a time, a response that does not force a false preference, and interaction traces retained separately for later testing. The voice example demonstrates richer expression, repeated participation, visible confirmation, and direct payment. Neither establishes that the measure is valid, accessible to every participant, or better than an existing instrument. Those questions require separate evaluation.</p>

    <p>Recruitment, consent, screening, participation, support, payment, feedback, and recontact form one service from the participant's point of view. That becomes especially visible when researchers return to the same people. Optimising the instrument while neglecting the surrounding service cannot create a relationship worth maintaining.</p>

    <h2 id="panel">3. Build panels with relationships</h2>

    <h3>When change is the claim</h3>

    <p>A fresh sample at each wave can estimate how an aggregate measure differs from one point to the next. It cannot tell us how the same people's views or circumstances changed. If 200 people answer in January and a different 200 answer in April, a flat average can conceal individuals moving in opposite directions, while an apparent shift can partly reflect who happened to enter each sample. If the decision depends on within-person change, the design needs within-person data.</p>

    <p>That makes longitudinal relationships especially valuable for commercial tracking. Instead of repeatedly drawing convenient samples from an opaque pool, researchers should be able to return to known participants and ask stable questions over time. Diary methods and experience sampling already show the value of repeated, in-the-moment evidence (<a class="cite" href="https://jaohar.com.au/human-data-collection.html#ref-csikszentmihalyi-1987">Csikszentmihalyi &amp; Larson, 1987</a>; <a class="cite" href="https://jaohar.com.au/human-data-collection.html#ref-stone-1994">Stone &amp; Shiffman, 1994</a>; <a class="cite" href="https://jaohar.com.au/human-data-collection.html#ref-bolger-2003">Bolger, Davis, &amp; Rafaeli, 2003</a>).</p>

    <p>Academic researchers often work under severe constraints. An undergraduate convenience sample, a prize draw, or a small one-off study may be the only feasible way to see whether an early idea has any promise. Commercial research that informs real expenditure has a different opportunity. There is money in the system, and some of it should be used to improve the sample, the relationship, and the experience. Not every improvement costs more, but the parts that do should be treated as research infrastructure rather than avoidable overhead.</p>

    <h3>A relationship worth maintaining</h3>

    <p>I am interested in how curated, longitudinal commercial panels can be designed, developed, and maintained: known and segmented groups of people whom researchers can return to, who understand the relationship they are entering, and who are paid fairly and predictably. Their profiles can become richer with consent, their participation can span scheduled studies and rapid pulses, and the provenance of the sample can be described instead of hidden behind a provider's assurance.</p>

    <p>An established relationship can also stop every study from beginning at zero. With permission, a panel can maintain verified attributes that are relevant across studies, such as age range, location, household circumstances, category participation, or previous research activity. A researcher can then request the participants who fit the study without asking each person to repeat the same screening and demographic questions. The study should receive only the attributes it needs, and participants should be able to see why those attributes are being used.</p>

    <p>Build a panel that can be described honestly rather than promising perfect representation. Recruit people from relevant communities, establish their characteristics with a small set of screening questions, and report results by known groups rather than implying that every result generalises to everyone. The limits of non-probability samples still apply <a class="cite" href="https://jaohar.com.au/human-data-collection.html#ref-baker-2013">(Baker et al., 2013)</a>, but those limits can be described instead of concealed. Participants should also be paid directly for the time and evidence they provide.</p>

    <p>Trust has to run in both directions. Researchers often design as though the participant is an adversary to be caught speeding, straightlining, or answering carelessly. Participants have their own reasons not to trust us: repetitive questions, unexplained purposes, weak screening, thoughtless wording, delayed payment, and no accessible way to correct a problem. Every invitation asks them to take a risk on whether this will be a good survey or a bad one. Onboarding has to earn trust rather than merely test compliance.</p>

    <p>A durable panel can also make participation habitual without making it extractive. Ask people when they can genuinely focus, perhaps a regular half-hour on Friday morning, and return at the time they chose. A predictable rhythm respects the fact that research invitations otherwise arrive between work, care, and household tasks. Rapid pulses then become part of a relationship, not an unexpected demand.</p>

    <h3>Give participants control of the relationship</h3>

    <p>A richer profile creates obligations as well as efficiencies. Participants should be able to inspect and correct what the panel holds about them, understand which attributes will be reused, choose whether they can be recontacted, and withdraw from the relationship. Withdrawal, profile correction, and deletion should be usable parts of the service rather than rights buried in a policy. Consent to join a panel is not unlimited consent for every later use.</p>

    <p>The operator also has to monitor what continuity changes. Repeated participation can condition answers, long-serving members may differ from people who leave, and the most burdened participants may be the first to disappear. Track participation frequency, refusals, attrition, profile completeness, and replenishment by relevant group. A longitudinal panel produces better evidence only when those changes remain visible.</p>

    <h3>What already exists</h3>

    <p>Pieces of this future already exist. <a href="https://www.prolific.com/longitudinal-studies" target="_blank" rel="noopener">Prolific's longitudinal projects</a> support multi-wave studies that return to the same participants, show them the schedule and payment upfront, and track retention across waves. <a href="https://www.dscout.com/platform/methods/diary-studies" target="_blank" rel="noopener">Dscout's diary tools</a> support repeated mobile activities with photographs, videos, and screen recordings. Australia's probability-based Life in Australia panel demonstrates that a repeatedly contacted national panel can be deliberately recruited, maintained, and evaluated <a class="cite" href="https://jaohar.com.au/human-data-collection.html#ref-kaczmirek-2019">(Kaczmirek et al., 2019)</a>. Those are important capabilities. The larger opportunity is to combine continuity, good instrumentation, richer modes, transparent sample quality, and a relationship participants would choose to maintain.</p>

    <p>Evidence about who participated can strengthen a panel without validating every answer. A receipt may support category membership, but it does not establish the truth of every response; repeated participation can condition participants; and detailed profiles create privacy obligations. Make these strengths, limits, and provenance visible, and charge only for the quality genuinely established.</p>

    <h2 id="practice">4. Set a higher standard for human research</h2>

    <p>Participants owe researchers nothing. Every invitation asks for time, attention, information, and trust, often from someone who has no prior relationship with the organisation asking. A higher standard begins by treating that contribution as something to earn rather than an input to extract. At minimum, a better research service should include:</p>

    <ul class="practice-list">
      <li><strong>Pointed surveys</strong> that test specific hypotheses rather than pushing broad, repetitive exploration onto participants.</li>
      <li><strong>Enough context</strong> to explain why the task matters without scripting the answer: an invitation to help with a decision, not a demand to surrender data.</li>
      <li><strong>Longitudinal measurement</strong> with the same people when the claim concerns individual change, alongside transparent repeated cross-sectional estimates where those are appropriate.</li>
      <li><strong>Rapid pulses</strong> for high-quality answers to small questions between scheduled waves, delivered at times participants have said work for them.</li>
      <li><strong>Interruption-safe continuity</strong> through progressive saving, visible checkpoints, exact resumption, and no repeated answer caused by a failure to retain or retrieve prior work.</li>
      <li><strong>Mobile-first tasks</strong> using large touch targets, short modules, the full device surface, and interaction modes matched to the participant.</li>
      <li><strong>Low cognitive burden</strong> through careful segmentation and prioritisation before fieldwork. If comparing twenty items is unpleasant for the research team, it should not outsource that burden to the participant.</li>
      <li><strong>Richer expression</strong> through voice, text, photographs, videos, diaries, receipts, and other relevant artefacts.</li>
      <li><strong>Unmistakable confirmation</strong> so people can see exactly what the system recorded or classified and amend it immediately.</li>
      <li><strong>Formal, separate validation</strong> of identity, category participation, attentive completion, consistency, and ability to use the response mode accurately.</li>
      <li><strong>Direct, predictable payment</strong> that reflects the time, sensitivity, and evidential value of what was requested.</li>
      <li><strong>Something returned</strong> where appropriate: a useful reflection, personal result, or clear account of how the contribution mattered.</li>
    </ul>

    <h3>Operate the complete journey as one service</h3>

    <p>Someone has to own the participant journey from invitation to final payment and recontact. Recruitment, screening, consent, the research task, support, correction, payment, feedback, withdrawal, and closure should be tested together before fieldwork begins. Outsourcing recruitment or hosting the instrument on another platform does not outsource responsibility for what the participant experiences.</p>

    <p>Maintaining a minimum standard requires operationalisation. “Respect participants” is a value, not yet a control. The service has to define what a participant must be able to observe, the threshold that must still hold when time or budgets are tight, the evidence that shows whether it happened, who owns the result, and how a failure is repaired. Progressive retention, visible confirmation, refusal without penalty, predictable payment, and correction that reaches every downstream use still under the operator's control are examples of standards that can actually be tested.</p>

    <p>Measurement validity and service quality need separate checks. A well-worded scale can still sit inside a confusing journey, while a polished interface can still collect the wrong construct. Test whether the instrument measures what the claim requires, and separately observe whether people understand the task, can provide the answer they intend, can recover from mistakes, and receive what they were promised.</p>

    <p>Completion is not enough evidence that the service worked. Record abandonment, forced or amended answers, repeated prompts, support requests, payment delays, complaints, refusals, and willingness to participate again. Review those signals by device, response mode, and relevant participant group after every wave. Participant corrections and complaints are not administrative noise; they show where the service and the resulting data may have failed together.</p>

    <p>Incentives often increase participation, but their effects depend on the mode, amount, timing, and population. In a meta-analysis of mail surveys, rewards included with the initial mailing increased response rates, while rewards conditional on return did not show the same effect <a class="cite" href="https://jaohar.com.au/human-data-collection.html#ref-church-1993">(Church, 1993)</a>. A broader review likewise found that the effects and costs of incentives vary across survey designs <a class="cite" href="https://jaohar.com.au/human-data-collection.html#ref-singer-2013">(Singer &amp; Ye, 2013)</a>. For a durable panel, payment also carries a relational message: careful participation is work, and the organisation values it.</p>

    <h2 id="future">5. Human research is a service-design opportunity</h2>

    <p>Humans belong wherever the evidence depends on lived experience, meaning, priorities, circumstances, or change within a person over time. Synthetic methods can narrow the field and help researchers ask sharper questions. They cannot provide the human evidence.</p>

    <p>Better human research data requires better-designed participation: sharper questions, richer ways to respond, and durable relationships with known participants.</p>

    <p>Organisations can continue buying access to a new sample for each study, asking the same screening questions, and accepting little visibility into the relationship behind the dataset. A higher standard is to understand and describe who is participating, use consented information without demanding it again, pay people properly, return to the same people when continuity matters, and give them ways to express more than a checkbox permits.</p>

    <p>Not every organisation can build the complete model immediately. It can still make the next study shorter, preserve an interrupted response, improve correction and payment, reuse an established attribute with permission, or return to the same participants for one important question. Each improvement can make participation more respectful and the resulting evidence more defensible.</p>

    <div class="article-cta">
      <strong>Want to improve how you collect human research data?</strong>
      <span>If you want to make participation clearer, more respectful, or better suited to the ways people can respond, I can assess the current experience and help design a different approach. <a href="https://jaohar.com.au/index.html#help">Send me a message.</a></span>
    </div>
  </div>
<section class="references">
    <h2>References</h2>
    <ol>
      <li id="ref-antoun-2017">Antoun, C., Couper, M. P., &amp; Conrad, F. G. (2017). Effects of mobile versus PC web on survey response quality: A crossover experiment in a probability web panel. <em>Public Opinion Quarterly, 81</em>(S1), 280&#8211;306. <a href="https://doi.org/10.1093/poq/nfw088" target="_blank" rel="noopener">https://doi.org/10.1093/poq/nfw088</a></li>
      <li id="ref-baker-2013">Baker, R., Brick, J. M., Bates, N. A., Battaglia, M., Couper, M. P., Dever, J. A., Gile, K. J., &amp; Tourangeau, R. (2013). Summary report of the AAPOR task force on non-probability sampling. <em>Journal of Survey Statistics and Methodology, 1</em>(2), 90&#8211;143. <a href="https://doi.org/10.1093/jssam/smt008" target="_blank" rel="noopener">https://doi.org/10.1093/jssam/smt008</a></li>
      <li id="ref-bolger-2003">Bolger, N., Davis, A., &amp; Rafaeli, E. (2003). Diary methods: Capturing life as it is lived. <em>Annual Review of Psychology, 54</em>, 579&#8211;616. <a href="https://doi.org/10.1146/annurev.psych.54.101601.145030" target="_blank" rel="noopener">https://doi.org/10.1146/annurev.psych.54.101601.145030</a></li>
      <li id="ref-church-1993">Church, A. H. (1993). Estimating the effect of incentives on mail survey response rates: A meta-analysis. <em>Public Opinion Quarterly, 57</em>(1), 62&#8211;79. <a href="https://doi.org/10.1086/269355" target="_blank" rel="noopener">https://doi.org/10.1086/269355</a></li>
      <li id="ref-csikszentmihalyi-1987">Csikszentmihalyi, M., &amp; Larson, R. (1987). Validity and reliability of the Experience-Sampling Method. <em>The Journal of Nervous and Mental Disease, 175</em>(9), 526&#8211;536. <a href="https://doi.org/10.1097/00005053-198709000-00004" target="_blank" rel="noopener">https://doi.org/10.1097/00005053-198709000-00004</a></li>
      <li id="ref-dejonckheere-2024">Dejonckheere, E., Verdonck, S., Andries, J., R&#246;hrig, N., Piot, M., Kilani, G., &amp; Mestdagh, M. (2024). Real-time incentivizing survey completion with game-based rewards in experience sampling research may increase data quantity, but reduces data quality. <em>Computers in Human Behavior, 160</em>, 108360. <a href="https://doi.org/10.1016/j.chb.2024.108360" target="_blank" rel="noopener">https://doi.org/10.1016/j.chb.2024.108360</a></li>
      <li id="ref-galesic-2009">Galesic, M., &amp; Bosnjak, M. (2009). Effects of questionnaire length on participation and indicators of response quality in a web survey. <em>Public Opinion Quarterly, 73</em>(2), 349&#8211;360. <a href="https://doi.org/10.1093/poq/nfp031" target="_blank" rel="noopener">https://doi.org/10.1093/poq/nfp031</a></li>
      <li id="ref-gill-2024">Gill, S. L. (2024). About research: Qualitative data collection: Photo elicitation. <em>Journal of Human Lactation, 40</em>(4), 503&#8211;505. <a href="https://doi.org/10.1177/08903344241273863" target="_blank" rel="noopener">https://doi.org/10.1177/08903344241273863</a></li>
      <li id="ref-kaczmirek-2019">Kaczmirek, L., Phillips, B., Pennay, D. W., Lavrakas, P. J., &amp; Neiger, D. (2019). <em>Building a probability-based online panel: Life in Australia</em> (Methods Paper No. 2/2019). ANU Centre for Social Research &amp; Methods and Social Research Centre. <a href="https://openresearch-repository.anu.edu.au/server/api/core/bitstreams/f8fe1125-02a4-441b-8ff9-317fbb171b91/content" target="_blank" rel="noopener">Open-access PDF</a></li>
      <li id="ref-kennedy-2019">Kennedy, C., &amp; Hartig, H. (2019). <em>Response rates in telephone surveys have resumed their decline</em>. Pew Research Center. <a href="https://www.pewresearch.org/short-reads/2019/02/27/response-rates-in-telephone-surveys-have-resumed-their-decline/" target="_blank" rel="noopener">pewresearch.org</a></li>
      <li id="ref-krosnick-1991">Krosnick, J. A. (1991). Response strategies for coping with the cognitive demands of attitude measures in surveys. <em>Applied Cognitive Psychology, 5</em>(3), 213&#8211;236. <a href="https://doi.org/10.1002/acp.2350050305" target="_blank" rel="noopener">https://doi.org/10.1002/acp.2350050305</a></li>
      <li id="ref-singer-2013">Singer, E., &amp; Ye, C. (2013). The use and effects of incentives in surveys. <em>The ANNALS of the American Academy of Political and Social Science, 645</em>(1), 112&#8211;141. <a href="https://doi.org/10.1177/0002716212458082" target="_blank" rel="noopener">https://doi.org/10.1177/0002716212458082</a></li>
      <li id="ref-stone-1994">Stone, A. A., &amp; Shiffman, S. (1994). Ecological momentary assessment (EMA) in behavioral medicine. <em>Annals of Behavioral Medicine, 16</em>(3), 199&#8211;202. <a href="https://doi.org/10.1093/abm/16.3.199" target="_blank" rel="noopener">https://doi.org/10.1093/abm/16.3.199</a></li>
      <li id="ref-westwood-2025">Westwood, S. J. (2025). The potential existential threat of large language models to online survey research. <em>Proceedings of the National Academy of Sciences, 122</em>(47), e2518075122. <a href="https://doi.org/10.1073/pnas.2518075122" target="_blank" rel="noopener">https://doi.org/10.1073/pnas.2518075122</a></li>
    </ol>
  </section>
]]></content:encoded>
    </item>
    <item>
      <title>Synthetic research: claims must match evidence</title>
      <link>https://jaohar.com.au/synthetic-research.html</link>
      <guid isPermaLink="true">https://jaohar.com.au/synthetic-research.html</guid>
      <pubDate>Wed, 29 Jul 2026 00:00:00 +1000</pubDate>
      <category>Synthetic research</category>
      <category>Evidence</category>
      <description>I built a synthetic-research prototype and tested it until I found the boundary between useful machine signals and claims about people.</description>
      <content:encoded><![CDATA[
<p class="standfirst">Over three weeks, I built a synthetic-research prototype and tested it until I found the boundary between useful machine evidence and claims about people.</p>
<p class="article-byline">First published <time datetime="2026-07-29">29 July 2026</time>. By Jonathan O’Hara.</p>
<nav class="article-summary-contents" aria-label="Article summary">
    <h2>Summary</h2>
    <ol>
      <li><a href="https://jaohar.com.au/synthetic-research.html#dune"><strong>1. Thinking machines are now research subjects</strong></a><p>AI systems influence how information is organised and decisions are made. We need ways to study what they produce, when their patterns hold, and where depending on them becomes unsafe.</p></li>
      <li><a href="https://jaohar.com.au/synthetic-research.html#definition"><strong>2. Synthetic systems are not synthetic respondents</strong></a><p>Studying a model supports claims about that system under recorded conditions. Treating generated answers as people makes a stronger claim that requires different evidence.</p></li>
      <li><a href="https://jaohar.com.au/synthetic-research.html#build"><strong>3. Testing stopped me from selling a convincing result</strong></a><p>I built a prototype that produced polished, plausible output. Testing showed that its central implication was not dependable enough, so I stopped preparing it as a product.</p></li>
      <li><a href="https://jaohar.com.au/synthetic-research.html#market-research"><strong>4. AI changes what market research must study</strong></a><p>Human research is under strain, while AI is entering information-heavy purchasing journeys. Models are becoming research objects, but studying them is not a shortcut to declaring what customers think.</p></li>
      <li><a href="https://jaohar.com.au/synthetic-research.html#evidence"><strong>5. Synthetic results are promising but bounded</strong></a><p>Some studies reproduce averages, effects, or individual patterns under demanding conditions. Variation, relationships, effect sizes, prompts, and representativeness can still fail in ways a plausible mean conceals.</p></li>
      <li><a href="https://jaohar.com.au/synthetic-research.html#standard"><strong>6. Claims must match methods</strong></a><p>Preserve the instrument, test sensitivity and failure, and operationalise the standard so unsupported claims can be stopped before they reach a reader.</p></li>
      <li><a href="https://jaohar.com.au/synthetic-research.html#triangulation"><strong>7. Hypotheses must leave the machine</strong></a><p>Synthetic work can generate propositions worth testing. Claims about people need independent data from people or behaviour, collected through faster and more respectful human research.</p></li>
      <li><a href="https://jaohar.com.au/synthetic-research.html#testing-appendix"><strong>8. A plausible ranking failed a simple reversal test</strong></a><p>The prototype completed thousands of comparisons and produced a tidy ranking. Reversing presentation order exposed a position effect strong enough to stop the proposed claim.</p></li>
    </ol>
  </nav>
<div class="body">
    <h2 id="dune">1. Thinking machines are now research subjects</h2>

    <p>Some people have <em>The Lord of the Rings</em>. I have <em>Dune</em>. I remember staying awake until four in the morning to finish the original novel. I loved the world, the films, and the miserable beauty of Arrakis.</p>

    <p>In <em>Dune</em>'s backstory, humanity created “Thinking machines”, suffered the consequences, and outlawed them after the Butlerian Jihad. Mentats, human beings trained as computational specialists, were developed to perform knowledge work that could no longer be entrusted to machines.</p>

    <p>Thinking machines are real now. There will be negative consequences. Nothing is gained by pretending AI will not influence almost everything. I am interested in the opportunity to make something positive with these tools without pretending the potential for harm is not there. My practical question is: how can we benefit from them, and in what contexts can we depend on them?</p>

    <h2 id="definition">2. Synthetic systems are not synthetic respondents</h2>

    <p>I was thinking of how to best describe what I do in plain language: <em>I survey thinking machines.</em> It reframes the work for me. Begin with the system, ask it structured questions, vary the conditions and learn whether its patterns hold.</p>

    <p><strong>Synthetic systems research</strong> is a term I use because it distinguishes this work from the synthetic-respondent research now attracting attention. Synthetic systems research treats the model or system itself as the thing being studied. Its basic claim is modest: this system produced this pattern under these conditions. Synthetic-respondent research makes the stronger claim that generated answers can stand in for what people, customers or a population would say.</p>

    <p>The distinction matters. A model can generate a plausible answer for a fictional Gen Z consumer in Melbourne. That answer may contain useful regularities from human language and behaviour. It does not contain that person's experience, obligations, private knowledge, or consequences. A generated backstory can become a persona, a synthetic respondent, an audience, and eventually a “digital twin” before anyone has established what is being twinned.</p>

    <h2 id="build">3. Testing stopped me from selling a convincing result</h2>

    <p>A model can answer thousands of questions in a day without recruitment, incentives or fieldwork. The headline result, supporting evidence and visualisations can exist before you have even engaged with a client. Human research is expensive, slow and burdensome to both the researcher and the person participating.</p>

    <p>Recently, over three weeks, I built a synthetic-research prototype that could run structured collections, repeat tasks, compare models and turn the output into an accessible plain-language report. Building it changed how I understood the benefits and limits of synthetic systems research. I intended it to generate hypotheses, not declare what customers thought.</p>

    <p>This prototype was intended to make claims such as: this ranking or pattern is stable enough to be worth checking. Several collection and reporting components worked as intended. The implied claims did not. Its output looked convincing, but the weakness was not apparent until I tested it. I could show that the system had produced an answer. I could not rely on what that answer implied strongly enough to support the full claimed proposition.</p>

    <p>Finding this limit was the consequence of doing the work. My background is in academic research, measurement, and applied data work. The danger would have been never finding it because the report looked finished. Once the central implication failed, I stopped preparing this new prototype as a product rather than make the evidence say more than it did.</p>

    <h2 id="market-research">4. AI changes what market research must study</h2>

    <p>Part of the attraction of synthetic research is that the survey methods market researchers have relied on are under increasing strain. There are more barriers to high-quality survey work than there used to be: falling participation <a class="cite" href="https://jaohar.com.au/synthetic-research.html#ref-kennedy-2019">(Kennedy &amp; Hartig, 2019)</a>, lengthy questionnaires that reduce both participation and later response quality <a class="cite" href="https://jaohar.com.au/synthetic-research.html#ref-galesic-2009">(Galesic &amp; Bosnjak, 2009)</a>, and incentives that do not always align with the researcher's need for careful, valid answers.</p>

    <p>Commercial panel members weigh the incentive, time and ease of completion when deciding whether to participate <a class="cite" href="https://jaohar.com.au/synthetic-research.html#ref-brosnan-2021">(Brosnan et al., 2021)</a>. Payment does not inherently produce bad data: research in a probability-based panel found little reason to worry that moderate incentives or incentive-motivated participation reduced response quality <a class="cite" href="https://jaohar.com.au/synthetic-research.html#ref-schwarz-2022">(Schwarz et al., 2022)</a>. The design problem is that payment is usually conditional on completion, while the quality we need is thoughtful participation.</p>

    <p>In one web-survey experiment, the fastest lower-education respondents were most vulnerable to response-order effects on unipolar rating scales. The same study warned that speed is only a proxy for satisficing among particular respondents and items, not a blunt exclusion rule <a class="cite" href="https://jaohar.com.au/synthetic-research.html#ref-malhotra-2008">(Malhotra, 2008)</a>.</p>

    <p>Researchers have built attention checks, timing thresholds and consistency tests around these problems. The defences are useful, but they are not perfect. A new problem now looms: AI agents create an attack surface because one person can use them to complete instruments repeatedly and at scale. Westwood's autonomous synthetic respondent passed 99.8% of 6,000 attention-check trials while producing coherent, persona-consistent answers and simulating ordinary interaction traces <a class="cite" href="https://jaohar.com.au/synthetic-research.html#ref-westwood-2025">(Westwood, 2025)</a>. Verifying that a panel member is a person does not establish that the answers were produced carefully by that person.</p>

    <p>The quality of the experience matters beyond one dataset. A confusing, repetitive or disrespectful survey teaches participants that surveys are not worth their attention. Poor instruments consume the trust on which better instruments also depend. Synthetic respondents look like a solution to this burden, but substituting generated answers does not repair the human claim. It changes what was measured.</p>

    <h3>Preference is often built, not retrieved</h3>

    <p>The familiar commercial story is that people hold stable preferences which careful questioning can retrieve. Consumer-research evidence has challenged that story for decades. People often construct a preference in response to the task, the available alternatives and the information in front of them <a class="cite" href="https://jaohar.com.au/synthetic-research.html#ref-bettman-1998">(Bettman, Luce, &amp; Payne, 1998)</a>. A brand tracker, conjoint task or attitude scale can therefore produce a consistent answer without recovering a durable object that existed before the interview.</p>

    <p>This matters most when the real decision is information-heavy. A buyer comparing a car, mortgage or business platform encounters explanations, alternatives, constraints and recommendations before choosing. The research task is not only to ask what is hidden inside the customer. It is also to understand the environment in which a preference takes shape.</p>

    <h3>AI is becoming part of that environment</h3>

    <p>AI is entering purchasing journeys, especially where decisions require more comparison and guidance. It is too early to call it a replacement for search: an analysis of 973 e-commerce sites found that ChatGPT referrals accounted for less than 0.2% of traffic and underperformed most established channels. The same study found stronger results in complex product categories, where people need more information, comparison and guidance <a class="cite" href="https://jaohar.com.au/synthetic-research.html#ref-kaiser-2026">(Kaiser &amp; Schulze, 2026)</a>. These information environments are becoming research objects in their own right.</p>

    <p>When a model is part of the decision context, its category account becomes observable. We can study what it makes prominent, what changes with the prompt, what survives repetition and where systems disagree. That is a real market-research object. It is evidence about the information environment, not a shortcut to declaring what customers think.</p>

    <h2 class="summary-target" id="evidence">5. Synthetic results are promising but bounded</h2>

    <p>The cleanest empirical questions concern the system itself. What does this model produce when asked this question under these conditions? Which patterns survive a change in wording, order, model or time? Those questions are real and useful. They do not require us to pretend the generated speakers are people.</p>

    <p>The model's fluency can obscure the distinction. The material looks credible because producing credible-looking material is the task. Unless we actively search for vulnerabilities, a confident answer can pass through the workflow without anyone seeing how much it depends on the context we supplied.</p>

    <p>There are credible positive results. Argyle and colleagues found that a language model conditioned on real survey backstories could reproduce nuanced patterns across several US human samples <a class="cite" href="https://jaohar.com.au/synthetic-research.html#ref-argyle-2023">(Argyle et al., 2023)</a>. Park and colleagues built agents for 1,052 people using extensive self-report data. Interview-only agents reached 83% of the participants' own two-week consistency benchmark. Agents combining interviews and surveys reached 86% <a class="cite" href="https://jaohar.com.au/synthetic-research.html#ref-park-2026">(Park et al., 2026)</a>.</p>

    <p>These were not strangers invented from a handful of demographic and behavioural descriptors. The stronger simulations began with extensive data from the actual individuals being modelled. Data from real people was the foundation of the method.</p>

    <p>Models can also predict some experimental results. Cui and colleagues attempted 156 replications from psychology and management. Models recovered 73% to 81% of main effects, although interactions were less reliable and effect sizes were consistently larger than in human studies <a class="cite" href="https://jaohar.com.au/synthetic-research.html#ref-cui-2025">(Cui et al., 2025)</a>. Ashokkumar and colleagues reported strong correlations between model predictions and 469 effects from 70 preregistered survey experiments. Their predictions still systematically overestimated effect sizes and performed less strongly in a second archive <a class="cite" href="https://jaohar.com.au/synthetic-research.html#ref-ashokkumar-2026">(Ashokkumar et al., 2026)</a>.</p>

    <p>This supports pilot testing, hypothesis generation and deciding which interventions are worth testing with people. It does not establish that an arbitrary synthetic sample represents a population.</p>

    <h3>Average resemblance is not enough</h3>

    <p>Bisbee and colleagues generated more than 3.6 million model responses matched to 7,530 participants in the American National Election Study. Some overall averages were close. The synthetic responses had less variation and frequently produced different relationships between variables. Among coefficients that differed significantly from their human counterparts, nearly a third changed sign <a class="cite" href="https://jaohar.com.au/synthetic-research.html#ref-bisbee-2024">(Bisbee et al., 2024)</a>.</p>

    <p>For me, a third of the materially different relationships changing direction is not an acceptable discrepancy. A reader can decide what level of error they will tolerate. I would not use that result to describe a market structure.</p>

    <p>A category average is rarely the whole commercial decision. We care about who differs, what predicts choice, where an effect reverses and how uncertain the estimate is. A plausible mean can conceal a false market structure. Claims about demographic or geographic representativeness remain hypotheses until they are tested against the people they claim to represent. If we do not look, we do not know.</p>

    <h3>The prompt helps make the person</h3>

    <p>The prompt does not merely describe the synthetic person. It helps create the answer. That makes persona construction part of the instrument, not an administrative preface. Wang, Morgenstern and Dickerson found that models used as human replacements could misportray identity groups, flatten differences within them and turn identity prompts into essentialised representations <a class="cite" href="https://jaohar.com.au/synthetic-research.html#ref-wang-2025">(Wang et al., 2025)</a>. Lutz and colleagues found that portrayals varied with role format, demographic priming and model <a class="cite" href="https://jaohar.com.au/synthetic-research.html#ref-lutz-2025">(Lutz et al., 2025)</a>.</p>

    <p>Wording, position, model, date and persona construction belong in the method, not behind the curtain. Pezeshkpour and Hruschka found large performance gaps when the same answer options were reordered <a class="cite" href="https://jaohar.com.au/synthetic-research.html#ref-pezeshkpour-2024">(Pezeshkpour &amp; Hruschka, 2024)</a>. Zhuo and colleagues found prompt sensitivity across models and tasks <a class="cite" href="https://jaohar.com.au/synthetic-research.html#ref-zhuo-2024">(Zhuo et al., 2024)</a>. Bisbee and colleagues found changes after small wording variations and across a three-month collection period <a class="cite" href="https://jaohar.com.au/synthetic-research.html#ref-bisbee-2024">(Bisbee et al., 2024)</a>. If these conditions are not tested and recorded, we do not know which part of the result belongs to the question and which part belongs to its presentation.</p>

    <h2 id="standard">6. Claims must match methods</h2>

    <p>Synthetic research will become cheap enough for almost anyone to feed in a brief and produce a polished result. The important question is what quality control sits between the ingredients and the thing handed to a client. Familiar research disciplines already provide much of what this new instrument needs.</p>

    <p>My own testing was exploratory: a fast attempt to find failure, not a formal validation study. It showed how much a polished final dataset can hide. Testing has to be designed into the work before the output starts looking authoritative. I advocate the following behaviours as a practical standard for synthetic research:</p>

    <ul>
      <li><strong>Define the claim, or label the exploration.</strong> “The model ranks these brands” and “customers rank these brands” are different studies.</li>
      <li><strong>Preserve the instrument and metadata.</strong> Keep prompts, examples, model identifiers, dates, parameters, transformations, exclusions and costs. Cost is not evidence, but it matters when deciding whether a method is usable.</li>
      <li><strong>Test the influence of wording and position.</strong> Use semantically equivalent prompts and rotate answer choices, labels and evidence.</li>
      <li><strong>Repeat across relevant conditions.</strong> One answer is not a measurement. Retest across fresh runs, model families, versions and time where stability matters, then report the degree of agreement. A more expensive model may provide a useful second opinion, but price is not a quality guarantee.</li>
      <li><strong>Use controls and leakage checks.</strong> Include invented alternatives, blinded labels and comparisons where no difference should exist. Check whether the prompt already contains the conclusion.</li>
      <li><strong>Track failure.</strong> Record failed responses, retries and the tasks on which they cluster. Retries can make a final table look complete while hiding how often the instrument failed along the way.</li>
      <li><strong>Inspect the structure of response values, not only the mean.</strong> Compare variance, distributions, subgroup differences, interactions and sign. Tidy data can still support the wrong substantive ranking.</li>
      <li><strong>Validate externally at the level of use.</strong> If the claim concerns people, purchases or markets, test it against independent human-research or behavioural data.</li>
    </ul>

    <p>Synthetic methods can expose assumptions, pretest instruments, search for failure cases and help decide where scarce human research should be spent. They need extensive testing, particularly when a method is extended beyond the people, domain or conditions in which it was evaluated.</p>

    <p>The minimum standards also have to be operationalised. Saying that a method is careful, respectful, transparent, or evidence-led does not maintain those qualities. The system has to turn each value into an observable behaviour, a threshold, an owner, evidence, a repair path, and a test at the point where the claim reaches a reader. For synthetic research, that means enforced claim boundaries, recorded provenance, visible caveats, refusal states, and checks that can stop an unsupported conclusion rather than merely warn about it in a methods note.</p>

    <p>With human grounding and statistical correction, synthetic methods may also extend limited datasets. Krsteski and colleagues found that synthetic generation alone produced substantial bias, while allocating human observations to rectification reduced that bias below 5% in their settings. Performance varied substantially across questions and datasets: a method that helped on one item could perform poorly elsewhere. That reinforces the need to test the method in the domain where it will be used <a class="cite" href="https://jaohar.com.au/synthetic-research.html#ref-krsteski-2026">(Krsteski et al., 2026)</a>.</p>

    <p>Use these systems for what they can demonstrate. Test the boundaries that matter. If the bridge fails, narrow the claim before the work reaches a decision.</p>

    <section class="article-breakout article-breakout-compact" aria-labelledby="claim-ladder">
    <h3 id="claim-ladder">A claim ladder for synthetic evidence</h3>

    <ol class="claim-ladder">
      <li><strong>System claim:</strong> this model produced this pattern under these recorded conditions.</li>
      <li><strong>Robust system claim:</strong> the pattern survived declared wording, order, repeat, model and time contrasts.</li>
      <li><strong>Externally grounded claim:</strong> the hypothesis was tested against independent human-research or behavioural data from the population, construct and setting it claims to describe.</li>
      <li><strong>Decision claim:</strong> the validated relationship is current, precise and relevant enough to guide this decision.</li>
    </ol>

    <p>Repeated generation can strengthen the first claim. Cross-model agreement can help with the second. Neither automatically reaches the third. A thousand samples from one underlying system do not become a thousand independently observed lives.</p>

    <p>A blended approach requires an explicit boundary. Before collecting anything, decide which claims belong to the system, which concern people, what evidence each method can provide, and what result would force the claim to narrow. Otherwise, a polished synthesis can quietly give one source authority it never earned.</p>
    </section>

    <h2 class="summary-target" id="triangulation">7. Hypotheses must leave the machine</h2>

    <p>Synthetic systems research can help us craft better hypotheses. But a hypothesis has to leave the machine. Putting it into another model may produce a better hypothesis, or simply hypotheses about hypotheses. If the claim concerns people, we need to ask real people real questions.</p>

    <p>No single research signal contains the whole truth. Surveys see declared answers. Conversation exposes language, context and contradiction. Behaviour shows what happened under actual constraints. A model shows what that system tends to produce from the information and framing available to it.</p>

    <p>These signals do not become interchangeable because they point in the same direction. In psychometrics, reliability concerns the consistency and precision of scores across relevant replications of a measurement procedure. Test-retest repetition is one form, not the whole concept <a class="cite" href="https://jaohar.com.au/synthetic-research.html#ref-standards-2014">(American Educational Research Association et al., 2014)</a>. Validity asks whether the evidence supports the interpretation and use we intend. A model can be highly reliable and still systematically wrong about the world.</p>

    <p>Hypotheses generated by synthetic results need a plan for what comes next. Triangulation is one way to sharpen and test them by taking more than one bearing. Use people, behaviour or a genuinely different method. If the signals disagree, investigate the disagreement instead of averaging it away. The synthesis becomes the actionable result, not whichever source produced the cleanest chart.</p>

    <p>Organisations will become accustomed to generating possibilities in hours. They will not want to wait weeks for a long conventional survey to validate every one. We need faster and more respectful ways to collect human research data: short quantitative decisions, better qualitative elicitation, voice and contextual methods, and hybrid designs that ask people only for the information the machine cannot credibly supply.</p>

    <p>The next challenge is to make human research faster, more respectful, and worth the time people give it. That is the subject of the companion article.</p>

    <div class="related-work-grid">
      <a class="related-work-card" href="https://jaohar.com.au/human-data-collection.html"><span>Methodology</span><strong>Where do humans fit in the future of research?</strong><small>Better participant experiences, longitudinal panels, mobile-first collection, and richer ways for people to respond.</small></a>
    </div>

    <div class="article-cta">
      <strong>Need an independent view of a synthetic-research claim?</strong>
      <span>If you are evaluating a report, method, product, purchase, or acquisition, I can test what it measured, identify unsupported claims, and ask the questions the decision requires. <a href="https://jaohar.com.au/index.html#help">Tell me about the work.</a></span>
    </div>

    <div class="testing-appendix" id="testing-appendix">
      <p class="eyebrow">Worked example</p>
      <h2>8. A plausible ranking failed a simple reversal test</h2>

      <p>This was an exploratory battery, not a formal validation study. I have withheld the subject, context and model names because the point is the failure mechanism, not the ranking the instrument produced.</p>

      <h3>The proposed claim</h3>
      <p>The new project was intended to report rankings associated with different outcomes and contexts. The output was complete and plausible. In the strongest run, a frontier model completed all 9,072 scheduled comparisons, and it used “insufficient basis” in only 2.58% of responses. Nothing about the finished ranking announced that it was unsafe.</p>

      <h3>The control</h3>
      <p>Each comparison was asked twice with the presentation positions reversed. That created 4,536 forward-and-back pairs. If the ranking reflected the responses rather than their placement, reversing the presentation should preserve the selected response.</p>

      <p>It did not. Only 68.35% of resolved pairs selected the same response after reversal. In the remaining 31.65%, the model changed its selection while retaining the same side of the screen. It was often following position while appearing to express a stable judgement.</p>

      <p>This problem was not specific to one model. A smaller model showed the same mechanism: among 4,025 resolved reversal pairs, 600 changed selection, and every one of those 600 switches preserved screen side. A more capable model did not remove the problem.</p>

      <h3>What I changed</h3>
      <p>I balanced presentation positions, repeated the comparisons in both directions, retained non-answers, moved from the smaller model to a stronger one and kept every raw response for rescoring. The ranking still failed the reversal threshold.</p>

      <p>I then designed a replacement that removed side-by-side choice entirely: one response target per prompt, exact repeats and neutral paraphrases, with the contrast calculated only after collection. I chose not to pursue it. At that point it would have been a new instrument measuring something different, not evidence that the failed ranking had become valid.</p>

      <h3>Why I stopped</h3>
      <p>This is the boundary that ended the three-week prototype experiment. The new protocol could collect a large volume of structured data and turn it into a convincing story. It could also reveal that the story depended on where a label appeared on the screen.</p>

      <p>The product I was trying to create during those three weeks needed sufficient stability to guide decisions or support hypotheses. This ranking could not survive its simplest hostile control. Continuing to tune prompts until it passed would have taught me how to manufacture a result, not whether the result deserved to be trusted. I stopped preparing this new prototype as a product and kept the more useful conclusion: <strong>a synthetic claim is only as strong as the testing it survives.</strong></p>
    </div>
  </div>
<section class="references">
    <h2>References</h2>
    <ol>
      <li id="ref-standards-2014">American Educational Research Association, American Psychological Association, &amp; National Council on Measurement in Education. (2014). <em>Standards for educational and psychological testing</em>. American Educational Research Association. <a href="https://www.testingstandards.net/uploads/7/6/6/4/76643089/standards_2014edition.pdf" target="_blank" rel="noopener">Open-access PDF</a></li>
      <li id="ref-argyle-2023">Argyle, L. P., Busby, E. C., Fulda, N., Gubler, J. R., Rytting, C., &amp; Wingate, D. (2023). Out of one, many: Using language models to simulate human samples. <em>Political Analysis, 31</em>(3), 337–351. <a href="https://doi.org/10.1017/pan.2023.2" target="_blank" rel="noopener">https://doi.org/10.1017/pan.2023.2</a></li>
      <li id="ref-ashokkumar-2026">Ashokkumar, A., Hewitt, L., Ghezae, I., &amp; Willer, R. (2026). Large language models can predict the results of social science experiments. <em>Nature</em>. Advance online publication. <a href="https://doi.org/10.1038/s41586-026-10742-x" target="_blank" rel="noopener">https://doi.org/10.1038/s41586-026-10742-x</a></li>
      <li id="ref-bettman-1998">Bettman, J. R., Luce, M. F., &amp; Payne, J. W. (1998). Constructive consumer choice processes. <em>Journal of Consumer Research, 25</em>(3), 187–217. <a href="https://doi.org/10.1086/209535" target="_blank" rel="noopener">https://doi.org/10.1086/209535</a></li>
      <li id="ref-bisbee-2024">Bisbee, J., Clinton, J. D., Dorff, C., Kenkel, B., &amp; Larson, J. M. (2024). Synthetic replacements for human survey data? The perils of large language models. <em>Political Analysis, 32</em>(4), 401–416. <a href="https://doi.org/10.1017/pan.2024.5" target="_blank" rel="noopener">https://doi.org/10.1017/pan.2024.5</a></li>
      <li id="ref-brosnan-2021">Brosnan, K., Kemperman, A., &amp; Dolnicar, S. (2021). Maximizing participation from online survey panel members. <em>International Journal of Market Research, 63</em>(4), 416–435. <a href="https://doi.org/10.1177/1470785319880704" target="_blank" rel="noopener">https://doi.org/10.1177/1470785319880704</a></li>
      <li id="ref-cui-2025">Cui, Z., Li, N., &amp; Zhou, H. (2025). A large-scale replication of scenario-based experiments in psychology and management using large language models. <em>Nature Computational Science, 5</em>(8), 627–634. <a href="https://doi.org/10.1038/s43588-025-00840-7" target="_blank" rel="noopener">https://doi.org/10.1038/s43588-025-00840-7</a></li>
      <li id="ref-galesic-2009">Galesic, M., &amp; Bosnjak, M. (2009). Effects of questionnaire length on participation and indicators of response quality in a web survey. <em>Public Opinion Quarterly, 73</em>(2), 349–360. <a href="https://doi.org/10.1093/poq/nfp031" target="_blank" rel="noopener">https://doi.org/10.1093/poq/nfp031</a></li>
      <li id="ref-kaiser-2026">Kaiser, M., &amp; Schulze, C. (2026). Frontiers: ChatGPT referrals to e-commerce websites: How do LLMs compare against traditional channels? <em>Marketing Science</em>. Advance online publication. <a href="https://doi.org/10.1287/mksc.2025.0489" target="_blank" rel="noopener">https://doi.org/10.1287/mksc.2025.0489</a></li>
      <li id="ref-kennedy-2019">Kennedy, C., &amp; Hartig, H. (2019). <em>Response rates in telephone surveys have resumed their decline</em>. Pew Research Center. <a href="https://www.pewresearch.org/short-reads/2019/02/27/response-rates-in-telephone-surveys-have-resumed-their-decline/" target="_blank" rel="noopener">pewresearch.org</a></li>
      <li id="ref-malhotra-2008">Malhotra, N. (2008). Completion time and response order effects in web surveys. <em>Public Opinion Quarterly, 72</em>(5), 914–934. <a href="https://doi.org/10.1093/poq/nfn050" target="_blank" rel="noopener">https://doi.org/10.1093/poq/nfn050</a></li>
      <li id="ref-krsteski-2026">Krsteski, S., Russo, G., Chang, S., West, R., &amp; Gligorić, K. (2026). Valid survey simulations with limited human data: The roles of prompting, fine-tuning, and rectification. In <em>Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)</em> (pp. 10887–10906). Association for Computational Linguistics. <a href="https://doi.org/10.18653/v1/2026.acl-long.498" target="_blank" rel="noopener">https://doi.org/10.18653/v1/2026.acl-long.498</a></li>
      <li id="ref-lutz-2025">Lutz, M., Sen, I., Ahnert, G., Rogers, E., &amp; Strohmaier, M. (2025). The prompt makes the person(a): A systematic evaluation of sociodemographic persona prompting for large language models. In <em>Findings of the Association for Computational Linguistics: EMNLP 2025</em> (pp. 23212–23237). Association for Computational Linguistics. <a href="https://doi.org/10.18653/v1/2025.findings-emnlp.1261" target="_blank" rel="noopener">https://doi.org/10.18653/v1/2025.findings-emnlp.1261</a></li>
      <li id="ref-park-2026">Park, J. S., Zou, C. Q., Kamphorst, J., Egan, N., Shaw, A., Hill, B. M., Cai, C., Morris, M. R., Liang, P., Willer, R., &amp; Bernstein, M. S. (2026). <em>LLM agents grounded in self-reports enable general-purpose simulation of individuals</em> (Version 3) [Preprint]. arXiv. <a href="https://doi.org/10.48550/arXiv.2411.10109" target="_blank" rel="noopener">https://doi.org/10.48550/arXiv.2411.10109</a></li>
      <li id="ref-pezeshkpour-2024">Pezeshkpour, P., &amp; Hruschka, E. (2024). Large language models sensitivity to the order of options in multiple-choice questions. In <em>Findings of the Association for Computational Linguistics: NAACL 2024</em> (pp. 2006–2017). Association for Computational Linguistics. <a href="https://doi.org/10.18653/v1/2024.findings-naacl.130" target="_blank" rel="noopener">https://doi.org/10.18653/v1/2024.findings-naacl.130</a></li>
      <li id="ref-schwarz-2022">Schwarz, H., Revilla, M., &amp; Struminskaya, B. (2022). Do previous survey experience and participating due to an incentive affect response quality? Evidence from the CRONOS panel. <em>Journal of the Royal Statistical Society: Series A (Statistics in Society), 185</em>(3), 981–1003. <a href="https://doi.org/10.1111/rssa.12857" target="_blank" rel="noopener">https://doi.org/10.1111/rssa.12857</a></li>
      <li id="ref-westwood-2025">Westwood, S. J. (2025). The potential existential threat of large language models to online survey research. <em>Proceedings of the National Academy of Sciences, 122</em>(47), e2518075122. <a href="https://doi.org/10.1073/pnas.2518075122" target="_blank" rel="noopener">https://doi.org/10.1073/pnas.2518075122</a></li>
      <li id="ref-wang-2025">Wang, A., Morgenstern, J., &amp; Dickerson, J. P. (2025). Large language models that replace human participants can harmfully misportray and flatten identity groups. <em>Nature Machine Intelligence, 7</em>(3), 400–411. <a href="https://doi.org/10.1038/s42256-025-00986-z" target="_blank" rel="noopener">https://doi.org/10.1038/s42256-025-00986-z</a></li>
      <li id="ref-zhuo-2024">Zhuo, J., Zhang, S., Fang, X., Duan, H., Lin, D., &amp; Chen, K. (2024). ProSA: Assessing and understanding the prompt sensitivity of LLMs. In <em>Findings of the Association for Computational Linguistics: EMNLP 2024</em> (pp. 1950–1976). Association for Computational Linguistics. <a href="https://doi.org/10.18653/v1/2024.findings-emnlp.108" target="_blank" rel="noopener">https://doi.org/10.18653/v1/2024.findings-emnlp.108</a></li>
    </ol>
  </section>
]]></content:encoded>
    </item>
  </channel>
</rss>
