Where do humans fit in the future of research?
Humans belong wherever research asks about priorities, circumstances, trade-offs, and lived experience. Synthetic methods can help decide what to ask; they cannot supply those answers. If hypotheses must leave the machine, the next question is what kind of human research deserves the time and trust we ask people to give it.
Summary
Jump to a section
- Participant experience is part of data quality
- Match the method to the question, person, and device
- Build panels with relationships
- Set a higher standard for human research
- Human research is a service-design opportunity
Reader’s noteAbout 3,600 words · 18-minute read (references excluded)
1Participant experience is part of data quality
I have conducted social research for years, and I participate whenever I have the opportunity. Seeing research from both sides keeps one fact visible: the instrument is also an experience, and the quality of that experience shapes what people provide.
Research practice often treats the instrument as the method and the participant experience as administration around it. The questionnaire has specifications. Recruitment, interruption, correction, payment, and closure are managed separately. That separation looks efficient, but it makes it easy for no one to own the conditions in which the data is produced.
What participation reveals
Participating in other people's research makes recurring problems difficult to miss. Instruments can be too long, repetitive, short on context, and weakened by careless screening and branching. I can understand a weak survey produced by someone doing so for the first time. What is harder to accept is the same experience from organisations fielding surveys at large scale. They may have strict requirements for what the instrument must collect, yet no equivalent requirement for the experience to be good for the participant.
One reason these surveys remain weak is that questions are crammed into one of many survey platforms without reconsidering the experience. We have largely replicated the paper questionnaire online rather than asking what an online-first experience should be. Online-first means device-first, and today that also means mobile-first.
The best questionnaires move naturally, make the context of each question clear, and prepare the participant for what comes next. They make it easy to provide the information the researcher actually needs. Their quality makes the prevailing standard harder to excuse.
The case for standardisation is real. Research teams need comparable measures, controlled costs, and instruments they can repeat. They cannot turn every survey into a bespoke service or preserve every form of expression. A well-designed participant journey can preserve those controls. The practical question is which parts need to stay fixed and which parts must adapt so a participant can understand the task and provide the answer intended.
What the data reveals
The participant experience directly affects the data. I have spent much of my time analysing large survey datasets, so I began to see where people checked out, where a task stopped working, and where an instrument began producing answers that were technically complete but difficult to believe. A finished report rarely shows how much unusable data was removed, or how much doubtful data remained because the study still had to produce a result.
I learned the same lesson earlier, while conducting computer-assisted telephone interviews. I was reprimanded for straying from the script when a question prompted someone to describe a difficult experience. The correct operational response was to acknowledge as little as possible and return them to option A or B. I was not especially good at that job because people do not naturally answer as instruments require. Older participants in particular often wanted to tell me what happened. The questionnaire demanded a fixed response.
An online form hides that encounter. If none of the answers fit, a participant who wants the incentive still has to give the system what it demands: an answer. Paper surveys occasionally reveal what the form excluded. While entering data from a regional public-health study, I found comments written in the margins: profane, profound, and sometimes more immediately relevant than the questions we had asked. People wanted their stories to be heard even when the instrument had nowhere to put them.
The interview and the handwritten margins revealed the same conflict. People were trying to tell us what had happened to them. The instrument was trying to turn that account into a permitted response. Good human research has to preserve enough structure to answer the research question without treating the part that does not fit as noise.
Design for interruption
Research is rarely the highest priority in a participant's life. People answer while travelling, working, caring for someone, waiting for an appointment, losing reception, or simply running out of attention. Interruption should be treated as expected behaviour, not participant failure.
A well-designed service saves each meaningful contribution as it is made, shows the participant what has been retained, and lets them resume from their last completed thought rather than merely reopening the last screen. Once somebody has given us an answer, our process should not make them give it again because we failed to hold or retrieve it. Re-asking can still be legitimate when a participant is revising an answer, reconfirming something that may have changed, resolving a contradiction, or contributing to an intentional time series. The service should say which of those it is doing and why.
The useful invitation is not “please start again”. It is “what you gave us has been retained; continue when you can”. That promise has to include correction, withdrawal, deletion, and recontact choices, with any limits explained before participation rather than discovered afterwards.
US telephone poll response rates have fallen for decades (Kennedy & Hartig, 2019). Respondents under cognitive load may satisfice rather than answer carefully (Krosnick, 1991). In a web-survey experiment, longer stated questionnaires reduced starts and completions, while questions placed later produced faster, shorter, and more uniform answers (Galesic & Bosnjak, 2009). Westwood showed that an autonomous AI respondent could pass standard attention checks while producing coherent, persona-consistent answers (Westwood, 2025). We cannot keep offering people weak incentives and poor experiences, then treat declining participation as something participants have done to us.
Spend human attention carefully
Where possible, do not push a raw question battery onto participants merely because the research team has not narrowed it. Synthetic methods can map a topic, compare structures, screen possibilities, and help sharpen hypotheses before human attention is spent. Client knowledge and existing evidence should do the same work.
A team that begins with a hundred possible concepts might use those signals to identify twelve worth investigating, then ask people to evaluate the twelve. The synthetic stage reduces burden; it does not become evidence of what people prefer. Questions about priorities, trade-offs, circumstances, and lived experience still need answers from people.
I develop that boundary in Synthetic research: claims must match evidence, including what synthetic work can reveal, how it should be tested, and what it cannot claim.
Once human attention is reserved for questions only people can answer, the next design choice is how to collect that evidence.
2Match the method to the question, person, and device
A long web form is only one way to collect human research data. Some questions are suited to a rating or a short choice. Others need a person's own words, a photograph, a short recording, a diary entry, a receipt, or another artefact from the experience. Asking someone to compress a complicated experience into a number can discard the very information we hoped to understand.
Photo elicitation, for example, can broaden and deepen a qualitative interview (Gill, 2024). The principle extends to videos, screenshots, proof of purchase, and material captured at the moment something happens. If entry into a category matters, verified evidence of that entry can be more valuable than another unsupported screening answer. That higher standard of evidence should also earn the participant a higher reward.
Use the device
These experiences should be mobile-first, not merely desktop questionnaires squeezed onto a smaller screen. A randomised crossover experiment found that smartphone respondents could provide careful answers, but small sliders and date pickers created input errors; the task has to be easy to perform on a touchscreen (Antoun, Couper, & Conrad, 2017). We do not have to inherit standard form controls. We can use the whole surface of the device to make one task unmistakable, show exactly what was recorded, and make correction easy.
That does not mean every participant should be expected to swipe, drag, or manipulate the same control. Dexterity, sensation, vision, and familiarity with a device differ. Offer a brief practice step, clear feedback, and an alternative way to answer. The purpose is to help people respond accurately through an interaction they can use.
Interface quality is data quality
I have seen what happens when an instrument is not fit for purpose. In an academic health-data role, I worked with records entered by numerous nurses through an interface that made year of birth easy to enter incorrectly. Implausible dates appeared systematically. For studies of children's growth, date of birth was not a minor field; it was essential to the analysis. The problem was well known and remained unresolved. Interface quality had become data quality.
Making a task easier or more satisfying does not require gamification. In one experience-sampling experiment, virtual rewards increased responses among the people who used them most, but made their momentary reports slightly less reliable (Dejonckheere et al., 2024). Points and badges can become another demand placed on the participant. The better aim is a satisfying interaction: clear progress, appropriate feedback, useful variety, and, where possible, something of value returned to the person.
Two examples of designed participation
These examples put the method into visible form. The first is an interactive prototype; the second is a visual exploration. Neither is a finished product or validated instrument. They are design hypotheses about how a familiar research task might change when participation is designed for the device and the person using it.
Designed for the device
A mobile-first choice task exploring how the device itself can support new response modes: full-screen swipes, a “too close to call” response, and interaction data retained for later validation.
Try the prototype →
A voice check-in
A visual exploration of short, paid voice reflections collected over time, with the participant's experience at the centre.
View the example →
The choice task demonstrates one question at a time, a response that does not force a false preference, and interaction traces retained separately for later testing. The voice example demonstrates richer expression, repeated participation, visible confirmation, and direct payment. Neither establishes that the measure is valid, accessible to every participant, or better than an existing instrument. Those questions require separate evaluation.
Recruitment, consent, screening, participation, support, payment, feedback, and recontact form one service from the participant's point of view. That becomes especially visible when researchers return to the same people. Optimising the instrument while neglecting the surrounding service cannot create a relationship worth maintaining.
When the research question concerns change over time, that service must become a relationship.
3Build panels with relationships
When change is the claim
A fresh sample at each wave can estimate how an aggregate measure differs from one point to the next. It cannot tell us how the same people's views or circumstances changed. If 200 people answer in January and a different 200 answer in April, a flat average can conceal individuals moving in opposite directions, while an apparent shift can partly reflect who happened to enter each sample. If the decision depends on within-person change, the design needs within-person data.
That makes longitudinal relationships especially valuable for commercial tracking. Instead of repeatedly drawing convenient samples from an opaque pool, researchers should be able to return to known participants and ask stable questions over time. Diary methods and experience sampling already show the value of repeated, in-the-moment evidence (Csikszentmihalyi & Larson, 1987; Stone & Shiffman, 1994; Bolger, Davis, & Rafaeli, 2003).
Academic researchers often work under severe constraints. An undergraduate convenience sample, a prize draw, or a small one-off study may be the only feasible way to see whether an early idea has any promise. Commercial research that informs real expenditure has a different opportunity. There is money in the system, and some of it should be used to improve the sample, the relationship, and the experience. Not every improvement costs more, but the parts that do should be treated as research infrastructure rather than avoidable overhead.
A relationship worth maintaining
I am interested in how curated, longitudinal commercial panels can be designed, developed, and maintained: known and segmented groups of people whom researchers can return to, who understand the relationship they are entering, and who are paid fairly and predictably. Their profiles can become richer with consent, their participation can span scheduled studies and rapid pulses, and the provenance of the sample can be described instead of hidden behind a provider's assurance.
An established relationship can also stop every study from beginning at zero. With permission, a panel can maintain verified attributes that are relevant across studies, such as age range, location, household circumstances, category participation, or previous research activity. A researcher can then request the participants who fit the study without asking each person to repeat the same screening and demographic questions. The study should receive only the attributes it needs, and participants should be able to see why those attributes are being used.
Build a panel that can be described honestly rather than promising perfect representation. Recruit people from relevant communities, establish their characteristics with a small set of screening questions, and report results by known groups rather than implying that every result generalises to everyone. The limits of non-probability samples still apply (Baker et al., 2013), but those limits can be described instead of concealed. Participants should also be paid directly for the time and evidence they provide.
Trust has to run in both directions. Researchers often design as though the participant is an adversary to be caught speeding, straightlining, or answering carelessly. Participants have their own reasons not to trust us: repetitive questions, unexplained purposes, weak screening, thoughtless wording, delayed payment, and no accessible way to correct a problem. Every invitation asks them to take a risk on whether this will be a good survey or a bad one. Onboarding has to earn trust rather than merely test compliance.
A durable panel can also make participation habitual without making it extractive. Ask people when they can genuinely focus, perhaps a regular half-hour on Friday morning, and return at the time they chose. A predictable rhythm respects the fact that research invitations otherwise arrive between work, care, and household tasks. Rapid pulses then become part of a relationship, not an unexpected demand.
Give participants control of the relationship
A richer profile creates obligations as well as efficiencies. Participants should be able to inspect and correct what the panel holds about them, understand which attributes will be reused, choose whether they can be recontacted, and withdraw from the relationship. Withdrawal, profile correction, and deletion should be usable parts of the service rather than rights buried in a policy. Consent to join a panel is not unlimited consent for every later use.
The operator also has to monitor what continuity changes. Repeated participation can condition answers, long-serving members may differ from people who leave, and the most burdened participants may be the first to disappear. Track participation frequency, refusals, attrition, profile completeness, and replenishment by relevant group. A longitudinal panel produces better evidence only when those changes remain visible.
What already exists
Pieces of this future already exist. Prolific's longitudinal projects support multi-wave studies that return to the same participants, show them the schedule and payment upfront, and track retention across waves. Dscout's diary tools support repeated mobile activities with photographs, videos, and screen recordings. Australia's probability-based Life in Australia panel demonstrates that a repeatedly contacted national panel can be deliberately recruited, maintained, and evaluated (Kaczmirek et al., 2019). Those are important capabilities. The larger opportunity is to combine continuity, good instrumentation, richer modes, transparent sample quality, and a relationship participants would choose to maintain.
Evidence about who participated can strengthen a panel without validating every answer. A receipt may support category membership, but it does not establish the truth of every response; repeated participation can condition participants; and detailed profiles create privacy obligations. Make these strengths, limits, and provenance visible, and charge only for the quality genuinely established.
These obligations become a standard only when one owner can test the complete journey.
4Set a higher standard for human research
Participants owe researchers nothing. Every invitation asks for time, attention, information, and trust, often from someone who has no prior relationship with the organisation asking. A higher standard begins by treating that contribution as something to earn rather than an input to extract. At minimum, a better research service should include:
- Pointed surveys that test specific hypotheses rather than pushing broad, repetitive exploration onto participants.
- Enough context to explain why the task matters without scripting the answer: an invitation to help with a decision, not a demand to surrender data.
- Longitudinal measurement with the same people when the claim concerns individual change, alongside transparent repeated cross-sectional estimates where those are appropriate.
- Rapid pulses for high-quality answers to small questions between scheduled waves, delivered at times participants have said work for them.
- Interruption-safe continuity through progressive saving, visible checkpoints, exact resumption, and no repeated answer caused by a failure to retain or retrieve prior work.
- Mobile-first tasks using large touch targets, short modules, the full device surface, and interaction modes matched to the participant.
- Low cognitive burden through careful segmentation and prioritisation before fieldwork. If comparing twenty items is unpleasant for the research team, it should not outsource that burden to the participant.
- Richer expression through voice, text, photographs, videos, diaries, receipts, and other relevant artefacts.
- Unmistakable confirmation so people can see exactly what the system recorded or classified and amend it immediately.
- Formal, separate validation of identity, category participation, attentive completion, consistency, and ability to use the response mode accurately.
- Direct, predictable payment that reflects the time, sensitivity, and evidential value of what was requested.
- Something returned where appropriate: a useful reflection, personal result, or clear account of how the contribution mattered.
Operate the complete journey as one service
Someone has to own the participant journey from invitation to final payment and recontact. Recruitment, screening, consent, the research task, support, correction, payment, feedback, withdrawal, and closure should be tested together before fieldwork begins. Outsourcing recruitment or hosting the instrument on another platform does not outsource responsibility for what the participant experiences.
Maintaining a minimum standard requires operationalisation. “Respect participants” is a value, not yet a control. The service has to define what a participant must be able to observe, the threshold that must still hold when time or budgets are tight, the evidence that shows whether it happened, who owns the result, and how a failure is repaired. Progressive retention, visible confirmation, refusal without penalty, predictable payment, and correction that reaches every downstream use still under the operator's control are examples of standards that can actually be tested.
Measurement validity and service quality need separate checks. A well-worded scale can still sit inside a confusing journey, while a polished interface can still collect the wrong construct. Test whether the instrument measures what the claim requires, and separately observe whether people understand the task, can provide the answer they intend, can recover from mistakes, and receive what they were promised.
Completion is not enough evidence that the service worked. Record abandonment, forced or amended answers, repeated prompts, support requests, payment delays, complaints, refusals, and willingness to participate again. Review those signals by device, response mode, and relevant participant group after every wave. Participant corrections and complaints are not administrative noise; they show where the service and the resulting data may have failed together.
Incentives often increase participation, but their effects depend on the mode, amount, timing, and population. In a meta-analysis of mail surveys, rewards included with the initial mailing increased response rates, while rewards conditional on return did not show the same effect (Church, 1993). A broader review likewise found that the effects and costs of incentives vary across survey designs (Singer & Ye, 2013). For a durable panel, payment also carries a relational message: careful participation is work, and the organisation values it.
5Human research is a service-design opportunity
Humans belong wherever the evidence depends on lived experience, meaning, priorities, circumstances, or change within a person over time. Synthetic methods can narrow the field and help researchers ask sharper questions. They cannot provide the human evidence.
Better human research data requires better-designed participation: sharper questions, richer ways to respond, and durable relationships with known participants.
This matters to research commissioners, panel providers, and anyone making decisions from the resulting data. A poor experience does more than frustrate a participant. It increases abandonment, encourages forced or hurried answers, weakens trust in later invitations, and makes the provenance of a finished dataset harder to defend.
Organisations can continue buying access to a new sample for each study, asking the same screening questions, and accepting little visibility into the relationship behind the dataset. A higher standard is to understand and describe who is participating, use consented information without demanding it again, pay people properly, return to the same people when continuity matters, and give them ways to express more than a checkbox permits.
Not every organisation can build the complete model immediately. It can still make the next study shorter, preserve an interrupted response, improve correction and payment, reuse an established attribute with permission, or return to the same participants for one important question. Each improvement can make participation more respectful and the resulting evidence more defensible.
References14 sources
- Antoun, C., Couper, M. P., & Conrad, F. G. (2017). Effects of mobile versus PC web on survey response quality: A crossover experiment in a probability web panel. Public Opinion Quarterly, 81(S1), 280–306. https://doi.org/10.1093/poq/nfw088
- Baker, R., Brick, J. M., Bates, N. A., Battaglia, M., Couper, M. P., Dever, J. A., Gile, K. J., & Tourangeau, R. (2013). Summary report of the AAPOR task force on non-probability sampling. Journal of Survey Statistics and Methodology, 1(2), 90–143. https://doi.org/10.1093/jssam/smt008
- Bolger, N., Davis, A., & Rafaeli, E. (2003). Diary methods: Capturing life as it is lived. Annual Review of Psychology, 54, 579–616. https://doi.org/10.1146/annurev.psych.54.101601.145030
- Church, A. H. (1993). Estimating the effect of incentives on mail survey response rates: A meta-analysis. Public Opinion Quarterly, 57(1), 62–79. https://doi.org/10.1086/269355
- Csikszentmihalyi, M., & Larson, R. (1987). Validity and reliability of the Experience-Sampling Method. The Journal of Nervous and Mental Disease, 175(9), 526–536. https://doi.org/10.1097/00005053-198709000-00004
- Dejonckheere, E., Verdonck, S., Andries, J., Röhrig, N., Piot, M., Kilani, G., & Mestdagh, M. (2024). Real-time incentivizing survey completion with game-based rewards in experience sampling research may increase data quantity, but reduces data quality. Computers in Human Behavior, 160, 108360. https://doi.org/10.1016/j.chb.2024.108360
- Galesic, M., & Bosnjak, M. (2009). Effects of questionnaire length on participation and indicators of response quality in a web survey. Public Opinion Quarterly, 73(2), 349–360. https://doi.org/10.1093/poq/nfp031
- Gill, S. L. (2024). About research: Qualitative data collection: Photo elicitation. Journal of Human Lactation, 40(4), 503–505. https://doi.org/10.1177/08903344241273863
- Kaczmirek, L., Phillips, B., Pennay, D. W., Lavrakas, P. J., & Neiger, D. (2019). Building a probability-based online panel: Life in Australia (Methods Paper No. 2/2019). ANU Centre for Social Research & Methods and Social Research Centre. Open-access PDF
- Kennedy, C., & Hartig, H. (2019). Response rates in telephone surveys have resumed their decline. Pew Research Center. pewresearch.org
- Krosnick, J. A. (1991). Response strategies for coping with the cognitive demands of attitude measures in surveys. Applied Cognitive Psychology, 5(3), 213–236. https://doi.org/10.1002/acp.2350050305
- Singer, E., & Ye, C. (2013). The use and effects of incentives in surveys. The ANNALS of the American Academy of Political and Social Science, 645(1), 112–141. https://doi.org/10.1177/0002716212458082
- Stone, A. A., & Shiffman, S. (1994). Ecological momentary assessment (EMA) in behavioral medicine. Annals of Behavioral Medicine, 16(3), 199–202. https://doi.org/10.1093/abm/16.3.199
- Westwood, S. J. (2025). The potential existential threat of large language models to online survey research. Proceedings of the National Academy of Sciences, 122(47), e2518075122. https://doi.org/10.1073/pnas.2518075122