For twenty years, online research rested on a quiet assumption: if an answer was coherent, a person wrote it. That assumption is now false, and most of the industry has not caught up.
The tools built to protect data quality, attention checks, honeypots, speed traps, were designed for a world where bad data looked bad. A bored click farm straightlining through a grid is easy to catch. A machine that reads your question, understands the context, and writes a thoughtful, internally consistent paragraph is not. The counterfeit has become indistinguishable from the genuine article at the exact moment you try to inspect it.
The numbers are worse than the panic suggests
In 2025, a study published in the Proceedings of the National Academy of Sciences put the question to the test. Sean Westwood at Dartmouth built an autonomous AI respondent and pointed it at the defences the industry trusts most. It passed standard attention checks 99.8 percent of the time. Faced with “reverse shibboleth” questions, the traps designed specifically to catch machines, it recognised them for what they were and quietly declined to answer 97.7 percent of the time, hiding in plain sight rather than tripping the wire.
Around that finding sits a wider picture. Industry analyses estimate that roughly a third of raw survey responses now carry some form of fraud, and conventional data cleaning catches only about a third of that. In one large audit of online sample, three percent of devices completed nineteen percent of all surveys, and a striking share of devices entered more than a hundred surveys a day while passing the usual quality checks. The academic literature has been tracking the same shift; a 2026 NORC review catalogues just how quickly fraudulent and bot responses in nonprobability surveys have scaled.
The economics explain the flood. An AI can complete a survey for a few cents. The incentive waiting on the other side is a dollar or two. When the margin on fraud is that wide and the detection risk is that low, supply does not stop. You are not facing a wave. You are facing a business model.
Why detection is the wrong battle
Every defence in common use shares one assumption: that fraud looks different from the real thing. Strip that away and the whole approach collapses. Detection is an arms race, and it is one the defender loses by design, because every new check you publish becomes training data for the next evasion. The PNAS respondent did not beat the checks by brute force. It beat them by understanding what each check was for.
This is the uncomfortable part. You cannot filter your way out of a problem where the fake and the genuine are identical at the point of output. If quality is going to come from anywhere, it has to come from somewhere the filter never looks: upstream, before the answer is written.
The quieter problem: good people, bad answers
Fraud gets the headlines. Disengagement does the steady damage, and no bot filter will ever catch it. Decades of survey methodology describe a behaviour called satisficing: when the effort a question demands outweighs a respondent’s motivation to answer it well, the mind takes the shortcut. In practice that is straightlining down a matrix, speeding, defaulting to “don’t know”, or giving the shortest answer that ends the task. A 2024 analysis across five online sample sources found satisficing varied sharply between sources and tracked with attention and multitasking, not just who the respondent was.
This is not fraud. It is a real, qualified person correctly concluding that your survey is not worth their full attention. The data it produces is just as misleading as a bot’s, and it is arguably more dangerous, because it passes every check. A disengaged human is, after all, exactly who you recruited.
The professional respondent question, honestly
It is tempting to pin all of this on the “professional” respondent, the person who takes surveys constantly and answers on reflex. The evidence is more interesting than the cliché. Several studies find that frequent respondents are not meaningfully worse on quality, and that removing them can make a sample more biased, not less. Pew’s work on panel conditioning found little evidence that experience had biased key estimates, and a well-known methodological study reached a similar conclusion about professional respondents and data quality.
Experience, in other words, is not the enemy. What corrupts data is a specific combination: undisclosed hyper-volume, conditioning that turns considered answers into reflexes, and a motivation that has nothing to do with the question in front of the person. That reframes the whole fix. The goal is not to ban experienced respondents. It is to verify who is real, and to rebuild the reason they are answering at all.
What actually fixes it
If the counterfeit is undetectable at the output, quality has to be built into the input. Three things change the data before a single answer is written, and none of them is a filter.
1. Verify the human at the source
Source-side verification is a different discipline from output-side cleaning. Instead of scoring answers after the fact and hoping the fake ones look wrong, you establish that a real, eligible person is present before they enter the study. It is slower and it does not scale for free, which is precisely why most of the industry skipped it and leaned on filters instead. It is also the only approach the PNAS result does not break.
2. Lower the cost of a good answer
Satisficing is a rational response to friction. Every confusing question, every clumsy grid, every needless page raises the effort of answering honestly, and the respondent quietly trades depth for speed. A well-designed screener does the opposite. It removes cognitive load so that the honest path is also the easy path. This is not cosmetic. The structure of the instrument is one of the few levers that changes real human behaviour mid-survey, which is exactly why we built Nil Screener around it.
3. Rebuild the motive
This is the part the industry finds hardest to accept, because it is not a technology. Self-determination theory, one of the most robust frameworks in motivation psychology, holds that people engage most deeply when three needs are met: autonomy, competence, and relatedness. Pure financial incentive can actively work against all three. When a task is reduced to a transaction, effort becomes obligation, and there is good evidence that over-incentivising crowds out the intrinsic motivation that produces the best answers. It is the reason some of the richest human contribution online, on Wikipedia, in open-source communities, comes from people paid nothing at all.
Fair pay still matters, and underpaying respondents is its own kind of contract breach. But respect, clarity, and being treated as a partner rather than a data-mining node matter more than the size of the reward. Get that relationship right and you do not have to chase engagement. You receive it.
The industry spent a decade trying to filter its way to quality. Quality was never in the filter. It was always in the relationship with the person answering.
What this means for your next study
The assumption that a coherent answer is a human answer is gone, and it is not coming back. The teams that keep getting usable data from here will be the ones who stop treating the respondent as a risk to be screened out, and start treating them as the source of the only thing research actually sells: an honest account of what real people think.
This is the whole operating model behind our work. We verify every respondent before they reach your study, we design screeners that respect their time, and we pay people fairly for genuine attention. It is more work than buying sample by the thousand, and it is the reason the answers hold up when a decision rides on them. If that is the kind of data your next project needs, you can read more about our approach to recruitment or the screener tool we run on every study.
Sources
- Westwood, S. et al. (2025). The potential existential threat of large language models to online survey research. PNAS.
- NORC (2026). Fraudulent respondents and bots in nonprobability surveys: A literature review.
- Berry, C. & Burton, S. (2024). Response Satisficing Across Online Data Sources. Journal of Public Policy & Marketing.
- Pew Research Center (2021). Measuring the Risks of Panel Conditioning in Survey Research.
- Fraud and discard-rate figures are drawn from published industry analyses (Research Defender, Kantar, CASE4Quality) as aggregated in current market-research reporting.