Market research has a quiet dependency that is rarely examined: almost everything the discipline knows about what people think comes from people agreeing to say so, in a structured setting, for a small payment.
Online panels made that arrangement cheap enough to become the default. They also made it fragile, and the fragility has now been exposed by something the panel model has no defence against.
The incentive was always there
A panel respondent is paid per completed survey. That has always produced a class of participant optimising for completion rather than accuracy — straight-lining, satisficing, clicking through screeners with whichever answer grants entry. The industry knew this and built defences: attention checks, timing thresholds, consistency traps, red herring brands that no honest respondent could claim to have used.
Every one of those defences assumes the respondent is a human being doing minimal work. None of them survives a respondent that is a language model doing the work properly.
A model completing a survey does not straight-line. It answers attention checks correctly, produces open-ended responses that are coherent and on-topic, takes a plausible amount of time, and maintains internal consistency across the instrument better than a tired human does. The quality-control apparatus was designed to catch carelessness. What has arrived is diligence without a person behind it.
What breaks, specifically
It would be convenient if contamination simply added noise, because noise is manageable: it widens confidence intervals and can be countered with sample size. This is worse than noise, in two respects.
It is biased. Synthetic responses do not scatter randomly around the truth; they cluster around the plausible — the modal, well-documented, unsurprising answer. A study contaminated in this way will tend to confirm what is already believed, and confirm it with unusual tidiness. The tell is not implausible data. The tell is data that is too clean.
It is invisible in the aggregate. Open-ended responses that read well are usually taken as a sign of engagement, which means the contaminated portion of the sample is the portion an analyst is least likely to discard.
What still works
The methods that resist contamination are the ones that observe behaviour rather than collect statements.
Purchase data cannot be fabricated by a respondent, because it is not supplied by the respondent. Neither can returns, churn, search volumes, or what happens when a price changes in one market and not another. These sources were always more reliable than stated preference — the gap between what people say they would pay and what they pay has been documented for half a century — and their advantage has now widened for a new reason.
Where stated data is genuinely necessary, the defences that survive are the expensive ones the industry moved away from: recruited samples with verified identity, interviews conducted live, smaller samples with real provenance. Qualitative research, long treated as the soft preliminary to the serious quantitative work, turns out to have been holding the more robust position, because a conversation with a person in a room has a property a panel response does not: you can see who is answering.
The correction that is coming
Research budgets have moved for two decades towards whatever produced the most responses per euro. That direction was defensible while responses were responses. It is now a direction that optimises for the input most vulnerable to fabrication.
The correction will not be a technical fix. Detection tools will improve, and so will the systems they are trying to detect, and there is no reason to expect the defenders to win that exchange permanently. What will actually change is the standard of evidence: the industry will have to say where its data came from and how the people in it were verified, and studies that cannot answer will carry less weight.
That is not a bad outcome. It is a return to a question research should never have stopped asking.