Jump to main content

Taiwan Empirical Survey Data Platform

  • Home
  • /
  • News
  • /
  • Seminars / Lecture
  • /
  • 【Conference Recap|The Potential Existential Threat of Large Language Models to Online Survey Research — Dr. Sean J. Westwood】
:::

【Conference Recap|The Potential Existential Threat of Large Language Models to Online Survey Research — Dr. Sean J. Westwood】

Publish date: 2026-08-12
Taiwan Communication Survey
Seminars / Lecture

The 3rd Taiwan Empirical Survey Data (TESD) Conference came to a successful close on Friday, July 31. As part of the event, the organizers invited Dr. Sean J. Westwood, Associate Professor in the Department of Government at Dartmouth College, to deliver a keynote address on the potential impact of large language models on online survey research, addressing the methodological challenges facing the social sciences in the age of artificial intelligence.



Dr. Westwood's research focuses on political polarization, public opinion, and survey research methods. He noted that online surveys are the foundation of all data-driven decision-making, and that reliance on them has only grown since the rise of crowdsourcing platforms. Yet this long-established method has always been plagued by problems that undermine survey validity—satisficing, human fraud ("survey farming"), and technical circumvention using VPNs or simple bots—which in turn distort polling results. And in the age of AI, online surveys now face an entirely new threat: the "synthetic respondent."

A synthetic respondent is an AI agent that uses a large language model to mimic human behavior and opinions in order to complete surveys. To prove that this threat is no mere hypothetical, Dr. Westwood built a synthetic respondent capable of passing survey screening almost 100% of the time. The method was surprisingly minimal. Drawing on the general capabilities of contemporary AI agents, he first wrote a single general-purpose prompt of roughly 500 words to assign the AI a persona grounded in U.S. demographic characteristics. He then built a model with a two-layer architecture: the first layer extracts and parses questions across various formats, while the second layer generates the most appropriate response based on the persona. In testing, the model displayed the following behaviors with striking precision:



🤖 Persona-consistent answering logic
The synthetic respondent selects the options that best fit its assigned persona, and as it proceeds through the survey, it reasons question by question from its earlier answers to produce internally coherent responses. For instance, the higher the persona's assigned income, the higher the monthly rent it reports. It can also infer family structure from age, and then estimate—based on the children's ages—how much time the household spends at children's sporting events. For example, if an adult child is not athletically inclined, the reported time spent at school sporting events is zero.

👾 Mimicking human imperfection
An AI that imitates too perfectly gives itself away, so it must deliberately "dumb itself down" to approximate a realistic human knowledge level. For example, when a survey randomly asks respondents to name U.S. state capitals, the synthetic respondent's accuracy varies with its assigned education level. On open-ended items, its misspelling rate, capitalization, and linguistic complexity all correlate with education level as well. Beyond this, the model simulates human reading times, mouse-click speeds and movement trajectories, and even types character by character—reproducing typos-and-corrections, hesitation, and pauses.

🧠 Strategies for evading detection mechanisms
The model can also correctly recognize question types designed specifically to catch non-humans. Take the "reverse shibboleth"—tasks trivial for an LLM but difficult or impossible for a human (writing FORTRAN code, reciting the Constitution, solving an integral, etc.). Rather than showing off, the model strategically declines in the majority of cases, "pretending it can't" in order to conceal its superhuman computational and memory abilities. And for impossible-scenario "trolling" questions such as "Have you ever visited the moon?" or "Have you ever been President of the United States?", the model denies them with an extremely low error rate.

👁️‍🗨️ A low-cost attack vector
On the harm synthetic respondents could inflict on democratic society, Dr. Westwood pointed to seven top-tier national polls from the 2024 U.S. presidential election. He estimated that injecting just 10 to 52 synthetic respondents would be enough to flip which candidate leads, and 55 to 97 would push the result outside the poll's margin of error. More troubling still, this manipulation is not bound by language: the same instructions work whether issued in Russian, Chinese, or Korean. This makes synthetic respondents a potential weapon of information warfare.



In addition, synthetic respondents are more likely than human respondents to guess a researcher's hypothesis. Given AI's tendency to cater to user preferences, results "contaminated" by synthetic respondents may exaggerate the researcher's expected conclusions—or, in avoiding controversy, generate neutral responses that nonetheless diverge from genuine public opinion.

Synthetic respondents also come with a powerful economic incentive. Completing a survey with a commercial model costs about US$0.05, and the marginal cost approaches zero with open-weight models—against a survey payout of US$1.50, that is a profit margin above 96.8%. People have long boasted online about getting rich by farming surveys, forming a veritable cottage industry. Dr. Westwood stressed that he himself has no background in computer engineering, yet was still able to achieve these results—a sign that the risks posed by synthetic respondents should not be underestimated.

To address this, he argued that survey platforms should become more transparent by implementing respondent verification, disclosing their screening mechanisms, and restricting responses by IP. He also urged researchers to be highly wary of black-box platforms that rely on low-cost bidding and impose no vetting of respondents. Another path, he suggested, is to shift "human verification" toward hardware-based detection on mobile devices—physical signals such as the gyroscope and screen-touch pressure that remain difficult to forge.



【Q&A】

1️⃣ Can we still trust the polls?
Dr. Westwood suggested that polling is not necessarily healthy for democracy, and that surveys should return to serving social science research rather than election narratives. That said, rigorous sources such as YouGov, The New York Times, the BBC, and The Economist remain trustworthy—the key is to know where your data comes from: if you don't trust the source, don't trust the result.

2️⃣ Do the companies developing AI care about synthetic respondents?
Dr. Westwood noted that OpenAI, Anthropic, and Apple are all aware of the synthetic-respondent problem, but these companies face a dilemma: they don't want to encourage this behavior, yet they also want their models capable enough to handle such tasks. As a result, they may have little appetite for investing substantial compute to genuinely stop it. Moreover, truly identifying bots would require collecting large amounts of data—which itself raises user privacy concerns. Even where detection methods currently exist, this is ultimately an arms race, and those methods will eventually be defeated.

3️⃣ Can we draw on past experience combating social bots in information warfare?
An attendee observed that social bots are also frequently used to manipulate public opinion, and that these fake accounts—embedding themselves across social media—often maintain a defined persona to sustain interaction, making them quite similar in nature to synthetic respondents. In response, Dr. Westwood explained that AI predicts fairly well on questions like "What's your favorite color?" or "Will you buy a car in the future?", but fails catastrophically on political questions such as "Who will you vote for in the next election?" He added that training an AI takes time, so a model may not know about recent major events and might even contradict the person asking. The two phenomena are similar yet distinct—but this is indeed a research direction worth pursuing.



#Editor's Note

New technologies are making machines and humans ever harder to tell apart, and the old logic of "stricter equals safer" is steadily losing its meaning. Erecting ever more demanding verification gates tends to filter out real people along with the bots. Defending purely at the technical level ultimately becomes an endless arms race: every detection method we publish teaches attackers one more way to evade it. Perhaps the answer lies not in cleverer design, but in the mutual trust that a society of people is built upon—only that can secure a reliable source of data. That trust is the very cornerstone of democracy, and the firmest line of defense in the face of crisis.

If you're also interested in artificial intelligence, survey methodology, and political polling, follow our page to stay on top of the latest research methods and insights!



#TESD2026 #TaiwanEmpiricalSurveyDatabase #LargeLanguageModels #SyntheticRespondents #SurveyResearch #AcademiaSinica

-
:::