A Penny for Your Prompts: Experiments Detecting and Mitigating LLM Usage by Survey Respondents
2026-07-01 • Human-Computer Interaction
Human-Computer InteractionCryptography and SecurityComputers and Society
AI summaryⓘ
The authors studied how often people use large language models (LLMs) like AI to help answer surveys on different crowdsourcing platforms. They found that AI use ranged from less than 10% to more than 80% depending on the platform. They tried methods like asking participants not to use AI and disabling copy-paste, which lowered AI use but didn't always make the survey data better. The authors suggest checking for AI by watching how people type and designing questions that discourage AI help.
large language modelscrowdsourcing platformssurvey data qualityAI detectionkeystroke analysisMechanical TurkProlificresponse validationcopy-paste controlmitigation strategies
Authors
Zane Xu, Nathan Malkin
Abstract
Large language models are increasingly used by participants on crowdsourcing platforms when responding to surveys, potentially undermining the validity of collected data. Our study aims to quantify the prevalence of this behavior and investigate methods to detect and prevent it. In a series of surveys (N = 250), we examined conditions such as platform choice, survey length, requests not to use AI, and disabling copy-paste functionality. We were able to identify distinct characteristics of LLM-assisted responses and found that their frequency varied widely, from under 10% on Prolific to over 80% on Mechanical Turk. Mitigation measures reduced LLM usage but did not necessarily improve data quality. No participants employed browser-use agents at the time of our survey, but we report on our own detection experiments. We recommend that researchers actively screen survey responses for LLM usage by recording and analyzing keystroke data and crafting instructions and questions aimed at AI.