Benchmarking Identity-Sensitive LLM Outputs for Surveillance and Security Robots
2026-08-17 • Robotics
RoboticsComputers and Society
AI summaryⓘ
The authors studied how large language models (LLMs) write descriptions of surveillance and security robots when given different identity-related prompts. They tested 236 different identity labels to see if these influenced how readable the robot design descriptions were. They found that readability did change depending on the prompt, design aspects, and the identities mentioned. While readability alone can’t tell if the descriptions are fair or appropriate, the authors suggest it as a simple way to start checking for differences in the generated outputs.
large language modelssurveillance robotssecurity robotsreadabilityprompt conditioningdemographic identityrobot designnatural language generationfairnesstext analysis
Authors
Nneka Hyman, Jasmine Khan, Raj Korpan
Abstract
Large language models (LLMs) are increasingly used to generate textual robot design specifications, interaction policies, and risk assessments during early-stage robot development. Such outputs may influence how surveillance and security robots are conceptualized, documented, and ultimately implemented. This paper evaluates whether identity-conditioned prompts produce systematic differences in LLM-generated surveillance and security robot design descriptions. Using 236 demographic identity labels across single-label and model-augmented prompt conditions, we analyze readability as an initial benchmark for evaluating accessibility and identity-conditioned variation in generated robot design descriptions. The results show significant differences in readability across prompt conditions, design dimensions, and demographic identities. Although readability cannot determine whether an output is fair or socially appropriate, it provides an interpretable baseline within a broader benchmarking framework that also includes lexical, semantic, sentiment, syntactic, and fairness-focused analyses.