Papers for

security and compliance teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

System prompts mostly manage tools not model ethics

Configuration, Not Conscience: A Large-Scale Empirical Study of LLM System Prompts

Abstract: Leaked system prompts are often treated as windows into the hidden values of commercial language models, yet their composition is rarely studied at scale. We analyze a merged corpus of 407 leaked, reconstructed, or officially published system prompts from 62 vendors across four community collections, identifying 29 near-duplicate clusters covering 66 files. Operational content rather than ethical statements dominates the corpus; a deliberately simple block-level classifier assigns roughly 58\% of classified words to tool/protocol and roughly 5\% to safety policy, while the strictest rule-lines guard tool use and file safety over harmful content by an 11:1 margin. Literal text transfer concentrates in a small set of cross-vendor pairs. Prompts also carry measurable maintenance debt, with version chains turning over thousands of words per release. The evidence supports treating leaked prompts as operational specifications, closer to configuration files than value statements, and treats reuse and prompt rot as engineering and supply-chain concerns. Because most documents are adversarial in origin and the detectors are deliberately simple, all magnitudes are directional; we audit the main classifier's error modes.

Fri 25 SeptCryptography and Security
The gist
People often think leaked instructions for AI language models reveal the models' hidden values or ethics. This study shows that these instructions mostly focus on how to operate tools and keep files safe, rather than on moral guidelines. The prompts also change a lot over time, similar to software configuration, and are reused between vendors sometimes. The authors suggest treating these instructions as technical setups instead of ethical statements.
Open → 2609.31575v1