Papers for

ai operations teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Representation engineering helps llm safety under specific conditions

When Do Model Internals Help? Exploring the Role of Representation Engineering in LLM Safety

Abstract: Reliable AI safeguards require both control mechanisms that reduce unsafe behavior and monitoring mechanisms that detect safety risks during model interactions. Established behavioral safeguards include alignment methods that optimize model outputs and text monitors that assess interaction text. Representation engineering instead reads or modifies internal model states, but the relative strengths of these approaches remain unclear because they are often evaluated under different settings. We present a matched evaluation across two tracks. For safety control, we compare DPO, a behavioral alignment method, with three representation steering methods across robustness, practicality, and granularity. DPO provides the strongest overall control and generally improves with increasing training data, although its safety can degrade after subsequent benign fine-tuning. Representation steering remains competitive primarily in low-data settings, particularly with high-quality contrastive data. For safety monitoring, we compare representation probes with fine-tuned and open-weight text monitors across full-response detection, early detection, and computational cost. Specialized text monitors achieve the strongest overall detection accuracy, while representation probes remain competitive at substantially lower marginal cost. Finally, monitor-guided interventions recover much of the safety lost by DPO after benign fine-tuning, with little additional over-refusal. Overall, representation engineering does not generally replace behavioral safeguards, but offers practical advantages under specific conditions and can provide complementary safety benefits.

Mon 28 SeptArtificial IntelligenceComputation and LanguageMachine Learning
The gist
Keeping AI models safe is important but tricky. This paper compares two ways to make AI safer: changing the model’s behavior directly or tweaking what the model internally thinks. The authors find that changing behavior usually works better overall, but adjusting internal representations can help when little training data is available or for cheaper risk detection during use. They also show how combining these methods can restore safety after updates, meaning both methods have roles in making AI safer.
Open → 2609.34771v1

AI agents gain software structure for improved reliability

Agents as Software: A Programming Languages Agenda for Agent Reliability

Abstract: AI agents increasingly resemble software systems: they call tools, remember facts, follow policies, delegate work, and take actions with real consequences. % Yet the ``program'' of an agent is scattered across prompts, tools, memories, workflows, and execution traces, making its behavior difficult to inspect through ordinary testing and debugging alone. % This essay argues that a programming-systems perspective offers a natural lens for making agents reliable. % We recast agents as programmable artifacts whose behavior can be specified over traces and state, checked before deployment, monitored during execution, and improved from observed failures. % The goal is not to make probabilistic agents behave like deterministic programs, but to give them enough structure that their behavior can be reasoned about, controlled, and repaired.

Sat 26 SeptProgramming LanguagesArtificial Intelligence
The gist
AI agents today act like complex software, using tools and making decisions that affect the real world. The authors point out that these agents are tricky to test and fix because their 'programs' are spread across many parts like prompts and memories. They propose treating agents like software programs that can be checked, monitored, and improved systematically. This approach doesn't make AI agents fully predictable but helps control and repair their behavior more effectively.
Open → 2609.32198v1

Adaptive human oversight reduces risks in ai task delegation

When Should a Human Take Back Control? Optimal Delegation under Turbulent AI Risk

Abstract: Deploying AI systems requires deciding when to delegate tasks and when humans should intervene to monitor and mitigate risk induced by AI operations. These decisions are challenging when failures cluster: a hallucination or harmful output can trigger further errors, creating periods of elevated risk. We introduce a continuous-time framework for learning adaptive human oversight under such turbulent AI risk. Existing oversight and delegation formulations condition on history but do not model incident clustering, or its suppression by supervision effort, jointly with the delegation decision and this study fixes this gap. The self-exciting dynamics capture how risk events increase the likelihood of subsequent events, making their timing and history central to decision-making. We formulate a stochastic control problem combining human actions, monitoring effort, and switching between human-AI-assisted operation and full AI delegation, balancing operational rewards against oversight costs, and cascading AI-failures and induced uncertainty. Human participation is an endogenous component of risk management: the policy determines both when oversight is needed and how much effort to allocate. We study a relaxed switching formulation and propose Hawkes-PPO, a policy-gradient method that uses a bank of exponential filters of observed incident times. In a synthetic environment it attains a higher risk-adjusted objective than either fixed regime and approaches an approximate full-information oracle. We illustrate our results with numerical simulations by examining how cascade risks influence intervention and delegation, connecting reinforcement learning with adaptive human oversight of AI systems. In particular, we illustrate the benefit of our switching strategy and Hawkes-PPO algorithm to monitor the project efficiently along time, reducing turbulent risks occurrences and costs.

Fri 25 SeptMachine LearningArtificial Intelligence
The gist
Deciding when humans should take control away from AI is hard, especially when AI errors tend to cause more errors in a chain. The authors introduce a mathematical framework that helps decide when to monitor AI closely and when to let it operate freely. They use a technique that learns from past errors to know when risk is higher and human intervention is needed. Their method better balances the benefits of AI with the costs and dangers of errors compared to simpler approaches.
Open → 2609.32083v1