Epistemic Norms for AI Safety and Alignment Research

2026-07-27Artificial Intelligence

Artificial Intelligence
AI summary

The authors highlight that normal AI research focuses on improving abilities and accepts some small failure rates if the overall performance is good, while AI safety research wants to avoid any catastrophic failures, especially in unpredictable or risky situations. They point out that these goals differ in two main ways: focusing on avoiding dangerous behaviors instead of just adding good features, and managing extreme risks rather than average outcomes. They found important gaps in current AI safety work, such as a lack of independent checks. To improve this, the authors propose ECAISA, a set of rules and processes to make AI safety research clearer, more verifiable, and harder to manipulate, aiming for transparency and auditability rather than guaranteeing safety.

AI safetyAI alignmentcapability profilerisk profilefat-tailed riskindependent verificationepistemic practicestransparencyinformation hazardauditability
Authors
Keivan Navaie
Abstract
Mainstream AI research emphasises capability growth and tolerates low failure rates when average-case performance is high. AI safety and alignment research has a different mission: to ensure that catastrophic failures never occur, under sparse evidence, adversarial dynamics, and fat-tailed risk. We argue that the two domains differ along two analytically independent axes---{\it capability profile}, demonstrating the absence of hazardous behaviours rather than the presence of positive capabilities, and {\it risk profile}, bounding worst-case outcomes under fat-tailed uncertainty rather than optimising average-case performance---and that mainstream epistemic practices are inadequate on both. Building on a structured synthesis grounded in a preregistered bibliometric baseline, we identify five cross-cutting gap dimensions in current alignment research, including the near-absence of institutionalised independent verification. To address these gaps, we propose {\sc ECAISA}, an Epistemic Code for AI Safety and Alignment comprising eight principles, a three-level scoring rubric, a four-level disclosure ladder that reconciles transparency with information-hazard and commercial-confidentiality constraints, a tiered applicability scheme, an information-hazard adjudication procedure, and seven anti-gaming mechanisms. {\sc ECAISA} does not certify that any AI system is safe; it constrains how safety-relevant research claims are documented, checked, and relied upon, with auditability rather than certification as its governance target.