Most changes to AI safety promises go unreported by developers
Silent Revision: Measuring Undisclosed Change in the Safety Frameworks of Frontier AI Developers
Computers and SocietyArtificial IntelligenceSoftware Engineering
Summary
AI companies create safety rules to explain how they keep their powerful models from causing harm. The authors found that when these companies update their safety rules, over half of the important changes are not clearly explained or listed, making it hard to track what actually changed. Changes that reduce safety commitments are especially likely to be hidden. The researchers suggest that rules should require companies to clearly list all changes so the public can understand and hold them accountable.
AI safety frameworksaccountabilityversion controlcommitmentschangelogtransparencymodel riskregulation
Authors
Louis Yiven Zhu
Abstract
Frontier AI developers publish safety frameworks that commit them to evidencing whether their models are dangerous. The European Union and California now treat these documents as instruments of accountability, and both already impose duties on their revision. Neither requires the revision to be legible, in the sense that a reader could learn from the developer's own account what changed. We introduce the silent revision rate, the share of material changes to a framework's commitments that the developer's published account does not identify, and we release the versioned, hash-pinned corpus needed to compute it. The corpus contains every public version of the safety frameworks of the twelve developers that have published one, together with each provider's changelog, redline or announcement. We trace 710 commitment instances across twelve consecutive version pairs, code them against a frozen codebook, and adjudicate 244 individually. Three findings follow. First, 67% of material changes (95% CI 62 to 72) are silent under a strict standard and 53% under a lenient one, falling to 49% at section granularity. Second, silence appears to track the form of the account, since narrative announcements run at 74% against 63% for itemised changelogs, whereas account length in words barely matters; on the test that respects nesting the difference is suggestive. Third, 77% of traced changes weaken or remove a commitment, and in seven of eight pairs weakenings are more often silent than strengthenings. The statutory remedy therefore exists and specifies the wrong artefact. A justification explains why a framework changed, an enumeration states what changed, and only the latter makes revision auditable. We argue that publication duties should carry an enumeration duty, which one provider already meets, voluntarily and incompletely.