Circuits replicate model successes but not most errors

Rethinking Circuit Evaluation: Do Circuits Explain Model Errors?

Machine LearningArtificial Intelligence

Summary

Understanding how AI models make decisions involves looking inside them to find small parts called circuits that explain their behavior. This paper finds that these circuits often match when the model gets answers right but mostly miss explaining the mistakes the model makes. By studying a specific task in GPT-2, the researchers show that fixing which circuits are included improves how well the circuits explain errors. This means that to truly explain how models work, it's important to also explain their errors, not just their correct answers.

What this means in practice

  • For machine learning engineers: Improve model debugging by using circuit analyses that better capture both successes and failures in AI behavior.
  • For ai safety teams: Design more reliable interpretability tests that ensure explanations account for AI errors as well as correct outputs.

Authors

Li Zhang, Chuqin Geng, Mark Zhang, Chen Yang, Luke Zhang, Haolin Ye, Xujie Si

Abstract

Mechanistic interpretability (MI) aims to explain a model's behaviour through analyzing its internal computations; circuit-based explanations aim to isolate these computations with compact subnetworks validated by ablating the rest of the model. We show that circuits validated this way may fail to recover the underlying mechanism of the model's behaviour by closely reproducing its successful decisions while failing to account for most of its errors. Such explanations should account for the model's particular errors as well as its successes. We evaluate this requirement by measuring exact answer agreement separately on model successes and failures, across circuit sizes and ablation settings, on IOI, Docstring, and six model-task settings from the Mechanistic Interpretability Benchmark. We discover that many tested circuits closely replicate correct behaviour while missing most of the model's errors. On indirect object identification (IOI) for GPT-2 small, under mean ablation, the manual circuit and tested automated circuits, including one trained against the model's full output distribution, agree with the model on 97.3-99.5% of prompts it answers correctly but only 11.4-41.7% of errors. An IOI case study shows that lost errors are recoverable by restoring omitted attention-heads which raise error reproduction from 14.2% to 75.1% on a separate held-out set with 0.41 percentage point decrease on correct agreement, exceeding matched random extensions and scalar-biased control. Intervention traces show how omitted computations produce specific wrong answers for a reproducible subset of errors. In all, these findings show circuits can preserve task success without adequately explaining model's failures, and support exact error reproduction as a necessary, but not sufficient, test of circuit-based explanations of model behaviour.