Concept based explanations for neural networks with guaranteed accuracy

A Unifying Framework of Concept-based Explainable AI with Completeness Guarantees

Machine LearningArtificial IntelligenceComputer Vision and Pattern Recognition

Summary

Neural networks often make decisions based on complex math that is hard to understand. This paper focuses on explaining those decisions using simple, human-friendly ideas called concepts. The authors created a math framework that links these concepts to the model’s internal workings and its decisions. They found a way to measure how well these concepts capture the model’s behavior and can even tell how much each concept contributes to a decision with guarantees on the accuracy of those explanations.

What this means in practice

  • For machine learning engineers: Design interpretable neural network components that provide guaranteed accurate concept explanations for model predictions.
  • For ai safety teams: Assess and improve trustworthiness of neural network models by measuring and bounding how well concepts explain decisions.

A theory result. No direct application yet.

Authors

Vojtěch Kůr, Adam Kukučka, Tomáš Brázdil, Vít Musil

Abstract

Concept-based explanations describe neural network predictions through human-understandable properties of inputs called concepts. The field encompasses approaches that differ in how they define and represent concepts and connect them to model predictions. We introduce a theoretical framework that describes these approaches in a common mathematical language and supports a shared analysis of their properties. For concept discovery, which identifies concepts automatically within a latent space of a trained model, we employ a concept autoencoder view. An encoder extracts concept representations from the model's latent space, and a decoder uses them to reconstruct the original latent representation. The autoencoder's reconstruction error measures how accurately its decoder recovers the original latent representation. We revisit model completeness: how well the concepts can reproduce the model's outputs. We show that model incompleteness of the concepts can be bounded by the autoencoder's reconstruction error. The autoencoder view also provides a common way to define individual concept attributions, which measure each concept's contribution to a prediction. We establish when these attributions sum to the model's prediction, and bound the discrepancy otherwise, thus providing attribution completeness guarantees.