Paired evaluations improve testing for visual recognition errors

What Paired Evaluations Reveal under Visual Perturbations

Computer Vision and Pattern Recognition

Summary

Robustness tests check how well image recognition systems work when pictures are changed or disturbed, but real-world checks are expensive. The authors show that comparing predictions on original and changed images together reveals more about errors. They prove that simple counts of matching correct answers can reliably estimate how systems lose confidence, while using detailed comparisons helps focus on which images to test physically. This approach finds more mistakes using fewer physical tests than traditional confidence scores or averaging methods.

What this means in practice

Authors

Yongda Wei, Chen Zhang, Yifei Wang, Xinyu Wang, Bosen Shao, Hanxi Li, Liping Di

Abstract

Robustness evaluation must examine diverse visual perturbations, while benchmarks cover only some real-world conditions and physical testing is costly. Paired evaluations link clean and perturbed predictions for the same image, capturing changes in correctness, confidence, and acceptance beyond aggregate accuracy. We investigate how this image correspondence supports two needs in robustness evaluation: interpreting paired evaluation results and prioritizing samples for physical testing. To interpret paired evaluation results, we fix both sets of prediction records and vary their correspondence within each class. We prove that classwise correct-correct counts give the same sharp bounds on lost acceptance and mean true-class probability decrease among retained-correct inputs as any feasible five-state refinement. Distinguishing persistent from changed wrong answers can further constrain accepted-error transitions, while shared correspondence can establish policy orderings left unresolved by separate cost intervals. To prioritize samples for physical testing, we retain each image's synthetic responses and rank clean-correct images by their mean true-class probability under corruption. Across 44 classifiers, testing the highest-risk 20% finds 67% and 45% of failures under mild screen and print recaptures, versus 58% and 36% for clean confidence and 60% and 37% for an equal-size natural-transformation average. With both probability averaging and an A3Rank scoring adaptation, the tested corruption set yields higher mean failure recall than the natural-transform set; differences between scores depend on the source and budget. Together, these findings show that the value of correspondence depends on the evaluation objective: classwise counts suffice for specified reliability bounds, while image-specific synthetic responses improve the allocation of physical tests within the evaluated pool.