Papers for

image recognition engineers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Bayesian mixed-prior score improves rejecting unknown classes in recognition

Mixed-Prior Decision Risk for Open-Set Recognition

Abstract: In open-set recognition (OSR), a probe must either be identified as one of the known gallery classes or rejected as unknown, so three error types coexist: false acceptance, false rejection, and misidentification. An uncertainty score for selective recognition should rank probes by the risk of the decision the system has made. Bayesian gallery-aware models such as Holistic Uncertainty Estimation (HolUE) summarize the posterior over known and unknown classes by Kullback--Leibler (KL) divergence components and map them to an uncertainty score with a supervised nonlinear calibrator. We show that the KL summary is not generally monotone in decision risk: linear fusion of the KL components tuned on validation data yields negative filtering quality on several benchmarks. We propose MPRisk, a mixed-prior posterior decision-risk score that keeps the same Bayesian posterior but directly scores the error events associated with the selected decision: false-acceptance, misidentification, and false-rejection risks, plus a non-specificity penalty for rejections, enabled by modeling unknown identities as a continuous component. Four nonnegative weights tuned on a validation set suffice for ranking; no nonlinear supervised model is required. Across nine image, audio, and text benchmarks, MPRisk achieves the best or tied-best Prediction Rejection Ratio at every operating point on the image and audio benchmarks and on most text operating points, with bootstrap-confirmed gains over HolUE on five benchmarks (up to $+0.19$ PRR) at comparable or lower runtime.

Mon 28 SeptComputer Vision and Pattern Recognition
The gist
In recognition tasks, a system must decide if something is known or unknown, which creates three kinds of mistakes: saying unknown things are known, rejecting known things, or mixing up known classes. The authors show that existing Bayesian methods use a complex measure that doesn’t always reliably reflect these mistakes. They propose a new scoring method called MPRisk that directly scores the risk of these three error types plus a penalty for rejecting something unspecific. MPRisk is simpler to tune and performs better or equally well across many image, audio, and text benchmarks compared to a prior method.
Open → 2609.35043v1

Structured frequency attacks reveal robust model weaknesses

Frame the adversary: a structure-aware attack methodology

Abstract: Frequency-based adversarial attacks have recently grown popular by exploiting spectral sensitivities shared across neural architectures. Unlike spatial perturbations, frequency-based attacks expose deeper vulnerabilities, making them especially valuable for robust evaluation of safety-critical and security-sensitive applications. Yet, existing approaches are typically not derived as solutions to an optimization problem that explicitly captures transform-domain structure. In this paper, we propose a methodology for crafting principled frequency-based adversarial attacks, via a dedicated optimization framework. A cornerstone of our method hinges on the introduction of a perturbation constraint set, tied to highly structured non-orthogonal transforms, well-known for their flexible, non-predefined frequency handling. We prove that the attacks emerge as weighted $\ell_2$-projections onto this set, yielding a general and controlled attack generation mechanism. By this, we provide a clear geometric attack characterization, ensuring alignment between the optimization objective and the perturbation constraint. We assess our framework on standardized datasets, for pretrained and adversarially robust models. Results highlight that our attacks, being solutions to an optimization problem, over a structured perturbation set, are highly effective, even across different, unseen architectures. Our methodology could serve as a theoretical baseline for designing and analyzing transformed-based attacks, targeting fundamental model vulnerabilities, instead of mere architecture-specific artifacts typically studied in the robustness literature.

Fri 25 SeptMachine Learning
The gist
Some tricks can fool AI by making tiny changes to pictures that humans barely notice. This paper shows a new way to create these tricky changes by thinking about how images behave in terms of their hidden frequency patterns, rather than just pixels. The authors use math to carefully shape these changes, making attacks that are strong and work across different AI designs. Their method helps us understand AI weaknesses better, beyond quick fixes tied to one specific system.
Open → 2609.31128v1