Task aware compression improves classification by matching distributions
The Hidden Perception Constraint in Task-Aware Compression
Information TheoryMachine Learning
Summary
Compression schemes often aim to keep data looking and sounding natural so people think they are good quality. This paper looks at compression designed not just to keep things looking good but also to help computers classify information accurately. The authors find that when the classification rules don't fit the original data well, changing the compressed data to look like a better suited target helps classification work better. This means the way data is compressed can secretly influence how well computers understand it afterward.
What this means in practice
- •For image compression engineers: Design compression algorithms that improve downstream image classification accuracy by adapting data distributions to decision boundaries.
- •For speech recognition developers: Enhance voice data compression schemes to better support speech classification tasks by incorporating task-aware distribution adjustments.
Authors
Sahan Liyanaarachchi, Semih Akkoc, Sennur Ulukus, Aylin Yener
Abstract
With the recent advancements of neural compressors, explicitly incorporating perception constraints into the design of compression schemes has gained significant attention. Traditionally, these perception constraints ensure that the distribution of the reconstruction does not significantly deviate from the distribution of the source, thus attesting to the perceptual quality of the reconstruction. In this work, we uncover several perception constraints that are naturally present in task-aware compression. In particular, we consider a problem where the primary task is reconstruction and the secondary task is classification (i.e., a statistical test). We study this problem at varying levels of domain information available to us and discuss how to utilize the naturally emerging perception constraints to design rate-minimal compression schemes that also maximize the utility of our secondary task. We show that in this setting, if the decision boundaries of the classifier are ill-defined (mismatch) for our source distribution, then matching onto a target distribution enhances our classification accuracy.