Active pixel selection cuts errors and storage in image measurement
Active Data Acquisition with Side Information via Discrete Diffusion Priors
Machine Learning
Summary
Collecting detailed data like images can take a lot of power and space, and sometimes you get more information than you really need. The authors created a way to choose which parts of an image to measure based on some easy-to-get info, making sure the data collected is the most useful for many tasks. They use a special tool called a discrete diffusion model to help decide which pixels to pick. This method reduces mistakes when recognizing images and works well on several image sets, including medical scans.
What this means in practice
- •For medical imaging technicians: Reduce scan time and data size by selectively measuring MRI pixels guided by a learned discrete prior model.
- •For computer vision engineers: Improve image recognition accuracy on low-data sensors by using side information to pick which pixels to record.
Authors
An Vuong, Thinh Nguyen
Abstract
Acquiring data is costly: higher measurement fidelity costs power and storage and risks collecting irrelevant content, while aggressive cost reduction can discard information that later analysis needs. We address this trade-off with an information-theoretic framework that acquires data relevant to a broad set of tasks rather than to one model. A mask policy, conditioned on side information, chooses which pixels to measure so as to maximize the mutual information between a discrete image and its partial observation under a budget; since the image entropy does not depend on the mask, this is equivalent to minimizing the conditional entropy. A frozen discrete denoising diffusion model (D3PM) supplies the posterior, and we use it in two ways: as an entropy surrogate for training a one-shot mask generator, and as the criterion for sequential greedy acquisition. The one-shot generator outperforms random masks only with care, including an unbiased gradient estimator for binary masks. With sequential acquisition, on MNIST the prior makes $8\times$ fewer errors than random at a $10\%$ budget, and on CIFAR-10 it gains $0.9$--$3.4$~dB. On fastMRI, our proposed technique using a static mask outperforms the well-known methods such as variable density and LOUPE.