Entropy-punctured bloom filters improve memory use in machine learning

Entropy-Punctured Bloom Filters for Memory-Efficient Machine Learning

Machine LearningArtificial Intelligence

Summary

Machine learning models often need to use lots of memory to store and process data. The authors studied a way to shrink how data features are stored using something called Bloom Filters, which are like compact digital summaries. They improved these by removing bits that don’t change much, keeping important information but saving space. Tests showed this approach keeps predictions accurate while using less memory, which helps when memory is limited.

What this means in practice

  • For mobile app developers: Create machine learning features that use less memory on devices with limited storage and bandwidth.
  • For iot device engineers: Implement compact machine learning models to run analytics efficiently on devices with constrained memory.

Authors

John Cartmell, Mihaela Cardei, Ionut Cardei

Abstract

Memory-efficient feature representations are increasingly important in machine learning settings where storage, transmission cost, bandwidth, or privacy constraints limit access to raw data. Bloom Filter (BF) encodings provide compact probabilistic representations of engineered features, but their behavior under structural compression and their applicability to regression tasks remain underexplored. In this work, we propose entropy-punctured Bloom Filters, a memory-aware encoding strategy that removes low-variability bit positions identified using empirical entropy. Starting from fixed-length BF encodings of quantized features, the proposed approach produces reduced representations that preserve predictive structure while improving predictive efficiency relative to encoded representation size. We evaluate the approach on diverse regression datasets, comparing raw features, Principal Component Analysis (PCA), Random Projection (RP), and Bloom Filter variants under leakage-free evaluation protocols and approximately matched representation sizes. Performance is assessed using ridge regression, XGBoost, and neural networks, with predictive efficiency measured as R2 relative to encoded representation size per sample. Results show that Bloom Filter encodings remain competitive with classical compressed representations while achieving substantial storage savings. Entropy-based puncturing further reduces representation size with minimal loss in predictive fidelity, yielding improved predictive efficiency. These findings demonstrate that entropy-punctured Bloom Filters provide an effective representation-level compression approach for memory-constrained machine learning.