GeoRVQ improves accuracy and quality of physiological signal tokens

GeoRVQ: Decoder-aware geometry for residual-token prediction in physiological signals

Machine Learning

Summary

Turning heart and other body signals into short sequences helps computers analyze them faster. The authors designed a new method called GeoRVQ that judges errors based on the actual impact on signal quality, not just whether tokens match perfectly. Their approach improves accuracy and better preserves important features like heartbeat peaks. They tested GeoRVQ on major medical datasets and showed it predicts signal details more faithfully than previous methods.

What this means in practice

  • For medical device developers: Improve signal compression in wearable monitors for more accurate heartbeat and physiological event capture.
  • For clinical data engineers: Enhance quality of physiological waveform data encoding in hospital databases while preserving clinically important features.

Authors

Bo Cui, Yaowen Zhang

Abstract

Residual vector quantization (RVQ) turns physiological waveforms into compact token sequences, but conventional masked modeling treats every incorrect token as equally costly. We propose GeoRVQ, a coarse-to-fine masked token model whose objective reflects the local response of a frozen waveform decoder. Decoder-induced costs define geometry-aware soft targets and expected distortion, while quantizer-causal prediction follows residual dependencies from coarse to fine levels. In a descriptive aggregate over MIMIC-IV Waveform, VitalDB, and CODE-15\%, GeoRVQ increases exact token accuracy from $.133\pm.004$ to $.143\pm.003$, reduces decoded distance from $.606\pm.006$ to $.393\pm.007$, and increases R-peak F1 from $.784\pm.004$ to $.837\pm.008$ under matched model and training conditions. Across 45 held-out code substitutions, decoder-induced cost has a Spearman correlation of $.85$ with realized decoded cost, compared with $.54$ for Euclidean codeword distance. These results indicate that decoder-aware objectives can improve waveform and event preservation without requiring a large increase in exact token accuracy.