Density-ratio rescoring improves rare class detection in imbalanced data

Density-Ratio Rescoring for Imbalanced Classification

Machine Learning

Summary

When trying to identify rare cases in data, usual classifiers can miss them because the data is unbalanced. The authors propose a method called Density-Ratio Rescoring (DRR) that adjusts scores after training so that rare class features are better represented. DRR combines these adjusted scores with the original classifier scores without needing to retrain the model. Tests on many datasets show this method improves ranking of rare cases more effectively than other similar techniques.

What this means in practice

  • For medical data teams: Improve detection and ranking of rare disease cases in standard diagnostic classifiers without retraining models.
  • For fraud detection engineers: Enhance rare fraud event ranking in imbalanced transaction data by rescoring outputs from existing classifiers.

Authors

Dongha Kim, Seunghwan Park

Abstract

Density-Ratio Rescoring (DRR) augments a classifier trained at the original class prior with a survey-raking dual score. Raking reweights the majority sample to match minority feature moments within a tolerance. DRR marginally standardizes the dual and base scores and combines them with a fixed weight of one half, using the fitted dual directly for prediction without resampling or refitting the base classifier. Under exact population matching and a correctly specified log-linear tilt model, the dual equals the log density ratio up to an additive constant. A class-separation analysis characterizes the signal strength and correlation conditions under which fusion improves separation under common within-class covariance. On 24 tabular benchmarks, evaluated over 30 trials and five base learners, DRR at the D=128 random-feature setting improves average precision over the standardized base on every dataset, with a mean gain of 0.034. It exceeds the shared-dual raking-and-relabeling resampler on 22 of 24 datasets, with a mean gain of $0.092$, and on all eight one-versus-rest tasks of a shared gene-expression cohort. These results demonstrate the effectiveness of using raking duals as reusable scores for improving rare-class ranking while retaining classifiers trained at the original prior.