Comparing retraining methods reduces error gaps over deployment time
Evaluating Model Retraining under Drift: Paired Comparisons of Cumulative Subgroup Disparity
Machine Learning
Summary
Deciding when to update a computer model that makes decisions can affect how fairly it treats different groups. The authors compare several ways to decide when to retrain models based on performance changes or fixed schedules. They find that all updating methods generally reduce differences in error rates between groups compared to keeping the original model. However, the best method depends on factors like how errors appear over time and how groups change. Their work helps better evaluate fairness over a model’s lifetime.
What this means in practice
- •For machine learning teams: Assess and schedule model updates to manage fairness and error differences across groups during deployment.
- •For data ethics compliance teams: Monitor cumulative subgroup disparities over time to meet fairness and anti-discrimination guidelines in deployed AI systems.
Tested on simulated data.
Authors
Aaron Ceross
Abstract
Choosing when to retrain a deployed classifier requires assessing subgroup error rates across the sequence of models used, including periods between updates. We compare complete scheduled, loss-triggered, and subgroup-gap-triggered policies with retaining the initial model on the same observations and delayed labels. For true-positive and false-positive rates separately, the outcome is the paired difference in absolute subgroup gaps summed over deployment windows. Population evaluation in simulation, action records, and alternative schedules assess how measurement and retraining behaviour affect these comparisons. In a follow-up sample of 400 new trajectories per condition across two simulated drift regimes, all three policies had lower mean cumulative disparity, equivalent to reductions of 0.04 to 0.88 percentage points in the average gap per window. Evaluating the unchanged models against the known generating distributions preserved all mean directions, but finite-window and population comparisons agreed on whether updating increased, reduced or left cumulative disparity unchanged in 69 to 92 percent of trajectories. Under subgroup-specific drift, smaller true-positive-rate gaps accompanied lower sensitivity in both groups. In an exploratory American Community Survey replay, person weighting reversed all three race false-positive-rate mean comparisons without changing predictions or actions; all three weighted intervals included zero. Policy comparisons require group-specific rates, action distributions, and an explicit evaluation population alongside mean disparity. These analyses are non-confirmatory. Shared replay requires policy-independent observations and complete labels after the specified delay.