Papers for

fraud detection engineers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Density-ratio rescoring improves rare class detection in imbalanced data

Density-Ratio Rescoring for Imbalanced Classification

Abstract: Density-Ratio Rescoring (DRR) augments a classifier trained at the original class prior with a survey-raking dual score. Raking reweights the majority sample to match minority feature moments within a tolerance. DRR marginally standardizes the dual and base scores and combines them with a fixed weight of one half, using the fitted dual directly for prediction without resampling or refitting the base classifier. Under exact population matching and a correctly specified log-linear tilt model, the dual equals the log density ratio up to an additive constant. A class-separation analysis characterizes the signal strength and correlation conditions under which fusion improves separation under common within-class covariance. On 24 tabular benchmarks, evaluated over 30 trials and five base learners, DRR at the D=128 random-feature setting improves average precision over the standardized base on every dataset, with a mean gain of 0.034. It exceeds the shared-dual raking-and-relabeling resampler on 22 of 24 datasets, with a mean gain of $0.092$, and on all eight one-versus-rest tasks of a shared gene-expression cohort. These results demonstrate the effectiveness of using raking duals as reusable scores for improving rare-class ranking while retaining classifiers trained at the original prior.

Sun 20 SeptMachine Learning
The gist
When trying to identify rare cases in data, usual classifiers can miss them because the data is unbalanced. The authors propose a method called Density-Ratio Rescoring (DRR) that adjusts scores after training so that rare class features are better represented. DRR combines these adjusted scores with the original classifier scores without needing to retrain the model. Tests on many datasets show this method improves ranking of rare cases more effectively than other similar techniques.
Open → 2609.23926v1

Detecting coordinated online manipulation by analyzing data distortions

Principled Detection of Coordinated Manipulation from Aggregate Distortion and Account Reuse

Abstract: Coordinated manipulation is collective: plausible accounts can jointly distort ratings, rankings, and engagement. Existing defenses primarily construct evidence from identities, graphs, content, or co-activity. We introduce an aggregate-first evidence layer that treats distortion of a context-level outcome distribution as the primary evidence object. The engine observes only a histogram, count, resolution, and reference distribution; identities are withheld until interval evidence is fixed. Because raw discrepancies have positive finite-sample expectation, we subtract a matched null expectation to obtain signed evidence and account for reference uncertainty. Participation logs then accumulate these fixed increments across accounts. We characterize matched-exposure divergence, bound self-influence, establish finite-horizon separation, and derive an exact linear reuse law for paired contexts. We evaluate the mechanism with controlled rotation experiments and paired counterfactual interventions on historical Amazon review streams. Historical reviews provide the behavioral background; synthetic identities provide known coalition membership, and exact clean twins provide counterfactual controls. In a fixed-attack sweep against historical non-donor comparison accounts, reassigning the same manipulated events across identities with increasing reuse raises account-score ROC-AUC from 0.500 to 0.797. With activity- and exposure-matched clean twins, frequency is at chance while counterfactual attribution achieves ROC-AUC 0.744. Under a mean-preserving shape intervention, Wasserstein-1 and Jensen-Shannon evidence achieve ROC-AUC 0.909 and 0.967, while frequency and mean-based attribution remain at chance. Aggregate evidence complements repeated co-activity, improving mixed-mechanism ROC-AUC from 0.750 to 0.874 with a simple untrained combination.

Fri 11 SeptCryptography and SecuritySocial and Information Networks
The gist
Online platforms face problems when groups of users work together to unfairly change ratings or rankings. The authors introduce a new method that looks at how the overall scores or distributions are changed, rather than just focusing on who made the changes. This method uses mathematical tools to measure unexpected differences and tracks which accounts contribute to suspicious shifts. They tested their approach on Amazon review data by simulating attacks and comparing manipulated accounts to clean ones, showing it can better spot coordinated manipulation than traditional methods.
Open → 2609.13407v1