Beyond Modern Asymptotics for Log-Likelihood Ratios in Logistic Regression

2026-08-03Information Theory

Information TheoryMachine Learning
AI summary

The authors study how the log-likelihood ratio statistic behaves in logistic regression when the sample size and number of features are finite, not infinite. They find exact formulas describing the worst-case values the statistic can take, depending on the number of features, sample size, and confidence level, without needing common assumptions about the data. Their results show that in very low dimensions, the behavior changes noticeably, while high dimensions match classical theory when data are Gaussian. These findings provide precise, uniform bounds that work for all parameter choices, improving on previous asymptotic results.

log-likelihood ratiobinary logistic regressionfinite sample theoryWilks theoremquantiledesign matrixhigh-dimensional statisticsnonasymptotic analysisGaussian designstatistical inference
Authors
Hugo Chardon, Reese Pathak, Nikita Zhivotovskiy
Abstract
We characterize the finite sample behavior of the log-likelihood ratio statistic in binary logistic regression, uniformly over both the design and the target parameter. For $n\geq d\geq 3$, we determine, up to universal constants, its worst case $(1-δ)$ quantile over all fixed collections of design vectors and all target parameters: \[ d\log\left(\frac{e n}{d}\right)+\log\left(\frac{1}δ\right). \] This is a nonasymptotic analogue of the Wilks $χ^2_d$ phenomenon and requires no regularity assumptions on the design. The low dimensional cases exhibit unusual behavior. The worst case quantile in dimension $d=2$ is sharply of order \[ \log\log\log n+\log\left(\frac{1}δ\right). \] The worst case quantile in dimension $d=1$ is of order $\log(1/δ)$, with no dependence on $n$. Finally, i.i.d. Gaussian design vectors recover the classical Wilks scale. In the regime $n\gtrsim d+\log(1/δ)$, we prove the sharp bound \[ d+\log\left(\frac{1}δ\right). \] Unlike existing asymptotic results, our bounds are uniform over the target parameter, which may depend on $n$, $d$, and $δ$.