Sharp risk bounds clarify challenges in estimating partial linear models

Sharp Structure-Agnostic Minimax Risk for Partial Linear Models

Machine Learning

Summary

Estimating relationships in partial linear models can be tricky when using two separate black-box methods to learn hidden factors. The authors solve an open problem by precisely describing how the estimation error depends on the balance between approximation errors and learning errors in each method. They introduce a new way to prove a lower bound on the error by testing different scenarios where errors interact in complex ways. Their findings show that combining information about both methods is crucial for the best estimation, improving on previous approaches that treated them separately. This work helps guide the choice of methods for better accuracy in such models.

partial linear modelminimax riskapproximation errorstochastic errorlocalized Rademacher complexitydouble machine learningcoefficient estimationlower boundblack-box learnersmixture testing

Authors

Haichen Hu, David Simchi-Levi

Abstract

We characterize the sharp structure-agnostic minimax risk for coefficient estimation in the partial linear model when the outcome and treatment nuisances are learned by two distinct black-box learners, which resolves the open problem in double machine learning posed by Gu (2025). For each nuisance \(q\in\{μ,π\}\), we characterize the available learner by an approximation-error budget \(a_q\) and a stochastic-error budget \(s_q\), with the latter controlled through localized Rademacher complexity. Writing \(\mathcal E_n\) for the minimax mean-squared error, we show that \[\mathcal E_n\asymp1\wedge\left\{\frac1n+\left(a_μa_π+\min\left\{a_πs_μ+s_π^2,\,a_μs_π+s_μ^2\right\}\right)^2\right\}.\] The main new ingredient is a novel lower bound for the general two-learner problem. Our proof constructs four finite-mixture testing experiments using orthogonal code functions. Across these experiments, the hidden perturbations are placed outside both learner classes, outside only the treatment learner class, outside only the outcome learner class, or inside both learner classes. These four configurations capture, respectively, the interaction between the two approximation errors, the two asymmetric interactions between one learner's approximation error and the other learner's learning error, and the joint estimation difficulty of learning both nuisances. Combining the four resulting lower bounds yields the displayed rate, which matches the latest upper bound in Gu (2026). Our result shows that standard double machine learning can overstate the intrinsic difficulty of target estimation and provides a target-specific principle for learner selection: approximation error and stochastic complexity must be jointly balanced across the two nuisance learners rather than optimized separately.