2026-09-16 | | Total: 14
Covariates missing not at random generally prevent point identification of regression coefficients without untestable restrictions. This paper studies linear regression when every missing covariate is restricted to a prespecified compact interval. The resulting population target is the set of best linear predictor coefficients compatible with the observed-data law. Under a non-atomic observed-data law and the stated regularity conditions, we can represent this identified set as the image, under the least-squares moment map, of the Aumann expectation of a random moment set. For the population set and its empirical analogue, we derive explicit bounds under bounded, sub-exponential and polynomial envelope conditions. The bounds show their dependence on sample size, dimension, confidence level and tail parameters. Under a Donsker condition and a uniform zero-set error bound, an oracle enlargement of the set-valued Z-estimator also converges at the root n rate. We separately study a Hadamard-design method for exploring the effect of admissible imputations without enumerating all vertices and build intuition to detect situations where the population target can be approximated by a zonotope. A numerical experiment illustrates how co-missingness and the width of the imputation intervals affect this diagnostic.
Two features intrinsic to survey sampling complicate semiparametric efficiency analysis: design-induced dependence among sampling indicators and the randomness of finite-population targets under the superpopulation law. For general semiparametric full-data models, we show that the observed-data experiment under a broad class of dependent designs is locally asymptotically normal with the tangent-space structure of a reference Poisson experiment. Under the joint superpopulation-design law, first-order efficiency depends on the design only through the limiting inclusion-probability function. Standard missing-at-random projection in the reference experiment characterizes the observed-data efficient influence function. Finite-population targets are treated through first-order asymptotic expansions, extending the analysis beyond exact random sums to nonlinear census characteristics. Superpopulation and finite-population centerings yield equivalent notions of local regularity and efficiency, with their bounds linked by a Pythagorean decomposition that gives a generalized finite-population correction. We then give general and design-specific conditions under which cross-fitted estimators with estimated nuisance functions attain both efficiency bounds. For the finite-population mean, the bound equals the large-sample limit of the Godambe-Joshi anticipated-variance lower bound. For scalar targets, we characterize optimal limiting inclusion probabilities. Simulations and California Academic Performance Index data illustrate the theory.
In this paper we study a linear drift perturbed by a superposition of $m$ independent fractional Brownian motions with known Hurst parameters and a common scale, observed at $N$ equidistant times. Inference for such models is usually asymptotic; we show that here it is exact. We derive the maximum likelihood estimators of the drift $θ$ and of the scale $α^{2}$ in closed form and obtain their exact finite-sample joint law: $\widehatθ$ is Gaussian, $N\widehatα^{\,2}/α^{2}$ is chi-square with $N-1$ degrees of freedom, and the two are independent. As this law is free of every model parameter, we deduce Student and chi-square confidence intervals and tests of exact level for every $N\ge2$, whatever the Hurst vector. We also prove that the estimators are uniformly minimum variance unbiased with $\widehatθ$ attaining the Cramér--Rao bound at every $N$, that both are strongly consistent and asymptotically normal, and that the drift estimators form, in law, a Brownian motion run along their own variance scale. A sharp non-asymptotic bound shows that the accuracy of the drift is governed by the length of the observation window and not by the mesh, and a Monte Carlo study confirms exact coverage, even at small sample sizes, and quantifies what is lost when the Hurst vector is misspecified.
For multivariate scale and location--scale models with independent components, we extend the univariate results of Zhou and Nayak (2012) and derive optimum equivariant estimators under the generalized Pitman closeness criterion. We first show, by a counterexample, that in the multivariate case the Pitman closeness comparison within the class of equivariant estimators is not transitive, so that a Pitman closest equivariant estimator does not exist in general. We then enlarge the transformation group by the coordinate permutations---equivalently, impose the formal equivariance principle of Berger (1985) across isomorphic component problems, in the spirit of the separable rules of Robbins' (1951) compound decision theory---and show that within the resulting restricted class an optimum is restored. A multivariate median lemma based on strictly convex losses then yields explicit Pitman closest equivariant estimators of the scale parameters, powers of the scale parameters, and the location parameters, given by median-adjusted versions of any given equivariant estimator. Applications to the multivariate uniform and multivariate normal distributions are worked out in detail. Monte Carlo experiments for the Rayleigh distribution and for a competing risks model with Rayleigh component lifetimes confirm that the proposed estimators dominate the maximum likelihood and Bayes estimators under the Pitman closeness criterion, and a real industrial data set on ball bearing failure times illustrates the feasibility of the method in practice.
Equivariance is increasingly used in machine learning and statistics, often without systematic justification. In a companion article, the equivariance criterion was applied to the normal linear model with a fixed design matrix (fixed-$X$), yielding the minimum risk equivariant (MRE) estimators of the coefficient vector and of the condensed diagonal covariance matrix under a multivariate invariant location--scale group. We extend these results to the random-$X$ case, with covariates sampled from a population. The extension hinges on a distinction vacuous for fixed-$X$ but fundamental for random-$X$: whether risk and unbiasedness are evaluated conditionally on the realized design or after averaging over the design distribution. Under conditional evaluation, the fixed-$X$ group applies given $X$: least squares remains the best equivariant estimator of the coefficient vector, and the MRE estimators of the population variances keep their fixed-$X$ forms with population sizes at the realized design. Under absolute evaluation with an i.i.d.\ design, the picture changes qualitatively: the natural scale group acting jointly on $(Y,X)$ fixes the coefficient vector, the induced parameter-space action is intransitive, equivariant risks are constant only along orbits indexed by the signal-to-noise ratio $ρ=\|β\|^2/σ^2$, and no uniformly minimum risk equivariant estimator exists. In the scalar case the optimal equivariant weight is the oracle shrinkage factor $w^*(ρ)=ρ/(ρ+E[T^{-1}])$, with least squares recovered as the infinite-signal limit $ρ\to\infty$---explaining and refining the known failure of the Gauss--Markov theorem with random regressors. For a centered design under location--scale transformations, least squares remains optimal within the natural invariant-contrast class, and the MRE estimator $S^2/(n-p+2)$ of the error variance is valid under both modes.
Protecting individual privacy has become a central and urgent concern in modern data analysis, given the vast quantities of data now generated and processed. In this paper, we develop a systematic theory of exact asymptotic efficiency for regular parametric estimation under central zero-concentrated differential privacy. In the privacy regime, the governing object is a diameter-constrained information region: the set of information matrices generated by statistics with diameter at most one. For weighted quadratic loss, we show that the exact local minimax risk is an inverse information variational functional over this region. More generally, a mixed information region yields a unified efficiency constant across different regimes, covering the privacy regime and classical Fisher efficiency. A matching estimator releases the empirical mean of a nearly optimal bounded statistic with Gaussian noise and locally inverts its population moment map. Our theory differs from classical efficiency theory in its set-valued information geometry and loss-dependent efficient estimator. As examples, we provide closed-form constants and optimal procedures for various concrete models, including one-dimensional regular families, Gaussian means, categorical probability vectors, and regression models among others. The theory also transfers to Gaussian differential privacy through an exact parameter rescaling.
This work addresses a longstanding gap in the statistical foundations of marginal maximum likelihood estimation for high-dimensional latent variable models. Marginal maximum likelihood estimation is widely used to fit latent variable models across the social sciences, ecology, and machine learning. Despite its broad use, rigorous asymptotic theory for nonlinear models remains limited when both the sample size and the number of observed variables diverge. The gap arises largely from the fact that integration over the latent variables creates a nonlinear objective that tightly couples the high-dimensional model parameters. To address this issue, we first show that the marginal likelihood exhibits multiple nearly flat directions even at the true parameter, in contrast to the behavior in the fixed-dimensional case. Building on this geometric characterization, we develop new techniques to establish consistency and asymptotic normality for the marginal estimator of the high-dimensional parameters. For the latent variables, we provide frequentist and Bayesian uncertainty quantification, proving asymptotic normality of the maximum a posteriori estimator and a Bernstein-von Mises-type result for the plug-in posterior. Together, these results provide rigorous foundations for marginal estimation and latent-variable inference in high-dimensional models.
Parameter estimation in finite mixture models can exhibit highly heterogeneous convergence behavior: locally isolated components may be estimated substantially faster than groups of competing components. Existing analyses based on Wasserstein distances typically characterize only the worst-case rate and therefore do not fully capture this local heterogeneity. In this paper, we introduce a Voronoi-based partial optimal transport (VPOT) framework for obtaining refined local and global convergence guarantees for the maximum likelihood estimator of the mixing measure. The key geometric idea is to localize the comparison of two mixing measures to extended Voronoi neighborhoods and use partial optimal transport to accommodate the unequal masses of their local restrictions. Within each neighborhood, the first-order POT discrepancy is raised to a power determined by the number of locally competing atoms, allowing the resulting loss to adapt to the local degree of singularity. Under suitable regularity and strong identifiability conditions, we establish uniform local and global upper bounds for a maximum likelihood estimator under the VPOT loss. These bounds reveal a configuration-dependent form of parameter estimation: less singular local configurations admit faster convergence, whereas the most singular configuration recovers the classical worst-case behavior characterized by Wasserstein-based analyses. We further establish a minimax lower bound showing that the convergence rate for estimating the mixing measure under the VPOT loss is optimal. Our results hold in arbitrary fixed dimension without requiring mixing proportions to be uniformly bounded away from zero or prior knowledge of the true number of mixture components. Overall, VPOT provides a configuration-adaptive framework for capturing heterogeneous parameter-estimation behavior in finite mixture models.
We show that, for a given dimension, the set of Spearman's rank correlation matrices and that of linear correlation matrices coincide if and only if the dimension is no larger than nine. For this, we construct an extreme rank-four counterexample in dimension ten and prove its incompatibility using moment identities and Cauchy-Schwarz. Appending unit directions produces counterexamples in every higher dimension. This, together with existing results, completes the dimensional classification and settles a long-standing open question in quantitative risk management.
For any fixed $k$, we prove a lower bound on the $k$th largest modulus of an eigenvalue of the non-backtracking matrix $B$. Specifically, consider any deterministic or random family of graphs that converges locally to the unimodular Galton-Watson tree with root degree distribution $D$, and set $κ:=\mathbb E[D(D-1)]/\mathbb E[D]$. Given $κ>1$ and an exponential-moment bound on the empirical degree distributions, we show that $|λ_k(B)|\geq\sqrtκ-o_N(1)$, where $N$ is the number of vertices. When restricted to locally tree-like regular graphs, this recovers a well-known consequence of the Ihara-Bass formula. In the specific case where the graph is generated through the Erdős-Rényi model with expected degree $d>1$, this proves a conjecture of Bordenave, Lelarge, and Massoulié. To do this, we show that the normalized log-determinant of the Bethe-Hessian of the graph is bounded by that of the Bethe-Hessian of its local limit. This bound is violated if the eigenvalues of the non-backtracking matrix are too small. We establish this using an effective-conductance interpretation of the tree Green's function recursion.
Large language model (LLM) distillation aims to transfer the capabilities of a powerful teacher to a smaller student. Direct imitation, however, can also transfer the teacher's systematic bias and errors. This challenge is particularly pronounced under covariate shift, when the teacher's reliability on target questions is uncertain and target-domain reward feedback is unavailable. We propose Coupled Calibration and Learning (CCL), an LLM distillation algorithm that couples teacher calibration with student updates through token-level branching, using reward feedback only on source questions. Each iteration calibrates the teacher using source feedback and then uses the calibrated teacher to train the student on target questions. The updated student, in turn, informs subsequent calibration. In an autoregressive policy framework, we prove that the output student's expected average Kullback-Leibler divergence to the oracle student converges to zero at a polynomial rate in the number of iterations. The oracle maximizes the true reference-regularized target reward within the student class, which need not represent the unrestricted optimal policy. Our analysis quantifies the progress of projected student gradient updates while controlling the error in teacher calibration. We further establish a separation from regularized direct matching: its error relative to the oracle student can remain bounded away from zero even when the teacher achieves higher regularized target reward than every student policy. These results demonstrate that LLM distillation can overcome persistent teacher bias and recover the optimal student through coupled calibration and learning, without target-domain reward feedback.
By consuming multiple copies of an unknown qubit state, one can modify its purity while preserving the direction of its Bloch vector. We determine the maximum linear rate at which qubit states of different purities can be interconverted, allowing a nonzero error, quantified, for instance, by the trace distance, provided that it vanishes in the limit of infinitely many copies. Interestingly, the optimal conversion rate is determined by the two eigenvalues of the complex right-logarithmic-derivative (RLD) Fisher information matrix associated with $\mathrm{SU}(2)$ rotations of the qubit state. When the output qubits have higher purity, corresponding to concentration, the optimal rate is given by the ratio of the maximum eigenvalues of the input and output RLD matrices. In contrast, when the output qubits have lower purity, corresponding to dilution, the optimal rate is given by the ratio of their minimum eigenvalues. Remarkably, both concentration and dilution can be implemented using SWAP tests as the only nontrivial two-qubit measurement primitive, together with ancillary qubits initially prepared in maximally mixed states, without requiring any additional two-qubit gates. Our work thus provides a novel operational interpretation of the full complex RLD Fisher information matrix. Crucially, its antisymmetric, purely imaginary part encodes geometric information beyond the statistical distance between density operators and plays an essential role in determining the optimal state-conversion rates.
Policy learning aims to determine who should be treated based on individual characteristics. In high-stakes settings such as medicine and public policy where safety is a central concern, improving the average outcomes alone may not be sufficient: decision makers may also seek to protect individuals from harm, in line with the Hippocratic principle of ``do no harm.'' In this paper, we propose \textit{conformal policy learning} (CPL), a policy learning procedure with a new distribution-free safety guarantee that controls the probability of assigning treatment to an individual who would be harmed relative to control. CPL views each treatment decision as testing a hypothesis of counterfactual harm and assigns treatment by thresholding conformal p-values. These p-values use observable proxies and selective calibration to address the challenge that the potential outcomes under comparison are never simultaneously observed. For randomized experiments, under standard exchangeability conditions, CPL provides finite-sample safety guarantee at a user-specified level, without imposing any outcome modeling assumptions. Moreover, when the outcome model is consistently estimated, CPL achieves asymptotically optimal welfare subject to the safety constraint. In observational studies, CPL with learn-then-balance weights achieves doubly robust safety guarantees. We evaluate CPL through extensive simulations and apply it to an empirical study of AI-powered interventions designed to reduce conspiracy beliefs.
Langevin-based Markov chain Monte Carlo (MCMC) algorithms use gradient information to improve sampling, particularly in high dimensions. Classical optimal scaling theory for these algorithms has largely focused on the Metropolis-Hastings (MH) acceptance rule. However, there has been a recent surge in acceptance rules beyond MH for applications spanning differential privacy, stochastic MCMC, diffusion models, and molecular dynamics. We develop optimal scaling results for Langevin proposals employed with generalized acceptance rules belonging to a suitable class. For high-dimensional targets, we recover the usual $O(d^{-1/3})$ scaling, while different acceptance rules lead to different optimal acceptance probabilities.