Statistics

2026-09-28 | | Total: 73

#1 First-Order Stationarity of Reverse Diffusions [PDF] [Copy] [Kimi] [REL]

Authors: Zhifeng Chen, Chenyang Jiang, Yazhen Wang

Recent literature has shown a strong connection between optimization and sampling. We develop the corresponding first-order theory for diffusion models. First, the SDE-based reverse-time flows of overdamped and underdamped Langevin diffusions contract relative Fisher divergences at explicit exponential rates whenever the stationary potential of the forward process is strongly convex---a condition on the noising process one chooses, not on the data. This is a unique advantage of SDE-based reverse diffusion, absent in the reverse process based on ODEs. Second, we incorporate discretization and establish averaged first-order stationarity bounds---the sampling analog of averaged gradient-norm guarantees in nonconvex optimization---for samplers of both overdamped and underdamped diffusion models. As in nonconvex optimization, the convexity-free certificate is local: it guarantees score consistency, not global mode weights.

Subjects: Machine Learning , Machine Learning

Publish: 2026-09-25 17:57:43 UTC


#2 Statistical attribute alignment for black-box generative AI via output post-processing [PDF] [Copy] [Kimi] [REL]

Authors: Kevin Jiang, Morgane Austern, Edgar Dobriban, Jason M. Klusowski

Generative AI systems are increasingly used, but aligning their outputs with user requirements poses a continuing challenge. Here, we aim to ensure that the distribution of an attribute of an AI-generated output aligns with a user-specified target. This is motivated by examples such as fairness, where we want to ensure that a protected attribute (e.g., gender, race, or age categories) follows a desired distribution, and synthetic data generation, where we want the generated data to be representative of a target distribution. We study the practically important black-box access setting, where a user can repeatedly query a generative AI model. The goal is to return $m\ge 1$ outputs whose joint attribute distribution is as close as possible to this target. For both exact and approximate alignment, we develop algorithms that minimize the expected number of queries to the generator, and we further demonstrate their optimality as the number of requested outputs $m \rightarrow \infty$. Experiments on text-to-image generation and geocoded persona generation tasks show that our post-processing algorithms improve statistical attribute alignment, complementing prompting-based interventions.

Subjects: Methodology , Artificial Intelligence , Machine Learning , Statistics Theory

Publish: 2026-09-25 17:55:49 UTC


#3 LAT: a Latinized aperiodic tiling for any sample size [PDF] [Copy] [Kimi] [REL]

Author: Pamphile T. Roy

Quasi-Monte Carlo methods allow computer experiments to be run with far fewer simulations than crude Monte Carlo. A digital net in base two, such as Sobol', is however only balanced when the number of samples is a power of two, and Latin Hypercube Sampling (LHS) accepts any sample size but only controls the one-dimensional margins. This work proposes a space- filling design defined for any sample size, referred to as LAT for Latinized aperiodic tiling. The unit hypercube is cut recursively across its longest edge following the golden section into N cells of equal volume. One point is then placed in each cell, and the margins are made Latin while every point stays inside its own cell. The construction only uses integer splits and costs O(dNlog N). The recursion is shown to follow the Fibonacci word and its aperiodicity is analysed. LAT is assessed with four L2-discrepancies and with the integration error on analytical functions and engineering emulators, and compared to Monte Carlo, LHS, Halton, Sobol' and a rank-1 lattice. LAT is better than Monte Carlo and LHS as soon as the integrand has interactions. Sobol' remains more accurate at the powers of two, but its error is one to two orders of magnitude larger at other sample sizes. The accuracy of LAT does not depend on the sample size. Finally, the cells form a partition of the hypercube for any N. This allows one to refine the design locally, to search for an optimum by splitting cells and to sample non-rectangular regions.

Subject: Methodology

Publish: 2026-09-25 17:51:55 UTC


#4 Uncertainty and Explainability in Deep Rough Volatility: A Neural Information-Theoretic Posterior Approach [PDF] [Copy] [Kimi] [REL]

Authors: Damiano Brigo, Raphaël Huser, Dan Leonte

Deep learning has substantially accelerated the calibration of complex stochastic-volatility models, but neural point calibration alone does not capture the uncertainty remaining after an implied-volatility (IV) surface has been observed. We develop a simulation-based inference framework for rough Heston (rHeston) calibration that learns the posterior distribution of the model parameters conditional on an IV surface. Using neural ratio estimation, we obtain calibrated posterior samples that can be propagated through heteroscedastic neural surrogate pricers for path-dependent exotic options. The resulting posterior-predictive distributions combine residual parameter uncertainty with conditional surrogate uncertainty and yield uncertainty-aware price intervals. We further introduce Hellinger-SHAP, an information-theoretic explainability method for posterior inference. Rather than attributing a single parameter point estimate, it applies local-background Kernel SHAP to a posterior-information functional measuring contraction from the prior to the posterior. This identifies maturity--moneyness regions associated with posterior information gain for individual rHeston parameters. In a simulation study, posterior-predictive intervals provide calibrated or conservative coverage across forward-start, barrier, and realized-variance claims, while point plug-in prices can be materially unreliable for selected contract regimes. Together, the UQ and XAI analyses provide a transparent framework for uncertainty-aware neural calibration and downstream exotic pricing under the specified prior-predictive model.

Subjects: Machine Learning , Machine Learning , Applications , Computation , Other Statistics

Publish: 2026-09-25 17:34:56 UTC


#5 Two Conformal Constructions for Adaptive Within-Document AI-Text Screening [PDF] [Copy] [Kimi] [REL]

Authors: Marco Mandap, Jerahmeel Hipolito, Arcel Galvez, Charlie Margaret Balagtas, Michael Joshua Buluran, Jeff Roel Durmiendo, Rizzette E. Lopez

We study false-alert control when screening for text generated by artificial intelligence (AI). The screening procedure selects document prefixes and detectors from observed evidence and may stop before exhausting its inspection budget. We give two finite-sample constructions under document-level exchangeability between human calibration documents and a new null document, with no restriction on dependence among tokens within a document. Construction A registers a finite family of prefix-detector scores and allocates a false-alert budget across their conformal ranks. A union bound protects any executed subset of that family. Construction B calibrates the complete-path maximum of a development-fixed adaptive policy. Each partial-path maximum is bounded by the complete maximum, so a terminal conformal rank protects early stopping without splitting the error budget. We prove marginal control of any false alert across the permitted inspection path and derive necessary calibration counts for rejection. We also state oracle testing, distribution-shift, and independent-audit bounds with their additional assumptions. Both constructions protect stopping within their specified scope; neither proof constructs an e-process or justifies multiplying conformal ranks. Detection power and computational savings remain questions for empirical evaluation.

Subjects: Methodology , Computation and Language

Publish: 2026-09-25 17:18:39 UTC


#6 Bi-invariant Geodesic Regression: Existence, Uniqueness, and Convergence [PDF] [Copy] [Kimi] [REL]

Authors: Martin Hanik, Christoph von Tycowicz

Bi-invariant geodesic regression generalizes linear regression to Lie groups. Its main feature is that it respects the symmetries of the group so that the resulting estimator is independent of arbitrary choices such as a reference frame. However, the local existence and uniqueness of the underlying estimator have not been shown until now. Furthermore, the convergence properties of the proposed algorithm for computing the estimator are not known. In this work, we investigate these questions. We prove that, locally, a unique estimator exists and give explicit bounds on the size of this neighborhood. We also show that the proposed iterative algorithm converges linearly to this estimator.

Subjects: Statistics Theory , Differential Geometry

Publish: 2026-09-25 17:03:24 UTC


#7 Statistical Foundations for a Google Play User-Review Sentiment Index: Signal Fusion, Shrinkage, Distributional Validation, and Dynamic Smoothing [PDF] [Copy] [Kimi] [REL]

Author: Marco Mandap

We develop a statistically explicit sentiment index for Google Play user reviews and establish the mathematical results supporting its construction. Normalized star ratings and text-sentiment scores are treated as noisy measures of latent review valence and fused by covariance-aware inverse-variance weighting. Review-level estimates are aggregated with bounded helpfulness and recency weights, then shrunk toward a population mean using estimated precision rather than an arbitrary review-count threshold. App-level rating histograms provide a distributional diagnostic for samples returned under different API sort orders; because star ratings are discrete, classical continuous Kolmogorov-Smirnov critical values are not used. A local-level state-space model and the Kalman filter provide a denoised temporal trend. Full proofs cover the BLUE and Gaussian maximum-likelihood result, Gaussian-conjugate shrinkage, the Glivenko-Cantelli and Donsker theorems, count transformations via the delta method, and exact Gaussian Kalman filtering. A worked three-review example shows how textual complaints can materially reduce an apparently perfect star-only score.

Subjects: Methodology , Computation and Language , Applications

Publish: 2026-09-25 16:54:09 UTC


#8 Beyond Empirical Support: Structured Outlier Generation via Sinkhorn Optimal Transport [PDF] [Copy] [Kimi] [REL]

Authors: Haixiang Sun, Andrew L. Liu

Outliers are essential for evaluating and improving the robustness of machine learning systems, especially when future distributions may differ significantly from historical training data. In high-stakes applications, robustness often depends on rare cases that finite datasets fail to capture, making simple resampling or perturbation insufficient for stress scenario generation. Existing outlier synthesis methods typically rely on sparse neighborhoods, low support latent regions, or classifier boundary crossings, which can be heuristic, unstable, and tied to specific modalities or architectures. We therefore propose Sinkhorn Boundary Outlier Generation (SBOG), a structured framework for latent-space outlier generation that couples Sinkhorn optimal transport geometry with distributionally robust boundary modeling. The resulting Sinkhorn-induced support cost guides the sampler toward weakly supported boundary regions, while semantic constraints prevent uncontrolled drift from the intended context, yielding controlled deviations from the in-distribution reference measure rather than arbitrary sparse-region samples. Experiments on time series anomaly generation and image outlier synthesis show that our framework produces informative, semantically controlled outliers and improves downstream robustness evaluation across modalities, providing a foundation for stress scenario generation beyond empirical support.

Subjects: Machine Learning , Machine Learning , Optimization and Control

Publish: 2026-09-25 16:18:02 UTC


#9 A likelihood-based coefficient for biomedical independence testing: the binomial-cut composite likelihood ratio [PDF] [Copy] [Kimi] [REL]

Author: Jing Qin

The standard dependence summaries used in biomarker studies -- Pearson's r, Spearman's rho, Kendall's tau -- take values in [-1, 1] with 0 indicating no linear or monotone association. Zero does not distinguish independence from non-monotone dependence, so the scale cannot represent threshold effects, heteroscedasticity, and tail shifts common in biomarker practice. We formulate independence testing as a composite Bernoulli likelihood ratio: at each threshold t, comparing the Bernoulli laws of 1(Y <= t) conditionally on X versus marginally, aggregated over cut points. The resulting coefficient xi_cut lies on [0, 1] with 0 iff X and Y are independent (under continuity of Y) and 1 iff Y is a measurable function of X. Fisher weighting arises at second order from the Bernoulli likelihood, and xi_cut equals twice the threshold-averaged mutual information between X and 1(Y <= t), giving a distribution-free lower bound on I(X; Y). A second-order expansion recovers the Fisher-weighted Dette-Siburg-Stoimenov measure, which coincides under continuity with Chatterjee's rank correlation. Estimation uses a Nadaraya-Watson plug-in with a max-over-grid bandwidth; inference is by exact permutation. In biomarker-motivated simulations T_cut substantially outperforms rank-based coefficients on W-shaped non-monotone and heteroscedastic alternatives. We illustrate on the Seattle cohort (n=70, ages 21-88) of the aging plasma proteome dataset, screening all 1,305 proteins for age dependence: under Benjamini-Hochberg control at q<0.05, T_cut rejects on 70 proteins, six of which are missed by Pearson, Spearman, and Chatterjee at the same FDR level.

Subjects: Methodology , Applications

Publish: 2026-09-25 16:17:20 UTC


#10 Nonparametric In-Context Learning under Growing Geometric Complexity: Minimax Optimality and Local Geometry-Adaptivity of Transformers [PDF] [Copy] [Kimi] [REL]

Authors: Jaehee Seo, Jisu Kim

Transformers have become a central architecture for in-context learning (ICL), particularly through their state-of-the-art performance in large language models. This success motivates understanding how transformers exploit task-relevant structure in geometrically heterogeneous data. However, existing nonparametric ICL theory has largely focused on Euclidean domains or single-manifold models. To address this gap, we study the prediction problem under unknown local geometry, modeled by sample size-dependent mixtures of manifolds with heterogeneous dimensions, smoothness, and sampling masses. Under local separation and small-perturbation conditions, we establish a minimax lower bound capturing the aggregate difficulty of the components and construct an oracle tangent local-polynomial estimator with a matching upper bound. This estimator is connected to a structure-informed, two-stage softmax transformer with a geometric preconditioner and chartwise reduced local-polynomial solvers. The transformer achieves negligible approximation error relative to the minimax rate with logarithmic depth and polynomial size. Finally, we derive an in-context generalization bound for near empirical risk minimizers over this class. Together, these results identify conditions under which the resulting predictor exploits local geometry and attains the aggregate minimax rate.

Subjects: Machine Learning , Machine Learning , Statistics Theory

Publish: 2026-09-25 16:10:53 UTC


#11 Equation discovery with Bayesian tree-adjoining grammars [PDF] [Copy] [Kimi] [REL]

Authors: Christopher A. Lindley, Nikolaos Dervilis, Keith Worden

Tree-Adjoining Grammars (TAGs) have recently been introduced to Nonlinear System Identification (NLSI) as a means of encoding an entire model class as a finite set of grammatical rules, from which candidate models are assembled as trees. Existing TAG-based identifiers rely on evolutionary optimisation and return point estimates of the model structure. This paper instead proposes the TAG framework within a Bayesian setting. A generative prior is defined over tree structures and their parameters, and a Reversible-Jump MCMC sampler with structure-preserving tree moves is used to infer the joint posterior over model structure, parameters and predictions. Two training objectives are considered; that is, a one-step-ahead objective with conjugate parameter proposals, and a simulation-based objective handled by likelihood-free inference. The approach is validated on a simulated polynomial NARX system, the Silverbox benchmark, and wave-loading data from the Christchurch Bay Tower, where embedding Morison's equation as a fixed initial tree yields a grey-box model that outperforms the physics-driven baseline. The results demonstrate that Bayesian TAGs are well suited to quantifying uncertainty in equation discovery for dynamical systems and to fitting physics-informed models.

Subjects: Machine Learning , Machine Learning , Systems and Control , Computation

Publish: 2026-09-25 15:04:40 UTC


#12 On the asymptotic shape of quantile surfaces [PDF] [Copy] [Kimi] [REL]

Authors: Florian Gach, Simon Hochgerner

This article is concerned with the asymptotic shape of quantile surfaces, defined as the set of quantiles at a given level $α$ generated by a controlled one-dimensional distribution. Specifically, when the distribution arises as a linear combination of log-normal random variables and the control is a vector of positive coefficients, we prove that quantile surfaces are globally concave in the left tail ($α\to0$) and globally convex in the right tail ($α\to1$). Moreover, these surfaces exhibit asymptotic separation of scale and shape.

Subjects: Statistics Theory , Probability , Mathematical Finance

Publish: 2026-09-25 14:44:45 UTC


#13 Modeling dependent degradation data considering inherent causal relationships for reliability analysis and remaining useful life prediction [PDF] [Copy] [Kimi] [REL]

Authors: Shi-Shun Chen, Xiao-Yang Li

Accurate modeling of dependent degradation processes is essential for credible reliability assessment and remaining useful life (RUL) prediction in complex systems. Existing dependent degradation models usually describe dependence using correlation-based methods with symmetric characteristics. However, they ignore inherent causal directionality between degradation paths, which may lead to biased reliability evaluation and RUL predictions when physical causality exists. To address this issue, this paper proposes a causality-driven framework for modeling dependent degradation data. Firstly, univariate degradation models with multi-source uncertainties are established for each performance indicator based on the Wiener process. Then, the stable Peter-Clark algorithm is employed to uncover inherent causal relationships between degradation processes, and an uncertainty-aware neural network is employed to quantify causal effects considering uncertainties. Next, univariate degradation predictions and causal predictions are integrated within a Bayesian framework to construct a causally dependent degradation model, and the corresponding loss function is derived for model training. Finally, system reliability and RUL predictions are derived via Monte Carlo simulation. The proposed methodology is validated on the C-MAPSS dataset. Results show that considering inherent causal directionality between degradation processes helps eliminate physically unrealistic degradation behaviors, yielding more accurate degradation and RUL predictions than independent and correlation-based dependent degradation models.

Subject: Applications

Publish: 2026-09-25 14:43:03 UTC


#14 Factorial Multivariate Bayesian Causal Forests: heterogeneous main and interaction effects of multiple treatments on correlated outcomes [PDF] [Copy] [Kimi] [REL]

Author: Danilo A. Sarti

Many studies expose units to several binary treatments at once and record several correlated outcomes, yet analysts usually estimate one treatment's average effect on one outcome at a time -- discarding how treatments interact and how their effects vary across units. We introduce the factorial multivariate Bayesian causal forest, a Bayesian nonparametric model that decomposes the factorial response surface over the treatment lattice into a prognostic sum-of-trees plus one sum-of-trees per main or interaction effect, each with correlated multivariate leaf parameters, estimated jointly with coherent uncertainty. A single indicator-weighted kernel samples every component -- the prognostic term is the special case whose indicator is unity -- so the two-treatment model, the single-treatment multivariate causal forest, and a general order-r truncation are one and the same sampler. We give the ANOVA/Möbius identification of each estimand, an efficient Rcpp engine, and an interpretability layer using value-suppressing uncertainty maps. In simulations the average-effect estimators are unbiased with nominal coverage, consistent, and robust under misspecification, failing only under unmeasured confounding, which we flag. We illustrate the method on clinical, agricultural and economic data and on a deeper application to the NHANES survey. Software is provided as an R package.

Subjects: Methodology , Applications , Computation

Publish: 2026-09-25 14:38:17 UTC


#15 Geometric Moment Contraction for Stochastic Nesterov Acceleration [PDF] [Copy] [Kimi] [REL]

Author: Wei Biao Wu

We study geometric moment contraction (GMC) of the constant-parameter stochastic Nesterov recursion \[ Y_k=Θ_k+β(Θ_k-Θ_{k-1}),\qquad Θ_{k+1}=Y_k-γG(Y_k,X_{k+1}). \] Under mean strong monotonicity and stochastic $L^p$ Lipschitz continuity, an explicit Perron comparison proves synchronous $L^p$ contraction when $βγL_p<(1-β)(1-q_{γ,p})$. This direct criterion includes infinite-variance gradients for $1<p<2$, but its small-step regime requires $β<μ/(μ+L_p)$. A complementary power-Lyapunov argument establishes a positive, generally much smaller, step-size interval for every fixed $β<1$ and every $p>1$, using only a finite $p$th gradient moment. At $p=2$, a simpler explicit certificate gives \[ 0<γ<\frac{2μ(1-β)^2}{L_2^2(1-β+2β^2)}. \] Its quadratic high-momentum scaling is a limitation of the chosen metric, not a sharp stability boundary. We quantify this loss, provide a general mean-only quadratic $S$-procedure, and exploit endpoint Lyapunov inequalities under stronger samplewise sector information. Verified endpoint certificates can be orders of magnitude less conservative than the explicit metric.

Subjects: Machine Learning , Machine Learning

Publish: 2026-09-25 14:14:14 UTC


#16 Rate-Preserving Shrinking-Support Gaussian Process Prediction [PDF1] [Copy] [Kimi] [REL]

Authors: Xiaopeng Xiang, Wenlin Dai, Marc G. Genton, Wenjia Wang

Gaussian process prediction is a central tool in spatial statistics, but standard implementations require dense matrix operations that become prohibitive for large datasets. We propose a scale-adjusted compactly supported working correlation for Gaussian process prediction, using the generalized Wendland family with a support radius $φ_n$ that is allowed to decrease with the sample size $n$. Under fixed-domain asymptotics with quasi-uniform designs, a fixed support radius does not yield asymptotic sparsity, whereas a shrinking support radius can make the covariance matrix sparse. We show that $φ_n$ and the regularization parameter can be jointly chosen so that the resulting predictor preserves the optimal integrated mean squared prediction error rate while reducing the number of nonzero covariance entries and the associated sparse matrix-vector multiplication cost. We also establish an analogous rate-preserving sparsification result for kernel ridge regression. Simulations and an ERA5 temperature application show that the proposed method achieves competitive prediction accuracy while retaining the computational efficiency provided by sparse linear algebra.

Subject: Methodology

Publish: 2026-09-25 13:54:39 UTC


#17 Dynamic factor and double PCA models for partially observed survival curves: Forecasting demand in short-term rental markets [PDF] [Copy] [Kimi] [REL]

Authors: Marthe Elisabeth Aastveit, Alex Lenkoski, Thordis Thorarinsdottir

This paper develops prediction models for population-level survival curves observed over time and sampled from a heterogeneous mix of populations. We consider a discrete-time setting where each curve is only partially observed and forecasts of the remaining trajectory are needed for downstream decision making. Our approach recasts cross-population heterogeneity into a multivariate sampling model. We propose two forecasting models for partially observed curves: a full factor analysis model that extends a general factor representation to incorporate the partially observed survival curve, and a double PCA model. The methodology is motivated by demand forecasting in short-term rental markets, where market-level occupancy paths can be viewed as survival curves over the booking horizon and where forecasts of future occupancy feed into dynamic pricing algorithms. We apply the models to the newly released Wheelhouse dataset, which contains time series of market occupancy curves for 500 markets from 2017 to 2022. Model performance is assessed using the integrated quadratic distance, and we compare the proposed PCA-based methods to Holt's linear trend model across multiple forecast horizons. The results show that the proposed models yield accurate and stable forecasts of the remaining survival trajectory and generally outperform Holt's method, particularly at longer horizons.

Subjects: Methodology , Applications

Publish: 2026-09-25 12:22:00 UTC


#18 Large and Moderate Deviations for Conservative Tail-Index Estimation [PDF] [Copy] [Kimi] [REL]

Authors: Martijn Gösgens, Bart P. G. van Parys, Bert Zwart

To design systems that are protected against events much rarer than the observational record, extreme-value methods are needed to extrapolate distribution tails. Tail-index estimators such as the Hill estimator are central to this extrapolation, but overestimating the tail exponent can lead to substantial underestimation of rare-event probabilities. Motivated by this, we derive large- and moderate-deviation asymptotics for the Hill estimator and use them to construct estimators whose probability of exceeding the true tail index decays at a controlled exponential rate (the decay rate). In the large-deviations regime, we show that a simple rescaled version of the Hill estimator achieves an optimal balance between bias and decay rate among scale-invariant estimators based on the same top $k$ order statistics. Under a second-order condition, we quantify the effect of the Hill bias, analyze a bias-corrected estimator, and identify sufficient conditions for moderate deviations in the boundary case where the second-order parameter $ρ$ equals zero.

Subjects: Statistics Theory , Probability

Publish: 2026-09-25 11:21:47 UTC


#19 Multiple change-point detection via bottom-up scanning [PDF] [Copy] [Kimi] [REL]

Authors: Jinhyeok Park, Hyeyoung Maeng, Hoseung Song

We study nonparametric multiple change-point detection for high-dimensional sequences, aiming to identify time points at which the underlying distribution changes. While many existing methods perform well when change-points are well-separated, their performance can deteriorate when structural breaks are densely clustered. To address this challenge, we propose gBottomup, a graph-based bottom-up framework for multiple change-point detection in high-dimensional settings. gBottomup constructs a hierarchical segmentation by proposing merges of adjacent segments and verifying them through an unmerge rule that combines absolute significance with relative local heterogeneity, thereby adaptively refining partitions and estimating the number of change-points. Simulation results demonstrate that gBottomup performs reliably across a range of structural configurations and is particularly effective in frequent change-point settings, where existing top-down procedures may lose sensitivity. Runtime experiments indicate favorable computational performance relative to graph-based top-down alternatives. We illustrate the proposed method through an analysis of a S&P 500 dataset.

Subject: Methodology

Publish: 2026-09-25 11:05:12 UTC


#20 EARL: Exposure- and Allocation-Reweighted Linear Estimator for Bipartite Experiments with Partial Assignment [PDF] [Copy] [Kimi] [REL]

Author: Alexey Kurennoy

Bipartite A/B tests are experiments in which treatment is randomised over one set of units, while outcomes are measured on another. For example, an online marketplace may test a new pricing algorithm on a random subset of items, while the outcome of interest, say, purchase satisfaction, is measured on customers, each of whom interacts with many items. Existing methods for analysing bipartite experiments assume that every randomisation unit is assigned to treatment or control. In practice, often only a subset participates: platforms cap rollout risk, reserve holdout groups, and split their population across concurrent tests. Ignoring the unassigned units biases estimation, while including them requires care. We construct an unbiased linear estimator for bipartite experiments with partial assignment. Observations must be reweighted not only by the (centred) share of treated connections among participating ones (the exposure) but also by the number of participating connections, each inverse-weighted by its participation propensity, so that units whose connections are well covered carry proportionally more weight. The resulting estimator, EARL (Exposure- and Allocation-Reweighted Linear), is unbiased, consistent, and asymptotically normal; we devise two asymptotic variance estimators and show that it has minimal variance in a natural class of linear estimators. EARL remains unbiased regardless of the experience the unassigned units receive, as long as they contribute to expected outcomes additively; in particular, they can be allocated to other, non-overlapping tests. Our theory is complemented with a simulation study on two public datasets, in which EARL attains up to six times lower error than the strongest existing baseline and, in some configurations, over an order of magnitude lower than Horvitz-Thompson-style alternatives.

Subject: Methodology

Publish: 2026-09-25 09:20:06 UTC


#21 Ensembles of Exactly Solved Subsamples for Clusterwise Regression: Trimming Without a Trimming Level [PDF1] [Copy] [Kimi] [REL]

Author: Samir Orujov

Clusterwise least squares partitions regression data into K groups with separate linear fits. We study an ensemble whose base learner is exact: solve the problem to global optimality on each of B random subsamples of size m << n, extend each solution by nearest-surface assignment, align the labels, and combine the replicates by vote or by selection. Each replicate is then an empirical K-quantizer on m points, and the ensemble admits an exact analysis. A vote is correct at a unit once the probability of a clean subsample times the clean-data replicate accuracy exceeds one half, whatever the contaminating values; iterating the ensemble on the units it has not flagged gives a variant that estimates the trimming level rather than requiring it. With up to 20% of gross outliers in the response its worst-case accuracy was 0.89, against 0.80 for trimmed alternation at the true contamination fraction and less at every fixed level tried. Conditionally on the data the replicates are i.i.d., so the vote converges exponentially fast in B to the plurality partition of the replicate law, agreeing with the criterion minimiser outside a boundary set whose size depends on m and the micro-solver, not on B. For the two-group location model, subsamples of order 1/pi_min make a replicate right more often than wrong above a separation threshold. An O(n^3) enumeration gives the exact minimiser for two groups and one covariate, against which the theory is checked. On clean data the ensemble loses to multistart alternation at equal cost.

Subjects: Computation , Methodology , Machine Learning

Publish: 2026-09-25 09:13:48 UTC


#22 Summary-powered prediction under distribution shift [PDF] [Copy] [Kimi] [REL]

Authors: Ivy Zhang, Dominik Rothenhaeusler

Prediction models can perform poorly when the deployment population differs from the training population. Data from the target population would help, but individual-level target data may be inaccessible because of access restrictions or reporting conventions. We consider a multi-resolution setting in which individual-level data are available from a source population, while the target population is observed only through subgroup summaries. We propose SAGE, a one-step estimator that updates a source-trained predictor using a gradient estimated from these summaries. Motivated by diagnostics consistent with the random distribution shift model, we choose SAGE's step-size to account for both sampling and distributional uncertainty. Under this model, SAGE reduces mean asymptotic target excess risk relative to the source-trained predictor. We also show that, under the model, entropy-balancing weighted empirical risk minimization (EB), which reweights source observations to match the target summaries, is asymptotically equivalent to a full-step SAGE update. SAGE with the optimal step-size has asymptotic mean squared error no larger than that of EB. Across real-world datasets, SAGE generally improves on the source-trained predictor and one-step updates that ignore distribution shift, including when the random shift model only partially captures the observed shifts. Compared to EB, SAGE improves prediction more consistently across the sample size and shift settings studied.

Subject: Methodology

Publish: 2026-09-25 07:14:18 UTC


#23 Conformal Prediction under Exponential-Tilt Joint Shift [PDF] [Copy] [Kimi] [REL]

Author: Seungjin Choi

Conformal prediction can lose coverage when the data distribution changes after deployment. We study adaptation using labeled source data and unlabeled target inputs, allowing both the input distribution and its relationship with outcomes to change. We use Exponential Tilt Reweighting Alignment (ExTRA), introduced for classification by Maity et al. (2023), to estimate structured distribution shifts. We compare using its estimated weights in conformal calibration with additionally tilting the source predictive distribution. Shared learned predictors, estimated weights, calibration samples, and test observations isolate the effect of tilting. Existing theory gives both procedures target coverage with true weights and a common coverage bound with estimated weights. Identification calculations and an analysis of how scoring interacts with weight estimation error help explain why their performance can nevertheless differ. In a synthetic regression setting where the assumed models match the data-generating process and target inputs are informative about the shift, tilting reduces mean set length by about $30\%$ relative to weighting alone, with both methods attaining coverage near nominal. Tilting can instead cause substantial coverage losses in synthetic classification and in regression when target inputs provide little information about the response shift. Real-data experiments also show no consistent benefit. Good coverage from weighted calibration alone does not ensure that adding predictive tilting will preserve coverage. Deciding when to apply this additional adjustment using only source labels and target inputs remains an open problem.

Subjects: Machine Learning , Machine Learning , Methodology

Publish: 2026-09-25 06:49:27 UTC


#24 Joint-Sparse Transfer Learning for High-Dimensional Multi-Output Regression [PDF] [Copy] [Kimi] [REL]

Authors: Sunwoo Lim, Mladen Kolar

Multitask linear models can improve estimation and prediction by exploiting structure shared across responses, bridging taskwise fitting and complete pooling. In many applications, however, the objective is estimation in a data-limited target domain, while data-rich but heterogeneous source domains are available. Borrowing from these sources can improve efficiency but introduce bias. We develop a joint-sparse transfer-learning framework for high-dimensional multi-output regression that combines shared predictor structure across responses with source-target similarity. The framework yields two complementary estimators: a fused estimator that aggregates jointly fitted domain-specific coefficients and a target-based debiased estimator that adjusts for source-induced shifts. Our error bounds show how transfer increases the available information and sharing predictors across responses reduces selection costs. They also reveal a tradeoff: the fused estimator benefits from larger sources but may retain bias if source shifts point in similar directions, whereas debiasing trades this bias for additional estimation error governed by the smaller target sample. Comparison with a minimax lower bound identifies regimes where the bounds match up to logarithmic factors, where matching remains unresolved, and where projection onto a target-based convex set closes the gap. Simulations and an analysis of single-cell RNA and surface-protein profiles across cell types support the theory.

Subject: Methodology

Publish: 2026-09-25 06:37:30 UTC


#25 Asymptotic Theory for Combining Dependent $p$-Values for Global Hypothesis Testing [PDF] [Copy] [Kimi] [REL]

Authors: Haoyi Yang, Lingzhou Xue

Combining $p$-values is a fundamental procedure in global hypothesis testing. In modern high-dimensional settings, however, component $p$-values often exhibit complex dependence and rely on asymptotic approximations rather than exact finite-sample uniform distributions. This paper establishes a unified asymptotic theory for weighted transformation statistics that decouples marginal finite-sample approximation error from the joint dependence structure. We also provide sufficient conditions based on conditional probability bounds to verify the joint-tail conditions. Utilizing this framework, we derive explicit dimension-growth and correlation rates for test statistics operating under asymptotic Gaussian and chi-square calibrations. For non-exact finite-sample statistics, we analyze standardized weighted sums, demonstrating how Cramér moderate deviations control relative tail error. Analytical examples demonstrate why both marginal and joint conditions are mathematically indispensable for valid global inference under dependence, and numerical experiments confirm that our asymptotic framework maintains accurate finite-sample size control at extreme significance levels.

Subject: Statistics Theory

Publish: 2026-09-25 03:15:29 UTC