Methodology

2026-07-15 | | Total: 16

#1 Factorial clinical trials in the presence and absence of plausible statistical interactions between treatment factors. A historical review of the methodological literature [PDF] [Copy] [Kimi] [REL]

Authors: Rebecca E. A. Walwyn, Pritaporn Kingkaew, Alan A. Montgomery, Alexandra Wright-Hughes, Steven G. Gilmour

Factorial trials can be conducted when statistical interactions between two or more treatment factors are not anticipated, but also when they are. A k-in-1 factorial trial answers k single-factor questions with the same number of units as one parallel-group trial if there are no interactions. Literature on k-in-1 factorial trials has dominated trialists understanding of factorial trials in the UK. However, factorial experiments originated from the Design of Experiments field to enable interactions to be robustly estimated. Seemingly conflicting guidance from these literatures poses a source of confusion and misunderstanding for trialists. We bring these literatures together to provide clarity on the arguments that have been used to recommend use of factorial trials in the presence and absence of interactions. We outline motivating examples. We summarise the rationales for using factorial trials, the treatment contrasts of interest, and the properties of their estimators, for the two schools of thought. We describe the debate, going back to 1935, and use an empirical example to illustrate the impact of different analysis approaches. We conclude that it is vital that trialists carefully and clearly specify their objectives. Estimands of interest and treatment contrasts follow, with properties of estimators dictated by this choice.

Subject: Methodology

Publish: 2026-07-14 17:42:28 UTC


#2 Higher-order Spillover Effects Under Partial Interference [PDF] [Copy] [Kimi] [REL]

Authors: Qixiang Xu, Laura Forastiere

Interference, under which a unit's outcome is affected by the treatment of other units through network connections, is often present when units interact on a network. When the network of interactions is measured, researchers are often interested in the spillover effect from first-order neighbors. When this is the case, the prevailing approach often involves the neighborhood interference assumption, which is oftentimes overly restrictive. In this paper, we instead rely on a generalized interference assumption, which allows one's potential outcomes to be influenced by the treatment of units from a wider area of the network, referred to as the "interference set". For instance, this can be a community detected through a community detection algorithm, or the set of units that can be reached through a finite network path. Under this assumption, we define new causal estimands to quantify spillover effects from first-order neighbors and, in general, from units at a specific network distance h. We employ two hypothetical Bernoulli distributions with different probabilities for the h-order neighborhood and for the rest of the units in the interference set. We first derive the bias of an approach that relies on a wrong interference set or incorrect exposure mapping function. We then develop new Horvitz-Thompson and Hajek estimators and corresponding weighted regression estimators under the generalized interference assumption. We conduct a series of simulations to assess the bias of OLS estimators -- which rely on restrictive interference assumptions and an exposure mapping function -- , and the performance of our estimators in different interference scenarios and random graphs. We then apply our estimators to a two-stage randomized trial implemented in Honduras to assess a maternal and child health intervention.

Subject: Methodology

Publish: 2026-07-14 15:07:52 UTC


#3 Hash-augmented adaptive multilevel splitting Monte Carlo algorithm for accurate estimation of two-sample permutation test p-values [PDF] [Copy] [Kimi] [REL]

Authors: Nikita Golikov, Vladimir Sukhov, Gennady Korotkevich, Alexey Sergushichev

Nonparametric permutation tests are widely used for statistical analysis. However, exact computation of test p-values can be algorithmically challenging, particularly for custom tests with complex test statistics. In contrast, Monte Carlo sampling can be easily applied to any test statistic, but it suffers from poor relative accuracy when estimating small p-values, interfering with multiple hypothesis testing correction and leading to other issues. In this work, we present a hash-augmented adaptive multilevel splitting Monte Carlo algorithm that enables accurate estimation of arbitrarily small p-values in two-sample permutation tests. Using the Kolmogorov-Smirnov and the Mann-Whitney U tests as examples, we highlight potential pitfalls related to the discreteness of the test statistic distribution and show how to address them. By comparing with an exact algorithm, we demonstrate the accuracy of the p-value estimates provided by the proposed algorithm and the validity of the associated confidence intervals. We provide a reference implementation of the proposed algorithm in the Python package hamstest, which allows p-value estimation for a user-defined statistic.

Subjects: Methodology , Statistics Theory , Computation

Publish: 2026-07-14 15:06:12 UTC


#4 Accounting for Preferential Sampling Using a Constructed Covariate [PDF] [Copy] [Kimi] [REL]

Authors: Andreia Monteiro, Isabel Natário, Ivone Figueiredo, Paula Simões

In geostatistics, it is commonly assumed that sampling locations are selected independently of the underlying spatial process. In practice, however, this assumption is frequently violated. In fisheries, for example, sampling sites are often chosen to maximize expected catches, creating a stochastic dependence between the abundance process and the sampling design. Such preferential sampling can introduce substantial bias and compromise statistical inference. This study investigates the use of constructed covariates, based on average distances from nearest neighbours observations, that are able to mitigate preferential sampling. The inclusion of such covariate in the geostatistical model might be able to account for the stochastic dependence of sampling locations on the spatial variable. If this inclusion sufficiently captures the dependence, conventional methods of inference may be applied without resorting to more complex models. The proposed methodology is evaluated through an extensive simulation study that explores a variety of sampling scenarios and spatial configurations. Additionally, we demonstrate the practical utility of the approach using two real-world datasets: one on fishery landings provided by the Instituto Português do Mar e da Atmosfera, and another concerning lead pollution biomonitoring in Galicia. Results show that incorporating the constructed covariate can substantially reduce the impact of preferential sampling, enabling reliable inference with standard geostatistical tools. We also discuss practical challenges, limitations, and paths for future methodological development.

Subject: Methodology

Publish: 2026-07-14 14:22:34 UTC


#5 Modification and extension of the Bayesian clinical trial design using external data for single-arm and hybrid-controlled trials [PDF] [Copy] [Kimi] [REL]

Authors: Wataru Murasaki, Tomohiro Ohigashi, Ryota Ishii, Kazushi Maruo, Masahiko Gosho

Limited patient availability complicates sample size determination in pediatric clinical trials. Although Bayesian methods incorporating external data offer a solution, rigorously controlling the type I error rate remains difficult. Psioda and Ibrahim (2019) proposed a simulation-based framework as a practical solution. However, although their framework was designed to relax the type I error control, this relaxation fails when the external data exhibit a large treatment effect, making it difficult to design clinical trials that incorporate external data. Furthermore, restricting the support of sampling priors can cause trial outcomes to fall outside of this support, leading to lower power. Additionally, their analytic prior formulation may induce bias, and their method is not applicable to hybrid-controlled trials involving two-group comparisons. Thus, we propose modifications to both the sampling and analytic prior specifications and extend the framework to hybrid-controlled trials. We redefine the null sampling prior as a normal distribution centered at the null boundary, ensuring a Bayesian type I error evaluation. For the analytic prior, we employ a weakly informative prior for the second component of a robust mixture prior to mitigate bias under prior-data conflict. Furthermore, we extend this methodology to hybrid-controlled trials. Simulation studies and a pediatric case study of cutaneous lupus erythematosus demonstrate that our method substantially reduces the required sample size compared with both frequentist and original Bayesian methods, while maintaining the target operating characteristics and controlling estimation bias under prior-data conflict. This framework provides a reliable and efficient approach for designing clinical trials that incorporate external information.

Subject: Methodology

Publish: 2026-07-14 08:54:25 UTC


#6 Symbolic Weak-form Recovery of 2-D Stochastic Generators [PDF] [Copy] [Kimi] [REL]

Authors: Sai Sathvik Gullipalli, Eshwar R A

Recovering two-dimensional Ito generators from trajectory data is difficult because drift increments have low signal-to-noise, bivariate weak designs can be ill-conditioned, and unconstrained tensor estimates need not be positive semidefinite. We study WG-SINDy estimator combining covariance-shaped spatial kernels, a ridge-stabilized local-polynomial projection, adaptive-LASSO/STLSQ selection, one in-sample per-component feasible diagonal GLS pass, and a PSD projection--Cholesky read-out with mild isotropic shrinkage. The released estimator uses a data-dependent full-cloud smoother and one in-sample per-component feasible diagonal GLS pass; accordingly, we do not claim exact finite-sample martingale cancellation or a feasible-GLS efficiency theorem for the reported implementation. We evaluate the estimator on 29 synthetic two-dimensional systems: 19 meet their declared per-system recovery contracts, eight are retained as named limits, and two remain scoped reviews. Across the 19 PASS rows, the median central-grid drift metric is 0.204 and the median tensor error is 0.0397. Among the six systems with a finite, non-degenerate off-diagonal target, the median $a_{12}$ cosine is 0.997. Positive-semidefinite validity is imposed by construction. These results are synthetic, in-sample sampled-region diagnostics and do not establish universal or real-data recovery.

Subjects: Methodology , Computational Engineering, Finance, and Science , Probability , Data Analysis, Statistics and Probability

Publish: 2026-07-14 08:33:53 UTC


#7 Does a Developed Comorbidity Index Really Add Value? A Selection-Aware Bootstrap for Post-Selection Concordance [PDF] [Copy] [Kimi] [REL]

Author: M. Ehsan Karim

Disease-specific comorbidity indices are routinely developed by building several candidate constructions and reporting the best-scoring one, then claiming it adds discriminative value over a fixed off-the-shelf comparator such as the Charlson or Elixhauser score. We show that the optimism correction in standard use does not make that claim valid. Because it corrects the selected model as if it were the only one ever fit, it omits the winner's-curse term from choosing the best of several candidates; so its confidence interval for the incremental concordance is not merely optimistic in small samples but structurally miscalibrated, and does not shrink as the sample grows. At a true null it inflates false claims of added value above the nominal level, increasingly so as more candidates are screened. We introduce a drop-in selection-aware bootstrap that re-runs the best-of-several selection inside each resample with the comparator held fixed, removing the structural bias. In a fully-known-truth simulation, 95% coverage under the standard correction falls from 0.94 with one candidate to 0.70 with a hundred, while the selection-aware interval holds near nominal; its coverage matches a calibrated cross-validation interval, and at a matched error rate it is at least as powerful. The results hold under Uno's concordance, and a semi-synthetic experiment on real survey data confirms when the correction matters. In practice, if several constructions were tried, report a selection-aware interval, most needed with many similar-quality candidates and few events per candidate. The scope is discrimination only; software and results reproduce every finding.

Subjects: Methodology , Applications

Publish: 2026-07-14 07:23:02 UTC


#8 Hierarchical Bayesian inversion using the Karhunen-Loève expansion with analytical eigenpairs of the squared exponential kernel [PDF] [Copy] [Kimi] [REL]

Authors: Tatsuya Shibata, Michael Conrad Koch, Kazunori Fujisawa

Hierarchical Bayesian inversion with Gaussian random field priors addresses uncertainty in covariance hyperparameters, such as the standard deviation and correlation length. When a Gaussian random field is represented by the Karhunen-Loève (KL) expansion, the basis functions depend on these hyperparameters through an integral eigenvalue problem (IEVP) associated with the covariance kernel. Consequently, the IEVP must be solved repeatedly whenever the hyperparameters are updated, leading to significant computational cost in hierarchical inference. In this paper, we focus on the squared exponential kernel and construct the KL expansion using the analytical solution to a Gaussian-weighted IEVP. This analytical KL expansion offers a computationally efficient alternative to the conventional KL expansion by eliminating the repeated numerical solutions of the IEVP during hyperparameter updates. While the analytical KL expansion is applicable to arbitrary domains and dimensions, it does not have the same mean-square optimality as the conventional KL expansion. To address this limitation, we employ an optimization-based approach that selects the standard deviation of the Gaussian weight function in the IEVP to effectively reduce the truncation error of the KL expansion. Numerical experiments in one- and two-dimensional settings show that this selection strategy provides sufficient accuracy for practical applications. Furthermore, the analytical KL expansion admits closed-form differentiation, enabling efficient posterior sampling via HMC. The proposed framework is applied to Bayesian inversion for a steady Darcy flow model, where the hydraulic conductivity field is successfully estimated using weakly informative hyperpriors.

Subjects: Methodology , Numerical Analysis

Publish: 2026-07-14 06:00:27 UTC


#9 Nonparametric Multi Change Point Detection for Markov Chains via Adaptive Clustering [PDF] [Copy] [Kimi] [REL]

Authors: Imon Banerjee, Jiaqi Lei, Sanjay Mehrotra

Offline change point detection tries to detect time points of distribution change in a given data sequence; and is now routinely used in signal processing, speech processing, climatology etc. Despite this broad applicability across economics, computer science, and planetary sciences, rigorous, nonparametric techniques for change point detection with non-independent and identically distributed (i.i.d.) datasets has remained elusive. This paper establishes such guarantees by proposing a non-parametric clustering algorithm which can accurately obtain the change points from a given Markovian dataset of length $n$. It does so by bridging together two different components of mathematical statistics; Rademacher complexities of Markov chains, and adaptive clustering via penalisation. Our first result uses recent advances in Rademacher complexities of regenerating Markov chains to derive a Dvoretzky Kiefer Wolfowitz (DKW) type inequality for the empirical distribution of the Markov chain. We then use this to show that an adaptive clustering algorithm recovers the correct change points for a Markovian sequence. We establish the tightness of our rates by showing that they essentially coincide with the best known rates for i.i.d. data. We end the paper by discussing the computational considerations of the problem.

Subjects: Methodology , Statistics Theory

Publish: 2026-07-14 05:41:59 UTC


#10 Fitting Generalized Linear Mixed-Effects Models using lme4 [PDF] [Copy] [Kimi] [REL]

Authors: Anna Ly, Rune Haubo Bojesen Christensen, Douglas Bates, Martin Maechler, Benjamin M. Bolker

The lme4 R package can be used to fit generalized linear mixed models (GLMMs), which extend the class of linear mixed models (LMMs). The two main extensions provided by GLMMs are (1) allowing for the conditional distribution of the response given the random effects to be non-Gaussian (e.g. binomial, Poisson) and (2) allowing the conditional mean to be a nonlinear function of a linear combination of the fixed and random effect coefficients, via an inverse link function. The conditional mode of the random effects given the observed data, the variance-covariance matrix of the random effects, and the fixed effect parameters are determined using penalized iteratively reweighted least squares. We compute an approximation of the integral over the distributions of the conditional modes to compute the maximum likelihood estimate for a given set of parameters (by default we use the Laplace approximation or, alternatively, the more computationally expensive adaptive Gauss-Hermite quadrature). The package provides all the standard features available for GLMs in base R, including the standard set of accessor functions as well as the possibility of user-specified distributions (within the exponential dispersion family) and link functions.

Subjects: Methodology , Computation

Publish: 2026-07-13 22:04:25 UTC


#11 Extracting Bayesian Evidence from Frequentist p-Values [PDF1] [Copy] [Kimi] [REL]

Authors: Frederik Aust, Samuel Pawel, Eric-Jan Wagenmakers

The $p$-value and the Bayes factor are measures of evidence that are often considered to be philosophically and mathematically incompatible: The $p$-value quantifies conflict between data and $H_0$ ("surprise"), whereas the Bayes factor quantifies the relative predictive accuracy of $H_0$ versus $H_1$ ("evidence"). We revisit Jeffreys's Approximate Bayes factor (JAB) -- a simple, largely overlooked approximation dating back to the 1930s -- which connects these two paradigms for objective hypothesis testing of the existence of an effect. Under a unit-information prior the approximation requires only the $p$-value and the effective sample size $n_\text{eff}$. We clarify the core assumptions and boundary conditions for the application of JAB and show across 704 published $t$-tests and 39 comparisons of proportions that JAB approximates objective Bayes factors remarkably well. The connection between $p$-values and JAB has a practical implication: The evidence implied by a $p$-value depends strongly on $n_\text{eff}$. Conventional verbal labels for $p$-values (e.g., "strong surprise" for .001 < $p$ < .01) correspond to similarly graded Bayes factors only around $n_\text{eff} \approx 8$; for larger samples the same $p$-value implies weaker evidence. In moderately sized to large samples, $p > .10$ can amount to moderate or even strong evidence for $H_0$. JAB offers a cheap, sample-size-sensitive supplement to $p$-values, computable from routinely reported statistics, that remains valid even under optional stopping.

Subjects: Methodology , Statistics Theory

Publish: 2026-07-13 20:30:51 UTC


#12 LatentFlow: A General Framework for Conditioning Stochastic Processes [PDF] [Copy] [Kimi] [REL]

Authors: Louis Sharrock, Lachlan Astfalck, Henry Moss

Stochastic-process models are, as a rule, far easier to simulate than to condition. Non-linear observations, non-Gaussian likelihoods, black-box information, and global constraints all induce intractable conditional laws, requiring bespoke, model-specific constructions. We introduce LatentFlow, a single framework for conditioning stochastic processes, with no learned neural approximations and no training. Our starting point is to write the stochastic process as the deterministic image of a tractable latent innovation, $f_0 = T_{\vartheta}(ξ_0)$, with $ξ_0$ sampled from a simple reference distribution. This reduces process-level conditioning to latent-space inference: pull the likelihood back through $T_{\vartheta}$, sample the resulting latent law with a tractable guided probability flow, and push the samples forward. This construction is provably exact at the level of the target law; in practice, approximation enters only through finite terminal noising, Monte Carlo guidance, and time discretisation of the continuous-time dynamics, each of which is explicit and systematically reducible. As LatentFlow is training-free, conditioning reduces to solving a single reverse-time SDE. This enables conditional sampling in seconds on a single desktop CPU across model classes that have never shared a scalable method: classical spatial priors, nonlinear stochastic dynamics, mechanistic models from the physical and life sciences, stochastic PDEs, heavy-tails and extremes, point and discrete-state processes, and neural or simulator-defined processes.

Subjects: Machine Learning , Machine Learning , Methodology

Publish: 2026-07-14 15:56:44 UTC


#13 Contrast-Free ICA and Causal Inference via Wasserstein Distances to the Gaussian [PDF] [Copy] [Kimi] [REL]

Authors: Félix Laplante, Christophe Ambroise, Pierre Humbert

We study the squared $2$-Wasserstein distance to the standard Gaussian as a non-Gaussianity criterion and use it for linear Independent Component Analysis (ICA) and causal inference in Linear Non-Gaussian Acyclic Models (LiNGAM). The analysis relies on a strict inequality between the Wasserstein non-Gaussianity of independent standardized sources and that of their linear combinations. When at most one source is Gaussian, any unit-norm linear combination involving at least two sources has strictly smaller squared Wasserstein distance than the corresponding weighted sum of source distances. At the population level, this yields exact identification of the ICA unmixing matrix, up to signed permutation, and gives an analogous characterization of causal orders through least-squares residuals. We then define empirical plug-in estimators and prove distribution-free uniform convergence bounds under finite-moment assumptions, before detailing three practical solvers: a Picard-style orthogonal optimizer for ICA, an exhaustive dynamic program for causal-order search, and a greedy order-search variant. Empirically, we demonstrate competitive performance for both tasks and provide open-source implementations for source separation and causal inference.

Subjects: Machine Learning , Methodology

Publish: 2026-07-14 14:49:31 UTC


#14 Forecasting Inflation with Microdata: An Adaptive Machine Learning Approach [PDF] [Copy] [Kimi] [REL]

Authors: Catherine Chen, Chen Gao, Jonathon Hazell, Lihua Lei, Chen Lian

Does microeconomic heterogeneity help to forecast aggregate inflation in a non-stationary environment? We develop a scan test for whether one forecast outperforms another, over an interval with unknown starting point and duration. To exploit any occasional forecasting power that the scan test detects, we design an adaptive machine learning pipeline. We encode the distribution of price changes into a high-dimensional vector, which we combine with a gradient boosted trees algorithm. We then combine this micro forecast with other benchmark forecasts, using an adaptive algorithm that makes use of the micro forecast only when it performs well. We apply the pipeline to UK microdata, with four main results. First, the micro forecast outperforms a univariate benchmark, but only in the volatile period after 2020. Second, the scan test detects periods of micro outperformance, so the micro forecast enters the combined forecast. Third, the combined forecast performs comparably to the univariate benchmark before 2020 and better at every horizon after 2020. Fourth, the value of microdata for the combined forecast materializes after 2020. We conclude that microdata are valuable for forecasting aggregate inflation, but only after large shocks.

Subjects: General Economics , Methodology , Machine Learning

Publish: 2026-07-14 04:46:43 UTC


#15 The Benjamini--Hochberg Procedure Can Fail to Control the FDR for Correlated Two-Sided Gaussian Tests [PDF] [Copy] [Kimi] [REL]

Author: Edgar Dobriban

We show that the Benjamini--Hochberg procedure can fail to control the false discovery rate (FDR) at its nominal level for correlated two-sided Gaussian $p$-values. We construct a factor model for which, at level $α=0.01$, a rigorous interval-arithmetic certificate proves $FDR>0.0104$ for all sufficiently large numbers of hypotheses. This disproves a conjecture widely believed to be true for twenty years. Monte Carlo experiments are consistent with the theoretical result. The proof was obtained by GPT-5.6 Pro and carefully checked by the author.

Subjects: Statistics Theory , Artificial Intelligence , Methodology

Publish: 2026-07-13 23:18:13 UTC


#16 Sensitivity to Subjective Expected Utility Maximization: A Methodological Study, with an Illustrative Application to LLM Decision-Making [PDF] [Copy] [Kimi] [REL]

Author: Jeff Helzner

Evaluating decisions made under uncertainty is hard when labeled outcomes are scarce, costly, or confounded with luck. We treat subjective expected utility (SEU) maximization as a stated standard and define a graded measure -- SEU sensitivity -- of an agent's conformity to it. The vehicle is a softmax choice model with a sensitivity parameter $α$ on SEU-valued alternatives; the contribution is a sequence of identifiability results for $α$ and for belief and utility parameters $(β, δ)$, validated in Stan via prior predictive checks, parameter recovery, and simulation-based calibration (SBC), with finite-sample caveats intact. In the uncertain-choice-only model $m_0$, $α$ is identifiable given the expected-utility vector $η$ and sharply recovered, while $(β, δ)$ are only weakly informed: the posterior barely contracts and concentrates on a $β$-$δ$ trade-off. In the extended model $m_1$, $δ$ becomes identifiable in principle via a $β$-free risky block, but its practical recovery gain at realistic sample sizes is negligible (matched-count CI-width reduction under 1%), and that block yields no detected $α$-precision gain at matched choice count. These are two distinct phenomena: for $δ$, identifiability does not imply precise estimability at realistic $n$; for $α$, identifiability is silent about what governs finite-$n$ precision. Marginal SBC passes for both models even where the joint posterior is weakly informed -- a demarcation we make precise. A two-by-two application (GPT-4o and Claude 3.5 Sonnet, each on insurance-claims triage and Ellsberg-style urns, with sampling temperature as the lever) runs end-to-end on real LLM choice data, detecting a structured comparative $α$ effect in two of four cells.

Subjects: Econometrics , Artificial Intelligence , Methodology

Publish: 2026-07-08 00:29:18 UTC