2026-07-28 | | Total: 115
Recent advances in spatial transcriptomics have enabled researchers to profile gene expression at the single-cell spatial resolution, often for multiple tissue samples in a single study. This high-dimensional molecular profile for each cell can be used to sort cells into cell types with distinct functions, or segment the tissue into biologically relevant spatial domains. Although many non-spatial and spatial clustering methods have been developed to cluster these cells into cell types or spatial domains, most have two main limitations: first, they perform dimension reduction and clustering separately; second, they cluster cells at a single scale, rather than treating cell type and spatial domain clustering as distinct tasks at two different scales. To overcome these limitations, we propose BayesClint, a Bayesian method that simultaneously performs factor analysis and spatial clustering on multiple samples, where the clustering is done jointly at the single-cell and tissue regional scale. To increase interpretability, we employ a feature selection mechanism within the estimation of the sparse factor loadings matrix, which detects active genes and differentially expressed genes that discriminate between cell type clusters. We illustrate the advantages of the method over alternative state-of-the-art approaches through simulation studies and two real data applications.
Early inquiries into the statistical properties of natural language found that words tend to occur with power-law frequencies. This observation, closely associated with Zipf's law, has spurred many investigations into why this power-law pattern emerges with such regularity. Rarely, however, has this property of text been leveraged in statistical inference. In this paper, we demonstrate that the Latent Dirichlet Allocation (LDA) model can accommodate power-law word frequencies. In particular, when the document length distribution is regularly varying, the word frequency distribution admits a hierarchy of power laws across documents and topics. We further leverage this finding to develop an efficient tensor decomposition algorithm for estimating the topic matrix via the moments of normalized extreme word frequencies. Applying our algorithm to the twenty newsgroups corpus reveals that the extreme-value methodology exhibits robustness to certain choices made in the pre-processing of the data. This work furthers the recent interest in adapting machine learning methods to the study of multivariate extremes.
We study imbalanced crowdsourcing with a focus on class-dependent annotator accuracy, a setting that, to the best of our knowledge, remains relatively underexplored despite its importance in real-world inspection systems where the labels of greatest operational importance are also the rarest ones. In this setting, annotators may be reliable on both classes, unreliable on both classes, majority-class specialists, or minority-class specialists. Existing models only partially address this problem: they either capture class-dependent errors but ignore item difficulty, or they model item difficulty without capturing class-dependent errors. To fill this gap for imbalanced datasets in crowdsourcing, we introduce a generative aggregation model combining item difficulty with class-dependent annotator competence. The model allows both annotator abilities and item difficulties to vary across classes. We then revisit Condorcet's Jury Theorem in the class-imbalanced setting. We also show that majority voting asymptotically preserves the underlying class proportion. We evaluate our model on $33$ real-world crowdsourcing datasets, covering multiclass tasks such as images and text, as well as two large-scale regimes: large-scale annotation datasets, with many annotations per item, and large-scale item datasets, with a large number of annotated instances. Across these diverse settings, our model consistently achieves the highest minority recall while remaining competitive in balanced accuracy, making it particularly relevant when rare-label recovery is the primary objective.
This paper proposes a stage independent multiple sampling plan (SIMSP) to improve inspection efficiency by reducing the number of samples required at each sampling stage. Unlike conventional multiple sampling plans (MSP), the proposed SIMSP eliminates the dependence of each sampling stage on the outcome of the preceding inspection. The proposed approach is developed for non-repairable products sold under a pro-rata warranty policy based on the generalized process capability index $C_{py}$. The SIMSP is designed under a Type-II hybrid censoring scheme (Type-II HCS), which provides greater flexibility in controlling test time and failure information in life testing experiments. The asymptotic distribution of the process capability index estimate is used to compute the operating characteristic (OC) function, and the exact Fisher information matrix (FIM) is obtained for further statistical analysis. A constraint optimization problem is formulated to determine the optimal design by minimizing the total cost subject to the manufacturer's and consumer's risk. Numerical investigations are conducted to examine the effects of model parameters, warranty policy characteristics, and cost factors on the optimal solution. The results demonstrate that the proposed approach provides an economically efficient and feasible method for lot acceptance while satisfying the tolerable risks requirements.
In linear mixed effects models (LMM), inference on fixed effects typically relies on either the generalized least squares (GLS) variance estimator obtained by assuming that the variance parameters are known, or on the Kenward-Roger method with a small-sample bias adjustment. A simulation study shows that both approaches can underestimate the variance of the treatment effect and inflate the type I error rate in large clinical trials under missing at random (MAR) mechanisms, with greater type I error inflation as the differences in the missingness mechanisms between treatment groups increase. Counterintuitively, the estimated fixed effects and variance parameters exhibit asymptotic dependence. Ignoring this dependence may lead to invalid inference. To account for this additional source of variability, we propose a new asymptotic variance estimator and a small-sample bias-corrected estimator to quantify uncertainty in estimation and prediction in LMM under MAR. The proposed methods are theoretically justified, and their performance is demonstrated by numerical examples.
Transportation network partitioning is essential for applications such as traffic analysis, simulation, and mobility pattern identification. However, transportation networks combine structural information with heterogeneous operational attributes, ranging from scalar indicators to temporal profiles. Existing approaches generally rely on predefined formulations to integrate these sources of information, limiting the ability to control their relative influence. This paper proposes a flexible framework for partitioning heterogeneous transportation networks represented as attributed graphs. The proposed methodology relies on a distance-based graph representation and an optimal transport formulation based on the semi-relaxed Fused Gromov-Wasserstein discrepancy, enabling joint consideration of network structure and attributes with explicit control over their trade-off. The proposed methodology is evaluated on two transportation systems with distinct characteristics: an urban road network for traffic-oriented partitioning and a bicycle-sharing system for identifying usage-based communities. Results demonstrate the ability of the framework to adapt the resulting partitions according to different structural and attribute preferences.
We study the task of estimating a $d-$dimensional quantum state $ρ$ under the constraint that the measurement is $α-$gentle. Such measurements $M$ do not collapse the state; they issue both a random variable $R^M = ω$ containing statistical information and a post-measurement state $ρ_{M \to ω}$ such that $\|ρ_{M\to ω} - ρ\|_{Tr} \leq α$. We describe gentle measurements and their connection to quantum differential privacy. Our results show that the optimal minimax estimation rate in Frobenius norm is of order $d^3/(n α^2)$, instead of $d^2/n$ for general measurements. Moreover, for rank $r$ states with $r\leq d$ we prove that the optimal minimax rate is $rd^2/(n α^2)$, instead of $rd/n$. Very surprisingly, the loss for gentleness $d/α^2$ scales with the ambient dimension of the Hilbert space, rather than the number of parameters $rd$, typically seen in classical differential privacy. We propose optimal gentle measurements and indicate how they can be physically implemented using an ancillary state and a CNOT gate to entangle it with the initial state. We notice that the resulting random variable has a likelihood that satisfies local differential privacy. Lower bounds are proven through a new quantum information-theoretic inequality applied to well chosen families of states in the manifold of (small-rank) quantum states.
Reservoir computing has emerged as an efficient machine learning framework for predicting time series generated by dynamical systems. In contrast to other machine and deep learning approaches, a reservoir computing trains only the output layer via linear regression, leaving the reservoir (recurrent layer) untrained. This simplification makes reservoir computers easier to train and more amenable to experimentation. However, because current reservoirs consist of networks of randomly connected nodes and require the optimization of numerous hyperparameters, a framework that precisely explains how reservoir computing operates and how it can be optimized remains missing. Here, we propose a frequency-based reservoir inspired by the brain's oscillatory dynamics and its hierarchy of timescales. The frequency-based reservoir can be interpreted as an ensemble of independent oscillatory units, each processing a portion of the input's frequency content. This allows us to understand the reservoir's internal behavior by modeling it as a single unit driven by an external input. Borrowing from the theory of a nonlinear oscillator forced by complex periodic inputs, we found that units of the frequency-based reservoir selectively amplify and store specific input frequencies, which are then used for prediction. The frequency-based reservoir performs as well as or better than equivalent random reservoirs. Furthermore, the frequency-based approach can be optimized to improve short-term prediction, a property that random reservoirs lack. Finally, we show that the frequency-based reservoir can also predict complex spatiotemporal dynamics. Our results show that reservoir computing can be designed using brain properties and theoretical insights borrowed from the physics of forced nonlinear oscillators.
Proxy outcomes (such as short-term behavioral signals, model predictions, or surrogate endpoints) are frequently used in place of primary outcomes that are too slow to mature, rare, or challenging to measure directly. But valid inference on a proxy does not guarantee valid inference on the primary estimate as proxy-based estimates can be systematically biased in ways that are difficult to predict, leading to improperly calibrated confidence intervals. We present proxymate, a framework and open-source Python package for proxy validation and adjustment. proxymate organizes into four levels: The Representativity Level (population validity), the Unit Level (measurement quality), the Estimate Level (decision validity), and the Domain Level (cross-domain transportability). Within each level, proxymate provides diagnostic checks, and targeted adjustment strategies that map specific failures to appropriate corrections. At Meta, proxymate has been adopted by many different use cases, spanning experimentation, prevalence estimation, and monitoring use cases, all facing different proxy challenges (limited human review time, long maturation window of outcomes, low detectability) and showcasing the modularity of the framework. Across all products, proxymate assessed and corrected millions of proxy, primary unit comparisons. It has facilitated launches across multiple work streams including enabling quick decision making on thousands of experiments.
We introduce a new notion of positive dependence, namely multivariate total positivity of order two for conditional distributions, denoted cMTP$_2$, that explicitly incorporates the conditioning variables. This property is weaker than multivariate total positivity of order two for densities, but stronger than multivariate total positivity of order two for distribution functions, stochastic monotonicity, and tail monotonicity. It thus provides a natural intermediate concept linking these classical notions of positive dependence. We verify the cMTP$_2$ property for several distributions and copula families, including Archimedean copulas, and establish stability under Markov product transformations. As a consequence, we derive comparison results for measures of directed dependence arising from Markov structures, including Chatterjee's rank correlation and a related sensitivity measure.
We develop kernel-based specification tests for semiparametric conditional moment models with high-dimensional nuisance parameters, extending existing conditional moment tests---which typically require asymptotically linear nuisance estimators---to accommodate modern machine-learning methods. The proposed locally robust kernel tests combine Neyman-orthogonal moments, cross-fitting, and reproducing kernel Hilbert space methods, yielding inference that is first-order insensitive to nuisance estimation error. We establish oracle equivalence between the feasible and infeasible test processes under local alternatives and weak nuisance-rate conditions, and characterize the resulting local power. A fast multiplier bootstrap avoids nuisance re-estimation. Applications include specification testing in high-dimensional linear and logistic regression, significance testing with machine-learning regressions, and tests of constant conditional treatment effects. Monte Carlo simulations and an application to the National Supported Work program illustrate the finite-sample performance of the proposed tests.
High-dimensional data with sparse structure and spatio-temporal dependence arise in many scientific domains. We develop a Bayesian feature-extraction framework for spatio-temporal settings that employs Gaussian and Diffused-gamma priors to induce structured sparsity. The modeling framework specifies a general likelihood via Bregman divergence, enabling compatibility with a range of loss functions and measurement models. Posterior computation is carried out via Markov chain Monte Carlo (MCMC), and we introduce a two-stage feature-extraction procedure based on posterior samples to stabilize selection across space and time. We illustrate the method with a multi-subject electroencephalography (EEG) case study examining the relationship between chronic alcohol exposure and activity in different brain regions. We first fit binary classification models at each time point, then use false discovery rate-controlled screening and subsequent clustering in a two-stage feature-extraction pipeline to identify active brain regions. The analysis demonstrates that our proposed priors improve recovery of sparse features and enhance interpretability in the presence of spatio-temporal dependence. The framework is broadly applicable to high-dimensional, structured problems where accurate feature selection and inference are required. The code to implement the model is publicly available via GitHub.
We introduce a class of goodness-of-fit tests for the spherical Cauchy model on the unit hypersphere. The proposed procedures exploit the invariance of the spherical Cauchy family under Möbius transformations: after estimating the parameter by the sample Möbius mean, the observations are transformed to approximate spherical uniformity, and a projection-based uniformity statistic is applied to the resulting sample. We show that, under mild conditions, the resulting tests are exactly distribution-free under the null hypothesis, so that exact critical values can be arbitrarily well approximated by Monte Carlo simulation. We study the Möbius mean as a population functional, establish its existence and uniqueness under mild conditions, and prove equivariance and asymptotic linearity of its empirical counterpart, which coincides with the spherical Cauchy maximum likelihood estimator. We derive the asymptotic null distribution of the test statistic and show that it coincides with that of the underlying uniformity statistic after removing the degree-one spherical-harmonic component, which corresponds to the tangent space of the spherical Cauchy model. We establish consistency against fixed alternatives and characterize local powers through the spherical-harmonic decomposition of contiguous alternatives. Monte Carlo experiments demonstrate the finite-sample accuracy of the asymptotic approximations and the empirical power of the proposed tests. A real data example is treated.
Self-reported numerical variables in sample surveys are frequently subject to coarsening, as respondents tend to report rounded or heaped values rather than the exact underlying quantity. While the statistical literature offers various model-based solutions, official statistics and public health surveillance routinely require design-based inference on finite-population indicators under complex sampling designs. This paper introduces a general framework that bridges this gap by treating the observed response as a coarsened manifestation of a latent variable, modeling reporting regimes of increasing coarseness jointly with the latent distribution via a survey-weighted pseudo-likelihood approach. The fitted model is used to generate posterior predictive replicates of the latent values, to which standard design-based estimators are applied, formally propagating the additional uncertainty through an explicit variance decomposition. Simulation studies assess the finite-sample performance and robustness of the proposed method under various misspecification and sampling scenarios. Finally, its practical relevance is demonstrated through an application to behavioural data from the Italian PASSI surveillance system, illustrating how coarsening affects different functionals, such as means, quantiles, and threshold-based prevalence indicators in heterogeneous ways.
Classical bootstrap methods can behave poorly in high-dimensional linear models: the pairs bootstrap often yields overly conservative inference, whereas the residual bootstrap can be anti-conservative, reflecting systematic failures in variance calibration. We propose a diffusion-based pairs bootstrap that replaces the empirical joint distribution with a learned generative law. We establish variance consistency under a score approximation assumption, using complementary SDE and PDE arguments. Counterexamples show that terminal $W_4$ convergence alone is insufficient for variance consistency. Experiments indicate that diffusion pairs bootstrap improves variance calibration and generally improves Type~I error calibration, including in settings not covered by our theory.
Classical early-warning signals assume that recovery slows near a tipping point, but in a well-studied CESM simulation of AMOC collapse, autocorrelation and variance decrease instead. We develop a preregistered diagnostic that uses only past data to estimate the slowest mode of a memory operator fitted to the evolving overturning structure. The operator ensemble produces an exploratory alarm 77 years before collapse, compared with 29 years for a physics-based freshwater-transport indicator. A blind test on the recovery branch does not validate the same timing rule, supporting interpretation of the diagnostic as a proximity measure rather than a universal clock. In nine additional simulations, it warns before all four finite-rate forced transitions but not before three rapid noise-induced collapses available only as scalar AMOC records. Matched ablations show that the signal depends on long-timescale memory and vertical overturning structure rather than trend, forcing, or multivariate information alone.
Fréchet regression is well developed for Euclidean predictors, but local linear methods remain limited for general manifold-valued predictors. We propose local constant and local linear estimators for predictors lying on a general Riemannian manifold and responses taking values in a general metric space. The proposed local linear estimator is the first local linear Fréchet regression method in this setting. Our construction uses geodesic neighborhoods, logarithmic-map coordinates, volume-density correction, and frame-invariant scalar equivalent weights. For both estimators, we establish not only pointwise consistency and convergence rates but also uniform consistency and convergence rates. Simulations and real data applications demonstrate the finite-sample performance and practical applicability of the proposed methods across diverse predictor and response geometries.
Over the past 20 years, kernel discrepancies have been leveraged as a highly powerful tool for quantifying the disagreement of distributions, with numerous successful applications in two-sample, goodness-of-fit, and independence testing, among others. Their fastest estimators are known to converge at a parametric rate---$n^{-1/2}$---under mild conditions. While this rate is known to be minimax optimal on $\mathbb R^d$ under strict assumptions with bounded kernels, little is known about its optimality beyond the finite-dimensional Euclidean setting with unbounded kernels. In this work, we prove that the minimax lower bound of estimation of the most popular kernel discrepancies (maximum mean discrepancy, Hilbert-Schmidt independence criterion and kernel Stein discrepancy; MMD, HSIC, KSD) is $n^{-1/2}$ on general topological spaces, and under mild assumptions on the kernel; the same rates are shown (as corollaries) to hold for the estimation of the mean embedding and the centered cross-covariance operator. Our results settle the question of optimal estimation of these kernel discrepancies.
Causal inference is often focused on average effects, which can hide important aspects of the effect distributions. Here we consider the entire posterior effects distribution by estimating full counterfactual outcome distributions. We propose a methodology for inference on counterfactual distributions which builds upon the martingale posterior framework of Fong et al. (2023). This provides a highly flexible approach to estimating densities, distribution functions, and derived quantities such as quantiles, which coherently quantifies the epistemic uncertainty on any target estimand of interest. As the predictive recursions are based on an underlying nonparametric model (a Dirichlet process mixture model), our method naturally inherits robustness with respect to restrictive parametric assumptions. In addition, implementation of our method is typically very fast. This approach can be applied to marginal or conditional counterfactual distributions and is easily extended to an instrumental variables setup. Using the concept of almost conditionally identically distributed random variables, we prove convergence of the martingale posterior inference on the counterfactual outcome distributions for the causal models considered in the paper. We illustrate our approach on both simulated and real data. Using the latter, we investigate the effect of zinc lozenges on common cold duration, the impact of vitamin A supplementation on children's survival rates with one-sided non-compliance (analysed in Imbens and Rubin, 1997a) and the effect of job training (LaLonde, 1986).
We study the contextual dynamic pricing problem under non-stationarity, where a firm sells products to $T$ sequentially arriving consumers that behave according to an unknown demand model that can change over time. The demand model is assumed to be a generalized linear model (GLM), allowing for a feature vector in $\mathbb{R}^d$ that encodes products and consumer information. To achieve optimal revenue (i.e., least regret), the firm needs to learn and exploit the unknown GLMs while monitoring for potential changes. We propose a multiscale change-point detection based algorithm that achieves a regret of order $\widetilde{O}(\sqrt{s_TdT}\wedge\{V_T^{1/3}d^{1/3}T^{2/3}+\sqrt{dT}\})$, where $s_T$ is the number of piecewise stationary segments and $V_T$ is a newly defined notion of design-adjusted variation budget of model parameters. Our algorithm is adaptive and does not require knowing $s_T$ or $V_T$. Moreover, to our knowledge, this is the first dynamic pricing algorithm that is adaptive to the nature of changes and achieves the best-of-both-worlds rate, thus closing a long-standing gap in the literature. We remark that, due to the varying contexts, existing works in the adaptive non-stationary bandit literature cannot be applied to achieve optimality for contextual dynamic pricing. The regret is further accompanied with a newly constructed minimax lower bound, confirming the optimality of our algorithm (up to logarithmic factors). Extensive numerical experiments are conducted to illustrate the efficiency and robustness of the proposed algorithm in non-stationary dynamic pricing.
In many applications, a diffusion trajectory is observed without distortion only while it remains inside a time-varying region; outside this region, only its nearest boundary point and an indicator of visibility are recorded. We formalize this setting through a censoring scheme for multidimensional stochastic differential equations, in which the latent state is observed via its Euclidean projection onto a random, time-dependent closed convex set. Classical arguments handle the one-dimensional case. In higher dimensions, however, the projection generates additional finite-variation terms, and it is not immediate that the local information required for drift estimation is preserved under censoring. We show that it is. Under a suitable regularity condition on the projection map, allowing the censoring region to be nonsmooth, the observed process is a continuous semimartingale and admits a generalized Itô decomposition, whose finite-variation correction term assigns no mass to visible times. Consequently, stochastic integrals over visible times coincide exactly for the observed and latent processes. Building on this identity, we construct a Nadaraya-Watson-type drift estimator with a data-driven bandwidth selection rule and establish oracle inequalities. We also derive anisotropic minimax lower bounds for both the full-observation and censored experiments.
Bayes factors provide a coherent Bayesian measure of evidence for competing hypotheses and have recently been used as the basis for single-arm phase II trial designs with binary endpoints. In contrast to classical power analyses based on test statistics and p-values, Bayes-factor based sample size calculations target high probabilities of obtaining compelling evidence either for a relevant treatment effect or for the null hypothesis, given pre-specified Bayes-factor thresholds. This paper explains how to design single-arm phase II binomial trials using Bayes factors with a focus on simulation-free calibration of Bayesian and frequentist power, type-I-error, and the probability of compelling evidence for the null. Two oncology-motivated examples illustrate the approach and are implemented in the bfbin2arm R package, with code provided in an appendix. The methodology fits naturally into current efforts to innovate and modernize clinical trial design through Bayesian and adaptive methods.
Count data models are ubiquitous in many fields, yet Bayesian data augmentation algorithms for such models frequently encounter challenges with Markov chain Monte Carlo efficiency. Posterior simulation is especially demanding when modeling data with a high proportion of zero outcomes. In this paper, we address this issue by introducing a marginal data augmentation approach for semi-parametric Bayesian count data regression models, based on a working parameter that rescales latent outcomes corresponding to zero-count observations. This strategy alleviates the strong posterior dependencies that typically reduce the efficiency of standard data augmentation schemes and leads to substantial gains in sampling efficiency. Synthetic data examples and simulation studies are used to demonstrate the improvements in mixing compared to conventional sampling methods. A demographic application using latent factor analysis to model subnational mortality counts in Austria further underscores the broader applicability of the proposed methodology.
Over-parameterized linear regression has been widely studied over the last decade. However, most existing works assume that the covariates are independent and that their covariance matrices are non-degenerate. In this paper, we relax both assumptions and derive deterministic equivalents for the prediction risk in a vanishing-ridge regime. We show that degeneracy of the covariance matrices and dependence can lead to multiple descent, and characterize where the corresponding peaks can occur. Our proofs use a novel graph representation of the variance profile. We show that maximum matchings and the Dulmage--Mendelsohn decomposition of the associated bipartite graph identify the configurations at which the variance becomes singular.
We propose a methodology for computing marginal likelihoods for block models. The proposed estimator computes the marginal likelihood from Markov chain Monte Carlo (MCMC) samples and is simulation-consistent, even when the size of the dataset is fixed. Moreover, it is asymptotically normal, of finite variance, invariant to label switching and can be computed efficiently, even for models with an arbitrarily large number of components. We evaluate the method through simulation studies in settings where the true marginal likelihood is available analytically. Finally, we apply the approach to a social network dataset based on the 2023 United Nations Climate Change Conference (COP28) and discuss the resulting insights.