2026-08-18 | | Total: 99
Electrophysiological brain signals are typically acquired through indirect and noisy measurements, providing transformed representations of the underlying neural activity. Source reconstruction---the inverse problem of resolving underlying neural signals from these measurements---is essential for mapping brain function but remains challenging because it is ill-posed and sensitive to noise. Deep learning methods have shown promise across a range of inverse problems, yet many do not explicitly incorporate the biophysical principles governing data generation, limiting data efficiency and adaptation across subjects. Here, we introduce DeepOp-Informed, a biophysics-informed geometric deep operator learning framework that embeds the biophysics of the sensing process into the model through a custom differentiable layer, enabling more efficient learning and improved reconstruction performance. This layer enables the neural network to adapt to subject-specific variations in the physics of signal generation, resulting from differences in brain anatomy and sensor positioning. In realistic magnetoencephalography simulations, DeepOp-Informed generalizes to forward models from held-out subjects, reducing reconstruction error relative to several neural-network and classical baselines. Applied to adolescent auditory-evoked recordings, it produces anatomically plausible reconstructions localized to the auditory cortex. While our application focuses on magnetoencephalography, the framework is general and may be adaptable to other imaging modalities.
In survival analysis the way covariates act on the risk of an event often differs between early and late failure times, yet hazard- and mean-based summaries collapse this variation into a single number. Quantile-based modeling instead describes the full conditional distribution on the original time scale, but existing censored-data methods are either inflexible or produce logically inconsistent crossing quantile curves. We propose a Censored Non-crossing Quantile (CNQ) framework for right-censored data that jointly estimates several conditional survival quantiles and guarantees valid ordering by construction, with flexibility supplied by Kolmogorov-Arnold and Transformer backbones, and we establish a finite-sample excess-risk bound holding jointly across all fitted quantile levels. Across 27 simulation settings and six cohorts the framework attains lower pinball loss than quantile-, hazard- and tree-based competitors whenever the conditional distribution is asymmetric, with interval coverage closer to nominal on all six. In two clinical case studies (METABRIC, breast cancer; FLCHAIN, population mortality) it recovers covariate effects that vary across the survival distribution and would be hidden by a single hazard ratio, and yields coherent individualized quantile milestones. Code: https://github.com/BIG-S2/deepcnq
Under the ICH E9 (R1) addendum, treatment policy strategies for intercurrent events target the treatment effect regardless of treatment discontinuation. Sequential multiple imputation (MI) models that condition each visit's imputation on discontinuation status or pattern reduce bias relative to mixed models and standard MI, but require every subject to contribute at least one post-baseline observation, an assumption violated by subjects who withdraw before any post-baseline assessment, a baseline-only early dropout pattern common in chronic-disease trials. We propose Extended Pattern-based Sequential Multiple Imputation (EPSMI), which reconstructs missing data for baseline-only early dropouts using covariate-matched, same-arm donors before applying an extended discontinuation-pattern indicator within eight sequential MI models. Two strategies were evaluated: EPSMI-Full, imputing the entire post-baseline trajectory from a donor, and EPSMI-Y1, imputing only the first visit and leaving later visits to the pattern-extended MI engine. A simulation study grounded in published Sjogren's syndrome trials evaluated bias, coverage, precision, power, and Type I error across 24 scenarios, comparing EPSMI against MMRM, standard MI, and sequential MI after excluding early dropouts (No Early). Under random early dropout, both EPSMI strategies reduced bias relative to No Early. Under informative early dropout the strategies diverged: EPSMI-Y1 remained robust, matching or exceeding No Early coverage with only mild Type I error inflation, whereas EPSMI-Full's deterministic reconstruction produced larger bias, narrower intervals from underestimated variance, lower coverage, and clear Type I error inflation. EPSMI-Y1 is recommended as the primary analysis strategy for baseline-only early dropout, keeping estimation faithful to the treatment policy estimand over the full randomized population.
We study empirical Bayes estimation of the prior in high-dimensional linear regression $\mathbf{y}=\mathbf{X}\mathbfβ+\mathbf{\varepsilon}$, where the regression coefficients are drawn independently from an unknown sub-Gaussian prior. In contrast to the sequence model, the design matrix couples the latent coefficients, so that recovering the prior requires deconvolving it from both the noise and copies of itself. We introduce the \emph{Empirical Bayes Method of Moments} (EBMoM), a computationally efficient procedure for general designs that recursively estimates the prior moments through a lower-triangular system of estimating equations and runs in time $O(np^2)$. Under mild design conditions, satisfied in particular by a broad class of correlated random designs, we show that EBMoM consistently estimates a growing number of moments and hence the prior itself, provided that $n\geq p^{1-o(1)}$. A matching information-theoretic lower bound, valid for a broad class of designs, shows that this sub-linear sample complexity is optimal for nonparametric prior estimation. This improves on existing results for likelihood-based methods whose consistency requires a linear sample size $n=Ω(p)$.
Instance-wise feature selection is a valuable tool for interpreting labeled data and the predictions of black-box models. In contrast to global feature selection techniques, instance-wise methods dynamically identify important features for each instance. A growing number of methods learn a selector, which identifies important features, and a predictor, which uses these to make predictions. However, these pioneering methods face challenges including information leakage and lack of differentiability, which can slow training. In this paper, we present Hide&Seek, an end-to-end differentiable model for instance-wise feature selection. We jointly learn feature selection and prediction under a single objective without information leakage. Hide&Seek outperforms existing state-of-the-art models across a range of experiments and is fast to train. We achieve this by reformulating feature removal as a differentiable operation where instead of discretely removing features, we replace a proportion of each feature. Training is further stabilized via a parsimony-weight annealing framework.
Bayesian dynamic borrowing (BDB) methods leverage historical data to reduce treatment effect uncertainty, yet existing approaches rely on parametric outcome models susceptible to misspecification. We propose the nonparametric latent exchangeability prior (NP-LEAP), an outcome-agnostic, assumption-lean framework to borrow information from historical data. The NP-LEAP performs individual-level exchangeability assessment, inducing Bayesian model averaging over all possible partitions of the historical data into exchangeable and nonexchangeable subsets. Although applicable to a variety of data types with choice of appropriate kernel, the NP-LEAP is particularly well-suited for studies with time-to-event outcomes, where parametric BDB is potentially triply misspecified - imposing a parametric baseline hazard, the proportional hazards structure, and blanket exchangeability. We establish posterior consistency under mild regularity conditions. Simulation studies demonstrate favorable operating characteristics relative to parametric borrowing methods and nonborrowing semiparametric frequentist methods. We illustrate the method by augmenting the control arm in a randomized trial of patients with non-small cell lung cancer.
We prove that every isotropic Gaussian mixture with finitely many components has finitely many modes. In one dimension, classical theory of Chebyshev systems gives the sharp bound of at most $n$ modes for an $n$-component mixture. In several dimensions, however, it has remained open whether every such mixture has finitely many modes. Our main result is the stronger statement that the entire critical set has finite cardinality, which is proven by combining real analytic curve selection theorem and Ax's functional-transcendence theorem. As an application, we show that, for Gaussian location mixtures, every nonparametric maximum likelihood estimator (NPMLE) based on a finite dataset is finitely supported. More specifically, all NPMLEs share the same finite set of allowable atom locations.
The Bethe-Hessian is a symmetric matrix for which the negative spectrum has been observed to encode the informative structure of sparse stochastic block models. We prove that, in the stochastic block model where all vertices have expected degree $d>1$, the number of negative eigenvalues of the Bethe-Hessian is exactly the number predicted by the eigenvalues of the planted model lying outside the bulk spectrum. The condition $d>1$ is optimal, and matches a regime in which existing spectral approaches based on larger non-Hermitian matrices apply. Our result extends a theorem of Stephan and Zhu, who established the same conclusion under the assumption $d\geq 2$. Our improvement relies on two main ideas. First, we construct test vectors on the $2$-core, where degree fluctuations are substantially smaller, and then extend them to the entire graph while controlling the quadratic form. Second, we construct the test vectors using an isotropic basis of the underlying Markov random field, with coefficients adapted to each relevant planted eigenvalue. This allows us to control the fluctuations of the test vectors throughout the sparse regime.
Generalized Bayesian inference uses weights of the form $\exp\{-β_n\ell_{n}(θ)\}$, even when the loss is available only through simulation, numerical integration, or subsampling. Exponentiating an unbiased loss estimate changes the target, and when $β_n\asymp n$ an ordinary Monte Carlo (MC) loss estimate with $M^{-1}$ variance needs a per-proposal budget of order $n^2$ to keep the leading log-weight variance bounded. We introduce the Sign-Corrected Bessel Debiasing (SCBD) algorithm, a signed pseudo-marginal method based on $K$ independent block estimates of the loss, and study its independently randomized Randomised Quasi-Monte-Carlo (RQMC) implementation. Under an iid Gaussian block model, a Bessel factor constructed from the block sample variance exactly removes the Gaussian exponential bias. For general finite MC or RQMC blocks, signed averages instead target a density proportional to $π_n(θ)w_M(θ)$, where $w_M$ is the finite-block target factor remaining after the signed correction; its posterior variation is assessed separately from sign efficiency. If an RQMC block estimator, based on $d$-dimensional randomized inputs, has variance $O\{B^{-α}(\log B)^{d-1}\}$, a sufficient budget for bounded leading log-weight variance has order $n^{2/α}$, up to logarithmic factors. We give a finite-block total-variation bound on the approximation and show the separate improvements from the debiasing correction and RQMC. The numerical examples show that, when the RQMC representation is favorable, the selected budget for stabilizing the estimated log weight can inherit this order, and that the Bessel correction can matter even when the RQMC variance rate is close to that of ordinary Monte Carlo.
Seasonal infectious-disease interventions are commonly evaluated with interrupted time-series or pre--post designs that align epidemics by calendar week. When epidemic onset, speed or peak timing differs between seasons, such comparisons confound a shift in epidemic phase with a change in disease burden. We propose a Bayesian causal count model in which season-specific affine transformations map calendar time to a latent epidemic clock, and intervention effects are estimated on that clock rather than on the calendar. The alignment is a model component rather than a preprocessing step, so uncertainty about epidemic timing propagates into every causal contrast. The model uses a negative-binomial observation distribution, hierarchical area, season and area-season effects, a shrunk Fourier epidemic curve, and a continuous programme-intensity exposure. Posterior g-computation yields prevented cases, prevented fractions, peak attenuation and epidemic displacement, under both a controlled contrast and a dynamic contrast that propagates disease history within each arm. A two-tier simulation study evaluates bias, root mean squared error, interval coverage and parameter recovery under stable timing, epidemic-clock variation, intensity-dependent ascertainment and area-level confounding. We illustrate the framework using open Catalan primary-care surveillance and respiratory syncytial virus immunisation data, with explicit attention to the overlap in programme intensity that identifies the effect.
The log-rank test and Kaplan--Meier plot are standard tools for analyzing time-to-event data in randomized clinical trials, yet neither provides a summary of the magnitude of the treatment effect. Practitioners typically fill this gap by reporting a hazard ratio from a Cox proportional-hazards model or an acceleration factor from an accelerated failure time (AFT) model, but both require assumptions beyond those needed for the log-rank test or Kaplan--Meier estimator. We propose two nonparametric confidence intervals for scalar effect-size summaries, an additive shift c and a multiplicative factor $ρ$, obtained by inverting the log-rank test under sharp null hypotheses of constant treatment effects. Building on the randomization-inference framework of Li and Small (2023), both intervals are valid under the randomization distribution alone, requiring no assumptions for the event-time distribution. We evaluate the proposed multiplicative interval via simulation, finding that it maintains nominal coverage across a range of censoring rates and sample sizes, including under data-generating processes that misspecify a parametric AFT model, while incurring only a modest efficiency loss compared to parametric AFT inference under correct specification. We illustrate the approach using data from a randomized trial of rhDNase for cystic fibrosis and provide R code and a Shiny application for ease of implementation.
Dataset alignment is a central step in data analysis across science and engineering, where the goal is to match observations between datasets. Entropic Optimal Transport (EOT) offers a computationally tractable framework for this task by encoding cross-dataset affinities in a transport plan. However, when two datasets are sampled from geometrically similar low-dimensional structures with substantially different sampling densities, the EOT plan may match points by relative sampling density rather than geometric proximity, yielding geometrically misleading correspondences. To address this issue, we propose a density-reweighted EOT framework in which the influence of sampling density on the transport plan can be discounted to a desired degree, ranging from standard EOT to alignment driven purely by underlying geometry. Under suitable regularity conditions, we establish convergence of the reweighted EOT plan to a family of population-level plans whose dependence on sampling density is made explicit. Through simulations, we show that our approach recovers geometrically faithful correspondences, improving over related EOT-based frameworks when datasets exhibit substantial sampling density disparity.
This paper studies the regret analysis for parallel Gaussian process (GP) bandit optimization. The known regret upper bounds for the widely used GP batched upper confidence bound and GP batched Thompson sampling (GP-BTS) suffer from a multiplicative factor with respect to the batch size $Q$. To avoid this degradation, existing analyses require a polynomial number of uncertainty sampling (US) for $Q$ at the beginning of optimization. However, this initial US phase is often ineffective in practice. This paper shows that the regret upper bound without the multiplicative factor on $Q$ can be achieved without the initial US phase, using GP-BTS as an example. Furthermore, we show much better regret upper bounds in the noiseless setting than in the noisy setting, as in the sequential GP bandit setting.
Bayesian optimal experimental design (BOED) aims to collect informative data by optimizing an expected utility reflecting the goals of an experiment. However, this optimization is computationally challenging for common utilities and complex models. This is especially so for sequential or adaptive designs, where design and data collection alternate, so that feedback from already observed data must be taken into account. Most existing BOED research employs information gain as the utility, leading to the expected information gain (EIG) criterion. While EIG is widely useful, it may not always adequately reflect experimental goals. EIG can be viewed as rewarding experiments that produce large positive evidence for the truth on average, but it does not directly control the risk of an experiment producing misleading evidence. Here we consider an alternative criterion, which we call bias against (BA), that prioritizes such control. To address computational challenges when applying this criterion for adaptive design, we consider a policy-based deep adaptive design framework, which has previously been used for the EIG criterion. Minimizing a tractable upper bound on the BA objective is equivalent to maximizing a variance-penalized EIG criterion, and we optimize the latter by approximating it by Monte Carlo and learning design policies using stochastic gradient methods. The differences between BA and EIG designs are demonstrated in several examples including the adaptive design of a complex discrete choice experiment.
Problem-solving log process data from computer-based assessments provide detailed information about how respondents approach and complete tasks. However, the resulting action sequences are complex and noisy, making it difficult to identify specific behaviors associated with successful performance. This paper proposes a representation-learning item response modeling (IRT) framework for identifying behaviorally important actions while accounting for respondent proficiency and item-level differences. Raw log sequences and timing information are first transformed into action representations that incorporate the hierarchical structure of action labels and the sequential and temporal context in which each action occurs. These respondent-specific representations are then entered as covariates in an extended IRT model, with spike-and-slab priors used to identify action-item combinations associated with response accuracy. The framework therefore evaluates actions contextually rather than as simple occurrence indicators and provides posterior uncertainty for their associations with performance. We apply the approach to problem-solving process data from the OECD Programme for the International Assessment of Adult Competencies (PIAAC). The analysis identifies a sparse set of actions associated with successful and unsuccessful performance and reveals differences across items in where behavioral information occurs within the problem-solving process.
Ornstein-Uhlenbeck processes with flexible multivariate drift matrices are powerful models for capturing coupled, asymmetric, and damped-oscillatory mean reversion. However, likelihood-based inference is challenging because the drift matrix must remain Hurwitz stable, while likelihood and gradient evaluations require repeated, costly computation and differentiation of the drift matrix exponential and of solutions to the associated Lyapunov equation. We introduce the Hurwitz smooth spectral block parametrization (H-SSBP), which represents the drift as a change of basis applied to independent one- and two-dimensional stable blocks. Simple scalar constraints enforce Hurwitz stability, while the two-dimensional blocks vary smoothly between real- and complex-eigenvalue regimes, avoiding discrete model selection. The H-SSBP represents every real Hurwitz matrix diagonalizable over the complex numbers and has dense, full-measure support within the Hurwitz cone. Its block structure reduces the OU transition matrix, stationary and innovation covariances, and their reverse-mode derivatives to constant-size block or block-pair computations plus change-of-basis multiplications. Numerical experiments demonstrate dramatic speed-ups over alternatives, particularly for matrix-exponential adjoints and Lyapunov-equation kernels. Real-data analyses of asynchronous financial data and multivariate phylogenetic traits illustrate the proposed framework in practice.
Interference occurs when one individual's treatment or exposure affects another individual's outcome. In particular, we assume partial interference, where individuals are divided into groups such that there is no interference between individuals in different groups. In observational studies, inverse probability weighting (IPW) based on propensity scores is often used for causal effect estimation. However, under partial interference, the group-level propensity score must be estimated, and it is more likely to take extreme values than the usual individual-level propensity score. As a result, IPW estimators may have large variances. This problem can become more serious when many covariates are available. In this study, we propose an Outcome-Adaptive Lasso based on a mixed-effects logistic regression model to stably estimate causal effects under partial interference. The proposed method performs covariate selection and estimation in the propensity score model simultaneously while accounting for unobserved group-level heterogeneity in treatment assignment. Under regularity conditions, we show that the proposed method has the oracle property and that the IPW estimators based on the proposed method are consistent and asymptotically normal. Through Monte Carlo simulations, we demonstrate that the proposed method tends to select confounders and prognostic factors at high frequencies, while excluding instrumental variables and spurious variables. The results further suggest that the proposed method improves the finite-sample efficiency of IPW estimators. We evaluate the performance of the proposed method using malaria data from the Democratic Republic of the Congo Demographic and Health Survey (DHS).
The combination of traditional statistical models and neural network (NN) components into semi-structured hybrid models is an intriguing approach to construct models that, ideally, combine traditional interpretability with the unprecedented flexibility of NNs. In order to preserve interpretability, it is usually necessary to restrict the NN components to prevent them from dominating the model. However, existing methods that enforce structural constraints on their NN components severely limit their models' flexibility; in contrast, methods that only enforce weak, indirect constraints lose meaningful interpretability. The method we propose therefore leverages invertible residual neural networks (i-ResNets) to equip generalized linear models with both nonlinear parameter estimation and a flexible correction of their distributional assumptions while always retaining stochastic monotonicity of the modeled distribution in the (formerly linear) predictor. The i-ResNets correspond to a controlled deviation from identity and by constraining their Lipschitz constant one can rigorously limit and quantify how far the hybrid model deviates from its traditional counterpart. This enables a user-specifiable compromise between flexibility and interpretability without limiting the structure of nonlinear and interaction effects that can be learned. Furthermore, we develop specific inherent interpretation techniques for our model and enforce model identifiability through an adapted post-hoc orthogonalization.
In multi-institutional studies, different parties hold distinct feature blocks for partially overlapping sets of individuals. Responses may also be missing for some records. In such settings, we propose Assisted Learning with Block-Missing Data (ALB) for sparse high-dimensional linear estimation and coordinatewise inference without pooling records or relying on a coordinating server. ALB minimizes a regularized available-case quadratic loss using cyclic block updates. Each cycle communicates $O(n)$ scalars through sample-level linear summaries, regardless of data dimension $p$, and the iterates converge geometrically to the centralized solution. We derive estimation rates that separate statistical and optimization errors. For inference on a target coefficient, ALB estimates the corresponding precision column and uses a sample-level variance estimator that accounts for dependence among moments computed from overlapping samples. Under sparsity and overlap conditions, the studentized estimator is asymptotically standard normal at the $\sqrt n$ rate, even when there are no complete cases. We also study one-time perturbed covariate and response releases that reduce direct disclosure by replacing unperturbed sample-level quantities with noisy versions. Simulations and an analysis of multimodal Alzheimer's Disease Neuroimaging Initiative data indicate that ALB approximates its centralized benchmark and improves upon complete-case Lasso by incorporating partially observed records.
LLM judges are often evaluated with a single prompt and only a few repeated calls. When their verdicts vary, it remains unclear whether the variation comes from sampling noise within a prompt or systematic differences across prompts. We formalize this distinction using a second-order response law: the distribution of prompt-conditioned verdict distributions induced by a declared prompt policy. For a quadratic measure of prompt instability, we show that the usual plug-in estimator is biased upward at finite repeat budgets because it confounds within-prompt noise with between-prompt variation. We derive unbiased estimators for both sampled prompts and declared fixed prompt censuses from the difference between within- and across-prompt agreement. Under a crossed prompt-by-answer-order design, the same framework separates prompt, order, interaction, and residual call variation, while retaining invalid completed outputs as outcomes. Known-law simulations and a byte-identical live null recover the predicted finite-$R$ inflation. In a matched Qwen study, corrected low-repeat estimates are closer to an independently acquired $R=16$ reference than plug-in estimates, with the largest gains at small repeat budgets. A matched panel across four frozen judge configurations exhibits configuration-specific inflation magnitudes and component profiles. Prompt robustness can therefore be estimated separately from finite-call noise.
We propose a regression model for the extreme tail of a response variable, in which covariates rescale the tail without changing its shape. A single covariate-dependent function then characterizes the entire conditional tail, in contrast to extreme quantile regression, which targets a quantile at a pre-specified level. The tail shape itself is left unrestricted: heavy-, light- and short-tailed responses are covered by the same framework. We specify the function through a link function and a linear combination of the covariates, which is in the spirit of a generalized linear model. In estimation, we match the parametric specification to the underlying tail function under a Bregman divergence, over a region localized at the largest observations. The resulting loss is convex, and an $\ell_1$-penalty allows the number of covariates to exceed the effective sample size. The tail localization makes the asymptotic theory deviate from that for classical penalized generalized linear models. Only the tail observations selected by a random threshold are used in the statistical analysis, making them dependent. We derive the convergence rate of the penalized estimator and propose a debiased estimator that is asymptotically normal, yielding confidence intervals for individual coefficients. Its asymptotic variance is determined by the covariance of the score, which under tail localization differs from the Hessian and must be estimated separately. We apply the method to automobile insurance claims data.
Identification of dominant polynomial-chaos modes is usually formulated as a sparse-regression problem on a sampled multivariate polynomial dictionary. We develop coded Hankel polynomial chaos (CH-PC), a complementary spectral formulation for dominant-mode identification. A finite generating transform converts PCE coefficients into a coefficient-generating polynomial, and evaluation along a geometric phase orbit produces a finite exponential sum. Its model order and spectral nodes are encoded by low-rank Hankel matrices, while coordinate phase shifts attach root-of-unity labels from which the full polynomial multi-indices are recovered. Coordinate-shifted probes are combined as common-node snapshots, and independent phase encodings provide redundant representations when a single spectral encoding is poorly conditioned. For finite observations, population, finite-data, and observed probes are kept distinct: sampling or quadrature error and observation error enter as separate Hankel perturbations, which are then connected to spectral stability, discrete decoding, and phase voting. For tensor-product candidate sets, the generating kernel factorizes into one-dimensional sums and can be evaluated without assembling the full multivariate PCE design matrix. Numerical experiments on sparse Legendre benchmarks and a stochastic Darcy problem illustrate exact recovery, noise stabilization, unknown-order identification by phase persistence, and dominant-mode recovery for a PDE-generated quantity of interest.
Coresets distill large datasets into small, representative subsets for efficient downstream learning. Yet Optimal Transport (OT)-based selection typically requires intensive computation of transport plans, limiting scalability. We introduce a scalable Sinkhorn coreset method that permits closed-form updates of the entropically regularized OT coupling by allowing non-uniform coreset weights. This produces centroids that generalize k-means via soft assignments. We establish asymptotic consistency of the selected measure and Lipschitz stability to data perturbations, providing accuracy and robustness guarantees. Across synthetic and real-world benchmarks, the proposed method achieves competitive or improved approximation quality while substantially reducing runtime compared to Wasserstein- and standard Sinkhorn-based coreset selection, especially at large scale.
Estimating high dimensional Generalized Structural Equation Models presents severe computational challenges. Traditional simultaneous estimators frequently suffer from numerical instability and prohibitive computational costs. Moreover, there are no tractable algorithms for families such as Poisson, negative binomial, and gamma. To overcome these limitations, this article introduces a Two Stage Quasi-Likelihood Expectation-Maximization framework. The proposed method isolates the structural model from the measurement model. First, it approximates the conditional distribution of the latent variables given the observed indicators. Second, it employs marginal quasi-likelihood estimating equations to evaluate the structural parameters, deriving the necessary conditional moments either exactly or through Monte Carlo integration. This approach completely avoids the need to evaluate the full joint likelihood. Extensive simulations demonstrate that our method drastically reduces computational runtime, providing a numerically stable framework that minimizes the mean squared error and structural bias to yield a scalable and flexible solution for analyzing complex latent variable models.
Accurate parameter estimation for atmospheric turbulence channels is challenging because the probability density functions of the Gamma-Gamma (GG) and Lognormal-Rician (LR) models involve special functions and numerical integrations. This paper proposes two Gaussianization parameter estimators for GG and LR turbulence channels, i.e., the quantile-transformation (QT) estimator and the Box-Cox estimator. The QT estimator employs bidirectional cross-transformation together with higher-order statistical matching, whereas the Box-Cox estimator constructs an approximate likelihood by incorporating the transformation Jacobian. Asymptotic analysis identifies skewness as the leading-order deviation from Gaussianity under weak-to-moderate turbulence and yields asymptotic expressions for the Box-Cox power parameters. In addition, a physics-informed regularizer based on the extended Rytov-theory mapping from the Rytov variance to the GG shape parameters is introduced to improve parameter identifiability. Simulation results demonstrate that both estimators provide robust performance under noiseless and noisy conditions, while the physics-informed regularization term can improve the parameter estimation performance by at least three orders of magnitude compared with the iterative moment-based estimator and the method-of-moments/convex-optimization estimator in specific turbulence scenarios.