UAI.2026 - Poster

| Total: 306

#1 EagleConv: Bio-inspired Dual-Foveated Convolution for Robust Small Object Detection [PDF] [Copy] [Kimi] [REL]

Authors: Jianwei Liu, Lifei Hao, Baoqi Huang, Bing Jia, Xuandong Zhao

Small object detection underpins wide-area vision tasks such as UAV and remote-sensing imagery. However, standard convolutions often exhibit low-pass smoothing behavior, which suppresses the sparse edge cues of tiny targets and may cause them to be overwhelmed by background clutter, leading to semantic information loss at the early stages of feature extraction. Inspired by the dual-fovea physiology of the eagle eye, we propose EagleConv, a bio-inspired operator that superimposes three Gaussian components into a center-excitation, surround-inhibition, peripheral-context response profile to selectively amplify small-object signals while attenuating noise. Architecturally, EagleConv adopts a sparse dual-pathway design that integrates partial-channel depthwise dual-foveated filtering, pointwise channel mixing, and adaptive residual gating, followed by anti-aliasing downsampling for robust dimensionality reduction. Experiments on three public benchmarks show that augmenting YOLOv11s with EagleConv delivers consistent and significant gains over multiple convolution-augmentation baselines, achieving state-of-the-art performance and confirming its transferability and generality as a plug-and-play module for wide-area small-object detection.

Subject: UAI.2026 - Poster


#2 Calibration-Aware Online Adaptation under Label Shift [PDF] [Copy] [Kimi] [REL]

Authors: Jiun Jeong, Byeongwoo An, Gi-Soo Kim, Kyubo Shin

We study online adaptation of a pre-trained base classifier to streaming unlabeled data under label shift, where the marginal label proportions in the stream differ from those in the offline training data. We consider the case where the base classifier’s model class may be misspecified, motivating a sep- arate, calibrated auxiliary classifier used solely to estimate the target label proportions. While many works have studied this setting, it is less understood how calibration quality affects the performance of the adapted base classifier in an online setting. In this paper, we thoroughly analyze the estimation error of the adapted base classifier after the deploy- ment of a novel algorithm that estimates the target label proportions in an online fashion and dynami- cally adapts the base classifier using the estimated proportions. We decompose the error into a term that vanishes over time and a term determined by calibration quality. Moreover, we characterize an explicit trade-off between calibration granularity and finite-sample calibration error, and propose a novel calibration strategy which effectively bal- ances this trade-off.

Subject: UAI.2026 - Poster


#3 Long term sequential decision making under risk [PDF] [Copy] [Kimi] [REL]

Authors: Irmaan Mirzanejad, Nadjet Bourdache, Abdel-illah Mouaddib

We study finite-horizon MDP planning under \emph{root-based} (resolute) risk objectives that apply a rank-dependent functional to the distribution of total returns. Such objectives are non-linear in the return distribution and generally break Bellman optimality, so direct optimization by scenario-tree enumeration is intractable. We propose \textbf{ERQDP}, an enumeration-free and sampling-free method that solves a rank--quantile surrogate via exact DP (Dynamic Programming), evaluates candidate policies exactly by DP over return Probability Mass Functions (PMFs) on a discretized return grid (with an explicit rounding bound), and refines the surrogate in an anytime loop that reports an explicit upper--lower gap (certificate) for the target objective up to discretization budgets. Across tested benchmarks, ERQDP returns certified solutions or explicit residual gaps, enables fast risk-parameter sweeps with substantial runtime gains, and supports both risk-averse and risk-seeking behaviors.

Subject: UAI.2026 - Poster


#4 Collapse-Aware Regularization for Reliable Reasoning Under Distribution Shift [PDF] [Copy] [Kimi] [REL]

Authors: Quynh Vo, Cong-Duy T Nguyen

Reliable reasoning requires models to generalize under distribution shifts, yet in-distribution validation loss often fails to identify checkpoints that remain reliable out of distribution. We study representation collapse in Transformer hidden states as an internal degeneration that may reduce the effective capacity needed for multi-step inference. We characterize two complementary collapse signals: spectral capacity collapse, measured by the effective rank of layerwise representation covariance, and geometric alignment collapse, measured by average cosine alignment. We then propose the Collapse Risk Criterion (CRC), an ID-computable diagnostic score estimated from in-distribution validation representations. CRC is not a formal OOD guarantee, but a practical surrogate for OOD-free checkpoint selection. Across four reasoning benchmarks and three Transformer backbones, CRC correlates more strongly with OOD reasoning error than ID validation loss and standard confidence-based reliability scores. CRC-aware checkpoint selection improves OOD performance under depth/difficulty and template/rule shifts with minimal ID degradation. As an extension, collapse-aware training with a CRC-based regularizer improves average OOD performance in all evaluated dataset--backbone settings. Our results suggest that monitoring representation collapse is a simple and useful tool for improving reasoning reliability without OOD validation data.

Subject: UAI.2026 - Poster


#5 Who Guards the Guardians? The Challenges of Evaluating Identifiability of Learned Representations [PDF] [Copy] [Kimi] [REL]

Authors: Shruti Joshi, Théo Saulus, Wieland Brendel, Philippe Brouillard, Dhanya Sridhar, Patrik Reizinger

Identifiability in representation learning is commonly evaluated using standard metrics (e.g., *MCC, $R^2$, DCI*) on synthetic benchmarks with known ground-truth factors. These metrics are assumed to reflect recovery up to the equivalence class guaranteed by identifiability theory. We show that this assumption holds only under specific structural conditions: each metric implicitly encodes assumptions about both the data-generating process (DGP) and the encoder. When these assumptions are violated, metrics become misspecified and can produce systematic false positives and false negatives. Such failures occur both within classical identifiability regimes and in post-hoc settings where identifiability is most needed. We introduce a taxonomy separating DGP assumptions from encoder geometry, use it to characterise the validity domains of existing metrics, and release an evaluation suite for reproducible stress testing and comparison.

Subject: UAI.2026 - Poster


#6 Uncertainty Quantification of Click and Conversion Estimates for the Autobidding [PDF] [Copy] [Kimi] [REL]

Authors: Ivan Zhigalskii, Andrey Pudovikov, Aleksandr Katrutsa, Egor Samosvat

Modern e-commerce platforms employ various auction mechanisms to allocate paid slots for a given item. To scale this approach to the millions of auctions, the platforms suggest promotion tools based on the autobidding algorithms. These algorithms typically depend on the Click-Through-Rate (CTR) and Conversion-Rate (CVR) estimates provided by a pre-trained machine learning model. However, the predictions of such models are uncertain and can significantly affect the performance of the autobidding algorithm. To address this issue, we propose the $\texttt{DenoiseBid}$ method, which corrects the generated CTRs and CVRs to make the resulting bids more efficient in auctions. The underlying idea of our method is to employ a Bayesian approach and replace noisy CTR or CVR estimates with those from recovered distributions. To demonstrate the performance of the proposed approach, we perform extensive experiments on the synthetic, iPinYou, and BAT datasets. To evaluate the robustness of our approach to the noise scale, we use synthetic noise and noise estimated from the predictions of the pre-trained machine learning model.

Subject: UAI.2026 - Poster


#7 One-Shot Federated Learning based on Random Feature Extractor [PDF] [Copy] [Kimi] [REL]

Authors: Luyuan Yang, Shayan Shafaei, Yiming Liu, Naeem Shahabi Sani, Yu CAI, Jun Huan, Chao Lan

The transition from multi-rounds federated learning to one-shot federated learning (OFL) markedly alleviates communication burden and represents a major step toward realistic deployment. Most existing OFL approaches require clients to perform local training, imposing a substantial computational burden on client side, while others rely on pre-trained models. In this paper, we propose a novel one-shot federated learning framework based on random feature extractor (FedRFE). Unlike existing approaches, it does not require any local model training or pre-trained model, featuring superior resource efficiency. Through comprehensive experiments, we show FedRFE achieves competitive performance while being robust in challenging scenarios including data heterogeneity and client scalability. Our code is available at\url{https://github.com/luyuanxyang/FedRFE}

Subject: UAI.2026 - Poster


#8 Finding the Signal in the Spam: Jointly Learning Rewards and Worker Reliability from Pairwise Comparisons [PDF] [Copy] [Kimi] [REL]

Authors: Kaustubh Shivshankar Shejole, Tanish Agarwal, Arpit Agarwal, Avishek Ghosh

The problem of learning from pairwise comparisons has been widely studied across many domains such as recommendation systems, social choice, and more recently, fine-tuning large language models. In this problem, the goal is to learn item rewards based on pairwise comparisons between them. In many scenarios, these comparisons are elicited from crowdworkers using platforms such as Amazon Mechanical Turk, Scale AI, etc. However, crowdworkers are often unreliable due to limited domain knowledge or revenue-maximizing (spamming) behavior. In this work, our goal is to understand whether worker reliability (competency) can be learned jointly with item rewards. To this end, we adopt the Boltzmann-rational model for pairwise comparisons, which extends the Bradley–Terry–Luce model by incorporating worker competencies. We derive an EM-based algorithm for learning under this model by introducing Polya-Gamma latent variables to transform the logistic likelihood into a conditionally Gaussian form, enabling tractable optimization and leading to a simplified $Q$ function in the E-step of the algorithm. This technique allows us to reduce our formulation to a matrix sensing problem, using which we establish theoretical convergence guarantees for our algorithm. We conduct extensive experiments on real-world and synthetic datasets. These experiments demonstrate the advantages of using our algorithm over several baselines and confirm its strong robustness to both spammers and adversarial workers, highlighting its practical effectiveness in realistic crowdsourcing and reward learning settings.

Subject: UAI.2026 - Poster


#9 Distributional Deep Gaussian Processes [PDF] [Copy] [Kimi] [REL]

Author: Sebastian Popescu

Deep Gaussian processes (DGPs) offer a principled Bayesian framework with hierarchical uncertainty propagation, but their reliable propagation of uncertainty and out-of-distribution (OOD) detection performance remains underexplored and often unreliable in safety-critical settings. In this work, we propose a novel kernel operating in both Euclidean and Wasserstein-2 space to better account for the geometry of representation learning spaces, thus circumventing a common pathology called feature collapse, whereby inliers and outliers get mapped to similar spaces. Empirically, our approach consistently improves OOD detection in convolutional image tasks and shows improved performance on tabular datasets.

Subject: UAI.2026 - Poster


#10 Efficient Federated Conformal Prediction with Group-Conditional Guarantee [PDF] [Copy] [Kimi] [REL]

Authors: Haifeng Wen, Osvaldo Simeone, Hong Xing

Deploying trustworthy AI systems requires principled uncertainty quantification. Conformal prediction (CP) is a widely used framework for constructing prediction sets with distribution-free coverage guarantees. In many practical settings, including healthcare, finance, and mobile sensing, the calibration data required for CP are distributed across multiple clients, each with its own local data distribution. In this federated setting, data can often be partitioned into, potentially overlapping, groups, which may reflect client-specific strata or cross-cutting attributes such as demographic or semantic categories. We propose \emph{group-conditional} federated conformal prediction (GC-FCP), a federated extension of conditional conformal calibration for a target mixture over prespecified groups. GC-FCP constructs mergeable, atom-stratified coresets from local calibration scores, enabling compact aggregation at the server when the number of active atoms is moderate. Experiments on synthetic and real-world datasets validate the performance of GC-FCP compared to centralized calibration baselines. The code of our work can be found at https://github.com/HaifengWen/GC-FCP.

Subject: UAI.2026 - Poster


#11 Canonical Domain Reduction for Partial Counterfactual Identification [PDF] [Copy] [Kimi] [REL]

Authors: Yesong Choe, Yeahoon Kwon, Min Woo Park, Sanghack Lee

Many counterfactual and causal queries are only partially identified from data, especially under unmeasured confounding. A common approach represents compatible nonparametric structural causal models (SCMs) on a finite canonical domain and computes sharp bounds via linear programming (LP) over the induced simplex. However, the canonical full counterfactual state space grows exponentially even for small graphs, making naive LP-based bounding computationally heavy. We propose a constraint-aware reduction that quotients out degrees of freedom irrelevant to the optimization problem. Because sharp bounds are determined jointly by the query functional and the data-implied information set, we aggregate full states into equivalence classes that are indistinguishable to every linear functional appearing in the LP objective and constraints. We show that optimizing over the induced push-forward distribution on the reduced domain preserves feasibility and yields the same sharp bounds as the full-domain.

Subject: UAI.2026 - Poster


#12 Mixture-Greedy for Online Generative Model Selection: Is UCB Necessary in Diversity-Aware Multi-Armed Bandits? [PDF] [Copy] [Kimi] [REL]

Authors: Bahar Dibaei Nia, Farzan Farnia

Efficient selection among multiple generative models is increasingly important in modern generative AI, where sampling from suboptimal models is costly. This problem can be viewed as a *multi-armed bandit (MAB)* task. Under diversity-aware evaluation scores, a non-degenerate mixture of generators can outperform any individual model, distinguishing this MAB setting from classical best-arm identification. Prior approaches incorporate an Upper Confidence Bound (UCB) exploration bonus into the mixture objective. However, across multiple datasets and evaluation metrics, we observe that the UCB term consistently slows convergence and reduces sample efficiency. In contrast, a simple *Mixture-Greedy* strategy without explicit UCB-type optimism converges faster and achieves even better performance, particularly for widely used metrics such as FID and Vendi where tight confidence bounds are difficult to construct. We provide theoretical insight explaining this behavior: under structural conditions, diversity-aware objectives induce *implicit exploration* by favoring interior mixtures, leading to sampling of all arms and sublinear regret guarantees for diversity-based objectives. These results suggest that in diversity-aware multi-armed bandits, e.g., for generative model selection, exploration can arise intrinsically from the objective's geometry. The project code is available at https://github.com/bhrdbn/Mixture-Greedy.

Subject: UAI.2026 - Poster


#13 LENS: Latent Precision Inference in Multi-LLM Routing [PDF] [Copy] [Kimi] [REL]

Authors: Juntao Liu, Lixing Yu, Kun Yue, Zhiwen Tang

Large language model (LLM) routing aims to select an appropriate model for each query under performance–cost trade-offs. A natural approach to improve adaptivity is to leverage interaction feedback to form behavioral signatures and continuously refine routing decisions as the environment changes. However, in realistic deployments, such feedback can have highly variable effective precision due to selective logging, imperfect evaluation signals, and temporal drift. As a result, behavioral signatures may be noisy or weakly informative, and treating them as uniformly precise can miscalibrate performance estimation and destabilize routing effectiveness. We propose the \textbf{L}atent pr\textbf{E}cisio\textbf{N} inference \textbf{S}ystem (\textbf{LENS}), a probabilistic routing framework that explicitly models the latent precision of interaction-derived signals. LENS formulates routing as posterior utility maximization under imprecise supervision, and marginalizes over latent precision to adaptively control how strongly behavioral signatures influence model selection. We instantiate LENS with an efficient variational inference procedure and evaluate it on multi-LLM routing benchmarks across diverse tasks and distribution shifts. Experimental results show that LENS consistently improves performance–cost trade-offs, with particularly strong gains under task and model shifts.

Subject: UAI.2026 - Poster


#14 Fast Best-in-Class Regret for Contextual Bandits [PDF] [Copy] [Kimi] [REL]

Authors: Samuel Girard, Aurelien Bibaut, Jill-Jênn Vie, Arthur Gretton, Nathan Kallus, Houssam Zenati

We study the problem of stochastic contextual bandits in the agnostic setting, where the goal is to compete with the best policy in a given class without assuming realizability or imposing model restrictions on losses or rewards. In this work, we propose \textit{Online Pessimistic Policy Learning} and establish the first fast rate for regret relative to the best-in-class policy. Our proposed algorithm updates the policy at every round by minimizing a pessimistic objective, defined as a clipped inverse-propensity estimate of the policy value plus a variance penalty. By leveraging entropy assumptions on the policy class and a Hölderian error-bound condition (a generalization of the margin condition), we achieve fast best-in-class regret rates, including polylogarithmic rates in the parametric case. Our analysis is driven by a novel sequential self-normalized maximal inequality for bounded martingale empirical processes, which yields uniform variance-adaptive confidence bounds and guarantees pessimism under adaptive data collection.

Subject: UAI.2026 - Poster


#15 First-Order Softmax Weighted Switching Gradient Method for Distributed Stochastic Minimax Optimization with Stochastic Constraints [PDF] [Copy] [Kimi] [REL]

Authors: Zhankun Luo, Antesh Upadhyay, Sang Bin Moon, Abolfazl Hashemi

This paper addresses the distributed stochastic minimax optimization problem subject to stochastic constraints. We propose a novel first-order Softmax-Weighted Switching Gradient method tailored for federated learning. Under full client participation, our algorithm achieves the standard $\tilde{\mathcal{O}}(\epsilon^{-4})$ oracle complexity to satisfy a unified bound $\epsilon$ for both the optimality gap and feasibility tolerance. We extend our theoretical analysis to the practical partial participation regime by quantifying client sampling noise through a stochastic superiority assumption. Furthermore, by relaxing standard boundedness assumptions on the objective functions, we establish a strictly tighter lower bound for the softmax hyperparameter. We provide a unified error decomposition and establish a sharp $\mathcal{O}(\log\frac{1}{\delta})$ high-probability convergence guarantee. Ultimately, our framework demonstrates that a single-loop primal-only switching mechanism provides a stable alternative for optimizing worst-case client performance, effectively bypassing the hyperparameter sensitivity and convergence oscillations often encountered in traditional primal-dual or penalty-based approaches. We verify the efficacy of our algorithm via experiment on the Neyman-Pearson (NP) classification, fair classification, and federated safe reinforcement learning tasks.

Subject: UAI.2026 - Poster


#16 Model Merging as Probabilistic Inference in Fine-Tuning Parameter Space [PDF] [Copy] [Kimi] [REL]

Authors: Long Minh Bui, Tuan Anh Le Van, Tung Phi Duc, Phi Le Nguyen, Jana Doppa, Trong Nghia Hoang

Model merging aims to combine existing single-task solutions into a multi-task solution without additional data-driven fine-tuning.~Most existing approaches achieve this using geometric properties of local solution spaces. However, such geometric views provide limited guidance for scoring how statistically useful each task-specific update direction is across tasks during merging. To address this, we formulate model merging from a new perspective of probabilistic inference under a product-of-experts (PoE) scenario where each single-task solution defines an energy-based expert model (EBM) over the merged parameters. We show that several existing model merging methods arise as special cases of our framework under energy designs that impose implicit Gaussian assumptions on directional residuals between merged and task-specific models. Empirically, we find that these residuals are often heavy-tailed which exposes a mismatch with the imposed light-tailed Gaussian structures. We address this with a heavy-tailed PoE design based on Cauchy experts, which better captures the observed residual behavior while admitting a provably convergent inference procedure. Experiments across multiple tasks and architectures show significant improvements over state-of-the-arts baselines. Our code is available at https://github.com/MinhLong210/PoE-EBM-Merging.git.

Subject: UAI.2026 - Poster


#17 Partially Observed Structural Causal Models [PDF] [Copy] [Kimi] [REL]

Authors: Turan Orujlu, Jordan Kyle Matelsky, Martin V. Butz, Charley M Wu, Konrad Kording

Here we introduce Partially Observed Structural Causal Models (POSCMs) as an extension of structural causal models (SCMs) to settings where upstream contexts co-determine both the interaction structure and downstream mechanisms on observed variables. POSCMs thus provide a self-contained causal modeling framework for endogenous graphs, allowing for an intervention hierarchy spanning node- and edge-level contexts and endogenous variable interventions. To define edge interventions, we separate node mechanisms into edge-local transmission channels that can be modified without changing the source node or the rest of the target mechanism. We provide an identifiability theory that clarifies which intervention families would suffice to disentangle structure formation from mechanisms. We then empirically validate these theoretical results in two external simulators: a biophysically detailed virtual human retina and a gene-regulatory analogue. The experiments reproduce non-identifiability under latent context, expose structure-mechanism confounding under latent edges, and recover pathway-level input-output relationships under targeted interventions, consistent with our positive Markov kernel identifiability results. Together, POSCMs provide an intervention-oriented framework for causal systems in which contexts, graph structure, mechanisms, and measurements are jointly generated and only partially observed.

Subject: UAI.2026 - Poster


#18 Optimal Conformal Prediction under Epistemic Uncertainty [PDF] [Copy] [Kimi] [REL]

Authors: Alireza Javanmardi, Soroush H. Zargarbashi, Santo M. A. R. Thies, Willem Waegeman, Aleksandar Bojchevski, Eyke Hüllermeier

Conformal prediction (CP) is a widely used frequentist framework to quantify uncertainty by constructing prediction sets with user-specified marginal coverage guarantees. In practice, CP is typically applied on top of probabilistic classifiers, which are able to express aleatoric but not epistemic uncertainty. In this paper, we consider the question of how to optimally employ CP on top of a more expressive formalism, namely credal sets, which can express both aleatoric and epistemic uncertainty. More specifically, we propose probabilistic Bernoulli prediction sets and derive a variant that achieves conditional coverage for valid credal sets while remaining minimal in expected size. We then address the more realistic scenario in which the validity of the credal sets is not guaranteed. Assuming access to calibration data with ground-truth distributions over labels, we apply conformal risk control to BPS and derive a PAC-style guarantee: with high probability over the data, the achieved conditional coverage is at least the desired level. We validate our theoretical findings empirically over various datasets.

Subject: UAI.2026 - Poster


#19 Estimating Interventional Outcomes over Time with Causal Normalizing Flow [PDF] [Copy] [Kimi] [REL]

Authors: Yoonseok Yeom, Jonghwan Kim, Taehui Yun, Juhyun Lyu, Jung-Hee Kim, Sangmin Lee, Jinseok Yang, Hyemin Jung, Woohyung Lim, Sanghack Lee

Estimating outcome distributions under time-varying treatments is an essential task for personalized decision-making, particularly in domains such as healthcare. Most prior work in this area focuses on point predictions, which fail to capture the inherent variability in outcomes. Recent efforts in causal inference have begun integrating generative models to address this limitation by estimating interventional distributions. However, existing approaches—including causal normalizing flows—are generally restricted to static settings and are not well suited to sequential, time-dependent data. In this work, we propose a novel framework that extends causal normalizing flows to time-series, enabling simulation-based interventional density estimation over time. Our method learns representations of treatment and covariate history that capture temporal dependencies. Conditioned on these representations and guided by a causal graph, our flow-based model generates interventional samples, allowing for the simulation of outcome trajectories under alternative treatment strategies. We evaluate our approach on both linear and non-linear synthetic time-series as well as on a simulated tumor growth dataset, demonstrating that it achieves performance competitive with state-of-the-art baselines, while accommodating a broader spectrum of causal queries.

Subject: UAI.2026 - Poster


#20 COBALT: Censored Optimization and Bayesian Active Learning Techniques [PDF] [Copy] [Kimi] [REL]

Authors: Andrea Karlova, Rishabh Kabra, Daniel Augusto de Souza, Brooks Paige

We target Bayesian Active Learning (AL) and Optimization (BO) for censored data regimes. While the Tobit likelihood accurately models such clipped observations, its mixed continuous-discrete nature impedes the analytical evaluation of information-theoretic acquisition functions. To address this, we investigate Censored Optimization via Bayesian Active Learning Techniques (COBALT). We establish rigorous theoretical guarantees for this framework, proving posterior consistency and the asymptotic normality of the MAP estimator under greedy maximization. Central to our framework is the derivation of a closed-form entropy for the Censored Normal distribution, enabling an analytical BALD (Bayesian Active Learning by Disagreement) score compatible with any Gaussian posterior approximations. We further underpin this method by deriving a numerically stable Evidence Lower Bound (ELBO) for censored atoms, utilizing robust approximations of the log-normal cumulative density. Empirical evaluations using our open-source implementation demonstrate COBALT’s best accuracy–compute trade-off among censored-likelihood methods in learning GP posteriors and effectiveness on a variety of benchmarks.

Subject: UAI.2026 - Poster


#21 Robustness Quantification for Discriminative Models: a New Robustness Metric and its Application to Dynamic Classifier Selection [PDF] [Copy] [Kimi] [REL]

Authors: Rodrigo F L Lassance, Jasper De Bock

Among the different possible strategies for evaluating the reliability of individual predictions of classifiers, robustness quantification stands out as a method that evaluates how much uncertainty a classifier could cope with before changing its prediction. However, its applicability is more limited than some of its alternatives, since it requires the use of generative models and restricts the analyses either to specific model architectures or discrete features. In this work, we propose a new robustness metric applicable to any probabilistic discriminative classifier and any type of features. We demonstrate that this new metric is capable of distinguishing between reliable and unreliable predictions, and use this observation to develop new strategies for dynamic classifier selection.

Subject: UAI.2026 - Poster


#22 On Transportability for Structural Causal Bandits [PDF] [Copy] [Kimi] [REL]

Authors: Min Woo Park, Sanghack Lee

The structural causal bandit (SCB) framework offers a graphical approach to identifying suboptimal actions by leveraging prior knowledge of the underlying causal structure. While theoretically appealing, there has been limited guidance on how to systematically transfer information across heterogeneous datasets from multiple environments. In this paper, we study the structural causal bandit under transportability, where prior knowledge from source environments is integrated to accelerate learning in a target deployment environment. We show that exploiting causal invariances across environments improves sample efficiency in identifying the optimal action.

Subject: UAI.2026 - Poster


#23 Beyond Bounds: Quantifying the Probability of Counterfactual Fairness [PDF] [Copy] [Kimi] [REL]

Authors: Teahan Kim, Minyoung Cho, Sanghack Lee

Counterfactual fairness is a rigorous criterion for algorithmic decision-making but remains fundamentally unidentifiable from observational data alone. Existing partial identification methods address this by deriving bounds for fairness measures; however, these intervals are often wide, limiting their practical utility. To address this limitation, we propose a framework that quantifies the probability that a black-box model satisfies counterfactual fairness, under a stated prior over the structural causal models compatible with the observed data. By exploiting conditional independencies to reduce the parameter space, we demonstrate that the set of causal parameters compatible with the observed data forms a convex polytope. We further show that, given domain-specific priors on exogenous distributions, this prior-dependent probability can be estimated via Hit-and-Run Monte Carlo integration. Our approach provides a practical tool that complements worst-case bounds, offering a prior-dependent probabilistic summary for auditing model fairness.

Subject: UAI.2026 - Poster


#24 Score-Based Diffusion Priors for Adaptive Conformal Inference under Distribution Shift [PDF] [Copy] [Kimi] [REL]

Author: XiangyuJiang

Conformal prediction provides a distribution-free coverage guarantee for predictive inference, yet its validity degrades under distribution shift, a common challenge in real-world deployment. We introduce DiffConf, a framework that uses score-based diffusion models as expressive priors over the data-generating process to enable adaptive conformal inference under temporal and covariate distribution shifts. The main idea is that the score function learned by a diffusion model encodes rich geometric information about the data manifold, which can be repurposed to construct nonconformity scores that are sensitive to distributional changes. We derive a diffusion-guided conformity score that integrates the learned score field with a lightweight online recalibration mechanism, providing finite-sample marginal coverage guarantees even when the data distribution evolves over time. Theoretically, we establish that DiffConf achieves asymptotic conditional coverage under mild regularity conditions on the drift rate, and we prove a regret bound that scales gracefully with the complexity of the distribution shift. Experiments on synthetic benchmarks, real-world tabular regression tasks, and high-dimensional image datasets demonstrate that DiffConf produces prediction sets that are simultaneously valid and more efficient than existing adaptive conformal methods, reducing average set size by 6--39% across the real-data benchmarks (and by over 50% under large synthetic mean shifts) while keeping coverage within one point of target.

Subject: UAI.2026 - Poster


#25 Analytic Planning under Uncertainty with Moment Closure [PDF] [Copy] [Kimi] [REL]

Authors: Shishir Sharma, Doina Precup

Effective model-based reinforcement learning in stochastic environments requires planning that accounts for predictive uncertainty. Propagating full state distributions analytically offers a principled way to do this, but has traditionally required restrictive policy or reward structures to remain tractable. Consequently, modern deep reinforcement learning has largely retreated to either stochastic sampling, which introduces significant target variance, or deterministic point estimates that ignore predictive covariance entirely. We investigate whether distribution-aware planning is possible without these constraints. Using a quadratic action-value parameterization, we first reduce the Bellman backup to an expectation over the state-value function alone; the key idea is then a compatibility principle between the predictive transition distribution and the value function class, under which this expectation is analytic in the distribution's moments. We instantiate this principle with a Gaussian transition model paired with a radial-basis value function, yielding a closed-form backup that propagates both predictive mean and covariance. Empirically, our approach reduces target variance and yields well-calibrated predictive uncertainty under stochastic observations in continuous control, providing a principled framework for planning with learned distribution models.

Subject: UAI.2026 - Poster