MICCAI.2026

| Total: 1165

#1 SPHARM-Mamba: Rotation-Invariant Multiscale Modeling for Brain Age Prediction [PDF2] [Copy] [Kimi] [REL]

Authors: Choi Junho, Joo Mingyu, Son Jiwon, Kim Won Hwa, Lyu Ilwoo, Choi Junho, Joo Mingyu, Son Jiwon, Kim Won Hwa, Lyu Ilwoo

Accurate brain age prediction is essential for understanding neurodevelopmental trajectories and detecting abnormal aging patterns. Morphological features derived from cortical shape provide informative structural representations for brain age prediction. Recent deep learning approaches for surface analysis primarily rely on local aggregation mechanisms. Transformer- and Mamba-based architectures model long-range dependencies through patch partitioning combined with self-attention or selective scan. However, these strategies introduce high computational cost, uncertainty at patch boundaries, and sensitivity to surface rotations that necessitate anatomical registration or extensive data augmentation. To address these limitations, we propose SPHARM-Mamba, a spherical harmonics-based generic backbone for genus-zero surface data that eliminates patch partitioning and avoids surface registration. To this end, dense cortical signals are projected onto spherical harmonics to construct compact rotation-invariant descriptors naturally ordered from global to local scales. Mamba is then employed to model dependencies across harmonic degrees and capture multiscale interactions efficiently. Experiments on cortical age prediction demonstrate superior performance over existing geometric and patch-based models with substantially fewer parameters. The software is available at https://github.com/Shape-Lab/SPHARM-Mamba.

Subject: MICCAI.2026


#2 AGGRNet: Selective Feature Extraction and Aggregation for Enhanced Medical Image Classification [PDF] [Copy] [Kimi] [REL]

Authors: Makwe Ansh, Agrawal Akansh, Jain Prateek, Agrawal Akshan, Bagade Priyanka, Makwe Ansh, Agrawal Akansh, Jain Prateek, Agrawal Akshan, Bagade Priyanka

Medical image analysis for complex tasks such as severity grading and disease subtype classification poses significant challenges due to intricate and similar visual patterns among classes, scarcity of labeled data, and variability in expert interpretations. Although deep learning models can capture complex visual patterns for medical image classification, many architectures still struggle to distinguish subtle classes because they do not adequately capture inter-class similarity and intra-class variability, which can lead to incorrect diagnoses. To address this, we propose the AGGRNet framework to extract informative and non-informative features to effectively understand fine-grained visual patterns and improve classification for complex medical image analysis tasks. Experimental results show that our model achieves state-of-the-art performance on 5 distinct medical imaging datasets, with the maximum improvement of 5% over SOTA models. The code is available at: https://github.com/AnshMakwe/AGGRNet-Feature-Extraction-and-Aggregation

Subject: MICCAI.2026


#3 Low-Rank Text-Guided Spectral Learning for Semi-supervised MHSI Segmentation [PDF] [Copy] [Kimi] [REL]

Authors: Zhang Siqi, Zhang Qing, Wang Yan, Li Qingli, Zhang Siqi, Zhang Qing, Wang Yan, Li Qingli

Microscopic hyperspectral image (MHSI) provides rich spectral information that enables the detection of subtle biochemical variations in tissue, making it a powerful modality for computational pathology. However, hyperspectral data exhibits substantial spectral redundancy, which may obscure critical discriminative cues. Current approaches rarely focus on identifying informative spectral components, thus failing to leverage the spectral invariance of pathological samples. Moreover, large-scale precise annotations are difficult to obtain, and textual description for MHSI remains largely unavailable, limiting supervised learning. To address these challenges, we propose a novel \textbf{Lo}w-rank \textbf{T}ext-guided \textbf{S}pectral learning \textbf{Net}work for semi-supervised MHSI segmentation, named as \textbf{LoTS-Net}. Specifically, we introduce a text-guided spectral channel selection mechanism to extract refined low-rank spectral representations and incorporate textual semantics to enable interpretable and redundancy-aware channel selection. Then, we leverage intrinsic spectral similarity across samples by constructing a spectral feature container that retrieves correlated prototypes to improve spectral-spatial fusion. This container mitigates noise during the interaction between labeled and unlabeled samples. To support future research, we construct the first microscopic hyperspectral vision-language benchmark and conduct comprehensive evaluations. Experimental results demonstrate that our method significantly outperforms the state-of-the-art methods, particularly under the challenging 5\% labeled data setting. Code is available at \url{https://github.com/ECNU-MultiDimLab/LoTS-Net}.

Subject: MICCAI.2026


#4 Frequency and Geometry Guided Graph Clustering for Weakly Supervised Skin Lesion Segmentation [PDF] [Copy] [Kimi] [REL]

Authors: Deng Zhaoxin, Wang Dongang, Shen Hualei, Deng Zhaoxin, Wang Dongang, Shen Hualei

Pixel-level annotation for skin lesion segmentation is costly and subjective. This limits large-scale clinical deployment. Weakly supervised methods based on image-level labels alleviate this burden. However, most existing methods rely on class activation maps (CAMs), which often fail to cover the entire lesion. To address this, we introduce a Frequency and Geometry Guided Graph Clustering framework, reformulating weakly supervised skin lesion segmentation as a differentiable node-clustering problem. First, a Frequency-aware Feature Modulation module injects local spectral priors into deep features. This enhances the model’s sensitivity to ambiguous boundaries. Second, we construct a semantic-spatial graph from these features. A graph neural network (GNN) then captures long-range dependencies to ensure structural completeness. Finally, we propose a tri-source fusion strategy that integrates three cues (GNN cluster, CAM, and the input image) via an improved side-window mean filtering. Experiments on ISIC 2017, ISIC 2018, and PH$^2$ show that our method achieves state-of-the-art performance. It significantly narrows the gap with fully supervised models, providing a reliable solution for label-efficient skin lesion analysis.

Subject: MICCAI.2026


#5 End2Reg: Learning Task-Specific Segmentation for Markerless Registration in Spine Surgery [PDF] [Copy] [Kimi] [REL]

Authors: Pettinari Lorenzo, El Hadramy Sidaty, Wehrli Michael, Cattin Philippe C., Studer Daniel, Hasler Carol C., Licci Maria, Pettinari Lorenzo, El Hadramy Sidaty, Wehrli Michael, Cattin Philippe C., Studer Daniel, Hasler Carol C., Licci Maria

Intraoperative navigation in spine surgery demands millimeter-level accuracy. Currently, this is achieved through radiation-intensive intraoperative imaging and bone-anchored markers that are invasive and disrupt surgical workflow. Markerless RGB-D registration methods offer a promising alternative. However, existing approaches rely on weak segmentation labels to isolate relevant anatomical structures, potentially propagating errors through the registration process. We present End2Reg, an end-to-end deep learning framework that jointly optimizes segmentation and registration, eliminating the need for segmentation labels and manual steps. The network learns task-specific segmentation masks optimized for registration, guided solely by the registration objective without explicit segmentation supervision. End2Reg achieves state-of-the-art performance on ex- and in-vivo benchmarks, reducing median Target Registration Error by 32% and mean Root Mean Square Error by 61%, while maintaining robust performance under partial occlusions. Ablation results confirm that end-to-end optimization significantly improves registration accuracy. Overall, End2Reg advances towards fully automatic, markerless intraoperative navigation. Code and interactive visualizations are available at: https://lorenzopettinari.github.io/end-2-reg/.

Subject: MICCAI.2026


#6 MedTriage-LM: Anatomically Grounded Visual Phenotype Synthesis for Interpretable ED Triage [PDF] [Copy] [Kimi] [REL]

Authors: Lu Zhixiang, Liu Xiwei, Wang Jinfeng, Zhou Mian, Nguyen Anh, Su Jionglong, Razzak Imran, Song Sifan, Lu Zhixiang, Liu Xiwei, Wang Jinfeng, Zhou Mian, Nguyen Anh, Su Jionglong, Razzak Imran, Song Sifan

Accurate emergency severity assessment fundamentally dictates patient survival and resource allocation. However, existing automated triage pipelines rely exclusively on tabular data and clinical text, neglecting the visual semantics of patient distress. A paramount barrier to anatomically informed reasoning is the absolute absence of visual modalities in public cohorts like the MIMIC-IV-ED dataset. To overcome this limitation, this study introduces MedTriage-LM, an anatomically grounded Multimodal Large Language Model (MLLM) that algorithmically injects visual priors without requiring real patient images. The architecture maps latent clinical representations onto a canonical human template via a weakly supervised Gaussian field generator, yielding continuous, interpretable Visual Phenotype Maps (VPMs). Synergizing these spatial representations with clinical embeddings via cross-attention shifts the paradigm from numerical risk scoring toward actionable three-class triage instruction prediction: Life-saving, High-Risk Assessment, and Resource Estimation. MedTriage-LM attains state-of-the-art performance among fine-tuned open-source models and approaches the capabilities of proprietary MLLMs, while additionally providing explicit, anatomically grounded interpretability. Crucially, explicit spatial grounding empowers an efficient 8B-parameter model to approach large proprietary multimodal foundation models, generating spatially grounded textual rationales that improve clinical trustworthiness.

Subject: MICCAI.2026


#7 TopoOR: A Unified Topological Scene Representation for the Operating Room [PDF] [Copy] [Kimi] [REL]

Authors: Wang Tony Danjun, Kim Ka Young, Birdal Tolga, Navab Nassir, Bastian Lennart, Wang Tony Danjun, Kim Ka Young, Birdal Tolga, Navab Nassir, Bastian Lennart

Surgical Scene Graphs abstract the complexity of surgical operating rooms (OR) into a structure of entities and their relations, but existing paradigms suffer from strictly dyadic structural limitations. Frameworks that predominantly rely on pairwise message passing or tokenized sequences flatten the manifold geometry inherent to relational structures and lose structure in the process. We introduce TopoOR, a new paradigm that models multimodal operating rooms as a higher-order structure, innately preserving pairwise and group relationships. By lifting interactions between entities into higher-order topological cells, TopoOR natively models complex dynamics and multi-modality present in the OR. This topological representation subsumes traditional scene graphs, thereby offering strictly greater expressivity. We also propose a higher-order attention mechanism that explicitly preserves manifold structure and modality-specific features throughout hierarchical relational attention. In this way, we circumvent combining 3D geometry, audio, and robot kinematics into a single joint latent representation, preserving the precise multimodal structure required for safety-critical reasoning, unlike existing methods. Extensive experiments demonstrate that our approach outperforms traditional graph and LLM-based baselines across sterility breach detection, robot phase prediction, and next-action anticipation.

Subject: MICCAI.2026


#8 StrokeTimer: Robust Representation Learning for Ischemic Stroke Onset-Time Estimation from Non-Contrast CT [PDF] [Copy] [Kimi] [REL]

Authors: Wang Weiru, Olthuis Susanne G. H., Lavrova Elizaveta, van Oostenbrugge Robert J., Majoie Charles B. L. M., van Zwam Wim H., Su Ruisheng, Wang Weiru, Olthuis Susanne G. H., Lavrova Elizaveta, van Oostenbrugge Robert J., Majoie Charles B. L. M., van Zwam Wim H., Su Ruisheng

Ischemic stroke is a major global disease. Treatment decisions are highly time-sensitive, as eligibility for reperfusion therapies relies on the interval between stroke onset and intervention. However, the true onset time is often uncertain in clinical practice, necessitating imaging-based assessment of tissue age as a surrogate marker. Early ischemic changes on routinely acquired non-contrast CT (NCCT) are often subtle, and real-world clinical datasets exhibit pronounced onset-time class imbalance and center-scanner-related heterogeneity. In this work, we propose StrokeTimer, a fully automated framework for onset-time estimation in acute ischemic stroke. StrokeTimer integrates self-supervised disentanglement learning with energy-guided contrastive learning to capture subtle ischemic patterns while addressing long-tailed data distributions under acquisition variability. Onset time is categorized into three clinically relevant windows (<4.5 h, 4.5–6 h, and >6 h). Experimental results on a large multi-center NCCT dataset from two national cohorts, show that StrokeTimer achieves a macro AUC of 0.69 and a macro F1-score of 0.57, improving the strongest baseline by nearly 50% (p < 0.005). In this realistic, challenging setting, representative baseline approaches exhibit near-chance macro performance. Model explanations further highlight subtle gray–white matter blurring and hypodense regions consistent with established radiological biomarkers. These findings demonstrate the potential of StrokeTimer to support treatment decision-making in acute ischemic stroke. We will make the code publicly available.

Subject: MICCAI.2026


#9 Physics-Inspired Continuous Transformer for Fast QDSA Reconstruction [PDF] [Copy] [Kimi] [REL]

Authors: Liu Yang, Liu Zehua, Liao Xiangyun, Duan Chuanzhi, Si Weixin, Liu Yang, Liu Zehua, Liao Xiangyun, Duan Chuanzhi, Si Weixin

Intra-operative endovascular treatment of intracranial aneurysms calls for rapid and reliable hemodynamic quantification. Angiographic Parametric Imaging (API) is derived from DSA time-density curves (TDCs), yet irregular frame rates and dose-reduction gaps often make conventional pixel-wise Gamma-variate fitting unstable and computationally expensive. We propose PIC-Former, a physics-inspired continuous transformer that reconstructs a physiologically plausible TDC function ĉ(t) from irregular DSA samples. PIC-Former uses Δt-aware causal attention for inter-frame gaps and physics-inspired regularization to encourage valid contrast arrival, peak, and washout dynamics. On 52 paired DSA-4D Flow MRI cases and an independent 11-case CFD cohort, QDSA velocities computed from the reconstructed TDCs achieve relative MAEs of 3.1% and 2.01%, respectively, outperforming Gamma-variate fitting (11.8% on DSA-4D Flow MRI). PIC-Former generates full-field parametric maps for a complete DSA run in 12.45 s and remains stable under a 20% frame-drop stress test. In a retrospective intra-operative risk stratification study, the proposed pipeline improves stratification accuracy from 87.5% to 95.0% compared to the Gamma-variate-based baseline. The code and reproducibility test set is available at https://github.com/LY-SUSTech/PIC-Former.git.

Subject: MICCAI.2026


#10 Anatomy-Structured Hierarchical MIL for Weakly-Supervised Thoracic Disease Detection in Chest X-Rays [PDF] [Copy] [Kimi] [REL]

Authors: Kim Jeongin, Ahn Sohyun, Kang Seo Young, Sung Jaeyi, Kim Soomin, Cho Sungho, Lee Rena, Kim Kwanchang, Noh Junhyug, Kim Jeongin, Ahn Sohyun, Kang Seo Young, Sung Jaeyi, Kim Soomin, Cho Sungho, Lee Rena, Kim Kwanchang, Noh Junhyug

Weakly-supervised thoracic disease detection in chest X-rays (CXR) is challenging due to subtle appearances and complex anatomical overlap, motivating anatomy-aware modeling for improved localization. However, prior anatomy-aware methods typically rely on coarse region proxies or static spatial priors, which may restrict dynamic instance discovery and limit precise localization of small abnormalities. We propose Anatomy-Structured Hierarchical Multiple Instance Learning (ASH-MIL), a framework that introduces parallel anatomy-structured observation branches (cardiac, pulmonary, and agnostic) combined with hierarchical MIL aggregation. Anatomical priors are injected as soft spatial biases into decoder cross-attention, enabling anatomically grounded evidence maps without disease bounding-box supervision. Instance localization is derived directly from MIL-weighted cross-attention maps. Experiments on CXR8 and cross-domain MIMIC-CXR held-out sets demonstrate consistent improvements over prior weakly-supervised and anatomy-aware approaches, particularly under stricter localization criteria. Our code is available at https://github.com/jn-kim/ash-mil.

Subject: MICCAI.2026


#11 CTTok: Voxel-Abulary for Autoregressive 3D CT Volume Generation with Large Language Models [PDF] [Copy] [Kimi] [REL]

Authors: Wang Jiayi, Reynaud Hadrien, Dombrowski Mischa, Wang Xiaoliang, Hamamci Ibrahim E., Shit Suprosanna, Er Sezgin, Menze Bjoern H., Kainz Bernhard, Wang Jiayi, Reynaud Hadrien, Dombrowski Mischa, Wang Xiaoliang, Hamamci Ibrahim E., Shit Suprosanna, Er Sezgin, Menze Bjoern H., Kainz Bernhard

Existing text-to-imaging generation methods inherit diffusion architectures from natural image synthesis, relying on explicit cross-attention to condition generation on text at every step. We hypothesize that this design is unnecessarily complex for medical tomographic volumes: unlike natural scenes, human anatomy exhibits strong structural regularity, with organs occupying predictable locations and varying far less across individuals than open-ended visual content. Given a text-aligned tokenization, sequential modeling alone should therefore suffice for text-conditional anatomical generation, without dedicated cross-attention layers. We propose CTTok, a discrete autoregressive approach that extends the vocabulary of a pre-trained language model with anatomical tokens representing 3D CT patches. Text conditioning is handled by the language model’s existing causal attention over these anatomically grounded tokens. Because the synthesized volume is itself constructed from such tokens, any mild boundary artifacts from patch-level discretization carry no semantic ambiguity and can be removed by a single-pass, unconditional GAN refinement with no access to the original prompt, since all text-to-anatomy correspondence is already resolved at the token level. CTTok outperforms state-of-the-art diffusion and flow-matching baselines in diversity, image quality, and text-image alignment, at a fraction of the training and inference compute. Source code and model weights: \url{https://github.com/WongJiayi/CTTok}.

Subject: MICCAI.2026


#12 SIRA: Reasoning-Aware Surgical Instrument Segmentation via Query-Anchored Alignment [PDF] [Copy] [Kimi] [REL]

Authors: Zhang Zhibo, Wang Qijie, Yan Zengqiang, Zhang Zhibo, Wang Qijie, Yan Zengqiang

Surgical instrument segmentation (SIS) plays a critical role in robotic assistance and surgical workflow analysis. However, most existing SIS methods formulate segmentation as a category-driven localization problem, limiting their ability to capture procedural context and task-dependent semantics in surgical workflows. We introduce Reasoning-Aware Surgical Instrument Segmentation (RA-SIS), a task formulation that frames segmentation as query-conditioned inference under surgical context. To benchmark this setting, we construct SurgRS, a surgical reasoning segmentation dataset consisting of 41,000 image–text pairs, which aligns instance-level masks with structured query–answer supervision to enable semantic grounding at the pixel level. Based on SurgRS, we propose Surgical Instrument Reasoning and segmentation Assistant (SIRA), a multimodal framework that disentangles target-level and query-level semantics and integrates them with visual features through query-anchored dual alignment. By aligning query semantics with spatial features and segmentation prompts, SIRA enhances semantic-visual consistency in mask prediction. Extensive experiments on SurgRS demonstrate improvements over existing reasoning-aware baselines. Code is available at https://github.com/linxir226/SIRA.

Subject: MICCAI.2026


#13 TraceCXR: Verifiable Chest X-ray Reports via Evidence-Addressable Graph Prompting [PDF] [Copy] [Kimi] [REL]

Authors: Dong Hang, Sun Chao, Yan Hao, Hu Wei, Du Bo, Dong Hang, Sun Chao, Yan Hao, Hu Wei, Du Bo

Automatic chest X-ray report generation can assist clinical reading, but vision–language models may hallucinate findings or miss subtle cues, and their localized evidence is often hard to inspect. We propose TraceCXR, an evidence-traceable framework that links generated clinical statements to concept-level node traces and localized visual evidence maps. TraceCXR constructs ProtoGraph, which makes clinical concepts evidence-addressable by associating them with disease-aware visual prototypes, so retrieved concepts can be traced back to candidate image evidence. Building on ProtoGraph, EviLoop Prompting retrieves patient-specific graph context from the image and grounds the retrieved cues to localized visual evidence; confidence-gated residual injection suppresses unreliable graph prompts. We further introduce Process-Supervised Node Planning to regularize clinically relevant node selection and concentrated evidence attribution during training. Experiments on IU X-ray, MIMIC-CXR, and CheXpert Plus improve report quality and clinical correctness, with stronger evidence-centric retrieval and competitive grounding and transfer performance. Code: https://github.com/Blaise-H02/TraceCXR.

Subject: MICCAI.2026


#14 Trustworthy Endoscopic Super-Resolution [PDF] [Copy] [Kimi] [REL]

Authors: Silva-Rodríguez Julio, Konukoglu Ender, Silva-Rodríguez Julio, Konukoglu Ender

Super-resolution (SR) models are attracting growing interest for enhancing minimally invasive surgery and diagnostic videos under hardware constraints. However, valid concerns remain regarding the introduction of hallucinated structures and amplified noise, limiting their reliability in safety-critical settings. We propose a direct and practical framework to make SR systems more trustworthy by identifying where reconstructions are likely to fail. Our approach integrates a lightweight error-prediction network that operates on intermediate representations to estimate pixel-wise reconstruction error. The module is computationally efficient and low-latency, making it suitable for real-time deployment. We convert these predictions into operational failure decisions by constructing Conformal Failure Masks (CFM), which localize regions where the SR output should not be trusted. Built on conformal risk control principles, our method provides theoretical guarantees for controlling both the tolerated error limit and the miscoverage in detected failures. We evaluate our approach on image and video SR, demonstrating its effectiveness in detecting unreliable reconstructions in endoscopic and robotic surgery settings. To our knowledge, this is the first study to provide a model-agnostic, theoretically grounded approach to improving the safety of real-time endoscopic image SR. Code is available: https://github.com/jusiro/Endoscopic-CFM

Subject: MICCAI.2026


#15 Real-Time Hardware-Free HIFU Interference Suppression via Teacher-Student Diffusion Framework [PDF] [Copy] [Kimi] [REL]

Authors: Cai Dejia, Abdollahi Ali, Wang Xi, Yang Kun, Guo Zhaohui, Zhou Xiaowei, Chen Hao, Cai Dejia, Abdollahi Ali, Wang Xi, Yang Kun, Guo Zhaohui, Zhou Xiaowei, Chen Hao

High-Intensity Focused Ultrasound (HIFU) is a non-invasive therapy, yet its safety is often degraded by severe acoustic interference during continuous ultrasound guidance. Conventional HIFU interference suppression methods heavily rely on proprietary raw Radio-Frequency (RF) data or complex hardware synchronization, limiting their clinical utility and preventing real-time implementation. To address this limitation, we propose Manifold-Constrained Hyper-Connections Diffusion (mHC-Diff), an image-domain diffusion framework for real-time interference suppression without specialized hardware synchronization, disentangling complex interference from anatomical structures while ensuring high reconstruction fidelity. To achieve clinical real-time application, our approach employs a two-stage strategy: (i) anatomy-aware prior acquisition, where a diffusion model is trained with multi-step UNet as a highfidelity Teacher; and (ii) efficiency distillation, where this prior is distilled into a one-step Student via knowledge distillation to achieve real-time throughput. Extensive validation on a clinically representative dataset across diverse therapeutic scenarios shows that mHC-Diff achieves superior restoration (26.65 dB PSNR), while enabling real-time inference (~20 FPS) on a single NVIDIA RTX 4090, providing a ~6.8x speedup over iterative diffusion baselines (e.g., HIFU-Diff). By eliminating the requirement for specialized hardware synchronization and proprietary RF access, this image-domain framework ensures compatibility and facilitates real-time interference suppression during ultrasound-guided HIFU interventions. The code will be released upon acceptance.

Subject: MICCAI.2026


#16 Representation Learning with Distance-Residualized Functional Connectivity for Cortical Parcellation from rs-fMRI [PDF] [Copy] [Kimi] [REL]

Authors: Zhu Jianfei, Wei Baichun, Liu Shaohui, Zhu Haiqi, Jiang Feng, Yi Chunzhi, Zhu Jianfei, Wei Baichun, Liu Shaohui, Zhu Haiqi, Jiang Feng, Yi Chunzhi

Functional parcellation of the cerebral cortex from resting-state functional magnetic resonance imaging (rs-fMRI) is a fundamental step for large-scale brain network analysis. Learning functional parcellation from rs-fMRI is essentially a dimensionality reduction of vertex-wise functional connectivity (FC), where previous studies replied linear dimensionality reduction techniques, failing to capture the nonlinearity nature of brain functional organization. In this work, we propose a self-supervised representation learning framework for whole-cortical, group-level parcellation to release the linear constraints in feature extraction. Specifically, we first decouple FC strength from spatial proximity via distance-based residualization, enabling the learning of functionally meaningful features beyond geodesic constraints. A variational autoencoder (VAE) is then employed to learn compact and uncertainty-aware FC representations, augmented with a neighborhood-based contrastive objective to explicitly promote local functional coherence. Spatial continuity is further encouraged through Laplacian positional encoding prior to clustering. Experiments demonstrate that the proposed method consistently achieves better performance than widely used cortical parcellations, achieving improvement in functional homogeneity, functional boundary alignment, clustering quality, and consistency with task-evoked activation patterns.

Subject: MICCAI.2026


#17 Multi-parametric MRI for Contrast Agent Free Breast DCE-MRI Synthesis [PDF] [Copy] [Kimi] [REL]

Authors: Yang Zhikai, Zhang Cristina, Chen Huijian, He Muzhen, Moreno Rodrigo, Yang Zhikai, Zhang Cristina, Chen Huijian, He Muzhen, Moreno Rodrigo

Dynamic contrast-enhanced Magnetic Resonance Imaging (DCE-MRI) is a critical tool for breast cancer detection and diagnosis. Yet, its reliance on contrast agents presents risks, increases cost, and is contraindicated for certain patients, necessitating a contrast-agent-free alternative. In this paper, we use a conditional Generative Adversarial Network (GAN) with a multi-scale enhancement consistency loss for contrast agent-free breast DCE-MRI synthesis from a multi-parametric MRI input (T1w, multi-b-value DWI, and ADC maps). The proposed loss explicitly enforces consistency of the predicted enhancement maps across multiple scales. The model generates multiple DCE post-contrast phases simultaneously, reducing inference time and computational overhead. The proposed framework achieved the best performance among previous methods on public and in-house datasets. The code is available at https://github.com/minnelab/DCE-MRI-Syn.

Subject: MICCAI.2026


#18 CrownFusion: 3D Dental Crown Generation Using Geometry Images and Latent Diffusion [PDF] [Copy] [Kimi] [REL]

Authors: Ye Johan Ziruo, Yan Xingguang, Hauberg Søren, Chang Angel X., Zhang Hao, Søndergaard Peter Lempel, Ye Johan Ziruo, Yan Xingguang, Hauberg Søren, Chang Angel X., Zhang Hao, Søndergaard Peter Lempel

We introduce CrownFusion, a latent diffusion model to generate 3D dental crown designs from geometry images, a 2D-grid encoding of surface coordinates. Departing from traditional approaches that use point clouds, our “CrownImage” representation not only unlocks the full power of image-based generative architectures, but also preserves fine geometric detail by operating at far higher spatial resolutions than point clouds to alleviate smoothing artifacts in the 3D generation. Conditioned on jointly learned embeddings of the surrounding dentition, our probabilistic diffusion model produces anatomically plausible crowns with morphological variations tailored to individual patient cases. To accommodate multiple valid designs, we propose an occlusal fit metric based on proximity to the antagonist teeth. We evaluate our method through quantitative experiments and qualitative studies involving expert testimonials. We also release the FDI 16 Crown Dataset, comprising 11,513 annotated intra-oral scans, to support future research in dental CAD.

Subject: MICCAI.2026


#19 RePCM: Region-Specific and Phenotype-Adaptive Bi-ventricular Cardiac Motion Synthesis [PDF] [Copy] [Kimi] [REL]

Authors: Yang Xuan, Yuan Xiaohan, Li Hao, Chen Lingyu, Liu Yanan, Li Lei, Yang Xuan, Yuan Xiaohan, Li Hao, Chen Lingyu, Liu Yanan, Li Lei

Cardiac motion over a cardiac cycle is crucial for quantifying regional function and is strongly affected by cardiovascular diseases. Since temporally dense cardiac mesh sequences are difficult to obtain in practice, we focus on leveraging the more accessible end-diastolic frame to infer a full-cycle sequence. Due to strong regional and disease-specific differences, traditional methods often oversmooth the data by relying on generative models that are optimized for global patterns. To address this problem, we propose Region-Aware and Phenotype-Adaptive Bi-Ventricular Cardiac Motion Synthesis (RePCM) for single frame Bi-ventricular mesh motion completion. In Stage I, a reconstruction network learns vertex-wise motion descriptors and clustering yields a data-driven functional partition, providing an explicit motion derived region structure. In Stage II, a Region-Specific Injection Module enforces masked, synchronized region exchange within a conditional VAE, preserving localized specific dynamics and restricting cross-region mixing. A Phenotype-Adaptive Mixture-of-Experts prior conditioned on ED shape uses anatomy-guided cues to model latent motion trends and capture inter-disease variability. Experiments on three datasets covering different cardiovascular diseases show consistent gains in geometric and functional metrics and improved preservation of region specific dynamics. Our code is available at https://github.com/yangxuan-nus/RePCM.

Subject: MICCAI.2026


#20 FedSwitch: From Standard to Balanced Aggregation via Private Histogram Estimation in Federated Medical Image Classification [PDF] [Copy] [Kimi] [REL]

Authors: Cacace Paolo, Taiello Riccardo, Protani Andrea, Molina Van den Bosch Marc, Santos Diogo Reis, Brutti Pierpaolo, Serio Luigi, Cacace Paolo, Taiello Riccardo, Protani Andrea, Molina Van den Bosch Marc, Santos Diogo Reis, Brutti Pierpaolo, Serio Luigi

Federated Learning (FL) enables collaborative model training across institutions without sharing raw data. Performance degrades, however, when label distributions are imbalanced across sites, a common scenario in medical settings where disease prevalence could vary drastically across institutions. Existing distribution-aware methods rely on label distribution statistics that clients cannot share in real-world federated settings due to privacy and regulatory constraints. We propose FedSwitch, a model-agnostic FL framework that estimates the global label histogram, under client-level epsilon-differential privacy, and uses the estimate to switch from FedAvg to distribution-aware aggregation. Each client adds partial discrete Laplace noise to its bounded label counts and secret-shares the result via Shamir’s scheme, designated decryptors reconstruct the noisy aggregate without ever observing individual contributions. The optimal switching round can be estimated via a closed-form variance decomposition that accounts for bounded contributions, client sampling, and differential-privacy noise components. We evaluate FedSwitch on medical imaging benchmarks under realistic heterogeneous settings, demonstrating consistent improvements over standard FL aggregation with negligible communication overhead.

Subject: MICCAI.2026


#21 MedEnv: Scaling Multimodal Virtual Medical Environments for Long-Horizon Diagnosis [PDF] [Copy] [Kimi] [REL]

Authors: Fan Zhiting, Ren Xuxiang, Wang Yuan, Chen Ruizhe, Meng Zijie, Hu Keli, Liu Jiaxiang, Wu Jian, Zhai Weiqi, Xu Hongxia, Liu Zuozhu, Fan Zhiting, Ren Xuxiang, Wang Yuan, Chen Ruizhe, Meng Zijie, Hu Keli, Liu Jiaxiang, Wu Jian, Zhai Weiqi, Xu Hongxia, Liu Zuozhu

Real-world clinical care is a long-horizon decision-making process in which patient state and evidence evolve over multiple turns. In contrast, static, single-turn medical VQA is insufficient for training and evaluating interactive medical agents. To bridge this gap, we introduce DiagWorld, a scalable multimodal virtual medical environment based on large-scale real-world clinical interactions—an EHR-grounded simulator that supports multi-turn interviewing and tool-augmented evidence acquisition under realistic clinical workflows. Built from real patient records in MIMIC-IV, our pipeline transforms longitudinal EHR data—admissions, laboratory tests, imaging, and clinical notes—into 500K+ interactive trajectories, enabling large-scale training and evaluation. Unlike prior EHR-only simulators, we integrate external disease-level medical databases for consistency-aware completion and constraint validation, reducing diagnostic confounding from missing or noisy EHR signals and enabling scaling in patient volume and interaction horizon without amplifying artifacts. Leveraging this environment, we train a diagnostic agent with confidence-based process supervision: rather than sparse terminal rewards or manually annotated key steps, we provide intermediate rewards tied to increases in confidence toward the correct diagnosis, reinforcing informative questioning, retrieval, test ordering, and CXR analysis. This shaping improves sample efficiency and stabilizes learning in large-scale interactive training.

Subject: MICCAI.2026


#22 SynHydro: Biomechanics-Driven Domain Randomization for Hydrocephalus-Agnostic, Generalizable Tissue-Ventricle Segmentation Across Age, Modality, and Resolution [PDF] [Copy] [Kimi] [REL]

Authors: Ren Zehua, Zhai Jingyan, Wang Haifeng, Bo Yixuan, Wang Fan, Zhao Cailei, Lian Chunfeng, Ren Zehua, Zhai Jingyan, Wang Haifeng, Bo Yixuan, Wang Fan, Zhao Cailei, Lian Chunfeng

Early diagnosis and quantitative assessment of pediatric hydrocephalus are critical, yet automated segmentation remains challenging due to severe anatomical deformation and substantial appearance variation across early neurodevelopment, imaging modalities, and resolutions. Current deep learning approaches are often limited to ventricle-only segmentation and tailored to specific acquisition protocols, hindering their clinical utility for comprehensive joint tissue analysis and longitudinal evaluation. We propose SynHydro, a novel biomechanics-driven simulation and domain-randomized augmentation framework. SynHydro synthesizes diverse hydrocephalus anatomies from readily available healthy tissue labels and generates corresponding MRIs with randomized image characteristics. This creates a large-scale, synthetic training dataset encompassing a wide spectrum of disease severity, MRI contrast, noise, resolution, and artifacts for joint tissue-ventricle segmentation. Extensive experiments on in-house and public datasets demonstrate that a segmentation model trained exclusively on this pure synthetic data achieves performance comparable to models trained with real hydrocephalus annotations. More importantly, SynHydro-trained models exhibit superior cross-modality and cross-resolution robustness, and reliably preserve pre-/post-operative volumetric trends. By eliminating the dependency on labor-intensive pathological annotations, SynHydro provides a powerful and practical solution for building robust, generalizable segmentation tools for pediatric hydrocephalus monitoring across the clinical imaging landscape.

Subject: MICCAI.2026


#23 A Neurosymbolic Framework for Interpretable Skeleton-Based Seizure Detection via Concept-Driven Logical Reasoning [PDF] [Copy] [Kimi] [REL]

Authors: Ilyas Talha, Mehta Deval, Ge Zongyuan, Ilyas Talha, Mehta Deval, Ge Zongyuan

Video-based seizure detection is essential for the management of epilepsy patients, offering a non-invasive complement to electroencephalography. While several deep learning approaches have been developed for video-based seizure detection, none are inherently interpretable, limiting their adoption and translation into clinical practice. We present, to our knowledge, the first exploration of a neurosymbolic framework for video-based seizure detection that directly addresses this gap. Our approach (1) extracts patient-centric skeleton sequences from epilepsy monitoring units via a prompt-guided foundation model, (2) predicts seizure semiology clinically grounded spatio-temporal concepts, and (3) composes them via differentiable logic into interpretable Boolean rules with auditable contributions. Furthermore, to mitigate false positives arising from the traditional binary formulation (seizure vs.\ non-seizure), we sub-classify non-seizure segments into clinically relevant normal activities, providing the model with fine-grained discriminative supervision. Evaluated on two public seizure video benchmarks, our framework achieves 89.78% sensitivity with 0.06 false detections per hour on SAHZU and 85.27% / 0.09 on IEEE, while producing complete three-level interpretability: every prediction decomposes into which motor primitives were detected, how they were logically composed, and how much each rule contributed to the clinical decision. We publicly release all annotations, extracted pose sequences, our data pipeline and code. (https://gitfront.io/r/luffy/rELWxdnvy6q2/CDSD/)

Subject: MICCAI.2026


#24 BioGait-VLM: A Tri-Modal Vision–Language–Biomechanics Framework for Interpretable Clinical Gait Assessment [PDF] [Copy] [Kimi] [REL]

Authors: Chen Erdong, Ji Yuyang, Greenberg Jacob K., Steel Benjamin, Arkam Faraz, Lewis Abigail, Singh Pranay, Liu Feng, Chen Erdong, Ji Yuyang, Greenberg Jacob K., Steel Benjamin, Arkam Faraz, Lewis Abigail, Singh Pranay, Liu Feng

Video-based Clinical Gait Analysis often suffers from poor generalization as models overfit environmental biases instead of capturing pathological motion. To address this, we propose BioGait-VLM, a tri-modal Vision-Language-Biomechanics framework for interpretable clinical gait assessment. Unlike standard video encoders, our architecture incorporates a Temporal Evidence Distillation branch to capture rhythmic dynamics and a Biomechanical Tokenization branch that projects 3D skeleton sequences into language-aligned semantic tokens. This enables the model to explicitly reason about joint mechanics independent of visual shortcuts. To ensure rigorous benchmarking, we augment the public GAVD dataset with a high-fidelity Degenerative Cervical Myelopathy (DCM) cohort to form a unified 8-class taxonomy, establishing a strict subject-disjoint protocol to prevent data leakage. Under this setting, BioGait-VLM achieves state-of-the-art recognition accuracy. Furthermore, a blinded expert study confirms that biomechanical tokens significantly improve clinical plausibility and evidence grounding, offering a path toward transparent, privacy-preserving gait assessment.

Subject: MICCAI.2026


#25 VM-NeXT UNet: Synergizing ConvNeXT and Visual Mamba for Robust Medical Image Segmentation [PDF] [Copy] [Kimi] [REL]

Authors: Cao Boxuan, Jiang Peilin, Cao Boxuan, Jiang Peilin

In recent medical image analysis, Convolutional Neural Networks (CNNs) and State Space Models (SSMs) have set benchmarks in segmentation tasks. While CNNs excel in capturing local fine-grained features, SSMs achieve remarkable global context understanding with linear complexity. However, Mamba-UNet, a pioneering pure SSM-based model, still exhibits deficiencies in feature representation, fusion efficiency, and spatial detail reconstruction. To address these limitations, we propose VM-NeXT UNet, an improved architecture that synergizes ConvNeXT with VSS Blocks.VM-NeXT UNet introduces four key optimizations: (1) a dual-encoder parallel structure for comprehensive feature extraction; (2) a channel-spatial attention gating module in skip connections for adaptive feature screening; (3) multi-scale convolution layers in the decoder to preserve spatial details; and (4) a combined FocalLoss and DiceLoss strategy to focus on hard samples. Experiments on the Synapse and ACDC datasets yielded Dice scores of 86.21% and 92.42%, respectively. The results demonstrate that VM-NeXT UNet significantly outperforms the original Mamba-UNet and achieves competitive performance against state-of-the-art methods, highlighting its potential for reliable clinical deployment.

Subject: MICCAI.2026