0231-Paper2284@2026@MICCAI

Total: 1

#1 CTTok: Voxel-Abulary for Autoregressive 3D CT Volume Generation with Large Language Models [PDF] [Copy] [Kimi] [REL]

Authors: Wang Jiayi, Reynaud Hadrien, Dombrowski Mischa, Wang Xiaoliang, Hamamci Ibrahim E., Shit Suprosanna, Er Sezgin, Menze Bjoern H., Kainz Bernhard, Wang Jiayi, Reynaud Hadrien, Dombrowski Mischa, Wang Xiaoliang, Hamamci Ibrahim E., Shit Suprosanna, Er Sezgin, Menze Bjoern H., Kainz Bernhard

Existing text-to-imaging generation methods inherit diffusion architectures from natural image synthesis, relying on explicit cross-attention to condition generation on text at every step. We hypothesize that this design is unnecessarily complex for medical tomographic volumes: unlike natural scenes, human anatomy exhibits strong structural regularity, with organs occupying predictable locations and varying far less across individuals than open-ended visual content. Given a text-aligned tokenization, sequential modeling alone should therefore suffice for text-conditional anatomical generation, without dedicated cross-attention layers. We propose CTTok, a discrete autoregressive approach that extends the vocabulary of a pre-trained language model with anatomical tokens representing 3D CT patches. Text conditioning is handled by the language model’s existing causal attention over these anatomically grounded tokens. Because the synthesized volume is itself constructed from such tokens, any mild boundary artifacts from patch-level discretization carry no semantic ambiguity and can be removed by a single-pass, unconditional GAN refinement with no access to the original prompt, since all text-to-anatomy correspondence is already resolved at the token level. CTTok outperforms state-of-the-art diffusion and flow-matching baselines in diversity, image quality, and text-image alignment, at a fraction of the training and inference compute. Source code and model weights: \url{https://github.com/WongJiayi/CTTok}.

Subject: MICCAI.2026