Total: 1
Accurate emergency severity assessment fundamentally dictates patient survival and resource allocation. However, existing automated triage pipelines rely exclusively on tabular data and clinical text, neglecting the visual semantics of patient distress. A paramount barrier to anatomically informed reasoning is the absolute absence of visual modalities in public cohorts like the MIMIC-IV-ED dataset. To overcome this limitation, this study introduces MedTriage-LM, an anatomically grounded Multimodal Large Language Model (MLLM) that algorithmically injects visual priors without requiring real patient images. The architecture maps latent clinical representations onto a canonical human template via a weakly supervised Gaussian field generator, yielding continuous, interpretable Visual Phenotype Maps (VPMs). Synergizing these spatial representations with clinical embeddings via cross-attention shifts the paradigm from numerical risk scoring toward actionable three-class triage instruction prediction: Life-saving, High-Risk Assessment, and Resource Estimation. MedTriage-LM attains state-of-the-art performance among fine-tuned open-source models and approaches the capabilities of proprietary MLLMs, while additionally providing explicit, anatomically grounded interpretability. Crucially, explicit spatial grounding empowers an efficient 8B-parameter model to approach large proprietary multimodal foundation models, generating spatially grounded textual rationales that improve clinical trustworthiness.