Total: 1
In recent medical image analysis, Convolutional Neural Networks (CNNs) and State Space Models (SSMs) have set benchmarks in segmentation tasks. While CNNs excel in capturing local fine-grained features, SSMs achieve remarkable global context understanding with linear complexity. However, Mamba-UNet, a pioneering pure SSM-based model, still exhibits deficiencies in feature representation, fusion efficiency, and spatial detail reconstruction. To address these limitations, we propose VM-NeXT UNet, an improved architecture that synergizes ConvNeXT with VSS Blocks.VM-NeXT UNet introduces four key optimizations: (1) a dual-encoder parallel structure for comprehensive feature extraction; (2) a channel-spatial attention gating module in skip connections for adaptive feature screening; (3) multi-scale convolution layers in the decoder to preserve spatial details; and (4) a combined FocalLoss and DiceLoss strategy to focus on hard samples. Experiments on the Synapse and ACDC datasets yielded Dice scores of 86.21% and 92.42%, respectively. The results demonstrate that VM-NeXT UNet significantly outperforms the original Mamba-UNet and achieves competitive performance against state-of-the-art methods, highlighting its potential for reliable clinical deployment.