Total: 1
The spatiotemporal dynamics and coordination of laryngeal movements remain incompletely characterized. This study introduces a larynx segmentation and analysis pipeline for mid-sagittal speech production real-time MRI using Mask2Former, combining supervised learning with semi-supervised refinement. Our results suggest sufficient segmentation performance with ~33-79 annotations per participant (~25-60% of 794 training samples / 6 participants) using a 5% marginal gain threshold. Additional annotations beyond this yield diminishing returns. Semi-supervised learning achieves modest improvements but sometimes degrades performance and cannot reach the fully supervised upper bound. Results from a sample phonetic study on Mandarin tones demonstrate the capability of mid-sagittal speech production real-time MRI to capture spatiotemporal laryngeal dynamics, including both intrinsic and extrinsic pitch control and laryngeal constriction mechanisms. The proposed pipeline opens new avenues for studying laryngeal behaviors in linguistic contrasts such as voicing, tone, and phonation types using real-time MRI. Code repository: https://github.com/pkuzyb/ larynx_segmentation.