Total: 1
Prosodic cues such as rhythm and lexical stress affect intelligibility; they stand out as signals that guide perception. While these cues are shared across many languages, lexical stress is realized differently in Arabic and English, making stress detection challenging in multilingual and transfer settings. We propose a framework that uses self supervised speech representations as fixed features and a two stage classifier: a syllable level Pre-net and a word context Post-net that models inter syllable dependencies. We compare monolingual training, joint Arabic English multilingual training, cross lingual transfer using different self-supervised learning (SSL) model's feature vectors input into Pre-net DNN and Post-net TDNN models. Multilingual training retains near monolingual accuracy of 97% and 90% on English and Arabic respectively, whereas cross-lingual transfer is direction-dependent; the Post-net improves robustness and multilingual pre-training yields the strongest transfer.