INTERSPEECH.2026 - Analysis and Assessment

| Total: 166

#1 Membership Inference Attacks against Large Audio Language Models [PDF] [Copy] [Kimi] [REL]

Authors: Jia-Kai Dong, Yu-Xiang Lin, Hung-yi Lee

We present the first systematic membership inference attack (MIA) evaluation of LALMs. Using Multi-modal Blind Baselines based on textual, spectral and prosodic features, we demonstrate that common audio datasets exhibit near-perfect train/test separability (AUC ≈ 1.0) even without model inference, thus MIA may primarily detect distribution shift. We therefore introduce a blind-baseline protocol to control for this confound. Under this protocol, we identify that the distribution-matched datasets enable reliable MIA evaluation without distribution-shift artifacts. We benchmark multiple MIA methods and conduct modality disentanglement experiments on these datasets. The results reveal that LALM memorization is cross-modal, arising only from binding a speaker's vocal identity with its text. These findings establish a principled standard for auditing LALMs beyond spurious correlations. Our codebase is available at https://github.com/snooow1029/ALM_MIA.

Subject: INTERSPEECH.2026 - Analysis and Assessment


#2 Listening Like a Judge: A Music-Aware Framework for Automatic Singing Performance Evaluation [PDF] [Copy] [Kimi] [REL]

Authors: Neelam Saini, Sourav Ghosh

Automatic singing quality assessment (SQA) requires evaluating lyrical correctness and musical fidelity while handling expressive variations. However, existing systems largely rely on either acoustic cues or lyric transcriptions exclusively, limiting holistic performance evaluation. Furthermore, their integration is non-trivial due to challenges in robust singing transcription amid melisma, vibrato, and tempo elasticity. To this end, we propose MUSICJUDGE, a modality-guided framework for automated SQA that performs block-aligned multimodal analysis by coupling lyric correctness with pitch-rhythm fidelity. It detects semantically meaningful lyric blocks using multi-signal matching that integrates semantic embeddings, lexical similarity, and phonetic alignment. To improve singing audio transcription, we introduce Modality-Guided LoRA for ASR fine-tuning. Experiments across datasets demonstrate strong agreement with human expert judgments and validate the generalizability of MUSICJUDGE.

Subject: INTERSPEECH.2026 - Analysis and Assessment


#3 ELSA: Acoustic Event-Level Semantic Alignment for Fine-Grained Reference-Free Text-to-Audio Evaluation [PDF] [Copy] [Kimi] [REL]

Authors: Shuntaro Suzuki, Kento Tokura, Daichi Yashima, Kanon Amemiya, Komei Sugiura, Shinnosuke Takamichi

Text-to-audio (TTA) generation, synthesizing audio from natural language, has been widely studied for its ability to capture precise user intent. To effectively advance TTA models, it is essential to reliably evaluate generated audio without relying on costly human subjective ratings, motivating the development of automatic evaluation metrics that correlate well with human judgments. While recent CLAP-based metrics provide practical reference-free solutions, their coarse-grained text–audio similarity matching often correlates poorly with human ratings. To address this, we propose ELSA, a reference-free evaluation metric for fine-grained text–audio alignment. ELSA decomposes generated audio guided by distinct acoustic events derived from the text query and assesses event-level alignment. Experiments across four TTA benchmarks show that ELSA reveals a higher correlation with human subjective ratings than prior metrics, highlighting its effectiveness for reliable TTA evaluation.

Subject: INTERSPEECH.2026 - Analysis and Assessment


#4 PrefSQA: Pairwise Preference Prediction for Speech Quality Assessment and the Critical Role of High Quality Datasets [PDF] [Copy] [Kimi] [REL]

Authors: Junyi Fan, Donald S. Williamson

Mean opinion scores (MOS) are widely used for speech quality assessment, yet scalar labels are sensitive to rater variability and listening test differences. This introduces labeling noise, which limits the reliability of MOS prediction. Preference prediction reduces this variability as listeners compare signals directly, producing cleaner labels. We study MOS-free preference prediction and propose PrefSQA, which incorporates uncertainty-aware logits, an impairment attention head, and a module based on non-matching-reference comparisons. We use and refine five datasets, including MOS-derived and low-noise simulated sets with matching and non-matching content, experiment with human preference sets, and test on unseen data. Experiments show small improvements on MOS-derived data, while other sets reveal clear improvement over the baselines, highlighting the value of high-quality preference data and demonstrating the effectiveness of the proposed method.

Subject: INTERSPEECH.2026 - Analysis and Assessment


#5 UG-Bench: A Comprehensive Benchmark for Evaluating Large Audio-Language Models [PDF] [Copy] [Kimi] [REL]

Authors: Jiaming Zhou, Haoqin Sun, Hui Wang, Jinghua Zhao, Yuhang Jia, Shiyao Wang, Enzhi Wang, Shiwan Zhao, Yong Qin

Evaluating Large Audio-Language Models (LALMs) is challenging due to the diverse speech and audio tasks. Existing benchmarks often lack a unified framework, focusing either on understanding or generation. We introduce UG-Bench, a comprehensive benchmark systematically assessing LALMs across four competencies: speech perception, audio perception, speech generation, and spoken language understanding. Its decoupled design enables standardized comparisons and supports new tasks. Evaluating 11 LALMs and 5 specialized speech generation models, we find significant gaps in instruction following and generation quality, underscoring the need for better semantic understanding. UG-Bench offers a unified, scalable evaluation framework to advance LALM research.

Subject: INTERSPEECH.2026 - Analysis and Assessment


#6 A Fine-Grained Acoustically-Aware Pre-training Encoder for Speech Quality Assessment [PDF] [Copy] [Kimi] [REL]

Authors: Subrina Sultana, Donald S. Williamson

Self-supervised learning (SSL) has become popular in speech processing because it generalizes well across downstream tasks. However, many SSL methods focus on capturing long-term contextual and speaker information, often making their representations invariant to background acoustics, which is an issue for speech quality assessment as it depends heavily on non-speech and fine-grained acoustic cues. These models also tend to be parameter heavy, limiting their use on resource constrained devices. In this work, we introduce an encoder that incorporates acoustic detail by combining local spectral-temporal modeling, a frame wise spectral relationship aggregator, and explicit noise and reverberation information alongside speech content. Using this pre-training framework for speech quality assessment across multiple datasets, we show that the encoder effectively extracts fine-grained acoustic features and achieves performance comparable to much larger models.

Subject: INTERSPEECH.2026 - Analysis and Assessment


#7 VoxEffects: A Speech-Oriented Audio Effects Dataset and Benchmark [PDF] [Copy] [Kimi] [REL]

Authors: Zhe Zhang, Yigitcan Özer, Junichi Yamagishi

Speech audio in the wild is often processed by post-production effects, but existing speech datasets rarely provide precise annotations of effects and parameters, limiting systematic study. We introduce VoxEffects, a speech audio effects dataset that pairs produced speech with exact effect-chain supervision at multiple granularities. VoxEffects supports speech-oriented audio effect identification: given a produced waveform, infer which effects are present and how they are applied. Built from minimally edited clean speech, it provides an extensible rendering pipeline for both offline synthesis and on-the-fly rendering for efficient training and evaluation. The audio effect identification benchmark includes effect presence detection, preset classification, number-of-active-effects prediction, and intensity prediction, with a robustness protocol covering capture-side and platform-side degradations. We provide an AudioMAE-based multi-task baseline and analyses of domain shift, robustness, input duration, and gender fairness.

Subject: INTERSPEECH.2026 - Analysis and Assessment


#8 Evaluating Objective Speech Quality Metrics for Neural Audio Codecs [PDF] [Copy] [Kimi] [REL]

Authors: Luca A. Lanzendöerfer, Florian Grötschla, Roger Wattenhofer

Neural audio codecs have gained recent popularity for their use in generative modeling as they offer high-fidelity audio reconstruction at low bitrates. While human listening studies remain the gold standard for assessing perceptual quality, they are time-consuming and impractical. In this work, we examine the reliability of existing objective quality metrics in assessing the performance of recent neural audio codecs. To this end, we conduct a MUSHRA listening test on high-fidelity speech signals and analyze the correlation between subjective scores and widely used objective metrics. Our results show that, while some metrics align well with human perception, others struggle to capture relevant distortions. Our findings provide practical guidance for selecting appropriate evaluation metrics when using neural audio codecs for speech.

Subject: INTERSPEECH.2026 - Analysis and Assessment


#9 Calibration-Reasoning Framework for Descriptive Speech Quality Assessment [PDF] [Copy] [Kimi] [REL]

Authors: Elizaveta Kostenok, Mathieu Salzmann, Milos Cernak

Explainable speech quality assessment requires moving beyond Mean Opinion Scores (MOS) to analyze underlying perceptual dimensions. To address this, we introduce a novel post-training method that tailors the foundational Audio Large Language Model for multidimensional reasoning, detection and classification of audio artifacts. First, a calibration stage aligns the model to predict predefined perceptual dimensions. Second, a reinforcement learning stage leverages Group Relative Policy Optimization (GRPO) with dimension-specific rewards to heavily enhance accuracy of descriptions and temporal localization of quality issues. With this approach we reach state-of-the-art results of 0.71 mean PCC score on the multidimensional Qual-iSpeech benchmark and 13% improvement in MOS prediction driven by RL-based reasoning. Furthermore, our fine-grained GRPO rewards substantially advance the model's ability to pinpoint and classify audio artifacts in time.

Subject: INTERSPEECH.2026 - Analysis and Assessment


#10 CAL-MOS: Bridging Layers with Adapters for Robust MOS Prediction Across Speech Foundation Models [PDF] [Copy] [Kimi] [REL]

Authors: Alef Iury Ferreira, Pedro Botelho, Fernanda Silva, Daniel Casanova, Rafael Faustino, Frederico Oliveira, Arlindo Galvão Filho, Anderson da Silva Soares

Speech Quality Assessment (SQA) is essential for modern speech technologies, and recent non-intrusive SQA predictors increasingly rely on Speech Foundation Models (SFMs). However, because SFMs expose representations from many layers, it remains unclear which depths are most informative for MOS prediction and how multi-layer information should be combined reliably across backbones and datasets. We benchmark ten SFMs on four MOS datasets under three regimes: full fine-tuning, last-layer probing with a frozen encoder, and naive cross-layer weighted aggregation. We find that the best layer is strongly backbone- and dataset-dependent, and that naive weighted fusion can be unstable across settings. We further evaluate a layer-calibrated aggregation variant that applies per-layer adapters before pooling, which improves the robustness of multi-layer fusion and narrows the gap to full fine-tuning while keeping the backbone frozen.

Subject: INTERSPEECH.2026 - Analysis and Assessment


#11 AnimeScore: A Preference-Based Dataset and Framework for Evaluating Anime-Like Speech Style [PDF] [Copy] [Kimi] [REL]

Authors: Joonyong Park, Jerry Li

Evaluating 'anime-like' voices currently relies on costly subjective judgments, yet no standardized objective metric exists. A key challenge is that anime-likeness, unlike naturalness, lacks a shared absolute scale, making conventional Mean Opinion Score (MOS) protocols unreliable. To address this gap, we propose AnimeScore, a preference-based framework for automatic anime-likeness evaluation via pairwise ranking. We collect 15,000 pairwise judgments from 187 evaluators with free-form descriptions, and acoustic analysis reveals that perceived anime-likeness is driven by controlled resonance shaping, prosodic continuity, and deliberate articulation rather than simple heuristics such as high pitch. We show that handcrafted acoustic features reach a 69.3% AUC ceiling, while SSL-based ranking models achieve up to 90.8% AUC, providing a practical metric that can also serve as a reward signal for preference-based optimization of generative speech models.

Subject: INTERSPEECH.2026 - Analysis and Assessment


#12 Exploiting Neural Audio Codec Latents for Adversarial Audio Attacks [PDF] [Copy] [Kimi] [REL]

Authors: Sameek Bhattacharya, Bharath Krishnamurthy, Ajita Rattani

Deep learning–based audio classification systems, including automatic speaker verification, are vulnerable to adversarial attacks. Realistic real-time threat assessment remains difficult because optimization-based methods, such as projected gradient descent (PGD) and Carlini–Wagner, require costly iterative updates in the high-dimensional waveform domain. Generative attacks allow single-shot synthesis but often introduce perceptible artifacts or depend on computationally intensive architectures, while diffusion and autoregressive approaches incur high inference latency. To address this gap, we propose a generative attack framework operating in the continuous latent space of a neural audio codec. A conditional generator synthesizes class-specific perturbations in a single forward pass and decodes them into adversarial waveforms. Our method achieves targeted attack success rates up to 99% with sub-7 ms inference, outperforming generative baselines while reducing latency by 24×.

Subject: INTERSPEECH.2026 - Analysis and Assessment


#13 AURA Score: A Metric for Holistic Audio Question Answering Evaluation [PDF] [Copy] [Kimi] [REL]

Authors: Satvik Dixit, Soham Deshmukh, Bhiksha Raj

Audio Question Answering (AQA) is a key task for evaluating Audio-Language Models (ALMs), yet assessing open-ended responses remains challenging. Existing metrics used for AQA such as BLEU, METEOR and BERTScore, mostly adapted from NLP and audio captioning, rely on surface similarity and fail to account for question context, reasoning, and partial correctness. To address the gap in literature, we make three contributions in this work. First, we introduce AQEval to enable systematic benchmarking of AQA metrics. It is the first bench-mark of its kind, consisting of 10k model responses annotated by multiple humans for their correctness and relevance. Second, we conduct a comprehensive analysis of existing AQA metrics on AQEval, highlighting weak correlation with human judgment, especially for longer answers. Third, we propose a new metric, AURA score, to better evaluate open-ended model responses. On AQEval, AURA achieves state-of-the-art correlation with human ratings, significantly outperforming all base-lines. Through this work, we aim to highlight the limitations of current AQA evaluation methods and motivate better metrics. We release both the AQEval benchmark and the AURA metric to support future research in holistic AQA evaluation.

Subject: INTERSPEECH.2026 - Analysis and Assessment


#14 Aligning Audio Captions with Human Preferences [PDF] [Copy] [Kimi] [REL]

Authors: Kartik Hegde, Rehana Mahfuz, Yinyi Guo, Erik Visser

Current audio captioning relies on supervised learning with paired audio-caption data, which is costly to curate and may not reflect human preferences in real-world scenarios. To address this, we propose a preference-aligned audio captioning framework based on Reinforcement Learning from Human Feedback (RLHF). To capture nuanced preferences, we train a Contrastive Language-Audio Pretraining (CLAP) based reward model using human-labeled pairwise preference data. This reward model is integrated into an RL framework to fine-tune any baseline captioning system without ground-truth annotations. Extensive human evaluations across multiple datasets show that our method produces captions preferred over baseline models, particularly when baselines fail to provide correct and natural captions. Furthermore, our framework achieves performance comparable to supervised approaches with ground-truth data, demonstrating effective alignment with human preferences and scalability in real-world use.

Subject: INTERSPEECH.2026 - Analysis and Assessment


#15 High-Precision Prosodic Boundary Anchors from Acoustic Cues under Weak Supervision [PDF] [Copy] [Kimi] [REL]

Authors: Hanyu Liao, Xiaoluan Liu

Prosodic boundary detection has traditionally relied on manual annotations such as ToBI labels, which can be resource-intensive and are not always available in large speech corpora. This paper presents a weakly supervised framework that derives high-confidence prosodic boundary anchors from acoustic cues, without depending on manually labeled prosodic boundaries. We adopted a hierarchical anchor construction procedure in which long pauses were first used to identify a conservative set of boundary candidates, and pitch reset and energy reduction were then used to refine these candidates into progressively stricter anchors. These anchors were designed to emphasize precision over coverage and served as reliable positive instances for positive-unlabeled (PU) learning. By framing prosodic boundary detection as a boundary strength estimation problem, we employed PU learning to infer continuous boundary strength scores over all candidate junctures under weak supervision. Experiments on a large-scale Japanese speech corpus demonstrated that meaningful prosodic boundary patterns can be recovered from acoustic cues using PU learning, thus providing an interpretable, data-efficient approach to prosodic boundary modeling.

Subject: INTERSPEECH.2026 - Analysis and Assessment


#16 ArtNet: A JEPA-Like Articulatory Predictive Framework for Robust Zero-Shot Phoneme Recognition [PDF] [Copy] [Kimi] [REL]

Authors: Zeqian Hu, Fuliang Weng, Shu Shang, Yaqian Zhou

Zero-shot cross-lingual phoneme recognition is often hindered by the fragility of direct acoustic-to-symbol mapping, which is susceptible to language-specific variations. Echoing joint-embedding predictive architecture (JEPA) work in vision, we propose ArtNet, a framework that explores a structured feature prediction task based on articulatory features to enhance acoustic robustness. Specifically, ArtNet integrates an articulatory predictor—designed to extract universal articulatory representations from self-supervised learning (SSL) features—with a variational information bottleneck (VIB) to suppress language-specific variations. Experiments on seven unseen languages demonstrate that ArtNet, particularly when synergized with the proposed vector-space inventory alignment (VSIA) strategy, significantly outperforms competitive baselines, achieving a 20.56% relative reduction in phoneme error rate (PER) and 7.01% in phoneme feature error rate (PFER).

Subject: INTERSPEECH.2026 - Analysis and Assessment


#17 wav2VOT: automatic estimation of voice onset time, closure duration, and burst realisation with wav2vec2 [PDF] [Copy] [Kimi] [REL]

Authors: James Tanner, Morgan Sonderegger, Jane Stuart-Smith, Tyler Kendall, Jeff Mielke

While automatic tools for speech annotation are now commonplace within phonetic research pipelines, many tasks require substantial manual correction or training sets to perform accurately. Simultaneously, large speech models such as wav2vec2 have been shown to perform well at speech classification tasks, raising the question of how these models may be applied to phonetic annotation tasks. We introduce wav2VOT: a tool for the automatic estimation of voice onset time, closure duration, and burst realisation using wav2vec2. We demonstrate that wav2VOT performs comparably with current approaches on unseen datasets, and can estimate with high accuracy with fine-tuning. Analysis of wav2VOT predictions demonstrate high fidelity across stop voicing and place of articulation. These results demonstrate that large speech models are capable of producing accurate annotations, and further motivate exploration of large speech models as tools in phonetic research pipelines.

Subject: INTERSPEECH.2026 - Analysis and Assessment


#18 Time-normalized spectrograms reveal segmental differences in English heterographic homophones [PDF] [Copy] [Kimi] [REL]

Authors: Yu-Hsiang Tseng, Harald Baayen

Homophones are words that are commonly assumed to sound the same but have different meanings. They are either homographic, sharing the same spelling, or heterographic, differing in spelling. This study focuses on heterographic homophones. Past studies have shown that the spoken word duration of such homophones is systematically related to their frequency of use. Building on this finding, we investigate whether and how the segmental realizations of homophones differ, using time-normalized spectrograms and derived phone logits. We analyzed 14,000 homophone tokens and found that words within homophone pairs differ in their segmental realizations. Further analysis indicates that the exact phonetic realization of a token is shaped by its meaning in the utterance context. These findings are more consistent with lexicon models that assume token-level alignment between form and meaning than with frameworks that posit segments as abstract formal symbols.

Subject: INTERSPEECH.2026 - Analysis and Assessment


#19 Minimum Token Thresholds and Stabilisation for Reliable Automatic Vowel Alignment: Empirical Study on TIMIT Vowels and MFA [PDF] [Copy] [Kimi] [REL]

Authors: Simon Gonzalez, Jason Littlefield, Tao Hoang, Chloe Dean, Hayden Ooi, Myung Kim, Bradley Donnelly, Latchman Singh, Jennifer Biggs, Tim Cawley

Automatic forced alignment is standard in phonetic research, but the minimum data needed for reliable acoustic measurements remains unclear. This study examines how token quantity affects alignment reliability using the MFA on the TIMIT corpus, with manual phoneme boundaries as the gold standard. Focusing on vowels and three acoustic features (Duration, F1, F2) token subsets were incrementally sampled and mixed-effects models fitted at each step to evaluate automatic versus manual measurements. Results show 85% of vowel-feature combinations improve significantly with increasing tokens, with F1 showing the most consistent gains. Most vowels stabilise around 50% of available tokens, though variability exists: Duration stabilises earliest (≈730 tokens) while F2 requires the most data (≈2,335 tokens). These findings offer empirically grounded guidelines for minimum token requirements, with practical implications for large-scale sociophonetic studies and low-resource language research.

Subject: INTERSPEECH.2026 - Analysis and Assessment


#20 Word Lengthening as a Function of Utterance Position: A Multi-Corpus Study [PDF] [Copy] [Kimi] [REL]

Authors: Mateo Cámara, José Luis Blanco, Juan Ignacio Godino-Llorente, Jeung-Yoon Choi, Stefanie Shattuck-Hufnagel

Efficient turn-taking requires interlocutors to predict turn endings within a few hundred milliseconds. Beyond syntactic and pragmatic completion, prosody (especially pre-boundary lengthening) supports projection. We test whether turn-final words are longer than mid-sentence words, whether this reflects prosodic modification rather than lexical choice, and where within the word it concentrates. We analyze four corpora spanning styles and two languages (English, Spanish): Switchboard, Columbia Games, BU Radio, and Glissando, with >500 speakers, 39,470 turn-final and 206,268 mid-sentence tokens across ~ 39,500 turns. Turn-final words are longer (mean ≈191 ms; d = 1.14). The effect persists in matched-word, within-speaker comparisons (80 ms; p < 0.001) and is localized mainly to the final syllable (d = 0.89). Turn-final lengthening thus emerges as a robust, localized cue to floor transfer.

Subject: INTERSPEECH.2026 - Analysis and Assessment


#21 Larynx segmentation in mid-sagittal speech production real-time MRI [PDF] [Copy] [Kimi] [REL]

Authors: Yubin Zhang, Xuan Shi, Kevin Huang, Prakash Kumar, Kevin Lee, Louis Goldstein, Krishna Nayak, Shrikanth Narayanan

The spatiotemporal dynamics and coordination of laryngeal movements remain incompletely characterized. This study introduces a larynx segmentation and analysis pipeline for mid-sagittal speech production real-time MRI using Mask2Former, combining supervised learning with semi-supervised refinement. Our results suggest sufficient segmentation performance with ~33-79 annotations per participant (~25-60% of 794 training samples / 6 participants) using a 5% marginal gain threshold. Additional annotations beyond this yield diminishing returns. Semi-supervised learning achieves modest improvements but sometimes degrades performance and cannot reach the fully supervised upper bound. Results from a sample phonetic study on Mandarin tones demonstrate the capability of mid-sagittal speech production real-time MRI to capture spatiotemporal laryngeal dynamics, including both intrinsic and extrinsic pitch control and laryngeal constriction mechanisms. The proposed pipeline opens new avenues for studying laryngeal behaviors in linguistic contrasts such as voicing, tone, and phonation types using real-time MRI. Code repository: https://github.com/pkuzyb/ larynx_segmentation.

Subject: INTERSPEECH.2026 - Analysis and Assessment


#22 Reconciling Dynamic Data Analysis with Linguistic Reality: Comparing Legendre Polynomial Modelling and GAMM Applied to Prosodic Contact [PDF] [Copy] [Kimi] [REL]

Authors: Angelo Dian, Mary Baltazani, Spyros Armostis, Elinor Payne

We examine continuation rises in Cypriot Greek (CYG) and Athenian Greek (ATG) as a case study for comparing and evaluating methods of dynamic data analysis applied to contact-induced intonational variation. We analyse contemporary data from young speakers using two approaches: Legendre Polynomial modelling and Generalized Additive Mixed Modelling (GAMM). Both methods converge in identifying two distinct CYG patterns: a low-nuclear ATG-like contour (CYG-nl), and a high-nuclear variant (CYG-nh) with reduced global slope and increased curvature. While polynomial modelling captures global geometric properties of the contour, GAMM specifies differences in relative time, especially around the nuclear high target. The comparison shows that complementary modelling systems can shed light on different aspects of the phonetic realization and phonological interpretation of intonation in contact scenarios.

Subject: INTERSPEECH.2026 - Analysis and Assessment


#23 Achieving voicelessness in coda stop contexts: Insights from combined electroglottography and laryngoscopy [PDF] [Copy] [Kimi] [REL]

Authors: Joshua Penney, Jae Hyun Kim, Dijana Dragicevich, Prue Gourley

Voicelessness in English coda stops can be produced with either glottal spreading or glottal constriction, with consequences for voice quality of preceding vowels. Recent evidence from electroglottography (EGG) suggests that in Australian English voicelessness is achieved via glottal constriction for /t/, whereas glottal spreading is the strategy for /k/ and no consistent pattern is identified for /p/. However, EGG measures glottal state indirectly via electrical impedance and provides no direct information on supraglottal configuration, which may play a role in constriction. In this paper, we combine EGG with laryngoscopic imaging to investigate glottal/laryngeal settings during voiceless coda stop production and assess the extent to which inferences from EGG align with observed glottal/supraglottal behaviour. Results provide support for recent findings and reveal that constriction is the preferred strategy for voiceless stops at all places of articulation in phrase final position.

Subject: INTERSPEECH.2026 - Analysis and Assessment


#24 Multilingual and Cross-lingual Lexical Stress Detection Using SSL Feature Vectors [PDF] [Copy] [Kimi] [REL]

Authors: Abdulrahman Alhabshi, Beena Ahmed, Mostafa Shahin

Prosodic cues such as rhythm and lexical stress affect intelligibility; they stand out as signals that guide perception. While these cues are shared across many languages, lexical stress is realized differently in Arabic and English, making stress detection challenging in multilingual and transfer settings. We propose a framework that uses self supervised speech representations as fixed features and a two stage classifier: a syllable level Pre-net and a word context Post-net that models inter syllable dependencies. We compare monolingual training, joint Arabic English multilingual training, cross lingual transfer using different self-supervised learning (SSL) model's feature vectors input into Pre-net DNN and Post-net TDNN models. Multilingual training retains near monolingual accuracy of 97% and 90% on English and Arabic respectively, whereas cross-lingual transfer is direction-dependent; the Post-net improves robustness and multilingual pre-training yields the strongest transfer.

Subject: INTERSPEECH.2026 - Analysis and Assessment


#25 Scaling Human and G2P Supervision for Robust Phonetic Transcription [PDF] [Copy] [Kimi] [REL]

Authors: Alexander Metzger, Aruna Srivastava, Ruslan Mukhamedvaleev

Expert phonetic annotation is costly, especially for non-standard dialects and atypical speech. A common alternative is using Grapheme-to-Phoneme (G2P) models to auto-generate phonetic labels from text transcripts at scale. We study how automatic phonetic transcription performance scales with human and G2P supervision in English. Using a curated 80-hour benchmark spanning native, non-native and post-stroke speech, we identify a supervision quality threshold: G2P supervision helps only when fewer than 20-30 hours of human annotation are available. Beyond this threshold, it provides no significant benefit and can reduce cross-dialect robustness. What is effective after this threshold is ASR pretraining which we use to achieve a 2.3× reduction in weighted phone feature error rate over prior systems, with strong gains on non-native and aphasic speech. These results suggest that quantity-driven G2P scaling may yield diminishing returns for robust generalization.

Subject: INTERSPEECH.2026 - Analysis and Assessment