INTERSPEECH.2026 - Modelling and Learning

| Total: 63

#1 TAD: Token-Adaptive Contrastive Decoding with Confidence-Guided Gating for Hallucination Mitigation in Large Audio-Language Models [PDF] [Copy] [Kimi] [REL]

Authors: Heyu Chang, Nianwen Si, Hao Zhang, Wenlin Zhang, Dan Qu

Large audio-language models (LALMs) can hallucinate audio objects, answering "yes" to absent sound events, thus undermining reliability in audio question answering. We propose Token-Adaptive Decoding (TAD), a training-free strategy for hallucination mitigation that grounds the initial yes/no decision by contrasting logits under real audio with a matched silent reference. TAD introduces a token-adaptive, confidence-guided gate that is decision-critical at the first decoding step and class-conditional on affirmative tokens, using the audio-silent margin to avoid overcorrection when evidence is weak or already sufficient. Experiments on AudioCaps-Hallucination show that, relative to Audio-Aware Decoding (AAD), a contrastive baseline with fixed contrast strength, TAD improves F1 for Qwen2 by 0.059 to 0.117 across Popular, Adversarial, and Random splits, and for Gemma by 0.025 to 0.064, while on Clotho-AQA it raises F1 from 0.810 to 0.816 on Qwen2 and remains comparable to AAD on Gemma.

Subject: INTERSPEECH.2026 - Modelling and Learning


#2 TinyGiantALM: A Compact Audio-Language Model for Intent-Aware Reasoning under Resource Constraints [PDF] [Copy] [Kimi] [REL]

Author: Vinh-Thuan Ly

Current advancements in Audio Reasoning rely on massive Large Audio-Language Models (LALMs), hindering deployment in resource-constrained environments. We introduce Tiny-GiantALM, a compact 1.5B efficiency-oriented alternative. Instead of brute-force scaling, we propose an Instruction-Aware Feature Refinement framework using a Query-guided Projector and Semantic Gating to filter acoustic signals based on user intent. On the MMAR benchmark, TinyGiantALM achieves 46.4% zero-shot accuracy, significantly outperforming 7B-13B baselines. While a reasoning gap in logical narrative remains versus 30B+ models and certain trade-offs exist in overly dense or spatial scenes, our approach notably surpasses models up to 8× larger in disentangling mixed-modality environments. These findings demonstrate that architectural precision offers a tangible pathway to secure robust perception capabilities on edge-friendly scales.

Subject: INTERSPEECH.2026 - Modelling and Learning


#3 Beyond Symmetric Interaction: Capability-Aware Asymmetric Multi-Agent Collaboration for Audio Deep Reasoning [PDF] [Copy] [Kimi] [REL]

Authors: Yan Rong, Jinting Wang, Tianxin Xie, Xiang He, Chenxing Li, Dong Yu, Li Liu

Audio deep reasoning demands expert-level perception and multi-step reasoning. Vanilla multi-agent paradigms struggle here due to three challenges: (1) underutilization of the complementary abilities of diverse Large Audio-Language Models (LALMs); (2) neglect of role differentiation and capability bias; and (3) sampling instability inherent to LALMs. To address these, we propose AsymAudio, a novel Capability-Aware Multi-Agent Audio Reasoning framework with three key components: (1) an Inter-Agent Collaborative Interaction Module that fuses complementary acoustic strengths across LALMs; (2) a Role-Adaptive Asymmetric Strategy that assigns hierarchical roles and tailored feedback to reduce synergistic hallucination from the bucket effect and textual compensation; and (3) an Intra-Agent Consistency Refinement Module that uses resampling and self-correction to enhance stability. AsymAudio ranked 3rd in the Interspeech 2026 Audio Deep Reasoning Challenge (Agent Track).

Subject: INTERSPEECH.2026 - Modelling and Learning


#4 VISA: A Visual Information Strengthened Audio-Reasoning System for the Interspeech 2026 ARC Agent Track [PDF] [Copy] [Kimi] [REL]

Authors: Wenming Tu, Jian Gao, Yanru Huo, Yixuan Wang, Jing Peng, Bohan Li, Ziyang Ma, Tao Liu, Shuai Fan, Kai Yu, Xie Chen, Zilong Zheng

Audio reasoning requires multi-step, evidence-grounded inference over temporally dynamic and acoustically mixed signals, exceeding conventional perception tasks such as ASR or captioning. We present VISA, our submission to the Interspeech 2026 Audio Reasoning Challenge (Agent Track), evaluated via the MMAR Rubrics for correctness and reasoning quality. Under a "LALM as a Tool" paradigm, VISA strengthens large audio language models with auxiliary multi-modal evidence while avoiding heavy orchestration. The system integrates three components: multi-modal feature extraction for complementary audio and acoustic-visual clues, model-voting inference with consistency checking for stable predictions, and fine-grained category-aware routing to resolve disagreements and select rubric-aligned reasoning chains. On the official leaderboard, VISA ranks 2nd overall in the Agent Track with a Rubrics score of 66.23%, while achieving the top Accuracy of 77.40% in both the Single Model and Agent tracks.

Subject: INTERSPEECH.2026 - Modelling and Learning


#5 The Interspeech 2026 Audio Reasoning Challenge: Evaluating Reasoning Process Quality for Audio Reasoning Models and Agents [PDF] [Copy] [Kimi] [REL]

Authors: Ziyang Ma, Ruiyang Xu, Yinghao Ma, Chao-Han Huck Yang, Bohan Li, Jaeyeon Kim, Jin Xu, Jinyu Li, Carlos Busso, Kai Yu, Eng Siong Chng, Xie Chen

Recent Large Audio Language Models (LALMs) excel in understanding but often lack transparent reasoning. To address this "black-box" limitation, we organized the Audio Reasoning Challenge at Interspeech 2026, the first shared task dedicated to evaluating Chain-of-Thought (CoT) quality in the audio domain. The challenge introduced MMAR-Rubrics, a novel instance-level protocol assessing the factuality and logic of reasoning chains. Featured Single Model and Agent tracks, the competition attracting 156 teams from 18 countries and regions. Results show agent systems currently lead in reasoning quality, utilizing iterative tool orchestration and cross-modal analysis. Besides, single models are rapidly advancing via reinforcement learning and sophisticated data pipeline. We details the challenge design, methodology, and a comprehensive analysis of state-of-the-art systems, providing new insights for explainable audio intelligence.

Subject: INTERSPEECH.2026 - Modelling and Learning


#6 Multi-Source Evidence Fusion for Audio Question Answering [PDF] [Copy] [Kimi] [REL]

Authors: Aivo Olev, Tanel Alumäe

Large audio language models (LALMs) can answer questions about speech, music, and environmental sounds, yet their internal reasoning is largely opaque and difficult to validate. We describe TalTech's solution to the Agent Track of the Interspeech 2026 Audio Reasoning Challenge, in which systems are evaluated on reasoning process quality, specifically the factual accuracy, logical soundness, and completeness of their reasoning chains. Our multi-source ensemble pipeline uses two LALMs that generate independent observations, while a separate text-only reasoning model cross-checks these against outputs from 25 acoustic tools organized into reliability tiers. By grounding every inference step in explicit, reliability-tagged evidence, the system produces dense, verifiable reasoning chains. Our system ranked first in the challenge, outperforming all competing systems by a wide margin in the challenge's reasoning quality metric.

Subject: INTERSPEECH.2026 - Modelling and Learning


#7 Structured Prompting vs. Self-Training for Audio Reasoning Under Limited Data and Compute: Lessons from Interspeech Audio Reasoning Challenge 2026 [PDF] [Copy] [Kimi] [REL]

Authors: Sujit Noronha, Steven Au, Kaushlendra Tripathi

This study consists of our approaches to the Interspeech 2026 Audio Reasoning challenge evaluating chain-of-thought-based reasoning traces on the Multi-Modal Audio Reasoning benchmark (MMAR). For audio reasoning with limited data and compute, practitioners and researchers must choose among various strategies ranging from prompt engineering at no cost to self-training requiring substantial compute. We compare these on the state-of-the-art Qwen3-Omni 30B model across the 1000 audio questions spanning four reasoning categories and 16 subcategories. We evaluate three strategies which include structured prompt engineering through iterative error analysis, automated prompt optimization using DSPy MIPROv2 and Reinforced Self Training (ReST) with qLORA finetuning. Our structured prompting approach helped achieve an accuracy increase of 5.5% compared to baseline, while ReST based training led to a decrease in performance compared to baseline.

Subject: INTERSPEECH.2026 - Modelling and Learning


#8 Audio-Cogito: Towards Deep Audio Reasoning in Large Audio Language Models [PDF] [Copy] [Kimi] [REL]

Authors: Longhao Li, Hongjie Chen, Zehan Li, Qihan Hu, Jian Kang, Jie Li, Lei Xie, Yongxiang Li

Recent advances in reasoning models have driven significant progress in text and multimodal domains, yet audio reasoning remains relatively limited. Only a few Large Audio Language Models (LALMs) incorporate explicit Chain-of-Thought (CoT) reasoning, and their capabilities are often inconsistent and insufficient for complex tasks. To bridge this gap, we introduce Audio-Cogito, a fully open-source solution for deep audio reasoning. We develop Cogito-pipe for high-quality audio reasoning data curation, producing 545k reasoning samples. Based on this dataset, we adopt a self-distillation strategy for model fine-tuning. Experiments on the MMAR benchmark, the only audio benchmark evaluating the CoT process, show that our model achieves the best performance among open-source models and matches or surpasses certain closed-source models in specific metrics. Our approach also ranks among the top-tier systems in the Interspeech 2026 Audio Reasoning Challenge.

Subject: INTERSPEECH.2026 - Modelling and Learning


#9 Audio-DeepThinker: Progressive Reasoning-Aware Reinforcement Learning for High-Quality Chain-of-Thought Emergence in Audio Language Models [PDF] [Copy] [Kimi] [REL]

Authors: Xiang He, Chenxing Li, Jinting Wang, Yan Rong, Tianxin Xie, Zeyu Xie, Wenfu Wang, Li Liu, Dong Yu

Large Audio-Language Models (LALMs) excel at perception but lack grounded reasoning. Existing methods rely on supervised chain-of-thought (CoT) data or coarse Reinforcement Learning (RL) rewards that do not directly evaluate reasoning quality, yielding chains logically ungrounded in audio. To bridge this gap, we propose Audio-DeepThinker with two ideas. First, a hybrid reasoning similarity reward combining an LLM evaluator assessing logical path alignment and key step coverage with embedding similarity enforcing semantic alignment with reference chains. Second, a progressive two-stage curriculum enabling CoT to emerge via pure RL exploration (RL-Zero). Stage 1 uses the hybrid reward on foundational audio QA, while Stage 2 shifts to boundary cases with an LLM-only reward for reasoning diversity. Audio-DeepThinker achieves state-of-the-art on MMAR (74.0%) and MMAU-Test-Mini (78.5%), winning 1st Place in the Interspeech 2026 Audio Reasoning Challenge (Single Model Track).

Subject: INTERSPEECH.2026 - Modelling and Learning


#10 EChO-Agent: Evidence Chain Orchestration Agent for Audio Reasoning [PDF] [Copy] [Kimi] [REL]

Authors: Siyuan Zhang, Jian Zong, Junyu Wang, Peiyuan Jiang, Jiahao Yan, Jingyu Zhang, Tianrui Wang, Xiaobao Wang, Longbiao Wang, Jianwu Dang

While LALMs show promise on audio question answering, they fail to focus on question-relevant segments of audio and provide a clear, checkable reasoning process when dealing with complex audio reasoning. Reinforcement learning and tool-augmented prompting can help models better relate questions to audio but lack a reliable way to understand, integrate, and self-verify audio segments. To address this gap, we present EChO-Agent, a modular agent framework that reformulates complex audio QA as a planning, tool execution, evidence integration, and answer verification workflow. Experiments on MMAR benchmark show EChO-Agent improves both accuracy and rubric scores over baseline and ablation studies show evidence integration is the key factor.

Subject: INTERSPEECH.2026 - Modelling and Learning


#11 MATA: A Training-Free Approach to Mitigate Cross-Modal Attention Imbalance in Large Audio Language Models [PDF] [Copy] [Kimi] [REL]

Authors: Junyu Wang, Jian Zong, Tianrui Wang, Zhengding Luo, Meng Ge, Xiaobao Wang, Longbiao Wang, Jianwu Dang

Large Audio Language Models (LALMs) often suffer from audio-textual attention imbalance, prioritizing text over acoustic information during multi-modal fusion. This bias limits the utilization of acoustic cues and degrades audio reasoning performance. To mitigate this, we propose MATA, a novel training-free method that dynamically pushes LALMs to pay More Attention To Audio tokens within the self-attention mechanism. Specifically, MATA intervenes after raw attention scoring, targeting only the last token in intermediate layers without adding parameters or computational overhead. Experiments on MMAU and MMAR benchmarks confirm consistent performance gains. Furthermore, integrating MATA with the Qwen3-Omni-Thinking model secured second place in the Single Model Track of the Interspeech 2026 Audio Reasoning Challenge. As the only training-free approach among top solutions, MATA offers a highly efficient strategy to mitigate attention bias and advance LALM reasoning.

Subject: INTERSPEECH.2026 - Modelling and Learning


#12 MAC-VAD: A Modality-Aligned Cross-Attentive Framework for Robust Voice Activity Detection [PDF] [Copy] [Kimi] [REL]

Authors: Bruhanth Mallik, Chintan Tundia, Kumud Tripathi, Shreyas Nagoor, Pankaj Wasnik

Recent studies indicate that the accuracy of Voice Activity Detection (VAD), a crucial component in speech processing systems, can be enhanced by leveraging both audio and visual cues. However, effectively fusing multimodal information across diverse content to achieve a highly accurate and robust VAD is still a challenge. In this paper, we present an audio-visual Modality-Aligned dual Cross-attention framework for VAD, MAC-VAD. The proposed model uses an audio encoder optimised to adaptively acquire spectral and temporal representations from raw audio. Whilst the visual encoder is designed to predict speech onsets, utilising face and lip features, accurately. The proposed model also employs knowledge distillation from a teacher network to steadily guide the multimodal framework in rigorous learning for the VAD task. Evaluations with the MMVAD dataset prove MAC-VAD's superior performance, validating the effectiveness of modality-aligned cross-attention and self-supervised distillation

Subject: INTERSPEECH.2026 - Modelling and Learning


#13 Position-Aware Target Speaker Extraction for Long-Form Multi-Party Conversations: A Diarization-Free Framework for ASR [PDF] [Copy] [Kimi] [REL]

Authors: Yichi Wang, Junzhe Chen, Wangjin Zhou, Tatsuya Kawahara

In long-form multi-party conversations, highly imbalanced speaker activity and frequent overlap make it difficult to identify "who spoke when and what". Sliding-window continuous speech separation (CSS) mitigates sparse supervision, but often suffers from cross-window speaker inconsistency and residual crosstalk, which in practice requires diarization for reliable speaker attribution. Motivated by the stability of speakers' directions of arrival (DOAs) in meetings, we propose PATSE, a multi-channel Position-Aware Target Speaker Extraction front-end that uses DOA as a spatial prior to directly extract the speech of each target speaker. PATSE combines a DOA-guided spatial encoder and conditioner to generate speaker-attributed streams, from which speaker activity can be inferred via simple post-processing (e.g., VAD) without explicit diarization. Experiments on both replayed and real conversations show consistent ASR gains outperforming CSS and diarization-based pipelines.

Subject: INTERSPEECH.2026 - Modelling and Learning


#14 SEAM: Shortcut-Aware Real-Time Detection of Scripted vs. Spontaneous Speech for Interview Guardrails [PDF] [Copy] [Kimi] [REL]

Authors: Vsevolod Kovalev, Pranay Manocha

Scripted vs spontaneous speech detection is appealing for interview guardrails, but benchmark performance can be inflated by shortcuts tied to corpus identity, channel conditions, and recording artifacts rather than speaking style itself. We present SEAM, a shortcut-aware framework for real-time scriptedness detection that combines uniform preprocessing, seam-aware sampling, non-speech augmentation, and a compact DistilHuBERT backbone. With 8s windows, the model achieves 0.971 ± 0.004 ROC-AUC on an external interview-domain evaluation set. Removing the shortcut-prevention components improves internal held-out metrics but sharply reduces external performance, indicating shortcut learning. Post-training quantization reduces the model footprint to 41.8 MB with little loss in external performance. The results demonstrate that robust real-time scriptedness detection depends not only on the backbone, but on shortcut-aware data design and evaluation. We release code and model checkpoints.

Subject: INTERSPEECH.2026 - Modelling and Learning


#15 MCA-DCF-DS: An Adaptive Framework for Unified Diarization and Separation with Spatial Information [PDF] [Copy] [Kimi] [REL]

Authors: Shutong Niu, Ruo-Yu Wang, Gao-Bin Yang, Ya Jiang, Tian Gao, Jia Pan, Jun Du

A Deep Cascade Fusion of Diarization and Separation (DCF-DS) framework was previously proposed to integrate speaker diarization with speech separation for handling the realistic multi-speaker scenarios. However, DCF-DS mainly relies on spectral cues to distinguish speakers, and its performance degrades in highly overlapped regions. In this paper, we first extend DCF-DS to multi-channel DCF-DS (MC-DCF-DS) by incorporating spatial information at the system level. We then further propose MCA-DCF-DS by introducing spatial information into the adaptation data simulation process to handle the trade-off between miss and confusion errors in diarization results. Experimental results demonstrate that incorporating spatial information at both the system and data levels improves downstream ASR performance. Moreover, under the same ASR backend, the proposed MCA-DCF-DS outperforms the CHiME-8 Task 2 champion system.

Subject: INTERSPEECH.2026 - Modelling and Learning


#16 Grammar-Guided Hierarchical Parsing for Long-form Audio Activity Recognition [PDF] [Copy] [Kimi] [REL]

Authors: Peng Zhang, Qingyu Luo, Philip J.B. Jackson, Wenwu Wang

Long-form audio exhibits an inherent hierarchy: fine-grained events form sub-activities, which in turn constitute higher-level activities. Prior work often models these levels separately, leading to cross-level inconsistencies and requiring supervision at multiple levels. We formulate the problem as hierarchical parsing from event-level evidence: given detected event segments with class posteriors, we infer an order-consistent Act-Sub-Event parse tree. We propose Hierarchical Activity Grammar, encoding hierarchical composition and temporal-order constraints, and perform grammar-guided decoding that combines event evidence with a grammar prior. This yields a temporally grounded parse tree from which sub-activity segmentation and activity classification are derived, without requiring sub-activity or activity labels for training. Experiments on the long-form MultiAct audio dataset demonstrate improved temporal-order consistency (Edit score) and produces interpretable hierarchies.

Subject: INTERSPEECH.2026 - Modelling and Learning


#17 SA-UAED: Joint Frame-Level Detection of Audio Events, Speaker Activities, and Speaker-Attributed Paralinguistic Events [PDF] [Copy] [Kimi] [REL]

Authors: Zekun Lan, Wangyou Zhang, Yanmin Qian

Existing sound event detection (SED) and speaker diarization (SD) systems typically fail to attribute paralinguistic events (e.g., laughter, coughing) to individual speakers, primarily due to the scarcity of fine-grained labeled data. To address this, we propose a dedicated simulation pipeline generating large-scale audio mixtures with precise frame-level annotations. Leveraging this, we introduce the Speaker-Attributed Unified Audio Event Detection (SA-UAED) framework for joint frame-level prediction of sound events, speaker activities, and paralinguistics. Experiments demonstrate that SA-UAED substantially improves speaker-attributed paralinguistic detection without compromising generic SED and SD accuracy.

Subject: INTERSPEECH.2026 - Modelling and Learning


#18 Music Artistic Captioning: Towards Translating Music into Expressive Language [PDF] [Copy] [Kimi] [REL]

Authors: Ubaid Ullah, Hyun-Chul Choi, Zied Bouraoui

Automated music captioning remains difficult for narrative-rich works (e.g., opera) where instrumentation, affect, and structure evolve over time. Existing LLM-based captioners depend on scarce paired audio-text data, and fine-tuning can overfit to dataset-specific phrasing, limiting transfer and controllable narration. We propose Music Artistic Captioning (MAC), a training-free framework that generates long-form, grounded descriptions by conditioning an instruction-following LLM on automatically extracted multi-scale audio evidence. MAC aggregates frame-, segment-, and song-level descriptors, from low-level acoustics (e.g., tempo) to predicted high-level semantic cues (e.g., instrumentation), into constrained pseudo-evidence for hierarchical captioning. To support user-preferred style, MAC uses a one-time iterative prompt optimization to balance narrative voice and evidence faithfulness. Experiments on an opera corpus and caption benchmarks show consistent gains over strong baselines.

Subject: INTERSPEECH.2026 - Modelling and Learning


#19 Content is What Remains: Invariant Speech Tokenization from Parallel Utterances [PDF] [Copy] [Kimi] [REL]

Author: Laurin Wagner

Discrete speech tokenizers aim to disentangle semantic from acoustic information, yet targets from self-supervised learning (SSL) models like HuBERT retain non-linguistic variation: speaker identity, prosody, and channel conditions leak into tokens, inflating entropy. Our insight: when enough speakers utter the same words under varying conditions, linguistic content is the only shared factor. We propose PINT (Parallel INvariant Tokenization), fine-tuning an SSL encoder with alignment losses across parallel utterances and augmentations to distill this shared residual. PINT collapses identical words onto consistent token sequences, drastically reducing conditional entropy. Unlike ASR text, PINT tokens preserve frame-level temporal grounding and serve as drop-in semantic targets for audio codecs. Experiments show a 98.7% relative reduction in speaker probe accuracy (93.1%→1.2%), 42% lower ABX error rate, and 27–30% lower LM perplexity versus baselines, confirming that the right invariance is key to efficient learning.

Subject: INTERSPEECH.2026 - Modelling and Learning


#20 Context-Adaptive Automated Audio Captioning with Symmetric Dual-MoE and Dynamic Reward Routing [PDF] [Copy] [Kimi] [REL]

Authors: Seyun Ahn, Joon-Hyuk Chang

Deep learning–based automated audio captioning (AAC) typically optimizes a uniform objective, limiting contextual adaptability and fine-grained audio–text alignment. Although reinforcement learning improves evaluation metrics, existing approaches rely on simple and static reward formulations. We propose a context-adaptive AAC framework based on a symmetric dual mixture-of-experts (MoE) architecture optimized with group relative policy optimization (GRPO). On the policy side, LoRA-based experts are integrated into a pretrained BART decoder and activated via cross-attention routing conditioned on acoustic context. On the reward side, parallel experts evaluate semantic relevance, grammatical correctness, lexical diversity, and audio–text alignment, while a context-aware router adaptively weights these signals. Experiments on Clotho and Audio-Caps show improvements in caption quality and human preference. Generated captions are available on the https://symmetric-moe.github.io/.

Subject: INTERSPEECH.2026 - Modelling and Learning


#21 Acoustic Prompting via Stage-wise Modulation for Few-Shot Learning in Audio Language Models [PDF] [Copy] [Kimi] [REL]

Authors: Hyebin Cho, Jaehyuk Jang, Changick Kim, Joon Son Chung

Audio-Language Models (ALMs) have shown remarkable success in zero-shot audio classification by aligning audio waveforms with text. Recent efforts to improve downstream performance focus on learning optimal text prompts. However, previous approaches focus on the text encoder, leaving the potential of learnable prompts within the audio encoder unexplored. In this paper, we propose a novel framework that introduces trainable prompts into the audio encoder to capture task-specific acoustic features. We demonstrate that integrating audio-side prompt learning with existing text-side approaches enhances few-shot adaptation. Through extensive experiments across 11 datasets show that integrating our method as a plug-and-play module alongside existing text prompt tuning generally leads to performance improvements. These findings suggest that explicitly modulating the audio representation space effectively complements text-only prompting approaches. The code is available at https://github.com/hyebin-c/aspl.

Subject: INTERSPEECH.2026 - Modelling and Learning


#22 Audio-Language Prompt Learning for Few-Shot Audio Classification [PDF] [Copy] [Kimi] [REL]

Authors: Qisheng Xu, Xiaoyi Tan, Wuyang Chen, Yutao Dou, Kele Xu

Audio-language models (ALMs) have shown strong generalization in standard audio classification tasks, yet their few-shot adaptation remains constrained by text-centric prompt learning, which under-adapts the audio encoder and struggles to distinguish subtle differences between accoustically similar classes. This text-dominated adaptation leads to imbalanced optimization across modalities and limits the discriminative capacity of ALMs under low-data regimes. To address this limitation, we propose MALP, a multi-modal prompt learning framework that jointly optimizes audio-specific, text-specific, and shared prompts. Specifically, MALP first enables modality-specific adaptation to capture complementary characteristics of audio and text, and then introduces shared prompts to strengthen cross-modal alignment while preserving modality-specific discriminability. Experiments on eleven benchmark datasets demonstrate consistent improvements over multiple strong baselines, achieving average gains of 7.21% over CoOp, 4.80% over CoCoOp, and 1.77% over PALM. Ablation studies further confirm the complementary roles of audio-specific and shared prompts, validating the effectiveness of multi-modal prompt learning for few-shot audio classification.

Subject: INTERSPEECH.2026 - Modelling and Learning


#23 Tone-Conditioned Curriculum Learning for Low-Resource Bantu Speech Recognition [PDF] [Copy] [Kimi] [REL]

Authors: Kesego Mokgosi, Vukosi Marivate, Sitwala Mundia, Unarine Netshifhefhe, Tsholofelo Mogale, Thapelo Sindane

Southern Bantu languages are spoken by over 80 million people, yet current foundation ASR models still produce zero-shot WER above 100%, which limits practical use in education and public services. We addressed this gap with a tone conditioned curriculum framework for 6 Southern Bantu languages that combined hybrid difficulty scoring, gated adapters driven by tonal statistics and staged curriculum training. We trained on a community corpus and tested transfer to NCHLT to measure robustness beyond matched evaluation. Results revealed clear interactions between architecture and language, with W2V-BERT outperforming Whisper on Nguni languages by 3 to 4 WER points whilst Whisper performed better on Sotho-Tswana languages. W2V-BERT with tone conditioning reached 28.41% average WER across datasets and 23.79% on Xitsonga transfer. No single model suited all 6 languages, so deployment should pair model selection per language with validation across corpora.

Subject: INTERSPEECH.2026 - Modelling and Learning


#24 Which Languages Transfer Best to Warlpiri? A Similarity-Based Study for Low-Resource ASR [PDF] [Copy] [Kimi] [REL]

Authors: Pravina Mylvaganam, Eliathamby Ambikairajah, Ting Dang, Vidhyasaharan Sethu, Tünde Szalay

This paper investigates how language similarity can improve cross-lingual transfer for automatic speech recognition (ASR) in extremely low-resource settings. Warlpiri, an Australian Aboriginal language, has very limited transcribed speech data, making transfer learning essential. We propose a framework combining acoustic similarity from pre-trained speech models with linguistic similarity based on typology, phoneme inventories, grammatical, and syntactic features to rank high-resource source languages and evaluate their effectiveness for ASR transfer to Warlpiri. Experiments with Whisper show that acoustically and typologically similar languages outperform monolingual and multilingual baselines. Assamese and Hindi achieve substantial reductions in word and character error rates. Correlation analysis further indicates that acoustic similarity is the strongest predictor of fine-tuning performance, while phoneme inventory and typological similarity better explain zero-shot transfer.

Subject: INTERSPEECH.2026 - Modelling and Learning


#25 From Academic Tool to Community Infrastructure: A Call for Indigenous Partnership in Speech Data Governance [PDF] [Copy] [Kimi] [REL]

Authors: Kaveri K. Sheth, Sebastien Christian

We present ELSI (ExELang Legacy Support Interface), a governance platform for managing sensitive child-centered audio datasets that includes recordings from 10 indigenous language communities across multiple continents. While ELSI was designed to balance open science with participant privacy, we identify a fundamental gap: custodianship concentrates control in the hands of European academic researchers, with no formal plan for indigenous communities to govern their own data. We describe the platform, map its current model against CARE and OCAP principles, and issue a concrete call for indigenous community partners to co-design a framework for identifying legitimate community custodians and implementing an enforceable access and control layer. We commit to the architectural and governance changes needed to make this real, and describe what meaningful co-design would look like in practice.

Subject: INTERSPEECH.2026 - Modelling and Learning