noronha26@interspeech_2026@ISCA

Total: 1

#1 Structured Prompting vs. Self-Training for Audio Reasoning Under Limited Data and Compute: Lessons from Interspeech Audio Reasoning Challenge 2026 [PDF] [Copy] [Kimi] [REL]

Authors: Sujit Noronha, Steven Au, Kaushlendra Tripathi

This study consists of our approaches to the Interspeech 2026 Audio Reasoning challenge evaluating chain-of-thought-based reasoning traces on the Multi-Modal Audio Reasoning benchmark (MMAR). For audio reasoning with limited data and compute, practitioners and researchers must choose among various strategies ranging from prompt engineering at no cost to self-training requiring substantial compute. We compare these on the state-of-the-art Qwen3-Omni 30B model across the 1000 audio questions spanning four reasoning categories and 16 subcategories. We evaluate three strategies which include structured prompt engineering through iterative error analysis, automated prompt optimization using DSPy MIPROv2 and Reinforced Self Training (ReST) with qLORA finetuning. Our structured prompting approach helped achieve an accuracy increase of 5.5% compared to baseline, while ReST based training led to a decrease in performance compared to baseline.

Subject: INTERSPEECH.2026 - Modelling and Learning