zhou26c@interspeech_2026@ISCA

Total: 1

#1 UG-Bench: A Comprehensive Benchmark for Evaluating Large Audio-Language Models [PDF] [Copy] [Kimi] [REL]

Authors: Jiaming Zhou, Haoqin Sun, Hui Wang, Jinghua Zhao, Yuhang Jia, Shiyao Wang, Enzhi Wang, Shiwan Zhao, Yong Qin

Evaluating Large Audio-Language Models (LALMs) is challenging due to the diverse speech and audio tasks. Existing benchmarks often lack a unified framework, focusing either on understanding or generation. We introduce UG-Bench, a comprehensive benchmark systematically assessing LALMs across four competencies: speech perception, audio perception, speech generation, and spoken language understanding. Its decoupled design enables standardized comparisons and supports new tasks. Evaluating 11 LALMs and 5 specialized speech generation models, we find significant gaps in instruction following and generation quality, underscoring the need for better semantic understanding. UG-Bench offers a unified, scalable evaluation framework to advance LALM research.

Subject: INTERSPEECH.2026 - Analysis and Assessment