dong26@interspeech_2026@ISCA

Total: 1

#1 Membership Inference Attacks against Large Audio Language Models [PDF] [Copy] [Kimi] [REL]

Authors: Jia-Kai Dong, Yu-Xiang Lin, Hung-yi Lee

We present the first systematic membership inference attack (MIA) evaluation of LALMs. Using Multi-modal Blind Baselines based on textual, spectral and prosodic features, we demonstrate that common audio datasets exhibit near-perfect train/test separability (AUC ≈ 1.0) even without model inference, thus MIA may primarily detect distribution shift. We therefore introduce a blind-baseline protocol to control for this confound. Under this protocol, we identify that the distribution-matched datasets enable reliable MIA evaluation without distribution-shift artifacts. We benchmark multiple MIA methods and conduct modality disentanglement experiments on these datasets. The results reveal that LALM memorization is cross-modal, arising only from binding a speaker's vocal identity with its text. These findings establish a principled standard for auditing LALMs beyond spurious correlations. Our codebase is available at https://github.com/snooow1029/ALM_MIA.

Subject: INTERSPEECH.2026 - Analysis and Assessment