Total: 1
Audio-language models (ALMs) have shown strong generalization in standard audio classification tasks, yet their few-shot adaptation remains constrained by text-centric prompt learning, which under-adapts the audio encoder and struggles to distinguish subtle differences between accoustically similar classes. This text-dominated adaptation leads to imbalanced optimization across modalities and limits the discriminative capacity of ALMs under low-data regimes. To address this limitation, we propose MALP, a multi-modal prompt learning framework that jointly optimizes audio-specific, text-specific, and shared prompts. Specifically, MALP first enables modality-specific adaptation to capture complementary characteristics of audio and text, and then introduces shared prompts to strengthen cross-modal alignment while preserving modality-specific discriminability. Experiments on eleven benchmark datasets demonstrate consistent improvements over multiple strong baselines, achieving average gains of 7.21% over CoOp, 4.80% over CoCoOp, and 1.77% over PALM. Ablation studies further confirm the complementary roles of audio-specific and shared prompts, validating the effectiveness of multi-modal prompt learning for few-shot audio classification.