Total: 1
Deep learning–based automated audio captioning (AAC) typically optimizes a uniform objective, limiting contextual adaptability and fine-grained audio–text alignment. Although reinforcement learning improves evaluation metrics, existing approaches rely on simple and static reward formulations. We propose a context-adaptive AAC framework based on a symmetric dual mixture-of-experts (MoE) architecture optimized with group relative policy optimization (GRPO). On the policy side, LoRA-based experts are integrated into a pretrained BART decoder and activated via cross-attention routing conditioned on acoustic context. On the reward side, parallel experts evaluate semantic relevance, grammatical correctness, lexical diversity, and audio–text alignment, while a context-aware router adaptively weights these signals. Experiments on Clotho and Audio-Caps show improvements in caption quality and human preference. Generated captions are available on the https://symmetric-moe.github.io/.