Total: 1
Automated music captioning remains difficult for narrative-rich works (e.g., opera) where instrumentation, affect, and structure evolve over time. Existing LLM-based captioners depend on scarce paired audio-text data, and fine-tuning can overfit to dataset-specific phrasing, limiting transfer and controllable narration. We propose Music Artistic Captioning (MAC), a training-free framework that generates long-form, grounded descriptions by conditioning an instruction-following LLM on automatically extracted multi-scale audio evidence. MAC aggregates frame-, segment-, and song-level descriptors, from low-level acoustics (e.g., tempo) to predicted high-level semantic cues (e.g., instrumentation), into constrained pseudo-evidence for hierarchical captioning. To support user-preferred style, MAC uses a one-time iterative prompt optimization to balance narrative voice and evidence faithfulness. Experiments on an opera corpus and caption benchmarks show consistent gains over strong baselines.