Total: 1
Current Large Audio Language Models (LALMs) can describe the sounds present in audio but cannot localize their occurrence-a critical gap for applications. Prior efforts to add temporal reasoning either rely solely on coarse multi-choice distinctions or depend on LLM-inferred timestamps whose accuracy is unverifiable. We address this with a time-aware audio instruction-tuning dataset, AudioGround-IT, that provides deterministic boundary supervision, yielding 49.9K instructions over 835 hours of audio across four temporal tasks. To leverage this supervision for temporal grounding in LALMs, we propose AudioGround, a lightweight extension of SALMONN that uses a sliding-window Q-Former to compress encoder features and receives timestamp conditioning and absolute time embeddings. Evaluations on multiple temporal grounding benchmarks show that AudioGround substantially outperforms prior LALMs, demonstrating that deterministic boundary supervision transfers effectively to real-world audio.