CLAIR-A: Leveraging Large Language Models to Judge Audio Captions

Automated Audio Captioning (AAC) aims to generate natural language descriptions of audio. Evaluating these machine-generated captions is a complex task, demanding an understanding of audio-scenes, sound-object recognition, temporal coherence, and environmental context. While existing methods focus on a subset of such capabilities, they often fail to provide a comprehensive score aligning with human judgment. Here, we introduce CLAIRA, a simple and flexible approach that uses large language models (lLMs) in a zero-shot manner to produce a “semantic distance” score for captions. In our experiments, ${CLAIR}_{A}$ more closely matches human ratings than other metrics, outperforming the domain-specific FENSE metric by 5.8% and surpassing the best general-purpose measure by up to 11% on the Clotho-Eval dataset. Moreover, CLAIRA allows the LLM to explain its scoring, with these explanations rated up to 30% better by human evaluators than those from baseline methods. The code for ${CLAIR}_{A}$ is made publicly available at https://github.com/DavidMChan/clair-a.

Paper

Similar papers

© 2026 NYSGPT2525 LLC