arXiv · 2024 · Preprint

Yang et al. · → Paper · Demo: ? · Code: ✓

AIR-Bench introduces the first benchmark for evaluating the generative instruction-following capabilities of large audio-language models across speech, sound, and music, using a GPT-4-based evaluation framework that assesses both single-task foundation abilities and open-ended chat comprehension.

Citation Stub

This paper is not a speech generation paper but is cited by the corpus. See Context in Speech Generation below for why it is relevant.

Context in Speech Generation

AIR-Bench provides a hierarchical evaluation resource covering 19 audio tasks across speech, sound, and music in a foundation benchmark (approximately 19k single-choice questions), plus a 2k-instance open-ended chat benchmark requiring free-form generation. Its key methodological contribution is a unified GPT-4-based evaluation framework that scores model hypotheses against reference answers derived from audio meta-information, demonstrating near-perfect alignment with human judgment and enabling fair comparison of models that produce diverse output formats. This benchmark enables the speech and audio research community to assess large audio-language models on instruction-following and comprehension capabilities beyond narrow task-specific metrics like WER. The chat benchmark component, which includes a novel audio mixing strategy with loudness control and temporal dislocation, provides a pathway for evaluating models on complex, real-world audio scenarios relevant to spoken conversational agent development.

Wiki Connections

  • evaluation-metrics — AIR-Bench replaces traditional automatic metrics (WER, ROUGE) with GPT-4-based scoring for evaluating open-ended generation quality in large audio-language models.
  • spoken-language-model — AIR-Bench directly benchmarks large audio-language models on their ability to follow audio-grounded instructions and engage in open-ended dialogue, providing a diagnostic tool for spoken language model capabilities.
  • 2311.07919 (Qwen-Audio) — Qwen-Audio-Chat is one of the nine LALMs evaluated on AIR-Bench, achieving top performance in the foundation benchmark.
  • 2310.13289 (SALMONN) — SALMONN is evaluated as a representative universal audio-language model and achieves the highest chat benchmark score on mixed audio among evaluated models.
  • 2305.11000 (SpeechGPT) — SpeechGPT is included in the AIR-Bench evaluation, revealing limitations in its instruction-following and audio comprehension capabilities.