arXiv · 2025 · Preprint
Kumar et al. · → Paper · Demo: ? · Code: ✓
MMAU-Pro introduces a 5,305-instance benchmark for evaluating holistic audio intelligence across 49 distinct skills spanning speech, sound, music, and their combinations, with multi-hop reasoning questions, long-form audio up to 10 minutes, spatial audio, and open-ended response formats.
Citation Stub
This paper is not a speech generation paper but is cited by the corpus. See Context in Speech Generation below for why it is relevant.
Context in Speech Generation
MMAU-Pro provides a comprehensive testbed for assessing how well large audio-language models and spoken conversational agents understand speech, environmental sounds, and music under realistic, demanding conditions. The benchmark extends evaluation beyond the short, single-source clips of prior benchmarks by including long-form audio (up to 10 minutes in four duration bins), multi-audio reasoning over simultaneous streams, spatial audio with binaural recordings, and multicultural music spanning eight regions. Its instruction-following subset uses synthesized speech to probe paralinguistic and constraint-based generation capabilities, making it directly applicable to TTS and speech LM evaluation. Results across 22 models reveal a substantial human-machine gap: the best-performing system (Gemini 2.5 Flash) achieves 59.2% accuracy against a 77.9% human ceiling, with near-random performance on spatial audio and multi-audio tasks for most models.
Wiki Connections
- evaluation-metrics — introduces an embedding-based MCQ evaluation strategy and LLM-as-a-judge open-ended scoring framework, extending standard accuracy-based audio benchmarking.
- 2503.01743 (Phi-4-Mini) — evaluated on MMAU-Pro as one of the 22 models included in the benchmark comparison.
- 2407.10759 (Qwen2-Audio) — evaluated on MMAU-Pro; its architecture family (Qwen2.5-Omni) represents the strongest open-weights omni model in the benchmark results.
- 2507.08128 (Audio Flamingo 3) — evaluated on MMAU-Pro and achieves the best fully open-source accuracy (51.7%); cited as a model that supports multi-audio multi-turn dialogue.
- 2504.18425 (Kimi-Audio) — evaluated on MMAU-Pro as part of the large audio-language model comparison.
- 2501.15368 (Baichuan-Omni-1.5) — evaluated on MMAU-Pro as an omni-language model.
- 2410.21276 (GPT-4o System Card) — GPT-4o-Audio is evaluated on MMAU-Pro and achieves 52.5% overall accuracy.
- 2306.12925 (AudioPaLM) — cited in related work as an early LALM with audio comprehension capabilities relevant to the benchmark’s scope.
- 2506.04779 (MMSU) — cited as a closely related benchmark covering 47 speech skills and 5,000 QA pairs, positioned as a predecessor that MMAU-Pro extends with multi-audio and spatial dimensions.
- 2503.20215 (Qwen2.5-Omni) — evaluated on MMAU-Pro; Qwen2.5-Omni-7B achieves 52.2% accuracy and is among the stronger open-weights omni models.
- 2505.09388 (Qwen3) — cited in the experimental setup as Qwen3-235B-A22B-Instruct used as a text-only LLM judge in the cascaded system evaluation.