arXiv · 2023 · Preprint
Chu et al. (Alibaba Group) · → Paper · Demo: ✓ · Code: ✓
Qwen-Audio scales audio-language pre-training to over 30 tasks and diverse audio types (speech, sound, music) using a hierarchical multi-task training framework built on a Whisper-initialized encoder and Qwen-7B language model backbone.
Citation Stub
This paper is not a speech generation paper but is cited by the corpus. See Context in Speech Generation below for why it is relevant.
Context in Speech Generation
Qwen-Audio connects a 640M-parameter Whisper-large-v2 audio encoder to the Qwen-7B language model via large-scale multi-task co-training across speech recognition, translation, audio captioning, sound classification, music analysis, and question answering tasks. To avoid interference between the heterogeneous label spaces of different datasets, the system conditions the decoder on a sequence of hierarchical tags specifying transcription mode, language, task type, and output format. TTS and SCA research references Qwen-Audio as a representative audio-language model that demonstrates how pre-training on broad audio understanding tasks enables instruction-following interaction without task-specific fine-tuning. The companion Qwen-Audio-Chat model, produced via supervised instruction fine-tuning, provides multi-turn dialogue over multiple simultaneous audio inputs and serves as a reference point for speech-capable conversational agent design.
Wiki Connections
Related concepts: spoken-language-model
In-corpus papers cited by this work: 2309.16609 · 2210.13438 · 1808.10583 · 2302.13971 · 2307.09288 · 2007.10310 · 2305.11000 · 2308.16692