arXiv · 2023 · Preprint

Ma et al. · → Paper · Demo: ? · Code: ✓

emotion2vec is a universal speech emotion representation model pre-trained via self-supervised online distillation on 262 hours of unlabeled emotion data, combining utterance-level and frame-level losses to capture both global and local emotional content.

Citation Stub

This paper is not a speech generation paper but is cited by the corpus. See Context in Speech Generation below for why it is relevant.

Context in Speech Generation

emotion2vec provides pre-trained speech emotion features that TTS and expressive synthesis papers use as an upstream representation model for emotion-conditioned generation and evaluation. The model is pre-trained using an online teacher-student distillation paradigm (based on data2vec 2.0 architecture) on 262 hours of open-source emotion data, and its frozen features can be probed with simple linear layers to achieve state-of-the-art results on speech emotion recognition (SER) across IEMOCAP and 10 languages. For the speech generation community, emotion2vec fills a practical gap: it enables reliable automatic emotion labeling and evaluation without requiring task-specific fine-tuning of large SSL models such as WavLM or HuBERT, which carry significant compute overhead. Its demonstrated generalization across speech SER, song emotion recognition, emotion prediction in conversation, and sentiment analysis makes it a broadly applicable feature extractor in emotion-conditioned TTS and voice conversion evaluation pipelines.

Wiki Connections

emotion-synthesis self-supervised-speech evaluation-metrics