arXiv · 2020 · Preprint
Desplanques et al. · → Paper · Demo: ? · Code: ?
ECAPA-TDNN enhances the x-vector TDNN architecture for speaker verification by introducing Squeeze-and-Excitation Res2Blocks, multi-layer feature aggregation, and channel-dependent attentive statistics pooling, achieving substantial EER reductions on VoxCeleb and VoxSRC 2019 benchmarks.
Citation Stub
This paper is not a speech generation paper but is cited by the corpus. See Context in Speech Generation below for why it is relevant.
Context in Speech Generation
ECAPA-TDNN is a speaker embedding extractor that produces fixed-length representations from variable-length utterances via an enhanced TDNN with channel attention, hierarchical feature aggregation, and Squeeze-and-Excitation modules. TTS and SCA systems use it as a speaker encoder for conditioning synthesis on a reference speaker: the 192-dimensional embeddings extracted from the final fully-connected layer serve as compact speaker representations that can be supplied as conditioning input to zero-shot or speaker-adaptive synthesis models. The architecture provides a strong off-the-shelf speaker encoder trained on VoxCeleb2 (5994 speakers), making it a practical backbone for speaker conditioning in settings where training a dedicated speaker model is not feasible. Its channel-dependent attentive pooling, which computes separate temporal attention weights per channel, allows the embedding to capture diverse speaker-specific features (vowel formants, consonant characteristics) rather than collapsing them into a single attention pass.