arXiv · 2023 · Preprint

Seamless Communication et al. (Meta FAIR) · → Paper · Demo: ? · Code: ✓

Introduces a family of models (SeamlessM4T v2, SeamlessExpressive, SeamlessStreaming, and Seamless) enabling end-to-end multilingual speech-to-speech translation that preserves vocal style and prosody in a streaming, low-latency fashion.

Citation Stub

This paper is not a speech generation paper but is cited by the corpus. See Context in Speech Generation below for why it is relevant.

Context in Speech Generation

Seamless builds on SeamlessM4T to address two gaps in speech translation: expressive preservation and streaming inference. SeamlessExpressive preserves speech rate, pauses, and sentence-level vocal style across languages using a combination of prosody modeling and voice style transfer, supporting six languages including English, French, German, Italian, Mandarin, and Spanish. SeamlessStreaming uses the Efficient Monotonic Multihead Attention (EMMA) mechanism to produce low-latency simultaneous speech-to-speech and speech-to-text translations for up to 100 source languages without waiting for complete source utterances. Its expressive prosody transfer approach, multilingual speech unit framework (UnitY2), and automatic prosody evaluation metrics (AutoPCP, rhythm evaluation toolkit) are relevant to TTS and spoken language systems research targeting expressive quality in cross-lingual synthesis.

Wiki Connections

Related concepts: multilingual-tts, prosody-control, streaming-tts, self-supervised-speech