Interspeech · 2025 · Conference

Zhipeng Li et al. (South China University of Technology / Alibaba Group) · → Paper · Demo: ✓ · Code: ✗

A Context-Aware Memory (CAM) block for autoregressive LM-based TTS that compresses, retrieves, and dynamically updates both speech and text memory across paragraph-length sequences, enabling more coherent and expressive long-form synthesis at lower inference cost than prior methods.

Problem

Current TTS systems generate speech sentence by sentence and concatenate the results, ignoring the cross-sentence contextual dependencies that make paragraph-level speech coherent. This produces inconsistencies in style, timbre, and prosody across a paragraph. Prior long-context TTS methods either feed excessively long speech prompts (increasing KV-cache cost) or compress too little context (losing distal information). There is no efficient mechanism that maintains a concise, dynamically updated summary of the preceding paragraph.

Method

The system builds on the CosyVoice-LM backbone (50Hz speech tokenizer) and extends it with a Context-Aware Memory (CAM) block and a prefix mask.

CAM block operates in three stages per sentence:

  • Compress: A Perceiver Resampler compresses the target text into a fixed-length (32-token) latent query.
  • Retrieve: The compressed query cross-attends over a fused representation of the long-term memory (Mem_{n-1}) and the immediately preceding sentence’s speech/text history (History_{n-1}), extracting the most contextually relevant information as Mem*_n.
  • Update: The updated memory Mem_n is a learned scalar blend of Mem*n and Mem{n-1}: Mem_n = sigmoid(α) * Mem*_n + (1-sigmoid(α)) * Mem_{n-1}. This allows the model to balance recent context with long-term paragraph structure.

Separate CAM-Speech and CAM-Text blocks (shared structure, independent weights) process acoustic and textual modalities, producing Mem-S_n and Mem-T_n. These memory tokens, together with paragraph-level X-vectors from Cam++, are prepended to the LM input for sentence-level speech token generation. Flow matching and a vocoder convert speech tokens to waveforms.

Prefix mask replaces the standard causal mask for the memory/text prefix tokens, applying bidirectional attention to them while maintaining unidirectional attention over generated speech tokens. This improves in-context learning without breaking autoregressive generation.

Training used ~15,000 hours of Chinese Mandarin audiobooks for 100M steps with 10,000 tokens/batch dynamic batching. Inference requires only one preceding context plus fixed-length 64-token memory, compared to 5 full preceding sentences (MMCE-Qformer) or variable-length prompts up to ~900 tokens (CLAP-RAG).

Overview of CAM block (a) and Long-Context LM (b)

Key Results

On the internal Chinese Mandarin audiobook test set (Table 1):

SystemMOSCoMOSSpeechBERTCERSPK-SIM (mean, var)
Proposed3.7963.99280.4484.140%85.685 (0.019)
MMCE-Qformer3.5573.88579.0315.075%85.110 (0.021)
CLAP-RAG3.4893.71778.8926.234%84.920 (0.037)
Baseline3.4683.46077.7765.850%85.051 (0.035)

The proposed method achieves the best naturalness (MOS), coherence (CoMOS), content accuracy (CER), and speaker consistency (lowest SPK-SIM variance) while using only fixed 64-token context — the same cost as MMCE-Qformer but without requiring 5 full preceding sentences.

Ablation confirms that both speech memory (Mem-S) and text memory (Mem-T) contribute, with speech memory having the larger effect (likely because speech captures richer prosodic information than text). The prefix mask provides a significant boost to both naturalness and coherence.

Novelty Assessment

The main contribution is the CAM block design: a perceiver-resampler-based compression fed into a cross-attention retrieval over a blended long-term/local-context memory, updated dynamically per sentence. This is architecturally inspired by Infini-Attention (Munkhdalai et al., 2024) but applied to TTS. The idea of maintaining separate speech and text memory streams is genuinely new. The prefix mask for LM-based TTS is an independent but useful contribution. Overall, the contribution is primarily architectural — the training data, backbone, and evaluation are fairly standard. The context budget reduction (from variable to fixed 64 tokens) is practically valuable.

Field Significance

Moderate — This paper introduces a structured memory mechanism for paragraph-level TTS that simultaneously reduces inference context cost and improves prosodic coherence, addressing a practical limitation of sentence-concatenation pipelines. The CAM block design provides a reusable pattern for long-context conditioning in autoregressive speech LMs, though evaluation is limited to a single language and proprietary data.

Claims

  • Maintaining dynamically updated compressed memory across sentences improves naturalness and coherence in paragraph-level TTS compared to methods that use fixed-window preceding sentences. (§3.4.1, Table 1)
  • Speech context representations are more effective than text context representations for guiding prosodic coherence in autoregressive LM-based TTS, due to the one-to-many relationship between text and speech. (§3.4.2, Table 1)
  • Applying bidirectional attention to prefix tokens via a prefix mask enhances in-context learning in decoder-only TTS LMs without compromising autoregressive generation consistency. (§2.3, §3.4.2, Table 1)
  • Excessively long variable-length inference prompts increase hallucination and content errors in autoregressive TTS, while fixed-length context representations mitigate this instability. (§3.4.1, Table 1)

Limitations and Open Questions

  • Evaluation is monolingual (Chinese Mandarin only); generalizability to other languages is unconfirmed.
  • The gap to Ground Truth (MOS 4.406 vs 3.796) remains substantial; paragraph-level naturalness is still an open problem.
  • No code release; replication requires re-implementing CAM on a CosyVoice backbone.
  • The model is evaluated only on single-speaker audiobooks; how it handles multi-speaker paragraphs is unknown.
  • The fixed 32-embedding Perceiver Resampler bottleneck may lose fine-grained phonetic detail.

Wiki Connections

  • autoregressive-codec-tts — the CAM block extends LM-based codec TTS with paragraph-level memory
  • prosody-control — cross-sentence coherence in style and timbre is the core problem addressed
  • neural-codec — the system relies on a 50Hz speech tokenizer for discrete speech representation