concept: neural-codec
last_updated: '2026-07-28'
paper_count: 183
papers:
- id: '1609.03499'
  published_date: "2016-09-12"
  entry_date: '2026-07-28'
  year: 2016
  venue: arXiv
  task:
  - TTS
  architecture:
  - autoregressive-LM
  relevance: medium
  evidence_role:
  - architecture_variant
  - acceleration_evidence
  - control_evidence
  current_role: foundational
  method_family:
  - autoregressive_codec_language_models
  claims:
  - claim_id: direct_generation_of_raw_audio_waveforms_without_intermediate_vocoder_parameters
    role: supports
    claim: Direct generation of raw audio waveforms, without intermediate vocoder parameters, produces substantially
      higher naturalness than parametric or concatenative synthesis pipelines as judged by human listeners.
    source: §3.2, Table 1
    evidence: Direct generation of raw audio waveforms, without intermediate vocoder parameters, produces substantially
      higher naturalness than parametric or concatenative synthesis pipelines as judged by human listeners.
    confidence: high
    relevance: medium
  - claim_id: dilated_causal_convolutions_enable_autoregressive_audio_models_to_achieve_receptive
    role: supports
    claim: Dilated causal convolutions enable autoregressive audio models to achieve receptive fields orders of
      magnitude larger than standard causal convolutions with comparable computational cost.
    source: §2.1, Figure 3
    evidence: Dilated causal convolutions enable autoregressive audio models to achieve receptive fields orders
      of magnitude larger than standard causal convolutions with comparable computational cost.
    confidence: high
    relevance: medium
  - claim_id: a_single_autoregressive_model_conditioned_on_speaker_identity_can_represent
    role: supports
    claim: A single autoregressive model conditioned on speaker identity can represent many voices with shared internal
      structure, and multi-speaker training improves per-speaker quality relative to single-speaker training.
    source: §3.1
    evidence: A single autoregressive model conditioned on speaker identity can represent many voices with shared
      internal structure, and multi-speaker training improves per-speaker quality relative to single-speaker training.
    confidence: high
    relevance: medium
  - claim_id: receptive_field_size_is_a_binding_constraint_for_prosodic_naturalness
    role: supports
    claim: 'Receptive field size is a binding constraint for prosodic naturalness: when the receptive field is insufficient
      to cover phrase-level F0 contours, prosody degrades even when segmental quality remains high.'
    source: §3.2
    evidence: 'Receptive field size is a binding constraint for prosodic naturalness: when the receptive field is
      insufficient to cover phrase-level F0 contours, prosody degrades even when segmental quality remains high.'
    confidence: high
    relevance: low
  - claim_id: autoregressive_raw_waveform_generation_achieves_high_naturalness_at_the_cost
    role: complicates
    claim: Autoregressive raw-waveform generation achieves high naturalness at the cost of sequential sample-level
      inference, creating a fundamental speed-quality trade-off that constrains deployment in real-time applications.
    source: §4, §3.2
    evidence: Autoregressive raw-waveform generation achieves high naturalness at the cost of sequential sample-level
      inference, creating a fundamental speed-quality trade-off that constrains deployment in real-time applications.
    confidence: high
    relevance: medium
  limitations:
  - Inference is strictly sequential at the sample level, requiring approximately one computation step per generated
    sample. At the reported generation rates (roughly 1.5× real-time compute), WaveNet is not suitable for real-time
    TTS deployment without hardware-specific optimisation or a parallel decoding approximation.
  - Evaluation is conducted on proprietary Google TTS databases, making direct replication by external researchers
    impossible. The MOS comparison is fair internally (same data, same test sentences for all systems) but cannot
    be directly compared to numbers from other published evaluations.
  - The receptive field of 240 milliseconds covers 3-4 phonemes. Long-range prosodic structure above the syllable
    and phrase level requires either an external F0 model or a substantially larger receptive field than dilated
    convolutions alone provide efficiently.
  - WaveNet as presented requires high-quality linguistic features derived from a text analysis front-end. It is
    not end-to-end trainable from text to waveform, deferring the alignment and duration prediction problems to
    external modules.
  caveats: []
- id: '2104.00355'
  published_date: "2021-04-01"
  entry_date: '2026-07-28'
  year: 2021
  venue: arXiv
  task:
  - TTS
  - VC
  architecture:
  - GAN
  - VAE
  relevance: high
  evidence_role:
  - core_evidence
  - control_evidence
  - evaluation_caution
  current_role: influential
  method_family:
  - gan_neural_codecs_and_decoders
  - vae_vector_quantized_codecs
  claims:
  - claim_id: ssl_content_representations_that_are_well_disentangled_from_speaker_identity
    role: supports
    claim: SSL content representations that are well-disentangled from speaker identity also exhibit stronger voice
      conversion performance, while representations that entangle speaker information perform worse at conversion
      but better at pitch reconstruction.
    source: §4, Table 1, Table 2
    evidence: SSL content representations that are well-disentangled from speaker identity also exhibit stronger
      voice conversion performance, while representations that entangle speaker information perform worse at conversion
      but better at pitch reconstruction.
    confidence: high
    relevance: medium
  - claim_id: discrete_speech_units_learned_by_ssl_models_can_form_the
    role: supports
    claim: Discrete speech units learned by SSL models can form the basis of an ultra-low-bitrate speech codec that
      outperforms classical parametric codecs in subjective quality.
    source: §4, Figure 2
    evidence: Discrete speech units learned by SSL models can form the basis of an ultra-low-bitrate speech codec
      that outperforms classical parametric codecs in subjective quality.
    confidence: high
    relevance: high
  - claim_id: among_self_supervised_content_encoders_hubert_units_carry_less_speaker
    role: supports
    claim: Among self-supervised content encoders, HuBERT units carry less speaker and pitch information than VQ-VAE
      units, making them better suited for downstream controllable synthesis.
    source: §4, Table 2
    evidence: Among self-supervised content encoders, HuBERT units carry less speaker and pitch information than
      VQ-VAE units, making them better suited for downstream controllable synthesis.
    confidence: high
    relevance: high
  - claim_id: pitch_and_speaker_identity_can_be_independently_conditioned_in_a
    role: supports
    claim: Pitch and speaker identity can be independently conditioned in a GAN vocoder through separate discrete
      token streams, enabling controllable F0 manipulation without retraining.
    source: §3, §4
    evidence: Pitch and speaker identity can be independently conditioned in a GAN vocoder through separate discrete
      token streams, enabling controllable F0 manipulation without retraining.
    confidence: high
    relevance: medium
  limitations:
  - The codec evaluation uses only 20 utterances from 5 VCTK speakers, all unseen during training but from the same
    corpus. Generalization to out-of-domain speech (conversational, noisy, or non-English) is untested.
  - The resynthesis MOS scores remain well below ground truth on both LJSpeech (3.66 vs. 4.33) and VCTK (3.41 vs.
    4.08), indicating a quality gap the system does not close. Disentanglement is evaluated indirectly through proxy
    metrics (EER, VDE, FFE) rather than a direct information-theoretic measure. The speaker encoder requires speaker
    embeddings from training-set speakers for the lookup-table variant; the d-vector approach generalizes but relies
    on a separately trained verification model. No ablation isolates the contribution of the F0 conditioning stream
    to final MOS. The MUSHRA scores in Figure 2 are visual only, making exact numerical comparison to baselines
    difficult to reproduce from the paper text alone.
  caveats: []
- id: '2209.03143'
  published_date: "2022-09-07"
  entry_date: '2026-07-28'
  year: 2022
  venue: arXiv
  task:
  - TTS
  - SCA
  architecture:
  - autoregressive-LM
  relevance: high
  evidence_role:
  - core_evidence
  - control_evidence
  current_role: foundational
  method_family:
  - autoregressive_codec_language_models
  claims:
  - claim_id: combining_self_supervised_semantic_tokens_with_codec_acoustic_tokens_in
    role: supports
    claim: Combining self-supervised semantic tokens with codec acoustic tokens in a hierarchical language model
      resolves the quality-versus-coherence tension that affects single-tokenizer audio language models.
    source: §III-B, §III-C, Table I
    evidence: Combining self-supervised semantic tokens with codec acoustic tokens in a hierarchical language model
      resolves the quality-versus-coherence tension that affects single-tokenizer audio language models.
    confidence: high
    relevance: high
  - claim_id: semantic_and_acoustic_tokens_in_speech_carry_complementary_information_semantic
    role: supports
    claim: 'Semantic and acoustic tokens in speech carry complementary information: semantic tokens primarily encode
      linguistic content and prosody, while acoustic tokens primarily encode speaker identity and recording conditions.'
    source: §IV-C, §IV-D, Tables II–III
    evidence: 'Semantic and acoustic tokens in speech carry complementary information: semantic tokens primarily
      encode linguistic content and prosody, while acoustic tokens primarily encode speaker identity and recording
      conditions.'
    confidence: high
    relevance: low
  - claim_id: autoregressive_language_modeling_over_discrete_audio_tokens_can_produce_speech
    role: supports
    claim: Autoregressive language modeling over discrete audio tokens can produce speech continuations indistinguishable
      from real speech to human listeners in an unpaired forced-choice test.
    source: §IV-G
    evidence: Autoregressive language modeling over discrete audio tokens can produce speech continuations indistinguishable
      from real speech to human listeners in an unpaired forced-choice test.
    confidence: high
    relevance: high
  - claim_id: the_semantic_to_acoustic_hierarchical_generation_pattern_transfers_across_audio
    role: supports
    claim: 'The semantic-to-acoustic hierarchical generation pattern transfers across audio domains: a model trained
      on piano music without symbolic notation also benefits from the two-tier tokenization.'
    source: §IV-I
    evidence: 'The semantic-to-acoustic hierarchical generation pattern transfers across audio domains: a model
      trained on piano music without symbolic notation also benefits from the two-tier tokenization.'
    confidence: high
    relevance: medium
  - claim_id: self_supervised_speech_representations_trained_with_masked_language_modeling_objectives
    role: supports
    claim: Self-supervised speech representations trained with masked language modeling objectives encode sufficient
      lexical and syntactic information to outperform earlier causal spoken language models on zero-resource linguistic
      benchmarks.
    source: §IV-E, Table IV
    evidence: Self-supervised speech representations trained with masked language modeling objectives encode sufficient
      lexical and syntactic information to outperform earlier causal spoken language models on zero-resource linguistic
      benchmarks.
    confidence: high
    relevance: medium
  limitations:
  - 'AudioLM is a continuation model only: it generates continuations of an audio prompt but cannot synthesise speech
    from a specified transcript. The paper explicitly frames TTS integration (encoder-decoder with text conditioning)
    as future work, which means the WER/CER results reflect acoustic fidelity to a given semantic token sequence,
    not instruction-following capability.'
  - The system requires three separately trained transformer models totaling approximately 0.9B parameters, plus
    frozen SoundStream and w2v-BERT models, making inference substantially more expensive than single-model TTS
    systems. Inference latency is not reported.
  - All speech experiments use English only; the model is trained on Libri-Light (audiobook speech), which is a
    clean, single-language, read-speech corpus. Generalisability to spontaneous speech, noise, or other languages
    is untested.
  - The paper does not evaluate unconditional generation quality against baselines with matched compute or training
    data, so the contribution of scale versus architecture is not fully disentangled.
  - The anti-spoofing classifier achieving 98.6% accuracy is trained on the same generative model it detects; its
    performance against other generators, or against an adversarially optimised version of AudioLM, is not assessed.
  caveats: []
- id: '2210.13438'
  published_date: "2022-10-24"
  entry_date: '2026-07-28'
  year: 2022
  venue: arXiv
  task:
  - codec
  architecture:
  - GAN
  relevance: medium
  evidence_role:
  - architecture_variant
  - acceleration_evidence
  - control_evidence
  current_role: foundational
  method_family:
  - gan_neural_codecs_and_decoders
  claims:
  - claim_id: a_multi_scale_stft_discriminator_alone_is_sufficient_for_perceptual
    role: supports
    claim: A multi-scale STFT discriminator alone is sufficient for perceptual quality in neural audio codecs, removing
      the need for waveform-domain discriminators.
    source: §4.5.1, Table 2
    evidence: A multi-scale STFT discriminator alone is sufficient for perceptual quality in neural audio codecs,
      removing the need for waveform-domain discriminators.
    confidence: high
    relevance: medium
  - claim_id: gradient_balancers_that_normalise_loss_contributions_by_expected_gradient_magnitude
    role: supports
    claim: Gradient balancers that normalise loss contributions by expected gradient magnitude substantially stabilise
      training when combining reconstruction, adversarial, and commitment losses with widely varying natural scales.
    source: §3.4, Table A.4
    evidence: Gradient balancers that normalise loss contributions by expected gradient magnitude substantially
      stabilise training when combining reconstruction, adversarial, and commitment losses with widely varying natural
      scales.
    confidence: high
    relevance: medium
  - claim_id: residual_vector_quantization_supports_variable_bitrate_operation_from_a_single
    role: supports
    claim: Residual vector quantization supports variable-bitrate operation from a single model by varying the number
      of active codebooks at inference, with each additional codebook yielding diminishing quality returns.
    source: §3.2, Table 1
    evidence: Residual vector quantization supports variable-bitrate operation from a single model by varying the
      number of active codebooks at inference, with each additional codebook yielding diminishing quality returns.
    confidence: high
    relevance: high
  - claim_id: auxiliary_transformer_language_models_over_rvq_codes_can_reduce_effective
    role: complicates
    claim: Auxiliary Transformer language models over RVQ codes can reduce effective bitrate by 25-40% through entropy
      coding without perceptual quality degradation, at the cost of increased latency.
    source: §3.3, §4.5
    evidence: Auxiliary Transformer language models over RVQ codes can reduce effective bitrate by 25-40% through
      entropy coding without perceptual quality degradation, at the cost of increased latency.
    confidence: high
    relevance: high
  - claim_id: neural_audio_codecs_outperform_traditional_dsp_codecs_at_low_bitrates
    role: supports
    claim: Neural audio codecs outperform traditional DSP codecs at low bitrates across both speech and music domains,
      with the quality gap widening as bitrate decreases.
    source: §4.5, Table 1, Figure 3
    evidence: Neural audio codecs outperform traditional DSP codecs at low bitrates across both speech and music
      domains, with the quality gap widening as bitrate decreases.
    confidence: high
    relevance: high
  limitations:
  - The 48 kHz model in non-streamable configuration operates slower than real time on CPU, limiting deployment
    without GPU acceleration or hardware-specific optimisation. Arithmetic coding precision issues (floating-point
    non-determinism across architectures) required a probability rounding workaround that the authors note may be
    insufficient for practical deployment.
  - The Transformer language model is small (5 layers, 200 channels) and neglects mutual information between codebooks
    at the same time step, leaving compression gains on the table. Music compression at 1.5 kbps still shows significant
    perceptual degradation. Training datasets mix multiple licenses, which may complicate commercial use. No evaluation
    of robustness to codec chaining (encode-decode-re-encode) or of downstream task performance degradation from
    quantization artefacts is provided.
  caveats: []
- id: '2301.02111'
  published_date: "2023-01-05"
  entry_date: '2026-07-28'
  year: 2023
  venue: arXiv
  task:
  - TTS
  - codec
  architecture:
  - autoregressive-LM
  relevance: high
  evidence_role:
  - core_evidence
  - control_evidence
  - evaluation_caution
  current_role: foundational
  method_family:
  - autoregressive_codec_language_models
  claims:
  - claim_id: treating_tts_as_conditional_language_modeling_over_discrete_codec_tokens
    role: supports
    claim: Treating TTS as conditional language modeling over discrete codec tokens enables zero-shot speaker generalisation
      as in-context learning, without speaker-specific fine-tuning or engineered speaker encoders.
    source: §4.1, §5.2
    evidence: Treating TTS as conditional language modeling over discrete codec tokens enables zero-shot speaker
      generalisation as in-context learning, without speaker-specific fine-tuning or engineered speaker encoders.
    confidence: high
    relevance: high
  - claim_id: training_on_large_scale_semi_supervised_speech_data_even_with
    role: supports
    claim: Training on large-scale semi-supervised speech data, even with noisy transcriptions and diverse acoustic
      conditions, yields stronger generalisation to unseen speakers than training on smaller clean corpora.
    source: §1, §5.2
    evidence: Training on large-scale semi-supervised speech data, even with noisy transcriptions and diverse acoustic
      conditions, yields stronger generalisation to unseen speakers than training on smaller clean corpora.
    confidence: high
    relevance: medium
  - claim_id: the_hierarchical_structure_of_residual_vector_quantization_supports_a_two
    role: supports
    claim: The hierarchical structure of residual vector quantization supports a two-stage AR+NAR generation pipeline
      in which first-codebook tokens carry speaker identity and subsequent codebooks refine fine acoustic detail.
    source: §4.2
    evidence: The hierarchical structure of residual vector quantization supports a two-stage AR+NAR generation
      pipeline in which first-codebook tokens carry speaker identity and subsequent codebooks refine fine acoustic
      detail.
    confidence: high
    relevance: high
  - claim_id: speaker_similarity_in_zero_shot_codec_tts_improves_monotonically_with
    role: supports
    claim: Speaker similarity in zero-shot codec TTS improves monotonically with acoustic prompt length, suggesting
      that speaker identity modelling does not saturate within a few seconds.
    source: §5.3, Table 6
    evidence: Speaker similarity in zero-shot codec TTS improves monotonically with acoustic prompt length, suggesting
      that speaker identity modelling does not saturate within a few seconds.
    confidence: high
    relevance: high
  - claim_id: stochastic_sampling_in_autoregressive_codec_generation_introduces_output_diversity_varying
    role: supports
    claim: Stochastic sampling in autoregressive codec generation introduces output diversity — varying speech rate,
      prosody, and accent realisation — that is both a feature for data augmentation and a complication for deterministic
      evaluation.
    source: §4.3, §5.4
    evidence: Stochastic sampling in autoregressive codec generation introduces output diversity — varying speech
      rate, prosody, and accent realisation — that is both a feature for data augmentation and a complication for
      deterministic evaluation.
    confidence: high
    relevance: high
  limitations:
  - 'Synthesis robustness is a material constraint: the autoregressive first-stage LM exhibits attention alignment
    failures that cause word deletions, insertions, and repetitions. WER on LibriSpeech test-clean is 5.9%, nearly
    three times the ground-truth rate of 2.2%. This limits deployment in applications requiring high intelligibility
    and is the principal limitation acknowledged by the authors.'
  - Training data is entirely English audiobook speech, limiting coverage of accented, spontaneous, or conversational
    speech styles. The two-model architecture (AR + NAR) adds inference complexity; the authors note that a single
    universal model is a natural future direction. Model parameter count is not reported, making compute comparisons
    with non-codec TTS systems difficult. The evaluation covers two benchmarks only; no noise-robustness or cross-domain
    experiments are included. Zero-shot voice cloning from 3 seconds of audio raises misuse risks that the paper
    acknowledges but does not experimentally mitigate.
  caveats: []
- id: '2301.11325'
  published_date: "2023-01-26"
  entry_date: '2026-07-28'
  year: 2023
  venue: arXiv
  task: []
  architecture:
  - autoregressive-LM
  relevance: medium
  evidence_role:
  - architecture_variant
  - evaluation_caution
  current_role: influential
  method_family:
  - autoregressive_codec_language_models
  claims:
  - claim_id: hierarchical_autoregressive_modeling_over_semantic_and_acoustic_tokens_enables_long
    role: supports
    claim: Hierarchical autoregressive modeling over semantic and acoustic tokens enables long-form music generation
      (several minutes) with temporal coherence at 24 kHz.
    source: §3.2, §6
    evidence: Hierarchical autoregressive modeling over semantic and acoustic tokens enables long-form music generation
      (several minutes) with temporal coherence at 24 kHz.
    confidence: high
    relevance: medium
  - claim_id: a_joint_audio_text_embedding_space_can_substitute_for_paired
    role: supports
    claim: A joint audio-text embedding space can substitute for paired text-audio supervision at training time,
      allowing generative models to be trained on audio-only corpora and conditioned on text at inference.
    source: §3.1, §4.2
    evidence: A joint audio-text embedding space can substitute for paired text-audio supervision at training time,
      allowing generative models to be trained on audio-only corpora and conditioned on text at inference.
    confidence: high
    relevance: medium
  - claim_id: semantic_token_intermediaries_improve_adherence_to_text_descriptions_in_hierarchical
    role: supports
    claim: Semantic token intermediaries improve adherence to text descriptions in hierarchical audio generation
      beyond what direct acoustic token prediction achieves.
    source: §5, Table 1
    evidence: Semantic token intermediaries improve adherence to text descriptions in hierarchical audio generation
      beyond what direct acoustic token prediction achieves.
    confidence: high
    relevance: medium
  - claim_id: for_text_conditioned_music_generation_perceptual_audio_quality_fad_and
    role: supports
    claim: For text-conditioned music generation, perceptual audio quality (FAD) and semantic text alignment (MCC,
      KLD) are complementary evaluation axes that do not always correlate with each other.
    source: §4.4, Table 1
    evidence: For text-conditioned music generation, perceptual audio quality (FAD) and semantic text alignment
      (MCC, KLD) are complementary evaluation axes that do not always correlate with each other.
    confidence: high
    relevance: medium
  - claim_id: large_autoregressive_audio_lms_trained_on_extensive_unlabeled_corpora_memorize
    role: supports
    claim: Large autoregressive audio LMs trained on extensive unlabeled corpora memorize only a small fraction
      of training sequences exactly, but approximate semantic matches affect a higher proportion of generated outputs
      under targeted prompting.
    source: §5, Figure 3
    evidence: Large autoregressive audio LMs trained on extensive unlabeled corpora memorize only a small fraction
      of training sequences exactly, but approximate semantic matches affect a higher proportion of generated outputs
      under targeted prompting.
    confidence: high
    relevance: medium
  limitations:
  - MCC, one of the two primary text-adherence metrics, is computed using MuLan itself, the same model used for
    conditioning. This circularity biases the metric in MusicLM's favour relative to baselines that do not use MuLan
    representations.
  - 'The system inherits MuLan''s known weaknesses: negations in text prompts are not well handled, and temporal
    ordering of described events is not captured. Vocal quality and lyrics generation are not supported and are
    noted as future work. The 280,000-hour training corpus is proprietary and not released, limiting reproducibility.
    The model weights are not released. Evaluation is conducted against two relatively weak baselines (Mubert, a
    rule-based tag-to-audio API; Riffusion, a fine-tuned image diffusion model on spectrograms); no contemporary
    neural music generation system of comparable scale is included in comparisons.'
  caveats: []
- id: '2303.03926'
  published_date: "2023-03-07"
  entry_date: '2026-07-28'
  year: 2023
  venue: arXiv
  task:
  - TTS
  architecture:
  - autoregressive-LM
  relevance: high
  evidence_role:
  - core_evidence
  - control_evidence
  current_role: influential
  method_family:
  - autoregressive_codec_language_models
  claims:
  - claim_id: large_scale_multilingual_codec_language_models_can_transfer_speaker_identity
    role: supports
    claim: Large-scale multilingual codec language models can transfer speaker identity, emotion, and acoustic environment
      across languages from a single source utterance without paired bilingual data.
    source: §3, §5.3, Table 3
    evidence: Large-scale multilingual codec language models can transfer speaker identity, emotion, and acoustic
      environment across languages from a single source utterance without paired bilingual data.
    confidence: high
    relevance: high
  - claim_id: language_id_conditioning_is_essential_for_native_sounding_accent_in
    role: supports
    claim: 'Language ID conditioning is essential for native-sounding accent in cross-lingual codec TTS: removing
      it significantly degrades accent scores even while modestly improving speaker similarity.'
    source: §5.5, Table 6
    evidence: 'Language ID conditioning is essential for native-sounding accent in cross-lingual codec TTS: removing
      it significantly degrades accent scores even while modestly improving speaker similarity.'
    confidence: high
    relevance: high
  - claim_id: in_context_learning_with_acoustic_token_prompts_provides_stronger_cross
    role: supports
    claim: In-context learning with acoustic token prompts provides stronger cross-lingual voice preservation than
      speaker embedding approaches across both TTS and speech-to-speech translation tasks.
    source: §5.3, §5.4, Tables 3, 5
    evidence: In-context learning with acoustic token prompts provides stronger cross-lingual voice preservation
      than speaker embedding approaches across both TTS and speech-to-speech translation tasks.
    confidence: high
    relevance: medium
  - claim_id: the_ar_nar_two_stage_codec_language_model_architecture_extends
    role: supports
    claim: The AR/NAR two-stage codec language model architecture extends naturally to cross-lingual generation
      by treating bilingual phoneme sequences as concatenated prompts.
    source: §3.2, §3.4
    evidence: The AR/NAR two-stage codec language model architecture extends naturally to cross-lingual generation
      by treating bilingual phoneme sequences as concatenated prompts.
    confidence: high
    relevance: high
  limitations:
  - Evaluations cover only English and Chinese, and test sets are small (40 speakers, 1373 samples for TTS; 14 speakers,
    350 utterances for S2ST on EMIME). Generalisation to typologically distant language pairs, more than two languages,
    or lower-resource settings is untested.
  - The EMIME dataset is bilingual by construction (same speakers recorded in both languages), making the ASV upper
    bound (tgt vs. src) unusually informative but also artificially favourable for a cross-lingual system. The ASV
    gap between the model and the upper bound for English-to-Chinese (0.48 vs. 0.58) suggests substantial room for
    improvement in voice transferability in that direction.
  - 'The language ID ablation reveals an inherent tension: stronger accent control comes at the cost of voice similarity.
    The right operating point depends on application, and the paper does not explore soft or learnable trade-off
    mechanisms. No model size figures are reported, making cost comparison difficult. S2ST quality is bottlenecked
    by translation quality; the paper does not disentangle acoustic and linguistic error sources beyond the oracle-text
    upper bound analysis.'
  caveats: []
- id: '2304.09116'
  published_date: "2023-04-18"
  entry_date: '2026-07-28'
  year: 2023
  venue: arXiv
  task:
  - TTS
  - VC
  - singing
  architecture:
  - diffusion
  - VAE
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: influential
  method_family:
  - diffusion_latent_codec_generators
  - vae_vector_quantized_codecs
  claims:
  - claim_id: latent_diffusion_models_operating_on_continuous_codec_vectors_avoid_the
    role: supports
    claim: Latent diffusion models operating on continuous codec vectors avoid the word-skipping and repetition
      errors that arise from autoregressive generation over long discrete token sequences.
    source: §2.3, §5.3, Table 7
    evidence: Latent diffusion models operating on continuous codec vectors avoid the word-skipping and repetition
      errors that arise from autoregressive generation over long discrete token sequences.
    confidence: high
    relevance: high
  - claim_id: speech_prompting_via_in_context_learning_during_training_enables_zero
    role: supports
    claim: Speech prompting via in-context learning during training enables zero-shot speaker adaptation without
      requiring speaker embeddings or multi-step speaker encoding pipelines.
    source: §3.3, §5.5
    evidence: Speech prompting via in-context learning during training enables zero-shot speaker adaptation without
      requiring speaker embeddings or multi-step speaker encoding pipelines.
    confidence: high
    relevance: medium
  - claim_id: prosody_adherence_in_zero_shot_tts_improves_monotonically_with_the
    role: supports
    claim: Prosody adherence in zero-shot TTS improves monotonically with the length of the reference speech prompt,
      at least up to 10 seconds.
    source: §5.5, Table 10
    evidence: Prosody adherence in zero-shot TTS improves monotonically with the length of the reference speech
      prompt, at least up to 10 seconds.
    confidence: high
    relevance: low
  - claim_id: non_autoregressive_tts_architectures_maintain_near_zero_error_rates_on
    role: supports
    claim: Non-autoregressive TTS architectures maintain near-zero error rates on adversarially difficult phoneme
      sequences where autoregressive models degrade significantly.
    source: §5.3, Table 7
    evidence: Non-autoregressive TTS architectures maintain near-zero error rates on adversarially difficult phoneme
      sequences where autoregressive models degrade significantly.
    confidence: high
    relevance: medium
  - claim_id: a_system_trained_jointly_on_speech_and_singing_data_can
    role: supports
    claim: A system trained jointly on speech and singing data can synthesise singing in novel timbres using only
      a speech reference prompt, demonstrating cross-modal timbre transfer within a shared latent space.
    source: §5.6
    evidence: A system trained jointly on speech and singing data can synthesise singing in novel timbres using
      only a speech reference prompt, demonstrating cross-modal timbre transfer within a shared latent space.
    confidence: high
    relevance: medium
  limitations:
  - The direct comparison with VALL-E is based on VALL-E demo page samples rather than a controlled shared test
    set — the 16 compared utterances are cherry-picked by the VALL-E authors and may not be representative. This
    limits the strength of the head-to-head quality claim.
  - The model is described as still underfitting at 300K training steps, meaning reported results are likely below
    the system's ceiling performance. Inference requires 150 diffusion steps (ODE solver), and 1000 steps for singing,
    which is slow for real-time deployment. The paper cites consistency models as future work for acceleration.
    Training and evaluation are English-only, so multilingual generalisation is uncharacterised. The singing dataset
    is approximately 30 hours of web-crawled data with no formal provenance or quality validation beyond alignment
    filtering, which raises questions about singing style coverage. Code and model weights are not publicly released,
    limiting reproducibility.
  caveats: []
- id: '2305.02765'
  published_date: "2023-05-04"
  entry_date: '2026-07-28'
  year: 2023
  venue: arXiv
  task:
  - codec
  architecture:
  - GAN
  - VAE
  relevance: medium
  evidence_role:
  - architecture_variant
  - evaluation_caution
  current_role: influential
  method_family:
  - gan_neural_codecs_and_decoders
  - vae_vector_quantized_codecs
  claims:
  - claim_id: grouping_residual_vector_quantization_into_parallel_chains_rather_than_a
    role: supports
    claim: Grouping residual vector quantization into parallel chains rather than a single sequential chain improves
      reconstruction quality per codebook, enabling competitive fidelity with fewer total quantizers.
    source: §3.3, Table 1
    evidence: Grouping residual vector quantization into parallel chains rather than a single sequential chain improves
      reconstruction quality per codebook, enabling competitive fidelity with fewer total quantizers.
    confidence: high
    relevance: high
  - claim_id: the_burden_that_codec_codebook_count_imposes_on_downstream_generation
    role: supports
    claim: The burden that codec codebook count imposes on downstream generation models is a practical constraint
      that drives codec architecture choices independently of raw reconstruction quality.
    source: §1 Introduction, §3.3
    evidence: The burden that codec codebook count imposes on downstream generation models is a practical constraint
      that drives codec architecture choices independently of raw reconstruction quality.
    confidence: high
    relevance: high
  - claim_id: objective_speech_quality_metrics_such_as_pesq_and_stoi_are
    role: supports
    claim: Objective speech quality metrics such as PESQ and STOI are insufficient alone to characterise codec reconstruction
      quality, and subjective evaluation is necessary but often omitted in codec research.
    source: §6 Limitations
    evidence: Objective speech quality metrics such as PESQ and STOI are insufficient alone to characterise codec
      reconstruction quality, and subjective evaluation is necessary but often omitted in codec research.
    confidence: high
    relevance: high
  - claim_id: publicly_available_training_code_and_pre_trained_baselines_for_neural
    role: supports
    claim: Publicly available training code and pre-trained baselines for neural audio codecs are necessary for
      reproducible research, as previously these were unavailable for EnCodec and SoundStream.
    source: §5 Conclusion, §6 Limitations
    evidence: Publicly available training code and pre-trained baselines for neural audio codecs are necessary for
      reproducible research, as previously these were unavailable for EnCodec and SoundStream.
    confidence: high
    relevance: medium
  limitations:
  - 'No subjective evaluation is included. The paper''s own §6 acknowledges this as a limitation: "Subjective evaluation
    is always the best choice, but this part is missed in this study." All quality comparisons rest solely on PESQ
    and STOI, which the authors themselves note may not accurately reflect perceptual quality.'
  - Beyond the missing subjective evaluation, the paper does not validate GRVQ on downstream generation tasks. The
    claim that 4 codebooks reduce burden on generation models is plausible and consistent with the motivation, but
    no TTS or audio LM experiments are included. The training data is limited to English and Chinese speech from
    public TTS corpora; generalisation to music, environmental sound, or noisy/spontaneous speech is not tested.
    Finally, the paper trains at 16kHz and 24kHz only; higher sample rates (44.1kHz, 48kHz) used in music and high-fidelity
    audio applications are not addressed.
  caveats: []
- id: '2305.07243'
  published_date: "2023-05-12"
  entry_date: '2026-07-28'
  year: 2023
  venue: arXiv
  task:
  - TTS
  architecture:
  - autoregressive-LM
  - diffusion
  - VAE
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: influential
  method_family:
  - autoregressive_codec_language_models
  - diffusion_latent_codec_generators
  - vae_vector_quantized_codecs
  claims:
  - claim_id: conditioning_a_diffusion_decoder_on_the_continuous_latent_activations_of
    role: supports
    claim: Conditioning a diffusion decoder on the continuous latent activations of an autoregressive model rather
      than its discrete token outputs substantially improves output quality in a cascaded AR-diffusion TTS pipeline.
    source: §2.2.2, Appendix B.4
    evidence: Conditioning a diffusion decoder on the continuous latent activations of an autoregressive model rather
      than its discrete token outputs substantially improves output quality in a cascaded AR-diffusion TTS pipeline.
    confidence: high
    relevance: medium
  - claim_id: contrastive_re_ranking_of_multiple_autoregressive_candidates_using_a_text
    role: supports
    claim: Contrastive re-ranking of multiple autoregressive candidates using a text-speech discriminator measurably
      improves the final output quality of a TTS system without requiring the expensive decoder to process every
      candidate.
    source: §2.3, §4
    evidence: Contrastive re-ranking of multiple autoregressive candidates using a text-speech discriminator measurably
      improves the final output quality of a TTS system without requiring the expensive decoder to process every
      candidate.
    confidence: high
    relevance: medium
  - claim_id: applying_image_generation_scaling_techniques_large_scale_self_supervised_data
    role: supports
    claim: Applying image-generation scaling techniques (large-scale self-supervised data, generalist transformer
      architectures, multi-stage AR-then-diffusion generation) to speech synthesis yields high-expressiveness multi-speaker
      TTS even when trained by a single researcher on commodity hardware.
    source: §7
    evidence: Applying image-generation scaling techniques (large-scale self-supervised data, generalist transformer
      architectures, multi-stage AR-then-diffusion generation) to speech synthesis yields high-expressiveness multi-speaker
      TTS even when trained by a single researcher on commodity hardware.
    confidence: high
    relevance: medium
  - claim_id: building_a_large_scale_tts_training_corpus_by_scraping_and
    role: supports
    claim: Building a large-scale TTS training corpus by scraping and filtering internet audio (audiobooks, podcasts)
      with automatic transcription is a viable path to tens-of-thousands-of-hours datasets without manual labelling.
    source: §5, Appendix A
    evidence: Building a large-scale TTS training corpus by scraping and filtering internet audio (audiobooks, podcasts)
      with automatic transcription is a viable path to tens-of-thousands-of-hours datasets without manual labelling.
    confidence: high
    relevance: medium
  limitations:
  - No formal listening test or MOS table is reported. The primary quality claim rests on informal sample comparisons;
    the paper's own evaluation suite (CLVP-FID) is not a standard benchmark. This makes it difficult to place TorToise
    on the same scale as systems evaluated under controlled conditions.
  - 'Additional limitations noted in the paper include: slow inference due to multi-pass DDIM sampling (64 steps,
    large candidate sets) making real-time use impractical; fixed positional encodings in the AR model limiting
    maximum utterance length; the CLVP re-ranking model was only trained on sequences up to 13 seconds, degrading
    on longer outputs; and the VQVAE codebook embedding dimension was not constrained, which subsequent work showed
    to be a missed performance gain. Training was resource-constrained to 8 consumer GPUs, so model scale is limited
    relative to what the paper''s own training curves suggest would still improve results.'
  caveats: []
- id: '2305.09636'
  published_date: "2023-05-16"
  entry_date: '2026-07-28'
  year: 2023
  venue: arXiv
  task:
  - TTS
  - SCA
  architecture:
  - autoregressive-LM
  relevance: high
  evidence_role:
  - core_evidence
  - acceleration_evidence
  - control_evidence
  current_role: influential
  method_family:
  - autoregressive_codec_language_models
  claims:
  - claim_id: parallel_iterative_masked_decoding_adapted_to_rvq_structure_enables_acoustic
    role: supports
    claim: Parallel, iterative masked decoding adapted to RVQ structure enables acoustic token generation two orders
      of magnitude faster than autoregressive generation at matched perceptual quality.
    source: §4.3, Figure 3
    evidence: Parallel, iterative masked decoding adapted to RVQ structure enables acoustic token generation two
      orders of magnitude faster than autoregressive generation at matched perceptual quality.
    confidence: high
    relevance: high
  - claim_id: non_autoregressive_rvq_level_by_level_decoding_maintains_better_voice
    role: supports
    claim: Non-autoregressive RVQ-level-by-level decoding maintains better voice and acoustic consistency over long
      sequences than autoregressive chunk-and-prompt approaches.
    source: §4.2, Table 1, Figure 2
    evidence: Non-autoregressive RVQ-level-by-level decoding maintains better voice and acoustic consistency over
      long sequences than autoregressive chunk-and-prompt approaches.
    confidence: high
    relevance: high
  - claim_id: fine_level_rvq_tokens_are_conditionally_independent_given_coarser_tokens
    role: supports
    claim: Fine-level RVQ tokens are conditionally independent given coarser tokens and can be decoded greedily
      in a single pass without measurable quality loss.
    source: §3.3, §4.3
    evidence: Fine-level RVQ tokens are conditionally independent given coarser tokens and can be decoded greedily
      in a single pass without measurable quality loss.
    confidence: high
    relevance: high
  - claim_id: confidence_based_iterative_decoding_provides_a_meaningful_quality_gain_over
    role: supports
    claim: Confidence-based iterative decoding provides a meaningful quality gain over greedy decoding at the coarsest
      RVQ level, but additional iterations at finer levels yield no significant improvement for speech.
    source: §4.3, Figure 4
    evidence: Confidence-based iterative decoding provides a meaningful quality gain over greedy decoding at the
      coarsest RVQ level, but additional iterations at finer levels yield no significant improvement for speech.
    confidence: high
    relevance: high
  - claim_id: coupling_a_text_to_semantic_token_model_with_an_efficient
    role: supports
    claim: Coupling a text-to-semantic token model with an efficient acoustic generator enables real-time synthesis
      of controllable multi-speaker dialogue at 30-second horizons.
    source: §5
    evidence: Coupling a text-to-semantic token model with an efficient acoustic generator enables real-time synthesis
      of controllable multi-speaker dialogue at 30-second horizons.
    confidence: high
    relevance: low
  limitations:
  - Evaluation uses a DNSMOS-style estimator rather than human listening tests for audio quality comparisons, and
    the subjective baseline is carried over from earlier AudioLM papers rather than re-run. Direct perceptual comparisons
    between SoundStorm and AudioLM on the same conditions by human raters are not reported.
  - The conditioning mechanism requires time-aligned semantic tokens, restricting drop-in use to AudioLM-family
    pipelines with compatible semantic tokenisers. Extension to cross-attention conditioning or unconditional generation
    is left as future work. The 350M acoustic model trained on LibriLight is English-only; no multilingual evaluation
    is presented. The dialogue synthesis component depends on a proprietary 100k-hour dialogue corpus, limiting
    reproducibility of that subsystem. The ablation on decoding iterations is limited to speech; the authors hypothesise
    that multiple fine-level iterations may matter for non-speech audio, but this is not tested.
  caveats: []
- id: '2305.11000'
  published_date: "2023-05-18"
  entry_date: '2026-07-28'
  year: 2023
  venue: arXiv
  task:
  - SCA
  - TTS
  architecture:
  - autoregressive-LM
  relevance: medium
  evidence_role:
  - architecture_variant
  current_role: influential
  method_family:
  - autoregressive_codec_language_models
  claims:
  - claim_id: expanding_an_llm_s_token_vocabulary_with_discretised_speech_units
    role: supports
    claim: Expanding an LLM's token vocabulary with discretised speech units enables a single model to perform both
      speech comprehension and speech generation without a cascade pipeline.
    source: §4.1
    evidence: Expanding an LLM's token vocabulary with discretised speech units enables a single model to perform
      both speech comprehension and speech generation without a cascade pipeline.
    confidence: high
    relevance: medium
  - claim_id: a_multi_stage_training_curriculum_separating_modality_adaptation_cross_modal
    role: supports
    claim: A multi-stage training curriculum, separating modality adaptation, cross-modal instruction tuning, and
      chain-of-modality alignment, is necessary to acquire reliable cross-modal instruction-following from an LLM
      backbone.
    source: §4.2
    evidence: A multi-stage training curriculum, separating modality adaptation, cross-modal instruction tuning,
      and chain-of-modality alignment, is necessary to acquire reliable cross-modal instruction-following from an
      LLM backbone.
    confidence: high
    relevance: medium
  - claim_id: the_chain_of_modality_pattern_generating_a_text_intermediate_before
    role: supports
    claim: The chain-of-modality pattern, generating a text intermediate before the speech response, is a practical
      mechanism for transferring LLM reasoning capability to speech output.
    source: §3.2, §4.2
    evidence: The chain-of-modality pattern, generating a text intermediate before the speech response, is a practical
      mechanism for transferring LLM reasoning capability to speech output.
    confidence: high
    relevance: medium
  - claim_id: large_scale_instruction_dataset_construction_via_gpt_4_generated_task
    role: supports
    claim: Large-scale instruction dataset construction via GPT-4-generated task descriptions applied to existing
      ASR corpora is a scalable approach to bootstrapping cross-modal training data.
    source: §3.1
    evidence: Large-scale instruction dataset construction via GPT-4-generated task descriptions applied to existing
      ASR corpora is a scalable approach to bootstrapping cross-modal training data.
    confidence: high
    relevance: medium
  limitations:
  - 'The paper provides no quantitative evaluation: no MOS, WER, or speaker similarity scores are reported, and
    no comparison to cascade baselines is made. All results are case studies. Claims about spoken dialogue quality
    and instruction-following capability cannot be independently verified from the paper alone.'
  - 'Additional limitations acknowledged by the authors:'
  - '- The system does not model paralinguistic information; it cannot generate responses with different emotional
    tones or speaking styles. - The chain-of-modality design requires generating a full text response before producing
    speech, introducing latency incompatible with real-time interaction. - Context length limitations (2048 tokens)
    prevent true multi-turn dialogue; only single-turn exchanges are demonstrated. - HuBERT units discard fine-grained
    acoustic detail (prosody, speaker identity beyond the vocoder''s speaker embedding), limiting expressiveness
    relative to audio codec approaches. - The unit vocoder is trained separately and is not end-to-end jointly optimised
    with the LLM, leaving a quality gap between unit-based and codec-based approaches.'
  caveats: []
- id: '2306.00814'
  published_date: "2023-06-01"
  entry_date: '2026-07-28'
  year: 2023
  venue: arXiv
  task:
  - TTS
  architecture:
  - GAN
  relevance: medium
  evidence_role:
  - architecture_variant
  current_role: influential
  method_family:
  - gan_neural_codecs_and_decoders
  claims:
  - claim_id: maintaining_constant_temporal_resolution_throughout_a_gan_vocoder_with_istft
    role: supports
    claim: Maintaining constant temporal resolution throughout a GAN vocoder, with ISTFT as the sole upsampling
      step, eliminates aliasing artefacts and dramatically reduces inference cost without sacrificing perceptual
      quality.
    source: §3.1, §4.3, Table 6
    evidence: Maintaining constant temporal resolution throughout a GAN vocoder, with ISTFT as the sole upsampling
      step, eliminates aliasing artefacts and dramatically reduces inference cost without sacrificing perceptual
      quality.
    confidence: high
    relevance: medium
  - claim_id: implicit_phase_wrapping_via_a_unit_circle_activation_is_essential
    role: supports
    claim: Implicit phase wrapping via a unit-circle activation is essential for stable GAN training of complex-valued
      spectrogram generators; alternatives that clamp or clip phase angles substantially degrade output quality.
    source: §3.2, §4.1.1, Table 1
    evidence: Implicit phase wrapping via a unit-circle activation is essential for stable GAN training of complex-valued
      spectrogram generators; alternatives that clamp or clip phase angles substantially degrade output quality.
    confidence: high
    relevance: medium
  - claim_id: fourier_domain_vocoders_reduce_periodicity_errors_more_effectively_than_time
    role: supports
    claim: Fourier-domain vocoders reduce periodicity errors more effectively than time-domain GANs, suggesting
      that modelling harmonics in the frequency domain provides a stronger inductive bias for voiced speech.
    source: §4.1.1, Table 1
    evidence: Fourier-domain vocoders reduce periodicity errors more effectively than time-domain GANs, suggesting
      that modelling harmonics in the frequency domain provides a stronger inductive bias for voiced speech.
    confidence: high
    relevance: medium
  - claim_id: convnext_blocks_with_isotropic_architecture_outperform_dilated_resblocks_in_the
    role: supports
    claim: ConvNeXt blocks with isotropic architecture outperform dilated ResBlocks in the Fourier-domain vocoder
      setting, even though dilated convolutions were motivated by the need to expand receptive fields in time-domain
      models.
    source: §4.1.1, Table 1
    evidence: ConvNeXt blocks with isotropic architecture outperform dilated ResBlocks in the Fourier-domain vocoder
      setting, even though dilated convolutions were motivated by the need to expand receptive fields in time-domain
      models.
    confidence: high
    relevance: medium
  - claim_id: a_fourier_domain_gan_vocoder_trained_as_a_neural_codec
    role: supports
    claim: A Fourier-domain GAN vocoder trained as a neural codec decoder can substantially improve perceptual quality
      over the original codec decoder across all bitrates without architectural changes to the upstream codec.
    source: §4.2, Table 5
    evidence: A Fourier-domain GAN vocoder trained as a neural codec decoder can substantially improve perceptual
      quality over the original codec decoder across all bitrates without architectural changes to the upstream
      codec.
    confidence: high
    relevance: high
  limitations:
  - Vocos's mel-spectrogram MOS scores are reported on LibriTTS using crowd-sourced listeners; the ground-truth
    MOS (3.81) is noticeably below what might be expected for studio speech, suggesting the evaluation pool or headphone
    compliance filtering may limit the discriminative power of the subjective test. Statistical equivalence with
    BigVGAN is shown, but the test may be underpowered for detecting small differences.
  - The speed benchmarks are conducted without hardware-specific optimisations and use batch size 16, which is unlikely
    to reflect latency-critical single-sample streaming scenarios. The CPU advantage (169x real-time) is particularly
    striking but is not validated in a streaming or low-latency deployment setting.
  - The MDCT variant explored in Appendix A performs worse than the ISTFT variant across all metrics, suggesting
    the overcomplete STFT representation provides a beneficial inductive bias. Whether MDCT-based generation could
    become competitive with improved training strategies remains open.
  - The model is trained exclusively on 24 kHz speech and music; performance at higher sample rates (44.1 kHz or
    48 kHz), which matter for high-fidelity TTS applications, is not reported.
  caveats: []
- id: '2306.12925'
  published_date: "2023-06-22"
  entry_date: '2026-07-28'
  year: 2023
  venue: arXiv
  task:
  - TTS
  - SCA
  architecture:
  - autoregressive-LM
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: influential
  method_family:
  - autoregressive_codec_language_models
  claims:
  - claim_id: initializing_a_speech_text_llm_from_a_pretrained_text_only
    role: supports
    claim: Initializing a speech-text LLM from a pretrained text-only checkpoint substantially outperforms training
      from scratch at equivalent model scale.
    source: §5.4.2, Table 6
    evidence: Initializing a speech-text LLM from a pretrained text-only checkpoint substantially outperforms training
      from scratch at equivalent model scale.
    confidence: high
    relevance: medium
  - claim_id: audio_tokenizer_quality_is_a_primary_bottleneck_in_llm_based
    role: supports
    claim: 'Audio tokenizer quality is a primary bottleneck in LLM-based speech generation and understanding: stronger
      semantic tokenizers yield large downstream gains independent of LM scale.'
    source: §5.4.3, Table 7
    evidence: 'Audio tokenizer quality is a primary bottleneck in LLM-based speech generation and understanding:
      stronger semantic tokenizers yield large downstream gains independent of LM scale.'
    confidence: high
    relevance: medium
  - claim_id: a_unified_multimodal_vocabulary_that_interleaves_text_and_audio_tokens
    role: supports
    claim: A unified multimodal vocabulary that interleaves text and audio tokens enables zero-shot speech translation
      to language pairs not seen during speech training, by inheriting translation capability from text pretraining.
    source: §5.2, Table 3
    evidence: A unified multimodal vocabulary that interleaves text and audio tokens enables zero-shot speech translation
      to language pairs not seen during speech training, by inheriting translation capability from text pretraining.
    confidence: high
    relevance: medium
  - claim_id: training_on_combined_tasks_that_decompose_complex_speech_operations_into
    role: supports
    claim: Training on combined tasks that decompose complex speech operations into intermediate text steps improves
      performance over direct end-to-end decoding.
    source: §5.4.4, Table 8
    evidence: Training on combined tasks that decompose complex speech operations into intermediate text steps improves
      performance over direct end-to-end decoding.
    confidence: high
    relevance: medium
  - claim_id: voice_identity_preservation_in_cross_lingual_speech_synthesis_can_exceed
    role: supports
    claim: Voice identity preservation in cross-lingual speech synthesis can exceed high-quality TTS-based references
      when an audio LM is conditioned on a short spoken prompt.
    source: §5.3, Table 4
    evidence: Voice identity preservation in cross-lingual speech synthesis can exceed high-quality TTS-based references
      when an audio LM is conditioned on a short spoken prompt.
    confidence: high
    relevance: medium
  limitations:
  - The entire system depends on the quality of the audio tokenizer, which is not released and requires access to
    Google-internal USM models. The best-performing configuration (AudioPaLM-2 with USM-v2 tokens) is not reproducible
    externally; the published ablations use the multilingual w2v-BERT tokenizer as the weakest condition, suggesting
    that reported performance at USM-v2 quality cannot be independently verified.
  - Adding S2ST tasks modestly degrades ASR and AST performance, suggesting that sharing model capacity across output
    modalities introduces trade-offs that are not fully resolved by the training mixture design (§5.4.5, Table 9).
    The paper evaluates primarily on speech translation tasks; generative speech quality at naturalness is only
    assessed in the S2ST with voice transfer setting, not for open-ended TTS. The subjective evaluations were conducted
    on an earlier version of AudioPaLM using AudioLM decoding rather than SoundStorm, meaning the best-performing
    decoder was not evaluated subjectively. Evaluation benchmarks for generative audio tasks more generally remain
    underdeveloped relative to text, a limitation the paper explicitly notes.
  caveats: []
- id: '2308.16692'
  published_date: "2023-08-31"
  entry_date: '2026-07-28'
  year: 2023
  venue: arXiv
  task:
  - TTS
  - codec
  architecture:
  - GAN
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: influential
  method_family:
  - gan_neural_codecs_and_decoders
  claims:
  - claim_id: separating_semantic_content_from_paralinguistic_information_across_rvq_layers_within
    role: supports
    claim: Separating semantic content from paralinguistic information across RVQ layers within a single codec improves
      both reconstruction quality and speech language model coherence compared to undifferentiated acoustic tokenisation.
    source: §4.4, Tables 2 and 4
    evidence: Separating semantic content from paralinguistic information across RVQ layers within a single codec
      improves both reconstruction quality and speech language model coherence compared to undifferentiated acoustic
      tokenisation.
    confidence: high
    relevance: high
  - claim_id: acoustic_tokens_from_standard_neural_codecs_encode_content_and_speaker
    role: supports
    claim: Acoustic tokens from standard neural codecs encode content and speaker identity in an entangled form
      that causes systematic word errors in autoregressive language model generation.
    source: §2.3, Table 3
    evidence: Acoustic tokens from standard neural codecs encode content and speaker identity in an entangled form
      that causes systematic word errors in autoregressive language model generation.
    confidence: high
    relevance: medium
  - claim_id: a_distillation_objective_computed_per_feature_dimension_d_axis_produces
    role: supports
    claim: A distillation objective computed per feature dimension (D-axis) produces stronger semantic guidance
      to a codec's first quantizer than the conventional per-timestep (T-axis) formulation.
    source: Appendix C, Table 7
    evidence: A distillation objective computed per feature dimension (D-axis) produces stronger semantic guidance
      to a codec's first quantizer than the conventional per-timestep (T-axis) formulation.
    confidence: high
    relevance: high
  - claim_id: the_first_layer_tokens_of_a_hierarchically_disentangled_codec_can
    role: supports
    claim: The first-layer tokens of a hierarchically disentangled codec can serve as a zero-shot voice conversion
      mechanism by swapping higher-layer tokens from a reference speaker, without requiring a separate conversion
      model.
    source: §5.2, Table 5
    evidence: The first-layer tokens of a hierarchically disentangled codec can serve as a zero-shot voice conversion
      mechanism by swapping higher-layer tokens from a reference speaker, without requiring a separate conversion
      model.
    confidence: high
    relevance: high
  - claim_id: codec_tokens_trained_without_explicit_content_supervision_exhibit_poor_codebook
    role: supports
    claim: Codec tokens trained without explicit content supervision exhibit poor codebook utilisation and weak
      phoneme-code correspondence, increasing the modelling burden on downstream language models.
    source: Appendix F, Table 8
    evidence: Codec tokens trained without explicit content supervision exhibit poor codebook utilisation and weak
      phoneme-code correspondence, increasing the modelling burden on downstream language models.
    confidence: high
    relevance: high
  limitations:
  - SpeechTokenizer is trained solely on English LibriSpeech. While preliminary results in Appendix G suggest cross-lingual
    token transfer is plausible, the codec is not validated for multilingual speech language models and the text-alignment
    properties of RVQ-1 may not hold for typologically distant languages.
  - The SLMTokBench mutual information metric relies on a variational upper bound and a fixed BLSTM downstream model;
    its absolute values are not directly comparable across evaluation frameworks. The USLM is evaluated only on
    VCTK with a single 3-second prompt per speaker, which does not represent the range of zero-shot conditions used
    in contemporaneous benchmarks. Model size is not reported for SpeechTokenizer itself. The voice conversion application
    (§5.2) is characterised as a single-shot demonstration rather than a full evaluation against dedicated VC baselines.
  caveats: []
- id: '2310.00704'
  published_date: "2023-10-01"
  entry_date: '2026-07-28'
  year: 2023
  venue: arXiv
  task:
  - TTS
  - VC
  - singing
  architecture:
  - autoregressive-LM
  - hybrid
  relevance: high
  evidence_role:
  - core_evidence
  - control_evidence
  - evaluation_caution
  current_role: influential
  method_family:
  - autoregressive_codec_language_models
  - hybrid_semantic_acoustic_tokenizers
  claims:
  - claim_id: training_a_single_audio_language_model_across_diverse_generation_tasks
    role: supports
    claim: Training a single audio language model across diverse generation tasks (TTS, voice conversion, sound
      synthesis, music, singing) produces consistent performance improvements over task-specific models trained
      on the same data.
    source: §3.4.1, Appendix C.1, Table 17
    evidence: Training a single audio language model across diverse generation tasks (TTS, voice conversion, sound
      synthesis, music, singing) produces consistent performance improvements over task-specific models trained
      on the same data.
    confidence: high
    relevance: medium
  - claim_id: the_autoregressive_property_is_critical_for_audio_generation_quality_parallel
    role: supports
    claim: 'The autoregressive property is critical for audio generation quality: parallel and delay-based codec
      prediction approaches yield measurably lower naturalness than fully autoregressive methods when codec quantization
      levels are held constant.'
    source: §3.4.2, Tables 4–5
    evidence: 'The autoregressive property is critical for audio generation quality: parallel and delay-based codec
      prediction approaches yield measurably lower naturalness than fully autoregressive methods when codec quantization
      levels are held constant.'
    confidence: high
    relevance: high
  - claim_id: hierarchical_factorisation_of_rvq_codec_token_sequences_into_inter_frame
    role: supports
    claim: Hierarchical factorisation of RVQ codec token sequences into inter-frame and intra-frame modeling substantially
      reduces training memory and time relative to flat-sequence autoregressive prediction, with comparable generation
      quality.
    source: §2.3, §3.4.2, Table 4
    evidence: Hierarchical factorisation of RVQ codec token sequences into inter-frame and intra-frame modeling
      substantially reduces training memory and time relative to flat-sequence autoregressive prediction, with comparable
      generation quality.
    confidence: high
    relevance: high
  - claim_id: pre_training_on_a_broad_multi_task_audio_corpus_enables
    role: supports
    claim: Pre-training on a broad multi-task audio corpus enables strong adaptation to unseen audio generation
      tasks via fine-tuning on small datasets, outperforming task-specific models trained from scratch on those
      tasks.
    source: §3.3, Appendix B.5–B.8, Table 17
    evidence: Pre-training on a broad multi-task audio corpus enables strong adaptation to unseen audio generation
      tasks via fine-tuning on small datasets, outperforming task-specific models trained from scratch on those
      tasks.
    confidence: high
    relevance: medium
  - claim_id: signal_level_metrics_such_as_pesq_are_poorly_suited_for
    role: supports
    claim: 'Signal-level metrics such as PESQ are poorly suited for evaluating generative audio models: systems
      achieving higher perceptual MOS scores routinely score lower on PESQ than discriminative baselines.'
    source: §3.2, §3.4.2, Table 11
    evidence: 'Signal-level metrics such as PESQ are poorly suited for evaluating generative audio models: systems
      achieving higher perceptual MOS scores routinely score lower on PESQ than discriminative baselines.'
    confidence: high
    relevance: medium
  limitations:
  - Model checkpoints are not released due to misuse concerns, limiting reproducibility. Only code and demos are
    public. Researchers cannot directly reproduce the full 165K-hour training run or perform ablations at scale.
  - 'UniAudio does not handle all known audio tasks: noise removal, noisy speech editing, and speech-to-speech translation
    are explicitly excluded. New modalities cannot be introduced during fine-tuning, only new combinations of modalities
    already seen at training time. The system relies entirely on labeled data; self-supervised or weakly supervised
    pre-training from unlabeled audio, which could substantially increase coverage, is left as future work.'
  - The multi-task benefit is empirically demonstrated but mechanistically underexplained. The paper offers informal
    hypotheses (shared codec token space, data augmentation equivalences between TTS and VC conditions) but no formal
    analysis. It is unclear whether the gains stem from better codec representations, increased effective data volume
    per task, or genuine cross-task transfer.
  caveats: []
- id: '2312.15821'
  published_date: "2023-12-25"
  entry_date: '2026-07-28'
  year: 2023
  venue: arXiv
  task:
  - TTS
  architecture:
  - flow-matching
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: influential
  method_family:
  - flow_matching_codec_decoders
  claims:
  - claim_id: a_single_generative_model_trained_across_speech_sound_and_music
    role: supports
    claim: A single generative model trained across speech, sound, and music modalities can match or surpass modality-specific
      models on dedicated benchmarks.
    source: §5.4, §6.4, §7.5, Tables 1, 5, 13, 14
    evidence: A single generative model trained across speech, sound, and music modalities can match or surpass
      modality-specific models on dedicated benchmarks.
    confidence: high
    relevance: medium
  - claim_id: self_supervised_pre_training_on_large_scale_unlabeled_audio_substantially
    role: supports
    claim: Self-supervised pre-training on large-scale unlabeled audio substantially improves multi-domain style
      generalisation in subsequent supervised fine-tuning, with gains most pronounced on out-of-domain test sets.
    source: §5.5, Table 3
    evidence: Self-supervised pre-training on large-scale unlabeled audio substantially improves multi-domain style
      generalisation in subsequent supervised fine-tuning, with gains most pronounced on out-of-domain test sets.
    confidence: high
    relevance: medium
  - claim_id: general_audio_language_embedding_models_trained_primarily_on_sound_events
    role: complicates
    claim: General audio-language embedding models trained primarily on sound events fail to capture fine-grained
      speech attributes, rendering them unreliable for evaluating description-conditioned speech generation.
    source: §7.3.1, Table 8
    evidence: General audio-language embedding models trained primarily on sound events fail to capture fine-grained
      speech attributes, rendering them unreliable for evaluating description-conditioned speech generation.
    confidence: high
    relevance: medium
  - claim_id: flow_matching_models_admit_post_training_inference_optimisation_via_learned
    role: supports
    claim: Flow-matching models admit post-training inference optimisation via learned ODE reparameterisation (Bespoke
      Solvers) that reduces function evaluations by 25x without measurable quality degradation.
    source: §8, Table 15
    evidence: Flow-matching models admit post-training inference optimisation via learned ODE reparameterisation
      (Bespoke Solvers) that reduces function evaluations by 25x without measurable quality degradation.
    confidence: high
    relevance: medium
  - claim_id: removing_trailing_silence_from_audio_context_prompts_substantially_improves_speaker
    role: supports
    claim: Removing trailing silence from audio context prompts substantially improves speaker similarity in zero-shot
      TTS, particularly on datasets with long prompt silences.
    source: §5.5, Table 3
    evidence: Removing trailing silence from audio context prompts substantially improves speaker similarity in
      zero-shot TTS, particularly on datasets with long prompt silences.
    confidence: high
    relevance: low
  limitations:
  - The speech description data is primarily English, and the description-based conditioning pipeline depends on
    LLM-generated captions derived from categorical attribute labels. The attribute vocabulary (age, gender, pitch,
    speaking rate, accent, emotion, environment) has limited granularity; fine-grained attributes like specific
    regional accents or breed-level sound events cannot be reliably generated if no paired training examples exist
    for those distinctions.
  - The gap between dedicated Audiobox Sound (FAD 0.77) and unified Audiobox (FAD 1.10) on the AudioCaps benchmark
    suggests that task-specific models still have an edge in their domain, even at this training scale. The description-based
    TTS evaluation relies on Joint-CLAP, which is introduced by the same paper and has not been independently validated
    as a general benchmark metric. The model and training data are not released publicly (no code repository linked),
    limiting reproducibility. The voice restylization capability, which requires disentangling vocal identity from
    environment and emotion, is evaluated subjectively on only two internal test sets, and the degree to which the
    voice prompt de-correlation training objective succeeds across diverse speakers remains unclear.
  caveats: []
- id: '2401.07333'
  published_date: "2024-01-14"
  entry_date: '2026-07-28'
  year: 2024
  venue: arXiv
  task:
  - TTS
  architecture:
  - autoregressive-LM
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: active_evidence
  method_family:
  - autoregressive_codec_language_models
  claims:
  - claim_id: interleaving_phoneme_tokens_with_their_corresponding_acoustic_frames_in_the
    role: supports
    claim: Interleaving phoneme tokens with their corresponding acoustic frames in the training sequence substantially
      reduces phoneme-level alignment errors in autoregressive codec LM TTS, including repetitions, transpositions,
      and omissions.
    source: §3.2, §4.2, Table 3
    evidence: Interleaving phoneme tokens with their corresponding acoustic frames in the training sequence substantially
      reduces phoneme-level alignment errors in autoregressive codec LM TTS, including repetitions, transpositions,
      and omissions.
    confidence: high
    relevance: high
  - claim_id: autoregressive_codec_language_models_that_concatenate_all_phoneme_tokens_ahead
    role: supports
    claim: Autoregressive codec language models that concatenate all phoneme tokens ahead of all acoustic tokens
      are prone to infinite-silence generation, with failure rates exceeding 80% under greedy decoding.
    source: §1, Table 1
    evidence: Autoregressive codec language models that concatenate all phoneme tokens ahead of all acoustic tokens
      are prone to infinite-silence generation, with failure rates exceeding 80% under greedy decoding.
    confidence: high
    relevance: high
  - claim_id: explicit_forced_alignment_supervision_at_training_time_enables_fine_grained
    role: supports
    claim: Explicit forced-alignment supervision at training time enables fine-grained phoneme-level control at
      inference, allowing deterministic truncation of abnormal synthesis and making greedy decoding viable.
    source: §3.2.1, §3.3, Figure 5
    evidence: Explicit forced-alignment supervision at training time enables fine-grained phoneme-level control
      at inference, allowing deterministic truncation of abnormal synthesis and making greedy decoding viable.
    confidence: high
    relevance: medium
  - claim_id: structural_alignment_constraints_in_the_token_sequence_provide_larger_accuracy
    role: supports
    claim: Structural alignment constraints in the token sequence provide larger accuracy gains than naturalness
      or speaker similarity improvements, suggesting that intelligibility and speaker identity are relatively easy
      to preserve while alignment robustness remains the primary challenge.
    source: §4.2, Table 2
    evidence: Structural alignment constraints in the token sequence provide larger accuracy gains than naturalness
      or speaker similarity improvements, suggesting that intelligibility and speaker identity are relatively easy
      to preserve while alignment robustness remains the primary challenge.
    confidence: high
    relevance: low
  limitations:
  - Results are compared only against a VALL-E baseline reproduced on LibriSpeech 960h, not against the original
    VALL-E trained on 60k hours of LibriLight. The restricted training data limits direct comparison to the published
    VALL-E numbers and leaves open whether the gains hold at scale.
  - MFA forced alignment is a preprocessing dependency that may fail or produce noisy alignments for short utterances,
    spontaneous speech, or languages without reliable MFA models. The paper does not analyze alignment quality or
    its effect on synthesis when MFA makes errors.
  - The evaluation is limited to LibriSpeech, a clean read-speech corpus in English. Generalization to noisy, spontaneous,
    or non-English settings is untested. The hard-case set of 100 sentences is internally designed rather than a
    standard benchmark, limiting reproducibility of the cross-speaker evaluation.
  - The local-advance hyperparameter (how many frames ahead to shift EOP and the next phoneme) is evaluated informally;
    systematic tuning across speech rates and phoneme types is left for future work.
  caveats: []
- id: '2402.01912'
  published_date: "2024-02-02"
  entry_date: '2026-07-28'
  year: 2024
  venue: arXiv
  task:
  - TTS
  architecture:
  - autoregressive-LM
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  - evaluation_caution
  current_role: influential
  method_family:
  - autoregressive_codec_language_models
  claims:
  - claim_id: automatic_acoustic_labeling_can_substitute_for_human_annotations_in_training
    role: supports
    claim: Automatic acoustic labeling can substitute for human annotations in training large-scale instruction-conditioned
      speech language models without a loss in attribute control accuracy relative to human-labeled systems.
    source: §3.1, §3.2, §4.1
    evidence: Automatic acoustic labeling can substitute for human annotations in training large-scale instruction-conditioned
      speech language models without a loss in attribute control accuracy relative to human-labeled systems.
    confidence: high
    relevance: medium
  - claim_id: including_a_small_proportion_of_high_fidelity_audio_approximately_1
    role: supports
    claim: Including a small proportion of high-fidelity audio (approximately 1%) in a predominantly noisy training
      corpus, combined with explicit recording-quality labels, enables a speech LM to generate professional-sounding
      speech on demand from text prompts alone.
    source: §3.1.2, §4.2, Table 1
    evidence: Including a small proportion of high-fidelity audio (approximately 1%) in a predominantly noisy training
      corpus, combined with explicit recording-quality labels, enables a speech LM to generate professional-sounding
      speech on demand from text prompts alone.
    confidence: high
    relevance: medium
  - claim_id: the_choice_of_neural_audio_codec_has_a_measurable_effect
    role: supports
    claim: The choice of neural audio codec has a measurable effect on perceptual audio quality in autoregressive
      TTS; higher-fidelity codecs translate directly to higher MOS and objective quality scores.
    source: §3.3, §4.2, Table 1–2
    evidence: The choice of neural audio codec has a measurable effect on perceptual audio quality in autoregressive
      TTS; higher-fidelity codecs translate directly to higher MOS and objective quality scores.
    confidence: high
    relevance: high
  - claim_id: natural_language_conditioning_on_accent_can_be_achieved_in_a
    role: supports
    claim: Natural language conditioning on accent can be achieved in a single TTS model covering dozens of accents,
      though classifier accuracy reflects the noise and imbalance inherent in automatic accent labeling of crowd-sourced
      data.
    source: §3.1.1, §4.1
    evidence: Natural language conditioning on accent can be achieved in a single TTS model covering dozens of accents,
      though classifier accuracy reflects the noise and imbalance inherent in automatic accent labeling of crowd-sourced
      data.
    confidence: high
    relevance: medium
  limitations:
  - The evaluation compares only against Audiobox. No standard TTS baselines (reference-based zero-shot systems,
    encoder-decoder models) are included, making it impossible to assess whether the MOS gains arise from the conditioning
    approach, the codec choice, or the training data mix.
  - The system is evaluated only on English audiobook speech; generalization to conversational, spontaneous, or
    non-English speech is stated as future work but untested. The accent accuracy of 68% indicates that discrete
    accent labels in crowd-sourced data are noisy, and the model's C50 control was found to be unreliable even after
    training. Model size and total compute are not reported. The automatic labeling pipeline requires training multiple
    specialized classifiers (accent, gender), each of which introduces its own noise floor. The demo website is
    the only verification source; no code or model weights are released.
  caveats: []
- id: '2402.08093'
  published_date: "2024-02-12"
  entry_date: '2026-07-28'
  year: 2024
  venue: arXiv
  task:
  - TTS
  architecture:
  - autoregressive-LM
  relevance: medium
  evidence_role:
  - architecture_variant
  - acceleration_evidence
  - control_evidence
  - evaluation_caution
  current_role: influential
  method_family:
  - autoregressive_codec_language_models
  claims:
  - claim_id: scaling_autoregressive_codec_tts_to_500m_parameters_and_10k_hours
    role: supports
    claim: Scaling autoregressive codec TTS to 500M+ parameters and 10K+ hours of training data produces qualitatively
      different prosody rendering on linguistically complex inputs compared to smaller models trained on less data.
    source: §4.3, Figure 4, Table 5
    evidence: Scaling autoregressive codec TTS to 500M+ parameters and 10K+ hours of training data produces qualitatively
      different prosody rendering on linguistically complex inputs compared to smaller models trained on less data.
    confidence: high
    relevance: high
  - claim_id: ssl_based_speech_representations_with_explicit_speaker_disentanglement_outperform_purely
    role: supports
    claim: SSL-based speech representations with explicit speaker disentanglement outperform purely acoustic codec
      representations for zero-shot TTS, particularly in lower-resource languages.
    source: §4.1, Table 3
    evidence: SSL-based speech representations with explicit speaker disentanglement outperform purely acoustic
      codec representations for zero-shot TTS, particularly in lower-resource languages.
    confidence: high
    relevance: high
  - claim_id: a_streamable_convolutional_decoder_can_match_or_exceed_a_diffusion
    role: supports
    claim: A streamable convolutional decoder can match or exceed a diffusion-based spectrogram decoder in subjective
      naturalness while reducing synthesis compute by approximately 3x and enabling low-latency streaming.
    source: §4.2, §4.5, Table 4
    evidence: A streamable convolutional decoder can match or exceed a diffusion-based spectrogram decoder in subjective
      naturalness while reducing synthesis compute by approximately 3x and enabling low-latency streaming.
    confidence: high
    relevance: low
  - claim_id: applying_bpe_to_discrete_speech_tokens_reduces_autoregressive_sequence_length
    role: supports
    claim: Applying BPE to discrete speech tokens reduces autoregressive sequence length by approximately 40% without
      degrading downstream synthesis quality, enabling longer-context training.
    source: §2.2.3
    evidence: Applying BPE to discrete speech tokens reduces autoregressive sequence length by approximately 40%
      without degrading downstream synthesis quality, enabling longer-context training.
    confidence: high
    relevance: high
  - claim_id: autoregressive_tts_trained_at_scale_generalises_to_a_wide_range
    role: supports
    claim: Autoregressive TTS trained at scale generalises to a wide range of textual phenomena without any explicit
      prosody annotation or task-specific supervision.
    source: §4.3, §6
    evidence: Autoregressive TTS trained at scale generalises to a wide range of textual phenomena without any explicit
      prosody annotation or task-specific supervision.
    confidence: high
    relevance: low
  limitations:
  - Model weights are not released, and evaluation uses proprietary test speakers. The MUSHRA baselines (YourTTS,
    Bark, TortoiseTTS) are not trained on comparable data or compute, making architecture-level conclusions difficult
    to separate from scale effects.
  - The speechcode decoder is tightly coupled to a specific frozen SpeechGPT checkpoint via hidden-state conditioning,
    preventing modular updates and complicating experimentation. The paper identifies this as a limitation requiring
    future work.
  - Hallucinations and cutoffs remain an inherent issue of the autoregressive formulation, worsened by misalignment
    between noisy web audio and ASR-generated transcripts. The authors avoid denoising during training to test robustness
    but acknowledge this makes the alignment problem harder.
  - Emotions and paralinguistics remain below ceiling even for BASE-large, suggesting that 100K hours and 980M parameters
    are not sufficient for reliable rendering of these categories. Formal scaling laws for TTS (analogous to Chinchilla
    for text LMs) are proposed as future work but not established here.
  - The "emergent abilities" phenomenon is characterised across only three data-scale points and assessed by a single
    expert linguist, leaving open whether it is a smooth or discontinuous function of scale and how sensitive it
    is to tokenization and architecture choices.
  caveats: []
- id: '2402.13236'
  published_date: "2024-02-20"
  entry_date: '2026-07-28'
  year: 2024
  venue: arXiv
  task:
  - TTS
  - SCA
  - codec
  architecture: []
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: active_evidence
  method_family: []
  claims:
  - claim_id: residual_vector_quantisation_is_the_dominant_quantisation_strategy_across_neural
    role: supports
    claim: Residual vector quantisation is the dominant quantisation strategy across neural audio codec models,
      with variation concentrated in discriminator design, bitrate, and semantic token integration rather than in
      the core compression mechanism.
    source: §II.A, Table II
    evidence: Residual vector quantisation is the dominant quantisation strategy across neural audio codec models,
      with variation concentrated in discriminator design, bitrate, and semantic token integration rather than in
      the core compression mechanism.
    confidence: high
    relevance: high
  - claim_id: codec_based_audio_language_models_increasingly_target_multi_task_coverage
    role: supports
    claim: Codec-based audio language models increasingly target multi-task coverage rather than single-task specialisation,
      with several systems spanning TTS, voice conversion, speech editing, speech enhancement, and translation in
      a single framework.
    source: §III.B, Table III
    evidence: Codec-based audio language models increasingly target multi-task coverage rather than single-task
      specialisation, with several systems spanning TTS, voice conversion, speech editing, speech enhancement, and
      translation in a single framework.
    confidence: high
    relevance: high
  - claim_id: integrating_semantic_tokens_from_self_supervised_speech_representations_into_the
    role: supports
    claim: Integrating semantic tokens from self-supervised speech representations into the codec quantisation process
      improves audio quality at low bitrates, with HuBERT-guided RVQ being the most common approach.
    source: §II.B
    evidence: Integrating semantic tokens from self-supervised speech representations into the codec quantisation
      process improves audio quality at low bitrates, with HuBERT-guided RVQ being the most common approach.
    confidence: high
    relevance: high
  - claim_id: discrete_units_derived_from_self_supervised_representations_enable_textless_speech
    role: complicates
    claim: Discrete units derived from self-supervised representations enable textless speech language modelling
      but sacrifice speaker and paralinguistic information relative to codec-based approaches.
    source: §III.A
    evidence: Discrete units derived from self-supervised representations enable textless speech language modelling
      but sacrifice speaker and paralinguistic information relative to codec-based approaches.
    confidence: high
    relevance: high
  limitations:
  - 'The survey covers only open-source codec models and does not include proprietary codecs used in industry systems.
    Evaluation methodology is not addressed: the paper does not compare codecs on shared benchmarks or report reproduction
    numbers, making it difficult to assess quality claims from the original papers in a unified way. The codec-based
    LM survey relies entirely on self-reported results from each system''s own paper, with no cross-system comparison
    on shared evaluation sets. Coverage is limited to models available as of early 2024; the field has moved substantially
    since, with systems such as SpeechTokenizer and more recent codec variants not included in the analysis.'
  - 'Open questions the survey surfaces: What is the quality impact of codec design choices (bitrate, discriminator
    type, semantic integration) on downstream speech generation systems? Can a universal codec serve all audio types
    equally well? How should speech LMs be prompted efficiently in zero-shot settings?'
  caveats: []
- id: '2403.03100'
  published_date: "2024-03-05"
  entry_date: '2026-07-28'
  year: 2024
  venue: arXiv
  task:
  - TTS
  - VC
  architecture:
  - diffusion
  - hybrid
  relevance: medium
  evidence_role:
  - architecture_variant
  - acceleration_evidence
  - control_evidence
  current_role: influential
  method_family:
  - diffusion_latent_codec_generators
  - hybrid_semantic_acoustic_tokenizers
  claims:
  - claim_id: explicit_disentanglement_of_speech_attributes_in_the_codec_representation_reduces
    role: supports
    claim: Explicit disentanglement of speech attributes in the codec representation reduces the complexity of zero-shot
      generation and improves speaker similarity, quality, and prosody simultaneously.
    source: §3, §4.2, Table 1, Table 2
    evidence: Explicit disentanglement of speech attributes in the codec representation reduces the complexity of
      zero-shot generation and improves speaker similarity, quality, and prosody simultaneously.
    confidence: high
    relevance: high
  - claim_id: gradient_reversal_combined_with_attribute_specific_supervised_losses_is_an
    role: supports
    claim: Gradient reversal combined with attribute-specific supervised losses is an effective mechanism for suppressing
      cross-attribute information leakage in neural codec quantization.
    source: §3.2.2, Appendix B.4
    evidence: Gradient reversal combined with attribute-specific supervised losses is an effective mechanism for
      suppressing cross-attribute information leakage in neural codec quantization.
    confidence: high
    relevance: high
  - claim_id: the_factorization_paradigm_for_codec_representations_is_architecture_agnostic_and
    role: supports
    claim: The factorization paradigm for codec representations is architecture-agnostic and improves both autoregressive
      and non-autoregressive generators when applied.
    source: §4.3.2, Table 6
    evidence: The factorization paradigm for codec representations is architecture-agnostic and improves both autoregressive
      and non-autoregressive generators when applied.
    confidence: high
    relevance: high
  - claim_id: discrete_masked_diffusion_over_disentangled_codec_tokens_is_faster_than
    role: supports
    claim: Discrete masked diffusion over disentangled codec tokens is faster than autoregressive LM-based codec
      generation at comparable or better quality.
    source: Appendix A.5, Table 10
    evidence: Discrete masked diffusion over disentangled codec tokens is faster than autoregressive LM-based codec
      generation at comparable or better quality.
    confidence: high
    relevance: high
  - claim_id: performance_on_zero_shot_tts_scales_predictably_with_both_training
    role: supports
    claim: Performance on zero-shot TTS scales predictably with both training data volume and model size when the
      underlying speech representation captures disentangled attributes.
    source: §4.4, Tables 7, 8
    evidence: Performance on zero-shot TTS scales predictably with both training data volume and model size when
      the underlying speech representation captures disentangled attributes.
    confidence: high
    relevance: high
  limitations:
  - FACodec requires phoneme-level transcriptions for content supervision during training, constraining its applicability
    to languages and settings where reliable alignments are unavailable. The zero-shot TTS evaluation is English-only;
    multilingual generalisation is stated as future work but not demonstrated.
  - 'Additional limitations: the attribute factorization is incomplete (background sounds, energy, and other fine-grained
    characteristics are not captured, as noted in Appendix C); the acoustic detail subspace retains some content
    and prosody leakage without gradient reversal (verified qualitatively in Appendix B.4); and the prosody evaluation
    relies on MCD and emotion classifiers on the RAVDESS dataset, which assesses a narrow range of acted emotions
    rather than naturalistic prosodic variation.'
  - Open questions include whether factorized disentanglement continues to improve at larger scales and whether
    supervision-free disentanglement is achievable without degrading reconstruction quality.
  caveats: []
- id: '2403.16973'
  published_date: "2024-03-25"
  entry_date: '2026-07-28'
  year: 2024
  venue: arXiv
  task:
  - TTS
  architecture:
  - autoregressive-LM
  relevance: high
  evidence_role:
  - core_evidence
  - control_evidence
  - evaluation_caution
  current_role: influential
  method_family:
  - autoregressive_codec_language_models
  claims:
  - claim_id: autoregressive_codec_language_models_can_perform_speech_infilling_with_naturalness
    role: supports
    claim: Autoregressive codec language models can perform speech infilling with naturalness approaching that of
      the original unedited recording when masked spans are relocated to the end of the sequence, enabling bidirectional
      context conditioning within a causal framework.
    source: §3.1, §5.3, Table 5
    evidence: Autoregressive codec language models can perform speech infilling with naturalness approaching that
      of the original unedited recording when masked spans are relocated to the end of the sequence, enabling bidirectional
      context conditioning within a causal framework.
    confidence: high
    relevance: high
  - claim_id: zero_shot_tts_and_speech_editing_can_be_unified_as
    role: supports
    claim: Zero-shot TTS and speech editing can be unified as a single autoregressive infilling operation without
      task-specific architectural components, at no cost to performance on either task.
    source: §3.4, §5.4, Table 6
    evidence: Zero-shot TTS and speech editing can be unified as a single autoregressive infilling operation without
      task-specific architectural components, at no cost to performance on either task.
    confidence: high
    relevance: medium
  - claim_id: wer_measured_by_asr_systems_is_an_unreliable_proxy_for
    role: supports
    claim: 'WER measured by ASR systems is an unreliable proxy for perceptual intelligibility when evaluating speech
      synthesis quality: systems can achieve lower WER than ground truth recordings while receiving substantially
      lower intelligibility ratings from human listeners.'
    source: §5.3, §5.4
    evidence: 'WER measured by ASR systems is an unreliable proxy for perceptual intelligibility when evaluating
      speech synthesis quality: systems can achieve lower WER than ground truth recordings while receiving substantially
      lower intelligibility ratings from human listeners.'
    confidence: high
    relevance: medium
  - claim_id: evaluation_of_speech_synthesis_exclusively_on_audiobook_data_underestimates_the
    role: supports
    claim: Evaluation of speech synthesis exclusively on audiobook data underestimates the performance gap between
      systems when applied to in-the-wild recordings with diverse accents, noise, and speaking styles.
    source: §5.3, §5.4
    evidence: Evaluation of speech synthesis exclusively on audiobook data underestimates the performance gap between
      systems when applied to in-the-wild recordings with diverse accents, noise, and speaking styles.
    confidence: high
    relevance: medium
  - claim_id: scaling_autoregressive_codec_lm_parameters_consistently_improves_objective_metrics_across
    role: supports
    claim: Scaling autoregressive codec LM parameters consistently improves objective metrics across intelligibility
      and acoustic fidelity measures, with larger gaps between larger model sizes suggesting further gains from
      continued scaling.
    source: §5.2, Table 3
    evidence: Scaling autoregressive codec LM parameters consistently improves objective metrics across intelligibility
      and acoustic fidelity measures, with larger gaps between larger model sizes suggesting further gains from
      continued scaling.
    confidence: high
    relevance: high
  limitations:
  - The inference-time artifact mitigation (generating 10 candidates and discarding the 4 longest) adds significant
    latency and compute cost. The strategy is acknowledged as inelegant, and the underlying cause (repetitive loop
    generation) is unresolved. This limits practical deployment in real-time or low-compute settings.
  - 'Additional limitations include:'
  - '- The model is trained and evaluated in English only; generalisation to other languages is untested. - Zero-shot
    speaker similarity, while competitive, still falls short of ground truth (MOS 4.34 vs 4.44), particularly for
    in-the-wild YouTube recordings where acoustic conditions are challenging. - REALEDIT covers only English and
    is limited to 310 examples, restricting statistical power for fine-grained analysis by edit type and domain.
    - The model does not incorporate explicit watermarking or deepfake detection, which the authors identify as
    an AI safety gap given the model''s voice cloning capability.'
  caveats: []
- id: '2404.03204'
  published_date: "2024-04-04"
  entry_date: '2026-07-28'
  year: 2024
  venue: arXiv
  task:
  - TTS
  architecture:
  - autoregressive-LM
  relevance: high
  evidence_role:
  - core_evidence
  - control_evidence
  current_role: influential
  method_family:
  - autoregressive_codec_language_models
  claims:
  - claim_id: autoregressive_codec_language_model_tts_can_be_substantially_stabilised_by
    role: supports
    claim: Autoregressive codec language model TTS can be substantially stabilised by predicting prosody tokens
      as explicit intermediate targets before speech token generation, without requiring reranking or a separate
      alignment model at inference.
    source: §3.2, Table 2
    evidence: Autoregressive codec language model TTS can be substantially stabilised by predicting prosody tokens
      as explicit intermediate targets before speech token generation, without requiring reranking or a separate
      alignment model at inference.
    confidence: high
    relevance: high
  - claim_id: duration_guided_attention_masking_which_restricts_each_speech_token_to
    role: supports
    claim: Duration-guided attention masking, which restricts each speech token to attend only to a local phoneme
      window based on predicted alignment, provides significant robustness improvements beyond prosody conditioning
      alone.
    source: §3.3, Table 4
    evidence: Duration-guided attention masking, which restricts each speech token to attend only to a local phoneme
      window based on predicted alignment, provides significant robustness improvements beyond prosody conditioning
      alone.
    confidence: high
    relevance: low
  - claim_id: the_robustness_deficit_of_ar_codec_tts_relative_to_non
    role: complicates
    claim: The robustness deficit of AR codec TTS relative to non-autoregressive methods is most pronounced on structurally
      unusual inputs (repetitive patterns, numeric sequences, code strings) where learned implicit alignment is
      most likely to fail.
    source: §4.4, Table 1
    evidence: The robustness deficit of AR codec TTS relative to non-autoregressive methods is most pronounced on
      structurally unusual inputs (repetitive patterns, numeric sequences, code strings) where learned implicit
      alignment is most likely to fail.
    confidence: high
    relevance: high
  - claim_id: reranking_over_multiple_samples_and_explicit_intermediate_prosody_prediction_address
    role: supports
    claim: Reranking over multiple samples and explicit intermediate prosody prediction address the same underlying
      alignment problem and can be combined for additive gain, but prosody CoT prompting reduces the dependency
      on reranking by improving single-sample quality.
    source: §4.2, Table 2
    evidence: Reranking over multiple samples and explicit intermediate prosody prediction address the same underlying
      alignment problem and can be combined for additive gain, but prosody CoT prompting reduces the dependency
      on reranking by improving single-sample quality.
    confidence: high
    relevance: low
  limitations:
  - The paper relies on an internal proprietary alignment tool for extracting phoneme-speech alignments during training.
    Alignment quality directly affects duration-guided masking effectiveness (§3.3 notes that alignment errors required
    loosening the masking window from k=0 to k=1). Reproduction requires either the same tool or a comparable open
    alternative, which the paper does not specify or release.
  - 'Evaluation is entirely English on LibriSpeech and the MLS English subset. Whether the CoT prosody approach
    transfers to languages with different phoneme-duration relationships or tonal languages is untested. The comparison
    with ELLA-V and VALL-T is not fully controlled: all three systems differ in training data size and evaluation
    subset, making it difficult to isolate the architectural contribution from data-scale effects. Speaker similarity
    (SIM 0.49) remains below the originally reported VALL-E score (0.58), a gap the authors attribute to differences
    in prompt resynthesis methodology rather than genuine capability differences but cannot rule out.'
  caveats: []
- id: '2406.00654'
  published_date: "2024-06-02"
  entry_date: '2026-07-28'
  year: 2024
  venue: arXiv
  task:
  - TTS
  architecture:
  - autoregressive-LM
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  - evaluation_caution
  current_role: influential
  method_family:
  - autoregressive_codec_language_models
  claims:
  - claim_id: standard_supervised_training_objectives_for_tts_produce_a_systematic_mismatch
    role: supports
    claim: Standard supervised training objectives for TTS produce a systematic mismatch with human perceptual evaluation
      metrics such as MOS and WER, and correcting this mismatch through preference-aware fine-tuning yields large
      performance gains.
    source: §1, §4.2, Table 1
    evidence: Standard supervised training objectives for TTS produce a systematic mismatch with human perceptual
      evaluation metrics such as MOS and WER, and correcting this mismatch through preference-aware fine-tuning
      yields large performance gains.
    confidence: high
    relevance: medium
  - claim_id: existing_rlhf_methods_requiring_pairwise_preference_data_from_the_same
    role: supports
    claim: Existing RLHF methods requiring pairwise preference data from the same input (DPO) are difficult to apply
      directly to autoregressive codec TTS because these models lack sufficient output diversity to form meaningful
      preference pairs from a fixed transcript-prompt combination.
    source: §4.2, Appendix B
    evidence: Existing RLHF methods requiring pairwise preference data from the same input (DPO) are difficult to
      apply directly to autoregressive codec TTS because these models lack sufficient output diversity to form meaningful
      preference pairs from a fixed transcript-prompt combination.
    confidence: high
    relevance: high
  - claim_id: uncertainty_in_human_speech_quality_annotations_is_not_noise_to
    role: supports
    claim: Uncertainty in human speech quality annotations is not noise to be discarded but an informative signal
      that, when incorporated into the optimization objective, improves the consistency of generated speech across
      listeners.
    source: §4.2, §6.3, Table 3
    evidence: Uncertainty in human speech quality annotations is not noise to be discarded but an informative signal
      that, when incorporated into the optimization objective, improves the consistency of generated speech across
      listeners.
    confidence: high
    relevance: medium
  - claim_id: rlhf_style_alignment_for_tts_can_be_achieved_with_a
    role: supports
    claim: RLHF-style alignment for TTS can be achieved with a small number of self-generated samples (hundreds)
      without access to ground truth speech, making it practical for post-training fine-tuning at low computational
      cost.
    source: §4.1, §5, Appendix D
    evidence: RLHF-style alignment for TTS can be achieved with a small number of self-generated samples (hundreds)
      without access to ground truth speech, making it practical for post-training fine-tuning at low computational
      cost.
    confidence: high
    relevance: medium
  - claim_id: alignment_objectives_designed_for_naturalness_mos_transfer_to_other_perceptual
    role: supports
    claim: Alignment objectives designed for naturalness MOS transfer to other perceptual dimensions such as emotion
      by substituting the selection criterion, demonstrating that preference-based fine-tuning generalises beyond
      a single quality axis.
    source: §6.4, Table 4
    evidence: Alignment objectives designed for naturalness MOS transfer to other perceptual dimensions such as
      emotion by substituting the selection criterion, demonstrating that preference-based fine-tuning generalises
      beyond a single quality axis.
    confidence: high
    relevance: low
  limitations:
  - 'The comparison with SpeechAlign is acknowledged by the authors to be partially unfair: SpeechAlign-DPO requires
    ground truth speech as positive samples during optimization, which is additional supervision not available to
    UNO. Presenting both as baselines without fully separating this distinction may understate SpeechAlign''s performance
    under matched conditions.'
  - UNO is evaluated exclusively on VoiceCraft, an autoregressive codec language model. Whether the training framework
    transfers to diffusion-based TTS (e.g., NaturalSpeech 3) or flow-matching models is discussed theoretically
    but untested. The paper notes that diffusion-based DPO works have been proposed in parallel for image generation,
    suggesting it is feasible, but speech-specific evidence is absent. Human evaluation uses only 10 listeners for
    120 samples, which is relatively small and may not capture the full distribution of listener preferences across
    accents, domains, or speaking styles. The uncertainty estimators (EDL, I-CNF) are trained on SOMOS, a dataset
    of 200 TTS systems with over 17 annotations per sample, and their behaviour on out-of-distribution speech styles
    is unknown. The practical annotation cost of 400 samples per optimization run is modest but still requires human
    or surrogate labelling infrastructure.
  caveats: []
- id: '2406.04904'
  published_date: "2024-06-07"
  entry_date: '2026-07-28'
  year: 2024
  venue: arXiv
  task:
  - TTS
  architecture:
  - autoregressive-LM
  - GAN
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: influential
  method_family:
  - autoregressive_codec_language_models
  - gan_neural_codecs_and_decoders
  claims:
  - claim_id: multilingual_zero_shot_tts_training_degrades_speaker_similarity_compared_to
    role: complicates
    claim: Multilingual zero-shot TTS training degrades speaker similarity compared to monolingual training on the
      same data, reflecting a fundamental trade-off in cross-lingual speaker conditioning.
    source: §4.1, Table 2, Table 3
    evidence: Multilingual zero-shot TTS training degrades speaker similarity compared to monolingual training on
      the same data, reflecting a fundamental trade-off in cross-lingual speaker conditioning.
    confidence: high
    relevance: low
  - claim_id: a_perceiver_resampler_based_speaker_conditioning_encoder_producing_multiple_fixed
    role: supports
    claim: A Perceiver Resampler-based speaker conditioning encoder, producing multiple fixed-length embeddings
      from variable-length reference audio, improves voice cloning robustness in massively multilingual autoregressive
      TTS over single-embedding approaches.
    source: §2
    evidence: A Perceiver Resampler-based speaker conditioning encoder, producing multiple fixed-length embeddings
      from variable-length reference audio, improves voice cloning robustness in massively multilingual autoregressive
      TTS over single-embedding approaches.
    confidence: high
    relevance: medium
  - claim_id: evaluating_multilingual_tts_models_against_monolingual_baselines_on_the_same
    role: supports
    claim: Evaluating multilingual TTS models against monolingual baselines on the same language produces misleading
      comparisons, because the multilingual model's per-language training data is substantially reduced.
    source: §3.2, §4.1
    evidence: Evaluating multilingual TTS models against monolingual baselines on the same language produces misleading
      comparisons, because the multilingual model's per-language training data is substantially reduced.
    confidence: high
    relevance: medium
  - claim_id: a_small_amount_of_target_speaker_fine_tuning_data_approximately
    role: supports
    claim: A small amount of target-speaker fine-tuning data (approximately 10 minutes) substantially improves speaker
      similarity in cross-lingual zero-shot synthesis, including extreme prosody styles such as whispering.
    source: §5
    evidence: A small amount of target-speaker fine-tuning data (approximately 10 minutes) substantially improves
      speaker similarity in cross-lingual zero-shot synthesis, including extreme prosody styles such as whispering.
    confidence: high
    relevance: low
  - claim_id: low_frequency_codec_codebook_entries_can_be_pruned_without_quality
    role: supports
    claim: Low-frequency codec codebook entries can be pruned without quality loss and improve expressiveness in
      multilingual discrete-token TTS.
    source: §2
    evidence: Low-frequency codec codebook entries can be pruned without quality loss and improve expressiveness
      in multilingual discrete-token TTS.
    confidence: high
    relevance: high
  limitations:
  - Speaker similarity lags behind monolingual specialists in English, and the multilingual evaluation uses cross-lingual
    prompting (English speaker references for non-English languages), which may understate true within-language
    similarity. No human listening test was conducted for multilingual outputs beyond subjective English comparisons.
  - The paper lacks ablations isolating the contribution of the Perceiver Resampler versus the larger reference
    representation alone. The VQ-VAE compression (21.53 Hz, 1 codebook) is highly compact compared to EnCodec at
    75 Hz with 8 codebooks; the quality ceiling this imposes is not characterised against higher-fidelity codecs.
    Arabic and CJK language results remain weakest by CER (Table 4), and the causes, whether limited training data,
    romanisation quality, or tokeniser coverage, are not analysed. The authors acknowledge future intent to disentangle
    speaker and prosody for cross-speaker prosody transfer, which the current architecture does not support.
  caveats: []
- id: '2406.05370'
  published_date: "2024-06-08"
  entry_date: '2026-07-28'
  year: 2024
  venue: arXiv
  task:
  - TTS
  architecture:
  - autoregressive-LM
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  - evaluation_caution
  current_role: influential
  method_family:
  - autoregressive_codec_language_models
  claims:
  - claim_id: adaptive_sampling_that_detects_and_breaks_token_repetition_loops_can
    role: supports
    claim: Adaptive sampling that detects and breaks token repetition loops can stabilise autoregressive codec LM
      decoding without requiring forced-alignment auxiliary data.
    source: §3.4.1, Table 1
    evidence: Adaptive sampling that detects and breaks token repetition loops can stabilise autoregressive codec
      LM decoding without requiring forced-alignment auxiliary data.
    confidence: high
    relevance: high
  - claim_id: grouping_codec_codes_into_multi_token_ar_steps_reduces_effective
    role: supports
    claim: Grouping codec codes into multi-token AR steps reduces effective sequence length and simultaneously improves
      long-context modelling quality at moderate group sizes.
    source: §3.1, §4.2.1, Table 1
    evidence: Grouping codec codes into multi-token AR steps reduces effective sequence length and simultaneously
      improves long-context modelling quality at moderate group sizes.
    confidence: high
    relevance: high
  - claim_id: autoregressive_codec_tts_can_match_or_exceed_ground_truth_speech
    role: supports
    claim: Autoregressive codec TTS can match or exceed ground-truth speech on robustness and speaker similarity
      metrics when evaluated on clean English audiobook benchmarks.
    source: §4.2.2, Table 2; §4.3.2, Table 5
    evidence: Autoregressive codec TTS can match or exceed ground-truth speech on robustness and speaker similarity
      metrics when evaluated on clean English audiobook benchmarks.
    confidence: high
    relevance: high
  - claim_id: prompt_availability_in_both_the_ar_and_nar_stages_is
    role: supports
    claim: Prompt availability in both the AR and NAR stages is independently necessary for preserving speaker identity;
      removing either prompt degrades speaker similarity substantially.
    source: §4.2.3, Table 3; §4.3.3, Table 6
    evidence: Prompt availability in both the AR and NAR stages is independently necessary for preserving speaker
      identity; removing either prompt degrades speaker similarity substantially.
    confidence: high
    relevance: low
  - claim_id: inference_time_multiple_sampling_followed_by_metric_based_selection_can
    role: supports
    claim: Inference-time multiple sampling followed by metric-based selection can substantially close the single-sample
      robustness gap, but at proportional computational cost.
    source: §4.1.3, Table 1
    evidence: Inference-time multiple sampling followed by metric-based selection can substantially close the single-sample
      robustness gap, but at proportional computational cost.
    confidence: high
    relevance: medium
  limitations:
  - Human parity is claimed solely from results on LibriSpeech test-clean and VCTK; both benchmarks are read speech
    from controlled or semi-controlled recording conditions. Generalisation to spontaneous, noisy, or low-resource
    speech is undemonstrated and the authors explicitly flag this caveat.
  - The model is English-only and trained on audiobook data (Libriheavy), leaving multilingual and conversational
    speech scenarios unexplored. No model size is reported, making it difficult to assess parameter efficiency against
    competing approaches. Code is not released, limiting reproducibility. The paper does not compare against contemporaneous
    diffusion or flow-matching zero-shot systems such as NaturalSpeech 3 or Voicebox on the same test sets, so cross-paradigm
    positioning is unclear. The stability benefit of Repetition Aware Sampling comes at the cost of a non-deterministic
    inference procedure, which may complicate deployment in latency-sensitive applications.
  caveats: []
- id: '2406.07855'
  published_date: "2024-06-12"
  entry_date: '2026-07-28'
  year: 2024
  venue: arXiv
  task:
  - TTS
  architecture:
  - autoregressive-LM
  relevance: medium
  evidence_role:
  - architecture_variant
  - acceleration_evidence
  - control_evidence
  current_role: active_evidence
  method_family:
  - autoregressive_codec_language_models
  claims:
  - claim_id: phoneme_monotonic_alignment_in_decoder_only_autoregressive_tts_can_close
    role: supports
    claim: Phoneme monotonic alignment in decoder-only autoregressive TTS can close most of the robustness gap caused
      by unconstrained attention, achieving near-ground-truth WER without encoder-decoder architectural changes.
    source: §3.2, Table 1
    evidence: Phoneme monotonic alignment in decoder-only autoregressive TTS can close most of the robustness gap
      caused by unconstrained attention, achieving near-ground-truth WER without encoder-decoder architectural changes.
    confidence: high
    relevance: medium
  - claim_id: downsampling_only_the_first_rvq_layer_of_a_neural_codec
    role: supports
    claim: Downsampling only the first RVQ layer of a neural codec at inference time reduces autoregressive steps
      and latency by more than half, with negligible impact on PESQ and STOI.
    source: §3.1, Table 5
    evidence: Downsampling only the first RVQ layer of a neural codec at inference time reduces autoregressive steps
      and latency by more than half, with negligible impact on PESQ and STOI.
    confidence: high
    relevance: high
  - claim_id: robustness_improvements_that_route_additional_phoneme_tokens_through_the_autoregressive
    role: complicates
    claim: Robustness improvements that route additional phoneme tokens through the autoregressive stream (as in
      ELLA-V) improve WER but increase inference time, illustrating a robustness-efficiency trade-off in codec LM
      TTS.
    source: §5.3, Table 4
    evidence: Robustness improvements that route additional phoneme tokens through the autoregressive stream (as
      in ELLA-V) improve WER but increase inference time, illustrating a robustness-efficiency trade-off in codec
      LM TTS.
    confidence: high
    relevance: high
  - claim_id: explicit_phoneme_level_alignment_in_a_codec_lm_enables_independent
    role: supports
    claim: Explicit phoneme-level alignment in a codec LM enables independent control of prosody and timbre by substituting
      preset phoneme sequences at inference, enabling a form of voice conversion.
    source: §3.2.3, Table 3
    evidence: Explicit phoneme-level alignment in a codec LM enables independent control of prosody and timbre by
      substituting preset phoneme sequences at inference, enabling a form of voice conversion.
    confidence: high
    relevance: high
  limitations:
  - All evaluations use LibriSpeech (clean English read speech). Robustness gains from monotonic alignment and codec-merging
    quality preservation have not been tested on noisy, expressive, or multilingual speech.
  - The model size is not explicitly reported, though the architecture (12-layer Transformer, 1024-dim hidden, 16
    heads) matches the VALL-E reference scale. Code and model weights are not publicly released, limiting reproducibility.
    The prosody control evaluation uses MCD-DTW-SL, a proxy metric; perceptual validation of prosody cloning quality
    is absent. The merging rate of 2x is validated by reconstruction metrics but its downstream effect on naturalness
    under diverse speaker and content conditions is not fully explored. RALL-E (chain-of-thought prompting for robustness)
    is included only in the efficiency comparison, not in the WER robustness comparison, making head-to-head robustness
    assessment with concurrent work incomplete.
  caveats: []
- id: '2407.05407'
  published_date: "2024-07-07"
  entry_date: '2026-07-28'
  year: 2024
  venue: arXiv
  task:
  - TTS
  architecture:
  - autoregressive-LM
  - flow-matching
  - hybrid
  relevance: medium
  evidence_role:
  - architecture_variant
  - acceleration_evidence
  - control_evidence
  current_role: influential
  method_family:
  - autoregressive_codec_language_models
  - flow_matching_codec_decoders
  - hybrid_semantic_acoustic_tokenizers
  claims:
  - claim_id: inserting_a_vector_quantizer_into_a_supervised_asr_encoder_yields
    role: supports
    claim: Inserting a vector quantizer into a supervised ASR encoder yields discrete speech tokens that preserve
      significantly stronger text-semantic alignment than unsupervised alternatives such as HuBERT or EnCodec tokens.
    source: §2.1, §5.1, Table 7
    evidence: Inserting a vector quantizer into a supervised ASR encoder yields discrete speech tokens that preserve
      significantly stronger text-semantic alignment than unsupervised alternatives such as HuBERT or EnCodec tokens.
    confidence: high
    relevance: high
  - claim_id: in_autoregressive_codec_lm_tts_both_the_text_tokenizer_and
    role: supports
    claim: In autoregressive codec-LM TTS, both the text tokenizer and the speech tokenizer independently contribute
      to content consistency, while speaker similarity is primarily controlled by the speaker embedding and acoustic
      model conditioning.
    source: §5.2, Table 7
    evidence: In autoregressive codec-LM TTS, both the text tokenizer and the speech tokenizer independently contribute
      to content consistency, while speaker similarity is primarily controlled by the speaker embedding and acoustic
      model conditioning.
    confidence: high
    relevance: high
  - claim_id: asr_re_ranking_is_an_effective_post_hoc_method_for
    role: complicates
    claim: ASR re-ranking is an effective post-hoc method for improving content consistency in autoregressive TTS
      without any model retraining, at the cost of increased inference-time compute.
    source: §5.3, Tables 8, 9
    evidence: ASR re-ranking is an effective post-hoc method for improving content consistency in autoregressive
      TTS without any model retraining, at the cost of increased inference-time compute.
    confidence: high
    relevance: medium
  - claim_id: instruction_fine_tuning_on_a_modest_amount_of_labelled_data
    role: supports
    claim: Instruction fine-tuning on a modest amount of labelled data (556h) enables a TTS system to control fine-grained
      paralinguistic features — including laughter, breath, and word emphasis — with substantially improved accuracy
      over the base model.
    source: §2.4, §5.4, Table 10
    evidence: Instruction fine-tuning on a modest amount of labelled data (556h) enables a TTS system to control
      fine-grained paralinguistic features — including laughter, breath, and word emphasis — with substantially
      improved accuracy over the base model.
    confidence: high
    relevance: medium
  - claim_id: high_quality_tts_synthesized_speech_can_serve_as_effective_training
    role: supports
    claim: High-quality TTS-synthesized speech can serve as effective training data augmentation for ASR, with text
      diversity of the synthesis prompts contributing more to downstream ASR gains than the raw duration of the
      synthetic corpus.
    source: §5.5, Table 11
    evidence: High-quality TTS-synthesized speech can serve as effective training data augmentation for ASR, with
      text diversity of the synthesis prompts contributing more to downstream ASR gains than the raw duration of
      the synthetic corpus.
    confidence: high
    relevance: medium
  limitations:
  - '- Only a single VQ codebook (4096 codes) is used; multi-level RVQ and its effect on quality vs. compression
    is left for future work. - The choice of VQ insertion layer (after layer 6 of 12) is not ablated — optimal placement
    is unresolved. - Cross-lingual cloning omits prompt prosody to prevent leakage, which may reduce naturalness
    in target language. - Instruction fine-tuning data amounts (556h) are modest; broader paralinguistic coverage
    remains open. - No subjective (MOS) evaluation in the main paper; relies entirely on objective WER/CER/SS metrics.'
  caveats: []
- id: '2407.08551'
  published_date: "2024-07-11"
  entry_date: '2026-07-28'
  year: 2024
  venue: arXiv
  task:
  - TTS
  architecture:
  - autoregressive-LM
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: influential
  method_family:
  - autoregressive_codec_language_models
  claims:
  - claim_id: continuous_mel_spectrogram_representations_preserve_more_speaker_relevant_acoustic_information
    role: supports
    claim: Continuous mel-spectrogram representations preserve more speaker-relevant acoustic information than vector-quantized
      codec codes at standard compression rates.
    source: §5.1, Table 1
    evidence: Continuous mel-spectrogram representations preserve more speaker-relevant acoustic information than
      vector-quantized codec codes at standard compression rates.
    confidence: high
    relevance: high
  - claim_id: autoregressive_tts_models_trained_to_predict_continuous_frames_can_achieve
    role: supports
    claim: Autoregressive TTS models trained to predict continuous frames can achieve naturalness comparable to
      human speech while avoiding the silence and repetition failures endemic to discrete codec language models.
    source: §5.2, Table 3
    evidence: Autoregressive TTS models trained to predict continuous frames can achieve naturalness comparable
      to human speech while avoiding the silence and repetition failures endemic to discrete codec language models.
    confidence: high
    relevance: high
  - claim_id: variational_sampling_in_the_continuous_latent_space_is_more_effective
    role: supports
    claim: Variational sampling in the continuous latent space is more effective than top-p discrete sampling for
      improving output diversity and speaker similarity in autoregressive TTS.
    source: §5.3, Table 4
    evidence: Variational sampling in the continuous latent space is more effective than top-p discrete sampling
      for improving output diversity and speaker similarity in autoregressive TTS.
    confidence: high
    relevance: low
  - claim_id: a_reduction_factor_that_predicts_multiple_frames_per_autoregressive_step
    role: supports
    claim: A reduction factor that predicts multiple frames per autoregressive step can substantially reduce inference
      time with only modest degradation in speaker similarity.
    source: §5.4, Table 5
    evidence: A reduction factor that predicts multiple frames per autoregressive step can substantially reduce
      inference time with only modest degradation in speaker similarity.
    confidence: high
    relevance: low
  limitations:
  - The subjective evaluation rests on only 40 samples from a single English corpus (LibriSpeech test-clean). The
    naturalness and speaker similarity advantages may not generalize to noisier prompts, non-native accents, or
    other languages.
  - The model's output quality is bounded by the HiFi-GAN vocoder, which was trained on only 585 hours of LibriTTS.
    Voicebox, which used a proprietary 60K-hour vocoder, showed higher SIM in part for this reason. Replacing or
    scaling the vocoder is identified as the most direct path to improvement.
  - The evaluation is English-only. Multilingual extension analogous to VALL-E X is deferred to future work. The
    paper also leaves open whether other continuous representations (VAE latent spaces, flow-based representations)
    would outperform mel-spectrograms as the target token. The model size is not reported, limiting cost comparisons
    with VALL-E 2 or Voicebox.
  caveats: []
- id: '2408.16532'
  published_date: "2024-08-29"
  entry_date: '2026-07-28'
  year: 2024
  venue: arXiv
  task:
  - codec
  architecture:
  - GAN
  - VAE
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: influential
  method_family:
  - gan_neural_codecs_and_decoders
  - vae_vector_quantized_codecs
  claims:
  - claim_id: a_single_large_codebook_quantizer_can_achieve_higher_perceptual_reconstruction
    role: supports
    claim: A single large-codebook quantizer can achieve higher perceptual reconstruction quality than multi-quantizer
      residual VQ systems at substantially higher bitrates, when supported by a strong decoder.
    source: §4.2, Table 1, Table 2
    evidence: A single large-codebook quantizer can achieve higher perceptual reconstruction quality than multi-quantizer
      residual VQ systems at substantially higher bitrates, when supported by a strong decoder.
    confidence: high
    relevance: high
  - claim_id: codec_tokens_produced_by_a_single_quantizer_enable_better_downstream
    role: supports
    claim: Codec tokens produced by a single quantizer enable better downstream autoregressive speech generation
      than tokens from multi-quantizer codecs, as measured by intelligibility and speaker similarity.
    source: §4.2, Appendix I, Table 12
    evidence: Codec tokens produced by a single quantizer enable better downstream autoregressive speech generation
      than tokens from multi-quantizer codecs, as measured by intelligibility and speaker similarity.
    confidence: high
    relevance: high
  - claim_id: incorporating_attention_mechanisms_in_the_codec_decoder_and_extending_the
    role: supports
    claim: Incorporating attention mechanisms in the codec decoder and extending the training context window improves
      the semantic richness of discrete audio tokens without requiring distillation from a semantic model.
    source: §3.3, §4.3, Table 9
    evidence: Incorporating attention mechanisms in the codec decoder and extending the training context window
      improves the semantic richness of discrete audio tokens without requiring distillation from a semantic model.
    confidence: high
    relevance: high
  - claim_id: expanding_the_vq_codebook_space_beyond_the_conventional_1024_entries
    role: supports
    claim: Expanding the VQ codebook space beyond the conventional 1024 entries improves reconstruction quality
      under extreme compression, but excessively large codebooks reduce codebook utilisation and yield diminishing
      returns.
    source: §3.2, §4.3, Table 5
    evidence: Expanding the VQ codebook space beyond the conventional 1024 entries improves reconstruction quality
      under extreme compression, but excessively large codebooks reduce codebook utilisation and yield diminishing
      returns.
    confidence: high
    relevance: high
  - claim_id: inverse_fourier_transform_decoding_significantly_outperforms_mirrored_transposed_convolution_upsampling
    role: supports
    claim: Inverse Fourier transform decoding significantly outperforms mirrored transposed-convolution upsampling
      in high-compression codec settings.
    source: §3.3, §4.3, Table 7
    evidence: Inverse Fourier transform decoding significantly outperforms mirrored transposed-convolution upsampling
      in high-compression codec settings.
    confidence: high
    relevance: high
  limitations:
  - The downstream TTS evaluation uses only LibriTTS (~960 hours) and a single model configuration; the claimed
    advantages of WavTokenizer over multi-quantizer codecs in generative modelling have not been validated at the
    scale of systems like VALL-E or Voicebox, where the codec is a fixed component in a much larger pipeline.
  - Acoustic codecs including WavTokenizer lack ASR-level speech understanding capabilities; the authors note this
    constrains use in unified multimodal understanding-and-generation frameworks (GPT-4o paradigm). The encoder
    design remains largely conventional, and the paper defers encoder improvement to future work. Training uses
    approximately 8K hours, which is substantially less than frontier systems, and the scalability of the codebook
    design with larger data and model scale is unexplored. The codebook analysis shows that even with 4000 hours
    of training, codebook utilisation plateaus below 2^12, leaving the upper bound of useful codebook expansion
    unclear.
  caveats: []
- id: '2408.16725'
  published_date: "2024-08-29"
  entry_date: '2026-07-28'
  year: 2024
  venue: arXiv
  task:
  - SCA
  - TTS
  architecture:
  - autoregressive-LM
  relevance: medium
  evidence_role:
  - architecture_variant
  - acceleration_evidence
  - control_evidence
  current_role: influential
  method_family:
  - autoregressive_codec_language_models
  claims:
  - claim_id: simultaneous_text_and_audio_generation_conditioned_on_text_tokens_generated
    role: supports
    claim: Simultaneous text and audio generation, conditioned on text tokens generated in parallel, enables streaming
      speech output without the latency penalty of sequential text-then-audio decoding.
    source: §3.2
    evidence: Simultaneous text and audio generation, conditioned on text tokens generated in parallel, enables
      streaming speech output without the latency penalty of sequential text-then-audio decoding.
    confidence: high
    relevance: low
  - claim_id: audio_reasoning_quality_in_end_to_end_speech_lms_lags
    role: supports
    claim: Audio reasoning quality in end-to-end speech LMs lags behind text reasoning quality when trained on similar
      data volumes, and batch inference strategies can partially bridge this gap.
    source: §3.2, §4.4
    evidence: Audio reasoning quality in end-to-end speech LMs lags behind text reasoning quality when trained on
      similar data volumes, and batch inference strategies can partially bridge this gap.
    confidence: high
    relevance: medium
  - claim_id: a_three_stage_adapter_based_training_curriculum_can_integrate_speech
    role: supports
    claim: A three-stage adapter-based training curriculum can integrate speech input and output into a frozen language
      model backbone with minimal degradation to text capabilities.
    source: §3.3
    evidence: A three-stage adapter-based training curriculum can integrate speech input and output into a frozen
      language model backbone with minimal degradation to text capabilities.
    confidence: high
    relevance: medium
  - claim_id: multi_codebook_audio_codecs_with_high_token_rates_require_parallel
    role: supports
    claim: Multi-codebook audio codecs with high token rates require parallel decoding schemes to maintain practical
      streaming throughput in autoregressive speech LMs.
    source: §3.1, §3.2
    evidence: Multi-codebook audio codecs with high token rates require parallel decoding schemes to maintain practical
      streaming throughput in autoregressive speech LMs.
    confidence: high
    relevance: high
  limitations:
  - The paper reports no MOS or naturalness metrics for speech output, making it impossible to quantitatively compare
    audio quality against TTS or SCA baselines. The claim that quality is "on par with common TTS systems" is unsupported.
  - Evaluation is restricted to ASR performance on LibriSpeech; there is no evaluation of conversational quality,
    response coherence, or latency. The 0.5B model size limits reasoning depth, and the paper acknowledges that
    audio reasoning remains weaker than text reasoning. The VoiceAssistant-400K dataset is entirely synthesized
    by GPT-4o, which may introduce systematic biases in prosody and topic coverage. The model supports only English.
    The paper is described as a work-in-progress technical report, with some experiments deferred to a future version.
  caveats: []
- id: '2409.00750'
  published_date: "2024-09-01"
  entry_date: '2026-07-28'
  year: 2024
  venue: arXiv
  task:
  - TTS
  - VC
  architecture:
  - autoregressive-LM
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: influential
  method_family:
  - autoregressive_codec_language_models
  claims:
  - claim_id: non_autoregressive_masked_generative_transformers_can_achieve_human_level_speaker
    role: supports
    claim: Non-autoregressive masked generative transformers can achieve human-level speaker similarity in zero-shot
      TTS without requiring explicit text-speech alignment or phone-level duration supervision.
    source: §4.2.1, Table 2
    evidence: Non-autoregressive masked generative transformers can achieve human-level speaker similarity in zero-shot
      TTS without requiring explicit text-speech alignment or phone-level duration supervision.
    confidence: high
    relevance: low
  - claim_id: replacing_k_means_quantisation_of_ssl_features_with_vq_vae
    role: supports
    claim: Replacing k-means quantisation of SSL features with VQ-VAE vector quantisation reduces information loss
      in tonal languages and improves downstream acoustic token prediction.
    source: §3.2.1
    evidence: Replacing k-means quantisation of SSL features with VQ-VAE vector quantisation reduces information
      loss in tonal languages and improves downstream acoustic token prediction.
    confidence: high
    relevance: high
  - claim_id: masked_generative_tts_substantially_outperforms_autoregressive_tts_on_hard_text
    role: supports
    claim: Masked generative TTS substantially outperforms autoregressive TTS on hard-text robustness (tongue twisters,
      repeating phrases) while maintaining competitive naturalness on standard benchmarks.
    source: §4.2.2, Appendix J, Table 13
    evidence: Masked generative TTS substantially outperforms autoregressive TTS on hard-text robustness (tongue
      twisters, repeating phrases) while maintaining competitive naturalness on standard benchmarks.
    confidence: high
    relevance: medium
  - claim_id: parallel_iterative_decoding_in_masked_generative_models_yields_constant_inference
    role: supports
    claim: Parallel iterative decoding in masked generative models yields constant inference cost regardless of
      output length, in contrast to autoregressive decoding whose cost scales linearly with utterance duration.
    source: §4.2.2
    evidence: Parallel iterative decoding in masked generative models yields constant inference cost regardless
      of output length, in contrast to autoregressive decoding whose cost scales linearly with utterance duration.
    confidence: high
    relevance: medium
  - claim_id: zero_shot_style_cloning_via_in_context_learning_extends_to
    role: supports
    claim: Zero-shot style cloning via in-context learning extends to accent and emotion transfer without task-specific
      architectural changes.
    source: §4.3, Tables 4–5
    evidence: Zero-shot style cloning via in-context learning extends to accent and emotion transfer without task-specific
      architectural changes.
    confidence: high
    relevance: low
  limitations:
  - Speech content editing is acknowledged as "not very robust" by the authors, who attribute this to a training
    objective mismatch (mask-and-predict vs. fill-in-mask). The editing capability is demonstrated qualitatively
    only, with no quantitative evaluation reported.
  - 'Training uses 100K hours of English and Chinese speech from Emilia, with multilingual extension at far smaller
    data budgets (2,500–8,200 hours per language). Multilingual performance is uneven: French and German show higher
    WER in cross-lingual dubbing, and the authors note limitations from insufficient retraining of all components
    on expanded data.'
  - Duration control requires either a ground-truth length or the flow-matching duration predictor; errors in predicted
    duration propagate to WER. The gap between predicted-length and ground-truth-length WER is measurable (e.g.,
    2.634 vs. 2.012 on LibriSpeech test-clean).
  - Inference steps of 25-50 for T2S plus the S2A step schedule add latency compared to single-pass systems, though
    the paper does not report real-time factor or wall-clock comparisons.
  - Emotion control requires post-training fine-tuning on labelled data; it is not available zero-shot from the
    base model alone.
  caveats: []
- id: '2409.03283'
  published_date: "2024-09-05"
  entry_date: '2026-07-28'
  year: 2024
  venue: arXiv
  task:
  - TTS
  architecture:
  - autoregressive-LM
  - flow-matching
  - hybrid
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: active_evidence
  method_family:
  - autoregressive_codec_language_models
  - flow_matching_codec_decoders
  - hybrid_semantic_acoustic_tokenizers
  claims:
  - claim_id: separating_the_waveform_generation_stage_into_a_low_sampling_rate
    role: supports
    claim: Separating the waveform generation stage into a low-sampling-rate Mel decoder and a super-resolution
      vocoder allows a system trained predominantly on low-sampling-rate data to produce high-fidelity 48kHz output.
    source: §3.3
    evidence: Separating the waveform generation stage into a low-sampling-rate Mel decoder and a super-resolution
      vocoder allows a system trained predominantly on low-sampling-rate data to produce high-fidelity 48kHz output.
    confidence: high
    relevance: medium
  - claim_id: few_shot_fine_tuning_of_a_large_foundation_tts_model
    role: supports
    claim: Few-shot fine-tuning of a large foundation TTS model substantially outperforms zero-shot in-context learning
      for highly expressive, distinctive target voices, even with only one hour of data.
    source: §5.2.1, Table 5
    evidence: Few-shot fine-tuning of a large foundation TTS model substantially outperforms zero-shot in-context
      learning for highly expressive, distinctive target voices, even with only one hour of data.
    confidence: high
    relevance: medium
  - claim_id: prompt_audio_enhancement_improves_voice_cloning_quality_for_noisy_prompts
    role: supports
    claim: Prompt audio enhancement improves voice cloning quality for noisy prompts but can slightly degrade performance
      when prompts are already clean.
    source: §5.2.2, Table 6
    evidence: Prompt audio enhancement improves voice cloning quality for noisy prompts but can slightly degrade
      performance when prompts are already clean.
    confidence: high
    relevance: medium
  - claim_id: instruction_tuning_with_a_small_domain_specific_dataset_dramatically_improves
    role: supports
    claim: Instruction tuning with a small domain-specific dataset dramatically improves emotion controllability
      in a pre-trained TTS language model, raising accuracy from near-chance to near-ceiling.
    source: §5.3, Table 7
    evidence: Instruction tuning with a small domain-specific dataset dramatically improves emotion controllability
      in a pre-trained TTS language model, raising accuracy from near-chance to near-ceiling.
    confidence: high
    relevance: low
  - claim_id: autoregressive_tts_systems_trained_on_predominantly_one_language_show_markedly
    role: supports
    claim: Autoregressive TTS systems trained on predominantly one language show markedly higher pronunciation error
      rates on under-represented languages, even at large data scales.
    source: §5.1.2, Table 3
    evidence: Autoregressive TTS systems trained on predominantly one language show markedly higher pronunciation
      error rates on under-represented languages, even at large data scales.
    confidence: high
    relevance: medium
  limitations:
  - All evaluations are conducted on proprietary internal test sets with no publicly released benchmarks, data,
    or model weights. This makes direct comparison with other systems difficult to reproduce and limits the generalisability
    of the reported numbers.
  - The streamable decoder incurs a measurable quality penalty (0.07 CoMOS) and the paper notes that Mel codec quality
    is a bottleneck, which the authors flag for future work. The English and code-switch pronunciation error rates
    remain high (12% and 8.5%), driven by limited language diversity in training data. The paralinguistic behaviour
    framework currently supports 13 types targeting primarily Chinese conversational speech; coverage of other languages
    and more complex prosodic phenomena is not addressed.
  caveats: []
- id: '2409.05377'
  published_date: "2024-09-09"
  entry_date: '2026-07-28'
  year: 2024
  venue: arXiv
  task:
  - codec
  architecture:
  - GAN
  - VAE
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: influential
  method_family:
  - gan_neural_codecs_and_decoders
  - vae_vector_quantized_codecs
  claims:
  - claim_id: at_very_low_bitrates_around_1_kbps_model_capacity_is
    role: supports
    claim: At very low bitrates (around 1 kbps), model capacity is a decisive factor in reconstruction quality,
      with larger models substantially outperforming architecturally sophisticated but smaller codecs.
    source: §III-C, §IV-B, Table I
    evidence: At very low bitrates (around 1 kbps), model capacity is a decisive factor in reconstruction quality,
      with larger models substantially outperforming architecturally sophisticated but smaller codecs.
    confidence: high
    relevance: medium
  - claim_id: low_dimensional_vector_quantisation_before_codebook_lookup_substantially_improves_codebook
    role: supports
    claim: Low-dimensional vector quantisation before codebook lookup substantially improves codebook utilisation
      in single-codebook, single-quantisation-step codec designs.
    source: §III-A, §IV-B
    evidence: Low-dimensional vector quantisation before codebook lookup substantially improves codebook utilisation
      in single-codebook, single-quantisation-step codec designs.
    confidence: high
    relevance: high
  - claim_id: adding_sequential_lstm_modelling_to_a_convolutional_codec_encoder_improves
    role: supports
    claim: Adding sequential (LSTM) modelling to a convolutional codec encoder improves both perceptual quality
      and speaker similarity at low bitrates, independently of the parameter count effect.
    source: §III-A, §IV-D, Table III
    evidence: Adding sequential (LSTM) modelling to a convolutional codec encoder improves both perceptual quality
      and speaker similarity at low bitrates, independently of the parameter count effect.
    confidence: high
    relevance: high
  - claim_id: scaling_codec_model_size_beyond_a_saturation_point_approximately_159m
    role: supports
    claim: Scaling codec model size beyond a saturation point (approximately 159M parameters in this setting) yields
      no further reconstruction benefit, analogous to scale-up behaviour observed in neural vocoders.
    source: §IV-D, Table III
    evidence: Scaling codec model size beyond a saturation point (approximately 159M parameters in this setting)
      yields no further reconstruction benefit, analogous to scale-up behaviour observed in neural vocoders.
    confidence: high
    relevance: high
  - claim_id: increasing_training_data_volume_from_960_hours_to_60k_hours
    role: supports
    claim: Increasing training data volume from 960 hours to 60k hours does not improve codec reconstruction quality,
      suggesting that model capacity rather than data quantity is the binding constraint at this bitrate.
    source: §IV-D, Table III
    evidence: Increasing training data volume from 960 hours to 60k hours does not improve codec reconstruction
      quality, suggesting that model capacity rather than data quantity is the binding constraint at this bitrate.
    confidence: high
    relevance: high
  limitations:
  - BigCodec is trained exclusively on clean English speech (LibriSpeech 960h), while competing codecs such as EnCodec
    and DAC train on diverse multilingual datasets including music and environmental sounds. The multilingual generalisation
    result is encouraging, but the clean-speech-only training domain limits applicability to noisy or music-heavy
    audio without fine-tuning.
  - The RTF of 1.1x on a high-end desktop CPU means BigCodec barely achieves real-time decoding; edge-device or
    streaming deployments would require hardware acceleration. The model's 159M parameter footprint is an order
    of magnitude larger than EnCodec (14M), imposing memory costs for downstream TTS systems that embed a codec.
    The paper does not report streaming latency or chunked-inference performance, which is relevant for spoken conversational
    agent use cases.
  - Whether the subjective superiority over ground truth reflects a genuine perceptual enhancement or a test artefact
    (e.g., listener anchoring, signal processing smoothing) remains unaddressed and warrants scepticism.
  caveats: []
- id: '2409.06666'
  published_date: "2024-09-10"
  entry_date: '2026-07-28'
  year: 2024
  venue: arXiv
  task:
  - SCA
  architecture:
  - hybrid
  relevance: medium
  evidence_role:
  - architecture_variant
  - acceleration_evidence
  - control_evidence
  current_role: influential
  method_family:
  - hybrid_semantic_acoustic_tokenizers
  claims:
  - claim_id: end_to_end_speech_llms_with_parallel_text_and_speech
    role: supports
    claim: End-to-end speech LLMs with parallel text and speech generation can achieve lower response latency than
      cascaded ASR-LLM-TTS pipelines without sacrificing prosody coherence under low-latency streaming conditions.
    source: §4.5
    evidence: End-to-end speech LLMs with parallel text and speech generation can achieve lower response latency
      than cascaded ASR-LLM-TTS pipelines without sacrificing prosody coherence under low-latency streaming conditions.
    confidence: high
    relevance: low
  - claim_id: aligning_llm_output_to_speech_interaction_conventions_through_targeted_instruction
    role: supports
    claim: Aligning LLM output to speech interaction conventions through targeted instruction data rewriting substantially
      improves response style suitability, independently of model architecture.
    source: §3, §4.4
    evidence: Aligning LLM output to speech interaction conventions through targeted instruction data rewriting
      substantially improves response style suitability, independently of model architecture.
    confidence: high
    relevance: medium
  - claim_id: non_autoregressive_ctc_decoding_from_llm_hidden_states_enables_streaming
    role: supports
    claim: Non-autoregressive CTC decoding from LLM hidden states enables streaming speech synthesis whose speech
      rate and naturalness are robust to chunk size variation, unlike word-level streaming TTS cascades.
    source: §4.5, Table 4
    evidence: Non-autoregressive CTC decoding from LLM hidden states enables streaming speech synthesis whose speech
      rate and naturalness are robust to chunk size variation, unlike word-level streaming TTS cascades.
    confidence: high
    relevance: low
  - claim_id: training_an_end_to_end_speech_interaction_model_on_a
    role: supports
    claim: Training an end-to-end speech interaction model on a small, carefully curated speech instruction dataset
      is sufficient to significantly close the gap with models trained on orders of magnitude more data, provided
      the LLM backbone is sufficiently capable.
    source: §4.4, §5
    evidence: Training an end-to-end speech interaction model on a small, carefully curated speech instruction dataset
      is sufficient to significantly close the gap with models trained on orders of magnitude more data, provided
      the LLM backbone is sufficiently capable.
    confidence: high
    relevance: medium
  limitations:
  - The ASR-WER of 10.82% is notably higher than cascaded baselines (3.78% for SALMONN+Orca), reflecting that the
    speech decoder is trained on only approximately 1K hours of response speech — far below industrial TTS scale.
    Intelligibility limitations restrict applicability in domains requiring precise spoken content.
  - The evaluation benchmark (InstructS2S-Eval) is derived from AlpacaEval with math and code questions removed,
    which skews toward conversational helpfulness and may not represent more demanding speech interaction tasks.
    The speech encoder relies on Whisper, which is optimised for ASR rather than general speech understanding, potentially
    limiting response to prosodic or para-linguistic cues in the user's speech. The current architecture does not
    support full-duplex interaction (interruption, turn-taking) — speech responses are generated after the full
    instruction is received. The training data is synthesised from text corpora, which may not capture the naturalness
    and variability of real spoken dialogue.
  caveats: []
- id: '2410.00037'
  published_date: "2024-09-17"
  entry_date: '2026-07-28'
  year: 2024
  venue: arXiv
  task:
  - SCA
  - TTS
  architecture:
  - autoregressive-LM
  - hybrid
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  - evaluation_caution
  current_role: foundational
  method_family:
  - autoregressive_codec_language_models
  - hybrid_semantic_acoustic_tokenizers
  claims:
  - claim_id: eliminating_the_text_bottleneck_in_spoken_dialogue_requires_modeling_acoustic
    role: supports
    claim: Eliminating the text bottleneck in spoken dialogue requires modeling acoustic tokens jointly with semantic
      tokens in a single generative model, as purely semantic approaches cannot capture paralinguistic information
      or generate in arbitrary voices.
    source: §3.4, §5.4
    evidence: Eliminating the text bottleneck in spoken dialogue requires modeling acoustic tokens jointly with
      semantic tokens in a single generative model, as purely semantic approaches cannot capture paralinguistic
      information or generate in arbitrary voices.
    confidence: high
    relevance: low
  - claim_id: predicting_time_aligned_text_tokens_as_a_per_frame_prefix
    role: supports
    claim: Predicting time-aligned text tokens as a per-frame prefix to audio tokens substantially improves the
      linguistic quality and factual accuracy of speech generated by audio language models, with minimal inference
      overhead.
    source: §3.4.4, §5.3, Table 6
    evidence: Predicting time-aligned text tokens as a per-frame prefix to audio tokens substantially improves the
      linguistic quality and factual accuracy of speech generated by audio language models, with minimal inference
      overhead.
    confidence: high
    relevance: medium
  - claim_id: modeling_conversation_as_parallel_autoregressive_streams_for_each_speaker_without
    role: supports
    claim: Modeling conversation as parallel autoregressive streams for each speaker, without explicit turn boundaries,
      enables full-duplex spoken interaction and allows training on naturally overlapping speech.
    source: §3.4.3, §5.6, Table 9
    evidence: Modeling conversation as parallel autoregressive streams for each speaker, without explicit turn boundaries,
      enables full-duplex spoken interaction and allows training on naturally overlapping speech.
    confidence: high
    relevance: medium
  - claim_id: adversarial_only_training_of_neural_audio_codecs_substantially_improves_subjectively
    role: supports
    claim: Adversarial-only training of neural audio codecs substantially improves subjectively rated audio quality
      relative to mixed reconstruction-adversarial objectives, despite degrading objective metrics such as VisQOL.
    source: §3.3, §5.2, Table 4
    evidence: Adversarial-only training of neural audio codecs substantially improves subjectively rated audio quality
      relative to mixed reconstruction-adversarial objectives, despite degrading objective metrics such as VisQOL.
    confidence: high
    relevance: medium
  - claim_id: standard_objective_audio_quality_metrics_visqol_mosnet_are_unreliable_proxies
    role: supports
    claim: Standard objective audio quality metrics (VisQOL, MOSNet) are unreliable proxies for perceived quality
      when the training objective changes, making human evaluation indispensable for codec comparison.
    source: §5.2, §5.8
    evidence: Standard objective audio quality metrics (VisQOL, MOSNet) are unreliable proxies for perceived quality
      when the training objective changes, making human evaluation indispensable for codec comparison.
    confidence: high
    relevance: high
  limitations:
  - 'Moshi''s spoken factual question answering performance lags substantially behind its Helium text baseline,
    particularly on multi-sentence or syntactically complex questions (TriviaQA: 22.8 vs. 56.4 for text-only Helium).
    This indicates that audio training causes significant forgetting of factual knowledge, and the instruct fine-tuning
    data does not cover the syntactic diversity needed to recover it.'
  - 'Signal-based watermarking (Audioseal) is ineffective against codec compression: Mimi''s own lossy coding removes
    the watermark to below detection threshold. The generative watermarking alternatives explored in §6.4 are blocked
    by the non-idempotence of audio codecs, leaving no robust content attribution mechanism available at release.'
  - The instruction fine-tuning pipeline relies heavily on synthetic TTS-generated speech for both conversation
    transcripts and user voice diversity. This introduces a distribution mismatch with real conversational speech
    that likely limits robustness to unusual acoustic conditions and speaking styles. The paper notes this but leaves
    more realistic instruct data collection as future work.
  - Quantization below 4-bit precision causes noticeable audio artifacts (repetitive generation, noisy voice) that
    current automatic metrics fail to detect, requiring entropy-spectrum analysis as a surrogate. This underscores
    a general gap in speech quality evaluation tooling for generative dialogue models.
  caveats: []
- id: '2410.03751'
  published_date: "2024-10-01"
  entry_date: '2026-07-28'
  year: 2024
  venue: arXiv
  task:
  - SCA
  - TTS
  architecture: []
  relevance: medium
  evidence_role:
  - architecture_variant
  - acceleration_evidence
  - infrastructure
  current_role: influential
  method_family: []
  claims:
  - claim_id: end_to_end_speech_generation_models_avoid_the_information_loss
    role: supports
    claim: End-to-end speech generation models avoid the information loss, latency, and cumulative error introduced
      by cascaded ASR-LLM-TTS pipelines, but require integrating speech tokenisation and synthesis into a unified
      training regime.
    source: §I, §II
    evidence: End-to-end speech generation models avoid the information loss, latency, and cumulative error introduced
      by cascaded ASR-LLM-TTS pipelines, but require integrating speech tokenisation and synthesis into a unified
      training regime.
    confidence: high
    relevance: low
  - claim_id: semantic_tokenizers_and_acoustic_tokenizers_impose_an_inherent_trade_off
    role: complicates
    claim: 'Semantic tokenizers and acoustic tokenizers impose an inherent trade-off: semantic tokens produce coherent
      content but poor acoustic quality, while acoustic tokens enable high-fidelity reconstruction but risk content
      inaccuracies.'
    source: §III-A, §IV-A1
    evidence: 'Semantic tokenizers and acoustic tokenizers impose an inherent trade-off: semantic tokens produce
      coherent content but poor acoustic quality, while acoustic tokens enable high-fidelity reconstruction but
      risk content inaccuracies.'
    confidence: high
    relevance: medium
  - claim_id: initialising_a_speech_lm_from_a_text_pretrained_checkpoint_accelerates
    role: supports
    claim: Initialising a speech LM from a text-pretrained checkpoint accelerates convergence and improves speech
      understanding, whereas initialisation from image-pretrained checkpoints yields worse results than random initialisation.
    source: §IV-B1
    evidence: Initialising a speech LM from a text-pretrained checkpoint accelerates convergence and improves speech
      understanding, whereas initialisation from image-pretrained checkpoints yields worse results than random initialisation.
    confidence: high
    relevance: medium
  - claim_id: interleaving_speech_and_text_tokens_during_pre_training_measurably_improves
    role: supports
    claim: Interleaving speech and text tokens during pre-training measurably improves cross-modal representation
      alignment compared to training on speech tokens alone.
    source: §IV-B1
    evidence: Interleaving speech and text tokens during pre-training measurably improves cross-modal representation
      alignment compared to training on speech tokens alone.
    confidence: high
    relevance: high
  - claim_id: post_alignment_techniques_rlhf_dpo_for_speech_lms_remain_substantially
    role: supports
    claim: Post-alignment techniques (RLHF, DPO) for speech LMs remain substantially underexplored relative to their
      established role in text LM development, leaving semantic consistency and acoustic quality gaps in deployed
      systems.
    source: §IV-B3, §VII
    evidence: Post-alignment techniques (RLHF, DPO) for speech LMs remain substantially underexplored relative to
      their established role in text LM development, leaving semantic consistency and acoustic quality gaps in deployed
      systems.
    confidence: high
    relevance: medium
  limitations:
  - The survey's arXiv version was submitted in October 2024 and the rapidly evolving SpeechLM landscape means several
    systems surveyed (notably Moshi, Mini-Omni, Llama-Omni) were still very recent preprints without peer-reviewed
    evaluation. The coverage of full-duplex systems and post-alignment techniques is acknowledged by the authors
    as incomplete.
  - 'The survey identifies several open questions: whether end-to-end joint training of all three components (tokenizer,
    LM, vocoder) outperforms separately trained pipelines; how to enable real-time speech generation with acceptable
    latency; the unique safety risks of SpeechLMs (toxicity, acoustic inappropriate content, speaker identity leakage)
    that differ from text LM safety challenges; and the potential of SpeechLMs for low-resource spoken languages
    where audio data is more available than text. The question of whether text alignment universally helps or degrades
    paralinguistic modelling (by anchoring the model too closely to textual semantics) is flagged but unresolved.'
  caveats: []
- id: '2410.11190'
  published_date: "2024-10-15"
  entry_date: '2026-07-28'
  year: 2024
  venue: arXiv
  task:
  - SCA
  architecture:
  - autoregressive-LM
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: active_evidence
  method_family:
  - autoregressive_codec_language_models
  claims:
  - claim_id: staged_adapter_training_encoder_alignment_before_language_model_fine_tuning
    role: supports
    claim: Staged adapter training (encoder alignment before language model fine-tuning) enables tri-modal extension
      of a compact language model with minimal data without catastrophic forgetting of the base model's text capabilities.
    source: §3.3
    evidence: Staged adapter training (encoder alignment before language model fine-tuning) enables tri-modal extension
      of a compact language model with minimal data without catastrophic forgetting of the base model's text capabilities.
    confidence: high
    relevance: medium
  - claim_id: using_continuous_encoder_features_whisper_rather_than_discrete_audio_tokens
    role: supports
    claim: Using continuous encoder features (Whisper) rather than discrete audio tokens for speech input yields
      more stable and semantically consistent representations, reducing ASR loss instability during training.
    source: §3.1, "Audio Encoder"
    evidence: Using continuous encoder features (Whisper) rather than discrete audio tokens for speech input yields
      more stable and semantically consistent representations, reducing ASR loss instability during training.
    confidence: high
    relevance: high
  - claim_id: adding_a_third_modality_vision_to_an_audio_text_spoken
    role: supports
    claim: Adding a third modality (vision) to an audio-text spoken conversational agent modestly degrades ASR performance,
      likely due to diluted training data proportion rather than architectural interference.
    source: §4.4, Table 2
    evidence: Adding a third modality (vision) to an audio-text spoken conversational agent modestly degrades ASR
      performance, likely due to diluted training data proportion rather than architectural interference.
    confidence: high
    relevance: medium
  - claim_id: command_based_semantic_interruption_intent_token_classification_provides_a_viable
    role: supports
    claim: Command-based semantic interruption (intent token classification) provides a viable alternative to VAD-based
      full-duplex detection, with the advantage of robustness to noise and unrelated background sounds.
    source: §3.4
    evidence: Command-based semantic interruption (intent token classification) provides a viable alternative to
      VAD-based full-duplex detection, with the advantage of robustness to noise and unrelated background sounds.
    confidence: high
    relevance: medium
  limitations:
  - 'Evaluation coverage is incomplete: no naturalness MOS, SMOS, or intelligibility metrics for speech output are
    reported in this version, and vision benchmark results are explicitly deferred. Claims about speech quality
    and vision understanding capability cannot be independently verified from this paper alone.'
  - The interruption mechanism is demonstrated on a single synthesised phrase ("Stop Omni") with a narrow distribution
    of noise conditions. Whether the approach generalises to arbitrary semantic interrupt commands or real conversational
    interruption patterns is an open question. The model is trained and evaluated exclusively on English data despite
    using multilingual Whisper; cross-lingual transfer to speech output is not evaluated. The 0.5B model scale is
    deliberately small, and the authors note that scaling data and compute would likely yield substantial capability
    gains, but this is not demonstrated. Synthetic data is used for spoken question-answering and interruption training,
    and the effect of synthetic-to-real domain mismatch on deployment robustness is not assessed.
  caveats: []
- id: '2410.17799'
  published_date: "2024-10-23"
  entry_date: '2026-07-28'
  year: 2024
  venue: arXiv
  task:
  - SCA
  architecture:
  - autoregressive-LM
  relevance: medium
  evidence_role:
  - architecture_variant
  - acceleration_evidence
  current_role: active_evidence
  method_family:
  - autoregressive_codec_language_models
  claims:
  - claim_id: full_duplex_spoken_dialogue_can_be_achieved_by_flattening_interleaved
    role: supports
    claim: Full-duplex spoken dialogue can be achieved by flattening interleaved speech and text token streams into
      a single autoregressive sequence, without modifying the backbone LLM architecture.
    source: §3, §3.3.2
    evidence: Full-duplex spoken dialogue can be achieved by flattening interleaved speech and text token streams
      into a single autoregressive sequence, without modifying the backbone LLM architecture.
    confidence: high
    relevance: low
  - claim_id: progressive_curriculum_training_modality_alignment_followed_by_half_duplex_then
    role: supports
    claim: Progressive curriculum training (modality alignment followed by half-duplex, then full-duplex) improves
      final full-duplex dialogue quality compared to training directly on full-duplex data.
    source: §4.3, Table 3
    evidence: Progressive curriculum training (modality alignment followed by half-duplex, then full-duplex) improves
      final full-duplex dialogue quality compared to training directly on full-duplex data.
    confidence: high
    relevance: low
  - claim_id: eliminating_intermediate_text_output_from_dialogue_models_substantially_reduces_response
    role: complicates
    claim: Eliminating intermediate text output from dialogue models substantially reduces response latency but
      causes a significant drop in semantic coherence, indicating a fundamental trade-off between speed and content
      quality in speech-to-speech generation.
    source: §3.3.2, §4.3, Table 3
    evidence: Eliminating intermediate text output from dialogue models substantially reduces response latency but
      causes a significant drop in semantic coherence, indicating a fundamental trade-off between speed and content
      quality in speech-to-speech generation.
    confidence: high
    relevance: low
  - claim_id: turn_taking_response_latency_in_full_duplex_speech_models_can
    role: supports
    claim: Turn-taking response latency in full-duplex speech models can be reduced by chunked interleaved sequence
      training, with practical response times under 200 ms achievable at 0.5B parameter scale.
    source: §4.3, Table 4
    evidence: Turn-taking response latency in full-duplex speech models can be reduced by chunked interleaved sequence
      training, with practical response times under 200 ms achievable at 0.5B parameter scale.
    confidence: high
    relevance: low
  limitations:
  - The 0.5B backbone is substantially smaller than comparators (LLaMA-Omni 8B, GLM-Voice 9B, Moshi 7B), making
    LLM-score comparisons in Table 3 not directly attributable to the method alone. The paper acknowledges GLM-Voice
    results may reflect test-set leakage. Dialogue quality scores remain well below the ground-truth ceiling.
  - Training data is entirely synthesised from text dialogues via a TTS pipeline; real conversational dynamics (natural
    prosody, disfluencies, real interruption patterns) are not represented. The model does not handle backchannels
    from either speaker, a basic feature of natural human conversation. User turn-taking accuracy at 25 tokens remains
    below 55% for both models evaluated, leaving interruption handling far from reliable. The paper does not report
    naturalness MOS, making direct quality comparison to TTS-oriented systems difficult. All evaluation uses simulated
    test data matching the training distribution, raising questions about real-world robustness.
  caveats: []
- id: '2411.00774'
  published_date: "2024-11-01"
  entry_date: '2026-07-28'
  year: 2024
  venue: arXiv
  task:
  - SCA
  architecture:
  - autoregressive-LM
  - hybrid
  relevance: medium
  evidence_role:
  - architecture_variant
  current_role: influential
  method_family:
  - autoregressive_codec_language_models
  - hybrid_semantic_acoustic_tokenizers
  claims:
  - claim_id: freezing_the_llm_backbone_during_speech_modality_alignment_reduces_the
    role: supports
    claim: Freezing the LLM backbone during speech-modality alignment reduces the intelligence gap between spoken
      and text question-answering performance compared to fine-tuned approaches.
    source: §3.4, Table 3
    evidence: Freezing the LLM backbone during speech-modality alignment reduces the intelligence gap between spoken
      and text question-answering performance compared to fine-tuned approaches.
    confidence: high
    relevance: medium
  - claim_id: a_three_stage_training_curriculum_using_large_asr_corpora_for
    role: supports
    claim: A three-stage training curriculum, using large ASR corpora for encoder pretraining followed by small-scale
      multi-modal Q&A fine-tuning, is sufficient to achieve competitive spoken dialogue quality without updating
      backbone LLM parameters.
    source: §2.2.2, §2.3.2
    evidence: A three-stage training curriculum, using large ASR corpora for encoder pretraining followed by small-scale
      multi-modal Q&A fine-tuning, is sufficient to achieve competitive spoken dialogue quality without updating
      backbone LLM parameters.
    confidence: high
    relevance: low
  - claim_id: chunk_level_state_classification_integrated_into_the_llm_s_prefill
    role: supports
    claim: Chunk-level state classification integrated into the LLM's prefill stage enables duplex interruption
      detection without requiring a separate monitoring model or additional LLM context.
    source: §2.4
    evidence: Chunk-level state classification integrated into the LLM's prefill stage enables duplex interruption
      detection without requiring a separate monitoring model or additional LLM context.
    confidence: high
    relevance: medium
  - claim_id: decoupling_encoder_and_llm_kv_cache_per_user_session_allows
    role: supports
    claim: Decoupling encoder and LLM KV-cache per user session allows a server-side pool of model replicas to handle
      concurrent users with chunk-granular scheduling.
    source: §2.4
    evidence: Decoupling encoder and LLM KV-cache per user session allows a server-side pool of model replicas to
      handle concurrent users with chunk-granular scheduling.
    confidence: high
    relevance: medium
  limitations:
  - The spoken Q&A benchmarks used for intelligence comparison (Web Questions, LlaMA Questions, Trivia QA) were
    synthesised from text using edge-tts rather than collected from real speakers. Results on naturally spoken or
    noisy input are not reported, limiting generalisability claims about real-world speech understanding.
  - Speech output quality is evaluated primarily through CER on 1000 utterances using a single speaker, not through
    subjective MOS ratings or speaker naturalness benchmarks. It is therefore difficult to assess voice quality
    relative to other systems. The system supports a limited number of output speakers and does not yet support
    style or voice instruct-following, which the authors flag as future work. Emotion understanding and audio captioning
    are also deferred to a planned encoder upgrade. The duplex state classifier operates at chunk boundaries (approximately
    160-320 ms non-statistical latency), which may be perceptible in fast-paced dialogue.
  caveats: []
- id: '2411.01156'
  published_date: "2024-11-02"
  entry_date: '2026-07-28'
  year: 2024
  venue: arXiv
  task:
  - TTS
  architecture:
  - autoregressive-LM
  - GAN
  relevance: medium
  evidence_role:
  - architecture_variant
  - acceleration_evidence
  current_role: active_evidence
  method_family:
  - autoregressive_codec_language_models
  - gan_neural_codecs_and_decoders
  claims:
  - claim_id: eliminating_grapheme_to_phoneme_conversion_by_directly_feeding_raw_text
    role: supports
    claim: Eliminating grapheme-to-phoneme conversion by directly feeding raw text to an LLM backbone is viable
      for multilingual TTS and can improve handling of context-dependent polyphonic words.
    source: §1, §3
    evidence: Eliminating grapheme-to-phoneme conversion by directly feeding raw text to an LLM backbone is viable
      for multilingual TTS and can improve handling of context-dependent polyphonic words.
    confidence: high
    relevance: medium
  - claim_id: hierarchical_decomposition_of_autoregressive_token_generation_into_semantic_level_and
    role: supports
    claim: Hierarchical decomposition of autoregressive token generation into semantic-level and acoustic-level
      stages improves codebook stability in grouped scalar quantization.
    source: §3.1, §3.1.1
    evidence: Hierarchical decomposition of autoregressive token generation into semantic-level and acoustic-level
      stages improves codebook stability in grouped scalar quantization.
    confidence: high
    relevance: high
  - claim_id: grouped_finite_scalar_vector_quantization_achieves_higher_codebook_utilisation_than
    role: supports
    claim: Grouped Finite Scalar Vector Quantization achieves higher codebook utilisation than residual vector quantization
      alternatives, mitigating dead-code collapse.
    source: §3.2.2, §3.2.3
    evidence: Grouped Finite Scalar Vector Quantization achieves higher codebook utilisation than residual vector
      quantization alternatives, mitigating dead-code collapse.
    confidence: high
    relevance: high
  - claim_id: real_time_tts_inference_with_low_first_packet_latency_is
    role: supports
    claim: Real-time TTS inference with low first-packet latency is achievable on consumer GPU hardware through
      standard inference optimisations without architectural compromise.
    source: §4.2
    evidence: Real-time TTS inference with low first-packet latency is achievable on consumer GPU hardware through
      standard inference optimisations without architectural compromise.
    confidence: high
    relevance: low
  limitations:
  - The entire experimental evaluation is conducted on a proprietary test set with undisclosed composition and size.
    No public benchmark is used, making it impossible to independently verify the claimed superiority over CosyVoice
    and F5-TTS or to compare against the broader literature.
  - The MOS evaluation uses "naive listeners" rather than trained raters or crowdsourced panels following standard
    listening test protocols (e.g. ITU-T P.800), which may inflate scores relative to conventional evaluations.
    The paper does not report model size, training compute, or inference memory requirements in full, limiting reproducibility.
    DPO training details are omitted from the main training description. The paper does not evaluate cross-lingual
    transfer or accent preservation, which are claimed motivations for the non-G2P design. It is also unclear how
    the system handles low-resource languages beyond the eight listed in the training data.
  caveats: []
- id: '2411.13577'
  published_date: "2024-11-15"
  entry_date: '2026-07-28'
  year: 2024
  venue: arXiv
  task:
  - SCA
  architecture:
  - autoregressive-LM
  - transformer-enc-dec
  - hybrid
  relevance: medium
  evidence_role:
  - architecture_variant
  - acceleration_evidence
  - evaluation_caution
  - infrastructure
  current_role: influential
  method_family:
  - autoregressive_codec_language_models
  - transformer_encoder_decoder_tokenizers
  - hybrid_semantic_acoustic_tokenizers
  claims:
  - claim_id: the_cascaded_and_end_to_end_paradigms_for_spoken_dialogue
    role: supports
    claim: The cascaded and end-to-end paradigms for spoken dialogue impose fundamentally different trade-offs between
      latency, paralinguistic fidelity, and intelligibility, and the appropriate choice depends on the target interaction
      scenario.
    source: §2.2, §2.3
    evidence: The cascaded and end-to-end paradigms for spoken dialogue impose fundamentally different trade-offs
      between latency, paralinguistic fidelity, and intelligibility, and the appropriate choice depends on the target
      interaction scenario.
    confidence: high
    relevance: low
  - claim_id: semantic_speech_representations_offer_higher_compression_rates_and_better_llm
    role: complicates
    claim: Semantic speech representations offer higher compression rates and better LLM compatibility than acoustic
      representations, but sacrifice expressiveness, timbre, and style fidelity, necessitating additional vocoders
      in pipeline-based generation.
    source: §3.3.1, Table 1
    evidence: Semantic speech representations offer higher compression rates and better LLM compatibility than acoustic
      representations, but sacrifice expressiveness, timbre, and style fidelity, necessitating additional vocoders
      in pipeline-based generation.
    confidence: high
    relevance: high
  - claim_id: achieving_genuine_full_duplex_spoken_dialogue_simultaneous_listening_and_speaking
    role: supports
    claim: Achieving genuine full-duplex spoken dialogue (simultaneous listening and speaking with interrupt handling)
      requires architecturally causal models throughout the full pipeline, a constraint that current systems largely
      satisfy only on the output side.
    source: §5.1, §5.2
    evidence: Achieving genuine full-duplex spoken dialogue (simultaneous listening and speaking with interrupt
      handling) requires architecturally causal models throughout the full pipeline, a constraint that current systems
      largely satisfy only on the output side.
    confidence: high
    relevance: low
  - claim_id: speech_text_modality_alignment_in_current_spoken_dialogue_systems_relies
    role: supports
    claim: Speech-text modality alignment in current spoken dialogue systems relies heavily on paired data, introducing
      catastrophic forgetting risk and creating a structural dependency on the availability of labelled speech corpora.
    source: §4.4.1
    evidence: Speech-text modality alignment in current spoken dialogue systems relies heavily on paired data, introducing
      catastrophic forgetting risk and creating a structural dependency on the availability of labelled speech corpora.
    confidence: high
    relevance: low
  - claim_id: evaluation_infrastructure_for_spoken_dialogue_lags_substantially_behind_system_capabilities
    role: supports
    claim: 'Evaluation infrastructure for spoken dialogue lags substantially behind system capabilities: interaction,
      streaming latency, and audio generation are either absent from or severely underrepresented in existing benchmarks.'
    source: §6.2, §6.3, Table 3
    evidence: 'Evaluation infrastructure for spoken dialogue lags substantially behind system capabilities: interaction,
      streaming latency, and audio generation are either absent from or severely underrepresented in existing benchmarks.'
    confidence: high
    relevance: low
  limitations:
  - The survey was produced concurrently with the primary wave of open-source spoken dialogue models it covers (late
    2024), meaning that some systems are described in early form and the field will have evolved by the time readers
    encounter the paper. The coverage of music and sound understanding and generation within dialogue systems is
    acknowledged as thin (the authors defer to an appendix), and security evaluation for spoken dialogue receives
    less treatment than its importance warrants. A number of prominent systems (Westlake-Omni, Hertz-dev, SpeechGPT2,
    Fish-Agent) lack published papers and are excluded from the timeline figure, which may leave gaps for practitioners
    interested in the deployed-systems landscape. The SuperCLUE benchmark, one of the more comprehensive interaction
    evaluations listed, is not open-source and focuses on Mandarin, limiting its utility for the broader research
    community.
  - 'Open questions surfaced include: whether speech tokenisers can be designed to enforce text-space alignment
    during encoding, eliminating the need for large paired corpora; what granularity of temporal alignment priors
    (sentence, word, phoneme level) is optimal for spoken dialogue training; and how preference optimisation techniques
    can be adapted for the joint text-speech output space.'
  caveats: []
- id: '2411.17607'
  published_date: "2024-11-26"
  entry_date: '2026-07-28'
  year: 2024
  venue: arXiv
  task:
  - SCA
  architecture:
  - autoregressive-LM
  - flow-matching
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: influential
  method_family:
  - autoregressive_codec_language_models
  - flow_matching_codec_decoders
  claims:
  - claim_id: synthetic_speech_text_interleaved_data_generated_by_converting_text_spans
    role: supports
    claim: Synthetic speech-text interleaved data, generated by converting text spans to speech tokens using a learned
      text-to-token model, enables effective cross-modal knowledge transfer from pre-trained LLMs to the speech
      domain.
    source: §2.2, §3.3.1, Table 5
    evidence: Synthetic speech-text interleaved data, generated by converting text spans to speech tokens using
      a learned text-to-token model, enables effective cross-modal knowledge transfer from pre-trained LLMs to the
      speech domain.
    confidence: high
    relevance: medium
  - claim_id: lower_speech_tokenizer_frame_rates_improve_speech_language_modelling_performance
    role: supports
    claim: Lower speech tokenizer frame rates improve speech language modelling performance within a fixed token
      budget, with gains plateauing around 12.5 Hz where information loss begins to outweigh efficiency benefits.
    source: §3.3.2, Figure 3a
    evidence: Lower speech tokenizer frame rates improve speech language modelling performance within a fixed token
      budget, with gains plateauing around 12.5 Hz where information loss begins to outweigh efficiency benefits.
    confidence: high
    relevance: medium
  - claim_id: supervised_speech_tokenization_derived_from_asr_model_fine_tuning_achieves
    role: supports
    claim: Supervised speech tokenization derived from ASR model fine-tuning achieves stronger semantic preservation
      at low frame rates than unsupervised codec-based tokenisers, while maintaining competitive speech reconstruction
      quality.
    source: §2.1, Table 1
    evidence: Supervised speech tokenization derived from ASR model fine-tuning achieves stronger semantic preservation
      at low frame rates than unsupervised codec-based tokenisers, while maintaining competitive speech reconstruction
      quality.
    confidence: high
    relevance: high
  - claim_id: speech_text_pre_training_with_interleaved_data_substantially_narrows_the
    role: supports
    claim: Speech-text pre-training with interleaved data substantially narrows the gap between speech-only and
      speech-to-text performance on spoken question answering, suggesting that cross-modal alignment transfers factual
      knowledge from text representations to speech decoding.
    source: §3.2, Table 3
    evidence: Speech-text pre-training with interleaved data substantially narrows the gap between speech-only and
      speech-to-text performance on spoken question answering, suggesting that cross-modal alignment transfers factual
      knowledge from text representations to speech decoding.
    confidence: high
    relevance: medium
  - claim_id: an_intermediate_text_response_in_speech_generation_text_guided_mode
    role: supports
    claim: An intermediate text response in speech generation (text-guided mode) provides meaningful accuracy gains
      for knowledge-intensive tasks, but a well-pre-trained model operating purely in the speech domain can still
      match text-guided baselines from prior work.
    source: §3.2, Table 4
    evidence: An intermediate text response in speech generation (text-guided mode) provides meaningful accuracy
      gains for knowledge-intensive tasks, but a well-pre-trained model operating purely in the speech domain can
      still match text-guided baselines from prior work.
    confidence: high
    relevance: medium
  limitations:
  - No code or model weights are publicly released with this paper, and the proprietary Chinese ASR dataset (10k
    hours) used in tokenizer training cannot be replicated by third parties, limiting reproducibility of the full
    pipeline.
  - The fine-tuning dataset (SpeechDialog-90K) is synthesised using MeloTTS for speech responses, so the chatbot's
    output speech may inherit MeloTTS quality characteristics rather than reflecting the model's own generative
    capacity. The evaluation of spoken chatbots relies on GPT-4 scoring, which is a reasonable but proxy measure
    that may not correlate perfectly with human judgements. The paper does not evaluate multilingual spoken QA or
    chatbot performance despite training on English and Chinese data. Full-duplex conversation capability, demonstrated
    by Moshi, is not explored. Scaling beyond 9B parameters is left as future work.
  caveats: []
- id: '2411.18803'
  published_date: "2024-11-27"
  entry_date: '2026-07-28'
  year: 2024
  venue: arXiv
  task:
  - codec
  architecture:
  - GAN
  relevance: medium
  evidence_role:
  - architecture_variant
  current_role: influential
  method_family:
  - gan_neural_codecs_and_decoders
  claims:
  - claim_id: transformer_architectures_can_match_convolutional_neural_audio_codecs_in_streaming
    role: supports
    claim: Transformer architectures can match convolutional neural audio codecs in streaming reconstruction quality
      while requiring substantially lower multiply-accumulate operations at similar parameter counts.
    source: §3.2, §5.1, Table 3
    evidence: Transformer architectures can match convolutional neural audio codecs in streaming reconstruction
      quality while requiring substantially lower multiply-accumulate operations at similar parameter counts.
    confidence: high
    relevance: low
  - claim_id: a_single_codebook_design_is_compatible_with_full_streaming_operation
    role: supports
    claim: A single-codebook design is compatible with full streaming operation, showing that the prior tradeoff
      between single-codebook simplicity and streaming capability is not fundamental.
    source: §2.3, §3.1
    evidence: A single-codebook design is compatible with full streaming operation, showing that the prior tradeoff
      between single-codebook simplicity and streaming capability is not fundamental.
    confidence: high
    relevance: high
  - claim_id: semantic_distillation_as_used_in_mimi_and_speechtokenizer_provides_consistent
    role: supports
    claim: Semantic distillation (as used in Mimi and SpeechTokenizer) provides consistent word error rate benefits
      over codecs trained without it, even when those codecs achieve higher perceptual quality scores.
    source: §5.1, §5.2, Table 3, Table 4
    evidence: Semantic distillation (as used in Mimi and SpeechTokenizer) provides consistent word error rate benefits
      over codecs trained without it, even when those codecs achieve higher perceptual quality scores.
    confidence: high
    relevance: medium
  - claim_id: at_equivalent_computational_budgets_transformer_based_codec_architectures_outperform_their
    role: supports
    claim: At equivalent computational budgets, transformer-based codec architectures outperform their causal convolutional
      counterparts across intelligibility, distortion, and naturalness metrics.
    source: §5.1, Figure 2, Figure 3
    evidence: At equivalent computational budgets, transformer-based codec architectures outperform their causal
      convolutional counterparts across intelligibility, distortion, and naturalness metrics.
    confidence: high
    relevance: high
  limitations:
  - All naturalness evaluations rely on UTMOS rather than human MOS. While UTMOS correlates well with human judgements
    on codec-reconstructed speech, the paper presents no subjective listening test to confirm its quality claims.
  - The codec is evaluated only on English speech (LibriSpeech). Generalisation to other languages, accents, and
    non-speech audio is untested. The training set (Libri-light) is entirely read speech; performance on conversational,
    emotional, or noisy speech is unknown. Code and checkpoints are not released (as of the preprint), limiting
    reproducibility. The paper does not evaluate latency (time-to-first-byte or algorithmic delay) in a real streaming
    deployment, only computational complexity in offline MACs. The effect of the large codebook sizes (65K, 131K)
    on downstream speech language model training and inference has not been demonstrated.
  caveats: []
- id: '2411.19842'
  published_date: "2024-11-29"
  entry_date: '2026-07-28'
  year: 2024
  venue: arXiv
  task:
  - codec
  architecture:
  - transformer-enc-dec
  - VAE
  relevance: medium
  evidence_role:
  - architecture_variant
  - evaluation_caution
  current_role: influential
  method_family:
  - transformer_encoder_decoder_tokenizers
  - vae_vector_quantized_codecs
  claims:
  - claim_id: scaling_transformer_architecture_parameter_count_in_neural_audio_codecs_produces
    role: supports
    claim: Scaling transformer architecture parameter count in neural audio codecs produces consistent quality improvements
      across objective and subjective metrics.
    source: §4.6, Table 4
    evidence: Scaling transformer architecture parameter count in neural audio codecs produces consistent quality
      improvements across objective and subjective metrics.
    confidence: high
    relevance: medium
  - claim_id: finite_scalar_quantization_achieves_near_perfect_codebook_utilization_without_explicit
    role: supports
    claim: Finite scalar quantization achieves near-perfect codebook utilization without explicit utilization regularization,
      simplifying downstream generative modeling compared to RVQ.
    source: §3.2, §A.8, Table 9
    evidence: Finite scalar quantization achieves near-perfect codebook utilization without explicit utilization
      regularization, simplifying downstream generative modeling compared to RVQ.
    confidence: high
    relevance: high
  - claim_id: a_neural_codec_trained_exclusively_on_english_speech_can_generalize
    role: supports
    claim: A neural codec trained exclusively on English speech can generalize effectively to unseen languages,
      outperforming multilingual-trained baselines of similar scale on most objective metrics.
    source: §A.5, Table 7
    evidence: A neural codec trained exclusively on English speech can generalize effectively to unseen languages,
      outperforming multilingual-trained baselines of similar scale on most objective metrics.
    confidence: high
    relevance: high
  - claim_id: perceptual_losses_derived_from_self_supervised_speech_models_wavlm_large
    role: supports
    claim: Perceptual losses derived from self-supervised speech models (WavLM-Large features) are critical for
      achieving intelligible reconstruction at low bitrates, beyond what adversarial and spectral reconstruction
      losses alone provide.
    source: §3.4, §A.1, Table 3
    evidence: Perceptual losses derived from self-supervised speech models (WavLM-Large features) are critical for
      achieving intelligible reconstruction at low bitrates, beyond what adversarial and spectral reconstruction
      losses alone provide.
    confidence: high
    relevance: medium
  - claim_id: systematic_spectral_bias_in_multi_resolution_stft_discriminators_arising_from
    role: supports
    claim: Systematic spectral bias in multi-resolution STFT discriminators, arising from power-of-two FFT configurations,
      causes periodic reconstruction artifacts that disproportionately affect large-capacity codec architectures.
    source: §3.3, §B.5
    evidence: Systematic spectral bias in multi-resolution STFT discriminators, arising from power-of-two FFT configurations,
      causes periodic reconstruction artifacts that disproportionately affect large-capacity codec architectures.
    confidence: high
    relevance: high
  limitations:
  - Training data is 16 kHz English audiobook speech only (105k hours). Multilingual generalization results are
    promising but the model was not trained or optimized for non-English data; claims about multilingual capability
    should be interpreted cautiously relative to models with dedicated multilingual training at scale.
  - The model has not been evaluated on noisy speech, overlapping speakers, or environmental audio, which are common
    real-world conditions. The large parameter count (950M) requires substantially more compute than lighter baselines
    (DAC at 76M, Mimi at 80M); while the RTF is acceptable on H100 GPUs for longer utterances, latency for short
    clips is roughly 3x that of smaller models, which matters for streaming applications. The post-hoc Residual
    FSQ decomposition is restricted to specific level configurations (L = 2^n + 1); arbitrary bitrate targets are
    not directly achievable without retraining. The systematic bias analysis in the discriminator (§B.5) raises
    open questions about whether similar biases appear in other convolutional discriminator architectures and how
    to address them in the general case.
  caveats: []
- id: '2412.02612'
  published_date: "2024-12-03"
  entry_date: '2026-07-28'
  year: 2024
  venue: arXiv
  task:
  - TTS
  - SCA
  architecture:
  - autoregressive-LM
  - flow-matching
  relevance: medium
  evidence_role:
  - architecture_variant
  - acceleration_evidence
  - control_evidence
  current_role: influential
  method_family:
  - autoregressive_codec_language_models
  - flow_matching_codec_decoders
  claims:
  - claim_id: speech_text_interleaved_pre_training_at_trillion_token_scale_enables
    role: supports
    claim: Speech-text interleaved pre-training at trillion-token scale enables an LLM to acquire speech understanding
      and generation capabilities that substantially close the gap between spoken and textual reasoning quality.
    source: §4.1, Table 4
    evidence: Speech-text interleaved pre-training at trillion-token scale enables an LLM to acquire speech understanding
      and generation capabilities that substantially close the gap between spoken and textual reasoning quality.
    confidence: high
    relevance: medium
  - claim_id: single_codebook_supervised_speech_tokenizers_derived_from_asr_models_can
    role: supports
    claim: Single-codebook supervised speech tokenizers derived from ASR models can achieve compact bitrates (below
      200bps) while retaining sufficient semantic fidelity for both downstream language modeling and speech synthesis.
    source: §3.1, Table 1
    evidence: Single-codebook supervised speech tokenizers derived from ASR models can achieve compact bitrates
      (below 200bps) while retaining sufficient semantic fidelity for both downstream language modeling and speech
      synthesis.
    confidence: high
    relevance: high
  - claim_id: streaming_interleaved_generation_templates_alternating_text_and_speech_token_output
    role: supports
    claim: Streaming interleaved generation templates, alternating text and speech token output, enable low-latency
      spoken responses without sacrificing content coherence by ensuring text generation consistently precedes its
      corresponding speech.
    source: §3.3
    evidence: Streaming interleaved generation templates, alternating text and speech token output, enable low-latency
      spoken responses without sacrificing content coherence by ensuring text generation consistently precedes its
      corresponding speech.
    confidence: high
    relevance: low
  - claim_id: end_to_end_speech_language_models_that_include_dedicated_speech
    role: supports
    claim: End-to-end speech language models that include dedicated speech pre-training produce measurably higher-quality
      and more stylistically controllable speech responses than LLMs fine-tuned solely on speech question-answering
      data.
    source: §5.2, Table 6
    evidence: End-to-end speech language models that include dedicated speech pre-training produce measurably higher-quality
      and more stylistically controllable speech responses than LLMs fine-tuned solely on speech question-answering
      data.
    confidence: high
    relevance: medium
  - claim_id: decoupling_the_text_and_speech_output_subtasks_during_fine_tuning
    role: supports
    claim: Decoupling the text and speech output subtasks during fine-tuning, through separate loss-masking passes
      at different epoch rates, addresses the discrepancy in learning dynamics between the two modalities.
    source: §4.2.2
    evidence: Decoupling the text and speech output subtasks during fine-tuning, through separate loss-masking passes
      at different epoch rates, addresses the discrepancy in learning dynamics between the two modalities.
    confidence: high
    relevance: medium
  limitations:
  - The paper reports no subjective listening test (MOS/MUSHRA) on the chat model output; UTMOS is used as a proxy
    for speech naturalness, and the chat evaluation relies on GPT-4o scoring of ASR transcriptions, introducing
    cascaded error from both the vocoder quality and the Whisper transcription step.
  - 'The 175bps tokenizer trades acoustic fidelity for compactness: VisQOL at 12.5Hz (2.52) is lower than SpeechTokenizer
    variants and the 50Hz variant of the same system. This may limit voice cloning quality and the fidelity of paralinguistic
    feature reproduction (accent, fine-grained emotion), though the paper does not directly evaluate these.'
  - Instruction-following for speech style (emotion, dialect, rate) is described and demonstrated qualitatively
    but not evaluated quantitatively; it is unclear how reliably the model follows complex or combined style instructions.
  - 'The streaming thoughts ratio (13 text : 26 speech tokens) and block size (b=0.8s) are empirically chosen hyperparameters;
    their sensitivity and generalisability to other tokenizer frame rates or LLM sizes is not studied.'
  caveats: []
- id: '2412.10117'
  published_date: "2024-12-13"
  entry_date: '2026-07-28'
  year: 2024
  venue: arXiv
  task:
  - TTS
  architecture:
  - autoregressive-LM
  - flow-matching
  - hybrid
  relevance: medium
  evidence_role:
  - architecture_variant
  - acceleration_evidence
  - control_evidence
  current_role: influential
  method_family:
  - autoregressive_codec_language_models
  - flow_matching_codec_decoders
  - hybrid_semantic_acoustic_tokenizers
  claims:
  - claim_id: finite_scalar_quantization_achieves_full_codebook_utilization_in_supervised_speech
    role: supports
    claim: Finite scalar quantization achieves full codebook utilization in supervised speech tokenizers, capturing
      substantially more semantic content than vector quantization at equivalent bitrates.
    source: §2.2, §4.1, Table 4
    evidence: Finite scalar quantization achieves full codebook utilization in supervised speech tokenizers, capturing
      substantially more semantic content than vector quantization at equivalent bitrates.
    confidence: high
    relevance: high
  - claim_id: replacing_a_randomly_initialized_custom_lm_with_a_pre_trained
    role: supports
    claim: Replacing a randomly initialized custom LM with a pre-trained LLM backbone improves content consistency
      in hybrid TTS systems without requiring a separate text encoder.
    source: §2.3, §4.3, Table 7
    evidence: Replacing a randomly initialized custom LM with a pre-trained LLM backbone improves content consistency
      in hybrid TTS systems without requiring a separate text encoder.
    confidence: high
    relevance: medium
  - claim_id: streaming_and_non_streaming_synthesis_can_be_unified_in_a
    role: supports
    claim: Streaming and non-streaming synthesis can be unified in a single autoregressive model through interleaved
      text-speech token sequences, with virtually lossless quality on typical inputs relative to offline mode.
    source: §2.3, §4.2, Table 8
    evidence: Streaming and non-streaming synthesis can be unified in a single autoregressive model through interleaved
      text-speech token sequences, with virtually lossless quality on typical inputs relative to offline mode.
    confidence: high
    relevance: low
  - claim_id: training_a_flow_matching_model_simultaneously_on_multiple_causal_mask
    role: complicates
    claim: Training a flow matching model simultaneously on multiple causal mask types — from non-causal to full-causal
      — enables a single model to span the latency-quality trade-off continuum at inference time, with masks providing
      implicit self-distillation.
    source: §2.4, §4.3, Table 8
    evidence: Training a flow matching model simultaneously on multiple causal mask types — from non-causal to full-causal
      — enables a single model to span the latency-quality trade-off continuum at inference time, with masks providing
      implicit self-distillation.
    confidence: high
    relevance: low
  - claim_id: differentiable_asr_reward_optimization_generalizes_better_to_out_of_domain
    role: supports
    claim: Differentiable ASR reward optimization generalizes better to out-of-domain and hard-case inputs than
      preference-based DPO in TTS speaker fine-tuning.
    source: §2.8, §4.7, Table 11
    evidence: Differentiable ASR reward optimization generalizes better to out-of-domain and hard-case inputs than
      preference-based DPO in TTS speaker fine-tuning.
    confidence: high
    relevance: medium
  limitations:
  - '- EN quality still lags CosyVoice 2 behind Seed-TTS and F5-TTS on SEED test-en (WER 2.57% vs. 2.25% and 1.83%),
    reflecting data imbalance toward Chinese. - Japanese synthesis degrades due to character set overlap with Chinese
    (CER 18.79% test-ja vs. 7.98% test-ko). - Cannot control timbre through text instructions. - Singing not supported.
    - Streaming still incurs a hard degradation on test-hard, suggesting that contextual information from future
    text is important for difficult patterns.'
  caveats: []
- id: '2412.15649'
  published_date: "2024-12-20"
  entry_date: '2026-07-28'
  year: 2024
  venue: arXiv
  task:
  - SCA
  architecture:
  - autoregressive-LM
  - flow-matching
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: active_evidence
  method_family:
  - autoregressive_codec_language_models
  - flow_matching_codec_decoders
  claims:
  - claim_id: decoupling_speaker_identity_from_semantic_content_in_spoken_dialogue_systems
    role: supports
    claim: Decoupling speaker identity from semantic content in spoken dialogue systems enables zero-shot timbre
      control without modifying the language model or adding speaker-conditioning layers.
    source: §3.4
    evidence: Decoupling speaker identity from semantic content in spoken dialogue systems enables zero-shot timbre
      control without modifying the language model or adding speaker-conditioning layers.
    confidence: high
    relevance: low
  - claim_id: grouping_audio_tokens_to_reduce_the_frequency_mismatch_between_text
    role: supports
    claim: Grouping audio tokens to reduce the frequency mismatch between text and speech representations substantially
      improves speech-text alignment and training efficiency in parallel audio-text dialogue models.
    source: §3.3, Table 5
    evidence: Grouping audio tokens to reduce the frequency mismatch between text and speech representations substantially
      improves speech-text alignment and training efficiency in parallel audio-text dialogue models.
    confidence: high
    relevance: low
  - claim_id: single_stage_fine_tuning_on_dialogue_data_can_match_or
    role: supports
    claim: Single-stage fine-tuning on dialogue data can match or outperform multi-stage pipelines that include
      ASR or TTS pre-training, because modality-specific pre-training degrades instruction-following and general
      knowledge retention.
    source: §5.3.2, Table 6
    evidence: Single-stage fine-tuning on dialogue data can match or outperform multi-stage pipelines that include
      ASR or TTS pre-training, because modality-specific pre-training degrades instruction-following and general
      knowledge retention.
    confidence: high
    relevance: low
  - claim_id: replacing_audio_token_history_with_text_only_history_in_multi
    role: supports
    claim: Replacing audio-token history with text-only history in multi-turn spoken dialogue models improves the
      system's ability to handle longer conversation contexts while leveraging pre-trained LLM in-context learning.
    source: §3.5
    evidence: Replacing audio-token history with text-only history in multi-turn spoken dialogue models improves
      the system's ability to handle longer conversation contexts while leveraging pre-trained LLM in-context learning.
    confidence: high
    relevance: low
  - claim_id: current_spoken_dialogue_models_consistently_underperform_text_only_llms_of
    role: supports
    claim: Current spoken dialogue models consistently underperform text-only LLMs of similar scale on semantic
      content quality, even after dialogue fine-tuning.
    source: §5.1, Table 3
    evidence: Current spoken dialogue models consistently underperform text-only LLMs of similar scale on semantic
      content quality, even after dialogue fine-tuning.
    confidence: high
    relevance: low
  limitations:
  - Historical text prompting discards all non-verbal information from prior dialogue turns (prosody, emotion, paralinguistic
    cues). In scenarios where voice-level context matters for dialogue coherence, this strategy may reduce response
    quality in ways not captured by the text-based evaluation metrics used.
  - The system is evaluated exclusively at 0.5B scale. The single-stage training advantage may not hold for larger
    LLMs, where the data volume required for joint audio-text modeling grows substantially.
  - Training data relies entirely on TTS-synthesized dialogue utterances rather than real recorded speech. Whether
    the system's timbre control and speech quality generalise to natural in-the-wild audio inputs at inference time
    is not tested.
  - The evaluation benchmark, while more comprehensive than prior SDM evaluations, is still largely built on synthetic
    TTS-generated test inputs. Evaluation on real end-to-end voice interactions would strengthen the claims.
  caveats: []
- id: '2502.04128'
  published_date: "2025-02-06"
  entry_date: '2026-07-28'
  year: 2025
  venue: arXiv
  task:
  - TTS
  architecture:
  - autoregressive-LM
  relevance: medium
  evidence_role:
  - architecture_variant
  - acceleration_evidence
  - control_evidence
  current_role: influential
  method_family:
  - autoregressive_codec_language_models
  claims:
  - claim_id: single_stage_autoregressive_tts_trained_with_next_token_prediction_over
    role: supports
    claim: Single-stage autoregressive TTS trained with next-token prediction over discrete speech tokens is competitive
      with multi-stage AR+NAR pipelines on intelligibility and speaker similarity in continuation mode, though SIM-o
      gaps remain due to codec acoustic reconstruction limits.
    source: §3.2.4, Table 3
    evidence: Single-stage autoregressive TTS trained with next-token prediction over discrete speech tokens is
      competitive with multi-stage AR+NAR pipelines on intelligibility and speaker similarity in continuation mode,
      though SIM-o gaps remain due to codec acoustic reconstruction limits.
    confidence: high
    relevance: high
  - claim_id: both_model_scale_and_training_data_volume_independently_improve_tts
    role: supports
    claim: Both model scale and training data volume independently improve TTS quality across naturalness, prosody,
      and text comprehension, consistent with scaling laws observed in text LLMs.
    source: §2.3, §3.2.2, Tables 2, 4
    evidence: Both model scale and training data volume independently improve TTS quality across naturalness, prosody,
      and text comprehension, consistent with scaling laws observed in text LLMs.
    confidence: high
    relevance: low
  - claim_id: inference_time_compute_scaling_via_speech_understanding_verifiers_can_substantially
    role: complicates
    claim: Inference-time compute scaling via speech understanding verifiers can substantially improve speaker similarity
      and emotional expressiveness beyond what train-time scaling alone achieves, at the cost of additional inference
      compute.
    source: §2.4, §3.2.3, Figure 2, Table 2
    evidence: Inference-time compute scaling via speech understanding verifiers can substantially improve speaker
      similarity and emotional expressiveness beyond what train-time scaling alone achieves, at the cost of additional
      inference compute.
    confidence: high
    relevance: low
  - claim_id: pure_process_reward_model_beam_search_for_tts_is_prone
    role: supports
    claim: Pure process reward model beam search for TTS is prone to mode collapse that degrades content accuracy
      (WER), and a hybrid partial-PRM strategy is needed to preserve both speaker similarity and intelligibility.
    source: §3.2.3, Figure 2
    evidence: Pure process reward model beam search for TTS is prone to mode collapse that degrades content accuracy
      (WER), and a hybrid partial-PRM strategy is needed to preserve both speaker similarity and intelligibility.
    confidence: high
    relevance: low
  - claim_id: single_vq_codecs_can_achieve_intelligibility_and_naturalness_competitive_with
    role: supports
    claim: Single-VQ codecs can achieve intelligibility and naturalness competitive with multi-layer RVQ codecs
      at the same token rate, but acoustic fidelity (speaker similarity) remains the limiting factor for single-VQ
      reconstruction.
    source: §3.1.2, Table 1
    evidence: Single-VQ codecs can achieve intelligibility and naturalness competitive with multi-layer RVQ codecs
      at the same token rate, but acoustic fidelity (speaker similarity) remains the limiting factor for single-VQ
      reconstruction.
    confidence: high
    relevance: high
  limitations:
  - 'The SIM-o gap between Llasa and RVQ-based baselines is intrinsic to the single-VQ design: acoustic reconstruction
    from a 65,536-entry single codebook at 50 Hz is weaker than 8-layer RVQ codecs, and this gap is only partially
    recovered by inference-time search. Systems requiring high timbre fidelity in a single inference pass would
    need a different codec design.'
  - Inference-time compute scaling requires running multiple candidates (beam search or Best-of-N) with auxiliary
    verifier models, which increases latency and compute cost substantially and makes the approach unsuitable for
    real-time or low-resource applications. The paper does not characterize latency or wall-clock overhead of the
    search strategies.
  - The text understanding evaluation uses expert-rated 3-point discrete scores, which are not directly comparable
    across papers and rely on a small number of sentences per category. The evaluation methodology is adapted from
    BASE TTS but the inter-rater reliability is not reported.
  - Models are trained on mixed Mandarin/English data, but the language coverage and balance are not fully documented.
    The internal data component of the 250k-hour corpus is not described, limiting reproducibility.
  caveats: []
- id: '2502.05512'
  published_date: "2025-02-08"
  entry_date: '2026-07-28'
  year: 2025
  venue: arXiv
  task:
  - TTS
  architecture:
  - autoregressive-LM
  - GAN
  relevance: medium
  evidence_role:
  - architecture_variant
  - acceleration_evidence
  - control_evidence
  current_role: active_evidence
  method_family:
  - autoregressive_codec_language_models
  - gan_neural_codecs_and_decoders
  claims:
  - claim_id: speaker_conditioning_via_a_multi_reference_conformer_perceiver_improves_zero
    role: supports
    claim: Speaker conditioning via a multi-reference Conformer Perceiver improves zero-shot voice cloning stability
      and timbre consistency over single-vector speaker embeddings.
    source: §2.3, Table 4
    evidence: Speaker conditioning via a multi-reference Conformer Perceiver improves zero-shot voice cloning stability
      and timbre consistency over single-vector speaker embeddings.
    confidence: high
    relevance: medium
  - claim_id: direct_waveform_decoding_from_lm_hidden_states_via_a_gan
    role: supports
    claim: Direct waveform decoding from LM hidden states via a GAN vocoder achieves competitive audio quality with
      faster inference than diffusion-based intermediate representation decoding.
    source: §2.4, Table 5
    evidence: Direct waveform decoding from LM hidden states via a GAN vocoder achieves competitive audio quality
      with faster inference than diffusion-based intermediate representation decoding.
    confidence: high
    relevance: high
  - claim_id: fsq_reaches_near_100_codebook_utilisation_with_less_training_data
    role: supports
    claim: FSQ reaches near-100% codebook utilisation with less training data than VQ, though VQ converges to similar
      utilisation with sufficient data scale.
    source: §3.3.2, Figure 2
    evidence: FSQ reaches near-100% codebook utilisation with less training data than VQ, though VQ converges to
      similar utilisation with sufficient data scale.
    confidence: high
    relevance: high
  - claim_id: hybrid_character_pinyin_tokenisation_enables_reliable_correction_of_chinese_polyphonic
    role: supports
    claim: Hybrid character-pinyin tokenisation enables reliable correction of Chinese polyphonic character mispronunciations
      at inference time without requiring a separate grapheme-to-phoneme module.
    source: §3.3.1, Table 2
    evidence: Hybrid character-pinyin tokenisation enables reliable correction of Chinese polyphonic character mispronunciations
      at inference time without requiring a separate grapheme-to-phoneme module.
    confidence: high
    relevance: medium
  limitations:
  - Model size is not reported, and the training data pipeline uses proprietary internet-sourced audio with pseudo-labels
    from commercial ASR; neither the data nor the code is released, limiting reproducibility.
  - The system is limited to Chinese and English, with acknowledged weak emotional expression replication. Instruction-based
    voice generation is explicitly unsupported. The MOS evaluation relies on 100 samples from an unspecified test
    set distribution, and the SPK-SIM metric uses ERes2Net rather than a standardised model, making direct comparison
    with published baselines difficult. The paper does not report streaming latency or real-time factor, despite
    positioning the hybrid architecture as streaming-capable.
  caveats: []
- id: '2502.06490'
  published_date: "2025-02-10"
  entry_date: '2026-07-28'
  year: 2025
  venue: arXiv
  task:
  - TTS
  - VC
  - codec
  architecture: []
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: influential
  method_family: []
  claims:
  - claim_id: acoustic_tokens_and_semantic_tokens_occupy_fundamentally_distinct_points_in
    role: complicates
    claim: Acoustic tokens and semantic tokens occupy fundamentally distinct points in a reconstruction-versus-semantics
      trade-off space, and no single tokenization strategy currently achieves strong performance on both axes simultaneously.
    source: §VI-C, Table I
    evidence: Acoustic tokens and semantic tokens occupy fundamentally distinct points in a reconstruction-versus-semantics
      trade-off space, and no single tokenization strategy currently achieves strong performance on both axes simultaneously.
    confidence: high
    relevance: medium
  - claim_id: speaker_disentanglement_in_acoustic_tokens_enables_voice_conversion_capability_but
    role: supports
    claim: Speaker disentanglement in acoustic tokens enables voice conversion capability but consistently reduces
      reconstruction quality metrics at equivalent bitrates.
    source: §VI-D, Table I
    evidence: Speaker disentanglement in acoustic tokens enables voice conversion capability but consistently reduces
      reconstruction quality metrics at equivalent bitrates.
    confidence: high
    relevance: medium
  - claim_id: k_means_clustering_on_ssl_model_embeddings_discards_prosody_information
    role: supports
    claim: K-means clustering on SSL model embeddings discards prosody information more severely than supervised
      or internally-quantized semantic tokenizers, making offline clustering ill-suited for tasks requiring prosody
      fidelity.
    source: §VI-C, §VI-D, Table I
    evidence: K-means clustering on SSL model embeddings discards prosody information more severely than supervised
      or internally-quantized semantic tokenizers, making offline clustering ill-suited for tasks requiring prosody
      fidelity.
    confidence: high
    relevance: low
  - claim_id: acoustic_byte_pair_encoding_achieves_greater_length_reduction_on_tokens
    role: supports
    claim: Acoustic byte-pair encoding achieves greater length reduction on tokens with lower information density,
      such as speaker-decoupled and semantic tokens, than on general-purpose acoustic tokens.
    source: §V-A, Figure 8
    evidence: Acoustic byte-pair encoding achieves greater length reduction on tokens with lower information density,
      such as speaker-decoupled and semantic tokens, than on general-purpose acoustic tokens.
    confidence: high
    relevance: medium
  - claim_id: single_codebook_tokens_at_very_low_frame_rates_improve_compatibility
    role: supports
    claim: Single-codebook tokens at very low frame rates improve compatibility with language model generation but
      currently exhibit measurable quality and intelligibility degradation relative to multi-codebook or higher
      frame-rate alternatives.
    source: §VIII.1, §VI-C
    evidence: Single-codebook tokens at very low frame rates improve compatibility with language model generation
      but currently exhibit measurable quality and intelligibility degradation relative to multi-codebook or higher
      frame-rate alternatives.
    confidence: high
    relevance: high
  limitations:
  - The experimental comparisons are conducted on English data only (LibriTTS, LibriSpeech), leaving multilingual
    tokenization trade-offs unexplored. The unified vocoder (CTX-vec2wav) is specifically designed for semantic
    tokens, which may introduce a systematic advantage for semantic token types in reconstruction experiments. Not
    all acoustic tokens support voice conversion in the paper's framework, so the VC comparison covers only a subset
    of systems.
  - 'Open questions identified by the survey include: the bitrate lower bound for single-codebook tokens with acceptable
    intelligibility; whether causal SSL architectures can match non-causal models for semantic token quality; how
    VFR tokens perform on generative tasks beyond ASR; and whether token vocoders trained at scale can match flow
    matching-based alternatives for timbre controllability.'
  caveats: []
- id: '2502.11946'
  published_date: "2025-02-17"
  entry_date: '2026-07-28'
  year: 2025
  venue: arXiv
  task:
  - TTS
  - SCA
  architecture:
  - autoregressive-LM
  - flow-matching
  - hybrid
  relevance: medium
  evidence_role:
  - architecture_variant
  - acceleration_evidence
  current_role: influential
  method_family:
  - autoregressive_codec_language_models
  - flow_matching_codec_decoders
  - hybrid_semantic_acoustic_tokenizers
  claims:
  - claim_id: a_dual_codebook_interleaved_tokenizer_that_combines_linguistic_and_semantic
    role: supports
    claim: A dual-codebook interleaved tokenizer that combines linguistic and semantic representations can achieve
      lower ASR error rates than either codebook alone, without sacrificing acoustic reconstruction quality.
    source: §4.4, §6.2.1
    evidence: A dual-codebook interleaved tokenizer that combines linguistic and semantic representations can achieve
      lower ASR error rates than either codebook alone, without sacrificing acoustic reconstruction quality.
    confidence: high
    relevance: high
  - claim_id: scaling_autoregressive_llm_backbone_size_from_3b_to_130b_parameters
    role: supports
    claim: Scaling autoregressive LLM backbone size from 3B to 130B parameters produces substantial gains in speech
      synthesis intelligibility on standard TTS benchmarks, suggesting speech generation quality is LLM-scale-sensitive.
    source: §6.2.2, Table 3
    evidence: Scaling autoregressive LLM backbone size from 3B to 130B parameters produces substantial gains in
      speech synthesis intelligibility on standard TTS benchmarks, suggesting speech generation quality is LLM-scale-sensitive.
    confidence: high
    relevance: medium
  - claim_id: rlhf_reward_models_trained_on_speech_interaction_data_can_exhibit
    role: supports
    claim: RLHF reward models trained on speech interaction data can exhibit systematic failure modes (such as rewarding
      evasive non-answers to unclear audio) unless explicit counter-examples are constructed during reward model
      training.
    source: §5.2.6
    evidence: RLHF reward models trained on speech interaction data can exhibit systematic failure modes (such as
      rewarding evasive non-answers to unclear audio) unless explicit counter-examples are constructed during reward
      model training.
    confidence: high
    relevance: medium
  - claim_id: speculative_response_generation_triggered_by_voice_activity_detection_can_reduce
    role: supports
    claim: Speculative response generation triggered by voice activity detection can reduce per-response latency
      by approximately 500 ms, with roughly 40% of pre-generated responses being usable, enabling practical real-time
      conversational systems.
    source: §3.4
    evidence: Speculative response generation triggered by voice activity detection can reduce per-response latency
      by approximately 500 ms, with roughly 40% of pre-generated responses being usable, enabling practical real-time
      conversational systems.
    confidence: high
    relevance: low
  - claim_id: synthetic_speech_data_generated_by_a_large_multi_modal_model
    role: supports
    claim: Synthetic speech data generated by a large multi-modal model can substitute for manually curated recordings
      in training TTS systems for low-resource dialects, emotions, and singing styles.
    source: §5.1.1
    evidence: Synthetic speech data generated by a large multi-modal model can substitute for manually curated recordings
      in training TTS systems for low-resource dialects, emotions, and singing styles.
    confidence: high
    relevance: medium
  limitations:
  - The StepEval-Audio-360 benchmark is proprietary and created by the same team; human evaluation results on it
    cannot be independently reproduced. Open-source benchmark comparisons mix locally re-run models with results
    taken from original publications, complicating direct numerical comparison.
  - 'Speaker similarity scores for the distilled Step-Audio-TTS-3B are noticeably lower than CosyVoice 2 on both
    Chinese and English SEED-TTS tests, suggesting that the dual-codebook approach trades some acoustic identity
    preservation for intelligibility gains. The AQTA+TTS design still relies on a cascade: errors in ASR transcription
    of history or in text generation propagate to the TTS stage. The paper''s future work section acknowledges that
    purely end-to-end audio-in/audio-out (AQAA) remains unsolved. Evaluation for singing, RAP, and dialect control
    is limited to instruction following scores without reference audio; absolute quality in these dimensions is
    difficult to assess from the reported numbers alone.'
  caveats: []
- id: '2502.17239'
  published_date: "2025-02-24"
  entry_date: '2026-07-28'
  year: 2025
  venue: arXiv
  task:
  - SCA
  - TTS
  architecture:
  - autoregressive-LM
  - flow-matching
  - hybrid
  relevance: medium
  evidence_role:
  - architecture_variant
  - evaluation_caution
  current_role: influential
  method_family:
  - autoregressive_codec_language_models
  - flow_matching_codec_decoders
  - hybrid_semantic_acoustic_tokenizers
  claims:
  - claim_id: multi_codebook_rvq_tokenizers_with_semantic_alignment_objectives_better_preserve
    role: supports
    claim: Multi-codebook RVQ tokenizers with semantic alignment objectives better preserve both acoustic and linguistic
      content than single-codebook designs, with each additional layer reducing ASR error substantially up to 8
      layers.
    source: §3.1, Table 1
    evidence: Multi-codebook RVQ tokenizers with semantic alignment objectives better preserve both acoustic and
      linguistic content than single-codebook designs, with each additional layer reducing ASR error substantially
      up to 8 layers.
    confidence: high
    relevance: high
  - claim_id: staged_pretraining_frozen_llm_first_then_joint_training_measurably_reduces
    role: supports
    claim: Staged pretraining (frozen LLM first, then joint training) measurably reduces intelligence degradation
      in end-to-end speech LMs relative to single-stage joint training.
    source: §3.3.1, Table 5
    evidence: Staged pretraining (frozen LLM first, then joint training) measurably reduces intelligence degradation
      in end-to-end speech LMs relative to single-stage joint training.
    confidence: high
    relevance: medium
  - claim_id: text_guided_aligned_generation_where_the_model_completes_text_tokens
    role: supports
    claim: Text-guided aligned generation, where the model completes text tokens before emitting the corresponding
      audio tokens, mitigates semantic incoherence in speech LM outputs.
    source: §3.3, §4.1
    evidence: Text-guided aligned generation, where the model completes text tokens before emitting the corresponding
      audio tokens, mitigates semantic incoherence in speech LM outputs.
    confidence: high
    relevance: medium
  - claim_id: flow_matching_decoders_trained_as_post_vq_refinement_stages_recover
    role: supports
    claim: Flow-matching decoders trained as post-VQ refinement stages recover significant audio quality lost during
      quantisation, with UTMOS improvements of 0.6 points possible without retraining the LLM.
    source: §3.2, Table 3
    evidence: Flow-matching decoders trained as post-VQ refinement stages recover significant audio quality lost
      during quantisation, with UTMOS improvements of 0.6 points possible without retraining the LLM.
    confidence: high
    relevance: medium
  - claim_id: end_to_end_speech_lms_evaluated_in_s_s_mode
    role: supports
    claim: End-to-end speech LMs evaluated in S→S mode exhibit notably lower benchmark performance than the same
      model evaluated in S→T mode, indicating that audio token generation itself introduces a quality penalty beyond
      the comprehension step.
    source: §4.3, Table 8
    evidence: End-to-end speech LMs evaluated in S→S mode exhibit notably lower benchmark performance than the same
      model evaluated in S→T mode, indicating that audio token generation itself introduces a quality penalty beyond
      the comprehension step.
    confidence: high
    relevance: medium
  limitations:
  - The TTS quality evaluation is limited to an in-house test set (MED-TTS); no comparison against standard TTS
    benchmarks (VCTK, LJSpeech, LibriTTS test-clean) or against dedicated TTS systems is reported. The naturalness
    and speaker similarity of the generated speech relative to state-of-the-art TTS systems is therefore unknown.
  - The intelligence gap to GPT-4o-Audio remains large (roughly 15-20 percentage points on QA tasks), and the paper
    does not explain what architectural or data factors account for this difference. The OpenAudioBench evaluation
    uses GPT-4o as judge, which may introduce evaluation bias. The full-duplex and interruption-handling capabilities
    common in deployed spoken conversational agents are not evaluated. Training data mix decisions (e.g. dropping
    INTLV audio loss, ITTS loss design) are motivated empirically but without systematic ablation.
  caveats: []
- id: '2502.18924'
  published_date: "2025-02-26"
  entry_date: '2026-07-28'
  year: 2025
  venue: arXiv
  task:
  - TTS
  architecture:
  - diffusion
  - VAE
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: influential
  method_family:
  - diffusion_latent_codec_generators
  - vae_vector_quantized_codecs
  claims:
  - claim_id: providing_coarse_stochastic_phoneme_anchors_rather_than_fully_expanded_forced
    role: supports
    claim: Providing coarse stochastic phoneme anchors rather than fully expanded forced alignments improves both
      naturalness and robustness simultaneously in latent diffusion TTS.
    source: §3.2, §4.3, Table 4, Table 7
    evidence: Providing coarse stochastic phoneme anchors rather than fully expanded forced alignments improves
      both naturalness and robustness simultaneously in latent diffusion TTS.
    confidence: high
    relevance: medium
  - claim_id: compact_continuous_latent_representations_at_very_low_token_rates_enable
    role: supports
    claim: Compact continuous latent representations at very low token rates enable higher zero-shot TTS quality
      than discrete codecs at higher bit rates when used as the target space for diffusion.
    source: §4.5, Table 5, Table 6
    evidence: Compact continuous latent representations at very low token rates enable higher zero-shot TTS quality
      than discrete codecs at higher bit rates when used as the target space for diffusion.
    confidence: high
    relevance: medium
  - claim_id: piecewise_rectified_flow_distillation_reduces_inference_steps_from_25_to
    role: supports
    claim: Piecewise rectified flow distillation reduces inference steps from 25 to 8 with negligible degradation
      in speaker similarity and intelligibility.
    source: §3.2, §4.2, Table 1
    evidence: Piecewise rectified flow distillation reduces inference steps from 25 to 8 with negligible degradation
      in speaker similarity and intelligibility.
    confidence: high
    relevance: low
  - claim_id: decoupled_text_and_speaker_guidance_scales_in_classifier_free_guidance
    role: supports
    claim: Decoupled text and speaker guidance scales in classifier-free guidance provide a continuous accent intensity
      control axis without requiring accent labels.
    source: §3.2, §4.4, Table 3
    evidence: Decoupled text and speaker guidance scales in classifier-free guidance provide a continuous accent
      intensity control axis without requiring accent labels.
    confidence: high
    relevance: medium
  - claim_id: latent_diffusion_tts_systems_exhibit_strong_data_and_model_scaling
    role: supports
    claim: Latent diffusion TTS systems exhibit strong data and model scaling behaviour, with both speaker similarity
      and intelligibility improving consistently as training data grows from 2k to 600k hours and model size grows
      from 0.5B to 7B parameters.
    source: Appendix D, Table 8
    evidence: Latent diffusion TTS systems exhibit strong data and model scaling behaviour, with both speaker similarity
      and intelligibility improving consistently as training data grows from 2k to 600k hours and model size grows
      from 0.5B to 7B parameters.
    confidence: high
    relevance: low
  limitations:
  - The main results are reported on LibriSpeech test-clean, a read-speech corpus recorded in controlled conditions.
    The scaling and cross-domain results (Appendix D) use an internal test set of 400 samples, limiting external
    reproducibility for those claims.
  - Language coverage is restricted to English and Chinese despite the 600k-hour multilingual training corpus. The
    sparse alignment mechanism still depends on an external forced aligner (Montreal Forced Aligner) at training
    time, which requires a transcription pipeline and does not eliminate the dependency on alignment tools, merely
    relaxing it. The relationship between alignment anchor density and generation quality is explored only qualitatively;
    no principled analysis determines the optimal sparsity level. Code and checkpoints are not publicly available
    at time of writing.
  caveats: []
- id: '2503.01710'
  published_date: "2025-03-03"
  entry_date: '2026-07-28'
  year: 2025
  venue: arXiv
  task:
  - TTS
  architecture:
  - autoregressive-LM
  relevance: medium
  evidence_role:
  - architecture_variant
  - acceleration_evidence
  - control_evidence
  current_role: influential
  method_family:
  - autoregressive_codec_language_models
  claims:
  - claim_id: disentangling_speech_tokens_into_linguistic_content_and_speaker_attributes_within
    role: supports
    claim: Disentangling speech tokens into linguistic content and speaker attributes within a single-stream codec
      enables a standard LLM to perform zero-shot TTS without a multi-stage pipeline.
    source: §3, §4.1
    evidence: Disentangling speech tokens into linguistic content and speaker attributes within a single-stream
      codec enables a standard LLM to perform zero-shot TTS without a multi-stage pipeline.
    confidence: high
    relevance: high
  - claim_id: small_llm_backbones_can_achieve_competitive_zero_shot_tts_intelligibility
    role: supports
    claim: Small LLM backbones can achieve competitive zero-shot TTS intelligibility when the codec reduces per-token
      modeling complexity through semantic alignment.
    source: §6.4, Table 4
    evidence: Small LLM backbones can achieve competitive zero-shot TTS intelligibility when the codec reduces per-token
      modeling complexity through semantic alignment.
    confidence: high
    relevance: high
  - claim_id: single_stage_autoregressive_tts_consistently_trails_multi_stage_or_non
    role: supports
    claim: Single-stage autoregressive TTS consistently trails multi-stage or non-autoregressive methods on speaker
      similarity metrics, even when intelligibility is comparable.
    source: §6.4, Table 4, Limitation
    evidence: Single-stage autoregressive TTS consistently trails multi-stage or non-autoregressive methods on speaker
      similarity metrics, even when intelligibility is comparable.
    confidence: high
    relevance: low
  - claim_id: fsq_based_global_token_quantization_with_learnable_cross_attention_queries
    role: supports
    claim: FSQ-based global token quantization with learnable cross-attention queries produces better speaker attribute
      representation than group-VQ at equivalent token lengths.
    source: §6.2, Table 2
    evidence: FSQ-based global token quantization with learnable cross-attention queries produces better speaker
      attribute representation than group-VQ at equivalent token lengths.
    confidence: high
    relevance: high
  - claim_id: attribute_controllable_tts_benefits_from_hierarchical_coarse_to_fine_prediction
    role: supports
    claim: Attribute-controllable TTS benefits from hierarchical coarse-to-fine prediction within the LM inference
      loop rather than requiring separate conditioning modules.
    source: §4.1, §6.3
    evidence: Attribute-controllable TTS benefits from hierarchical coarse-to-fine prediction within the LM inference
      loop rather than requiring separate conditioning modules.
    confidence: high
    relevance: medium
  limitations:
  - Speaker similarity in zero-shot cloning is meaningfully lower than multi-stage methods (SIM 0.672 vs. 0.774
    for MaskGCT on test-zh). The paper attributes this to AR variability without explicit disentanglement constraints
    between semantic and global tokens, and no solution is evaluated in this work.
  - The VoxBox training data and BiCodec codec are trained on separate, relatively limited datasets (3k hours for
    BiCodec; 102.5k hours for the LM). The BiCodec training data is English-only (LibriSpeech + Emilia EN/CN), which
    may limit acoustic reconstruction quality for languages outside this distribution.
  - The attribution control is evaluated primarily for gender, pitch, and speed; there is no evaluation of emotion
    control despite VoxBox containing emotion annotations for many source datasets.
  - Fine-grained numerical pitch control is evaluated on figures (Figs. 4–5) but no quantitative accuracy metric
    against target pitch values is reported, making it difficult to assess precision.
  caveats: []
- id: '2503.14345'
  published_date: "2025-03-18"
  entry_date: '2026-07-28'
  year: 2025
  venue: arXiv
  task:
  - TTS
  architecture:
  - autoregressive-LM
  - flow-matching
  relevance: high
  evidence_role:
  - core_evidence
  - control_evidence
  - evaluation_caution
  current_role: influential
  method_family:
  - autoregressive_codec_language_models
  - flow_matching_codec_decoders
  claims:
  - claim_id: spontaneous_scripting_from_an_llm_is_roughly_as_important_as
    role: supports
    claim: Spontaneous scripting from an LLM is roughly as important as the acoustic modeling choice for perceived
      spontaneity in long-form dialogue TTS.
    source: §4.2.2, Table 3
    evidence: Spontaneous scripting from an LLM is roughly as important as the acoustic modeling choice for perceived
      spontaneity in long-form dialogue TTS.
    confidence: high
    relevance: low
  - claim_id: full_sequence_interleaving_of_text_and_speech_codes_extended_to
    role: supports
    claim: Full-sequence interleaving of text and speech codes, extended to 40,000-token context windows with speaker-change
      tokens, enables coherent long-form zero-shot multi-speaker synthesis that turn-level concatenation cannot
      match.
    source: §3.2.1, Tables 1–2
    evidence: Full-sequence interleaving of text and speech codes, extended to 40,000-token context windows with
      speaker-change tokens, enables coherent long-form zero-shot multi-speaker synthesis that turn-level concatenation
      cannot match.
    confidence: high
    relevance: high
  - claim_id: curriculum_learning_progressively_exposing_a_codec_lm_to_increasing_dialogue
    role: supports
    claim: Curriculum learning, progressively exposing a codec LM to increasing dialogue complexity, is an effective
      strategy for developing long-context and spontaneous generation capability without requiring matched long-context
      data from the outset.
    source: §3.2.1
    evidence: Curriculum learning, progressively exposing a codec LM to increasing dialogue complexity, is an effective
      strategy for developing long-context and spontaneous generation capability without requiring matched long-context
      data from the outset.
    confidence: high
    relevance: high
  - claim_id: automatic_speaker_similarity_metrics_cosine_embedding_similarity_can_disagree_with
    role: supports
    claim: Automatic speaker similarity metrics (cosine embedding similarity) can disagree with subjective speaker
      similarity ratings in long-form generation settings, particularly when the acoustic model attends to prosodic
      rather than purely timbral features.
    source: §4.2.1
    evidence: Automatic speaker similarity metrics (cosine embedding similarity) can disagree with subjective speaker
      similarity ratings in long-form generation settings, particularly when the acoustic model attends to prosodic
      rather than purely timbral features.
    confidence: high
    relevance: low
  - claim_id: chunk_wise_autoregressive_decoding_with_a_causal_chunk_mask_provides
    role: supports
    claim: Chunk-wise autoregressive decoding with a causal chunk mask provides a practical solution to the continuity
      and memory constraints of mel-spectrogram reconstruction from long semantic code sequences.
    source: §3.2.2
    evidence: Chunk-wise autoregressive decoding with a causal chunk mask provides a practical solution to the continuity
      and memory constraints of mel-spectrogram reconstruction from long semantic code sequences.
    confidence: high
    relevance: medium
  limitations:
  - All evaluation is conducted on a small internal test set (4 knowledge sources for podcast, 7 podcasts for the
    script ablation), and training data is entirely proprietary (~515K hours). Results cannot be independently reproduced,
    and generalisability to other domains or languages beyond Chinese and English is unverified.
  - The notably lower SIM-O score for English (0.53 vs. 0.75 baseline) reveals unresolved speaker consistency challenges
    in the English long-context setting, which the authors attribute to an imbalanced training curriculum favouring
    audiobook over conversational English data. Hallucinations in speaker attribution (utterances assigned to the
    wrong speaker) emerge from the interplay of semantic token timbre leakage, diarization errors in training data,
    and ambiguous filler-word interpretations; no mitigation is proposed beyond discussion. The two-speaker restriction
    (host + guest) is a deliberate scope limitation; extension to three or more speakers is left as future work.
    Evaluation is subjective-only for multi-speaker interactions; no standardised zero-shot TTS benchmark (LibriSpeech,
    VCTK) is used, limiting direct comparison to single-speaker zero-shot systems.
  caveats: []
- id: '2504.08528'
  published_date: "2025-04-11"
  entry_date: '2026-07-28'
  year: 2025
  venue: arXiv
  task:
  - SCA
  architecture: []
  relevance: medium
  evidence_role:
  - architecture_variant
  - acceleration_evidence
  - control_evidence
  - infrastructure
  current_role: influential
  method_family: []
  claims:
  - claim_id: models_that_excel_at_content_level_spoken_language_understanding_tend
    role: supports
    claim: Models that excel at content-level spoken language understanding tend to underperform on tasks requiring
      reasoning about acoustic and paralinguistic details, indicating that these capabilities require separate optimisation.
    source: §7.2, Table discussion; VoiceBench vs. MMAU comparison
    evidence: Models that excel at content-level spoken language understanding tend to underperform on tasks requiring
      reasoning about acoustic and paralinguistic details, indicating that these capabilities require separate optimisation.
    confidence: high
    relevance: medium
  - claim_id: cascaded_speech_pipelines_and_end_to_end_spoken_language_models
    role: supports
    claim: 'Cascaded speech pipelines and end-to-end spoken language models have complementary strengths: cascades
      are stronger on semantic tasks, while end-to-end models better handle speaker and paralinguistic information.'
    source: §7.2 "Comparisons of cascaded vs. end-to-end SLMs"
    evidence: 'Cascaded speech pipelines and end-to-end spoken language models have complementary strengths: cascades
      are stronger on semantic tasks, while end-to-end models better handle speaker and paralinguistic information.'
    confidence: high
    relevance: medium
  - claim_id: jointly_training_a_speech_encoder_with_an_llm_backbone_improves
    role: supports
    claim: Jointly training a speech encoder with an LLM backbone improves instruction following and paralinguistic
      understanding but introduces risk of catastrophic forgetting of pre-trained text capabilities.
    source: §7.2, §4.3
    evidence: Jointly training a speech encoder with an LLM backbone improves instruction following and paralinguistic
      understanding but introduces risk of catastrophic forgetting of pre-trained text capabilities.
    confidence: high
    relevance: low
  - claim_id: full_duplex_spoken_dialogue_requires_architectural_support_beyond_turn_taking
    role: supports
    claim: Full-duplex spoken dialogue requires architectural support beyond turn-taking assumptions, and current
      approaches (dual-channel and time-multiplexing) each involve significant trade-offs in latency, naturalness,
      and modelling complexity.
    source: §6
    evidence: Full-duplex spoken dialogue requires architectural support beyond turn-taking assumptions, and current
      approaches (dual-channel and time-multiplexing) each involve significant trade-offs in latency, naturalness,
      and modelling complexity.
    confidence: high
    relevance: low
  - claim_id: the_tokenisation_choice_phonetic_tokens_versus_audio_codec_tokens_determines
    role: complicates
    claim: The tokenisation choice (phonetic tokens versus audio codec tokens) determines the trade-off between
      linguistic coherence and speaker or acoustic fidelity in spoken language model outputs.
    source: §3.1.2 "Comparison of token types"
    evidence: The tokenisation choice (phonetic tokens versus audio codec tokens) determines the trade-off between
      linguistic coherence and speaker or acoustic fidelity in spoken language model outputs.
    confidence: high
    relevance: high
  limitations:
  - The survey is primarily a snapshot of English-centric, high-resource SLM research. Multilingual, low-resource,
    and accessibility-focused SLMs receive minimal coverage, and the authors acknowledge that SLM research has largely
    not addressed dialects, accents, or speech-related medical conditions.
  - 'Beyond that scope limitation, several structural gaps are identified:'
  - '- Scaling behaviour for speech+text LMs and speech-aware text LMs is entirely unknown; existing scaling studies
    apply only to pure speech LMs. - Most SLM evaluations use non-overlapping benchmarks, making direct performance
    comparisons across model families impossible. - Very few SLMs are fully open-source (code, weights, and data),
    which prevents controlled ablations of design choices. - The representation of spoken output in SLMs is unsettled:
    discrete tokens, continuous flows, and hybrid approaches have not been systematically compared under controlled
    conditions. - Safety and trustworthiness evaluation for SLMs is nascent; non-verbal toxicity and speaker-type
    bias remain largely unstudied.'
  caveats: []
- id: '2504.10344'
  published_date: "2025-04-14"
  entry_date: '2026-07-28'
  year: 2025
  venue: arXiv
  task:
  - codec
  - TTS
  architecture:
  - autoregressive-LM
  - VAE
  relevance: high
  evidence_role:
  - core_evidence
  - control_evidence
  current_role: active_evidence
  method_family:
  - autoregressive_codec_language_models
  - vae_vector_quantized_codecs
  claims:
  - claim_id: frame_level_quantization_without_cross_frame_context_limits_codec_semantic
    role: supports
    claim: Frame-level quantization without cross-frame context limits codec semantic richness and increases the
      difficulty of autoregressive LM training on the resulting tokens.
    source: §1, Figure 1
    evidence: Frame-level quantization without cross-frame context limits codec semantic richness and increases
      the difficulty of autoregressive LM training on the resulting tokens.
    confidence: high
    relevance: high
  - claim_id: initialising_vq_codebook_entries_from_semantic_priors_derived_from_self
    role: supports
    claim: Initialising VQ codebook entries from semantic priors derived from self-supervised models improves downstream
      recognition accuracy without additional distillation overhead.
    source: §3.2, Table 6
    evidence: Initialising VQ codebook entries from semantic priors derived from self-supervised models improves
      downstream recognition accuracy without additional distillation overhead.
    confidence: high
    relevance: high
  - claim_id: lower_token_sequence_length_lower_hz_frame_rate_improves_both
    role: supports
    claim: Lower token-sequence length (lower Hz frame rate) improves both training and inference efficiency for
      audio language models, independent of bitrate.
    source: Appendix B, Table 8
    evidence: Lower token-sequence length (lower Hz frame rate) improves both training and inference efficiency
      for audio language models, independent of bitrate.
    confidence: high
    relevance: high
  - claim_id: reconstruction_quality_alone_is_an_insufficient_predictor_of_a_codec
    role: supports
    claim: 'Reconstruction quality alone is an insufficient predictor of a codec''s suitability for autoregressive
      language modelling: codecs with stronger semantic content produce more robust LM-based TTS even at higher
      bitrates.'
    source: §4.6, Table 4
    evidence: 'Reconstruction quality alone is an insufficient predictor of a codec''s suitability for autoregressive
      language modelling: codecs with stronger semantic content produce more robust LM-based TTS even at higher
      bitrates.'
    confidence: high
    relevance: high
  - claim_id: a_masked_autoencoder_auxiliary_loss_during_codec_training_increases_semantic
    role: supports
    claim: A masked autoencoder auxiliary loss during codec training increases semantic information in learned representations
      at a modest reconstruction cost.
    source: §3.3, Table 6
    evidence: A masked autoencoder auxiliary loss during codec training increases semantic information in learned
      representations at a modest reconstruction cost.
    confidence: high
    relevance: high
  limitations:
  - All downstream LM experiments use the same 1B-parameter LLaMA backbone with limited training data (2000 hours
    speech, ~500 hours each sound/music). Gains observed may not transfer to larger-scale audio LM systems, where
    baseline tokenizers may reach ceiling performance.
  - The two-stage training adds complexity and discards large sub-networks (MAE encoder/decoder, AR prediction transformer)
    after training, increasing resource cost without reuse. The authors note this explicitly and flag it as future
    work. Sound and music reconstruction remains challenging at 0.41 kbps, with VISQOL scores substantially lower
    than speech. Semantic retention still lags SSL models for ASR (18.3% WER vs. 6.2% for WavLM), indicating the
    codec does not fully substitute for purpose-built SSL representations. Code and model weights were not released
    at submission time, limiting reproducibility.
  caveats: []
- id: iclr-2025-868masI331
  published_date: "2025-04-24"
  entry_date: '2026-07-28'
  year: 2025
  venue: ICLR
  task:
  - TTS
  architecture:
  - autoregressive-LM
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: active_evidence
  method_family:
  - autoregressive_codec_language_models
  claims:
  - claim_id: reducing_the_first_layer_rvq_frame_rate_of_a_neural
    role: supports
    claim: Reducing the first-layer RVQ frame rate of a neural audio codec through hierarchical multi-resolution
      distillation enables autoregressive TTS to generate minute-long speech with stable intelligibility.
    source: §4, §7.3, Table 3, Table 4
    evidence: MReQ-Encodec at 8 Hz achieves WER 9.79% on MinutesSpeech test-90s where standard Encodec at 8 Hz produces
      100% WER and naive VALL-E at 48 Hz with long training data yields 16.14%.
    confidence: high
    relevance: high
  - claim_id: lowering_the_codec_frame_rate_in_autoregressive_tts_improves_temporal
    role: complicates
    claim: Lowering the codec frame rate in autoregressive TTS improves temporal coherence and intelligibility for
      long utterances but degrades speaker similarity.
    source: §7.3, Table 4, Table 6
    evidence: HALL-E consistently lags VALL-E by 0.025-0.042 SIM points on MinutesSpeech tests; the paper attributes
      this to acoustic information loss when compressing from 48 Hz to 8 Hz in the first RVQ layer.
    confidence: high
    relevance: high
  - claim_id: post_training_from_a_pre_trained_lm_checkpoint_is_essential
    role: supports
    claim: 'Post-training from a pre-trained LM checkpoint is essential for hierarchical codec TTS: training from
      scratch without pre-trained weights collapses quality.'
    source: §7.4, Table 10
    evidence: Removing VALL-E pre-training from HALL-E increases WER from 9.79% to 49.8% on MinutesSpeech test-90s,
      and removing MReQ pre-training similarly degrades codec reconstruction.
    confidence: high
    relevance: high
  - claim_id: the_frame_rate_at_which_autoregressive_speech_token_generation_becomes
    role: refines
    claim: The frame rate at which autoregressive speech token generation becomes unstable is approximately 8 Hz,
      consistent with average phoneme durations of around 100 ms.
    source: Appendix C.1, Table 18
    evidence: Ablation at 4 Hz first-layer rate raises WER to 20.07%, while 8 Hz yields 9.79%; the paper notes that
      phoneme duration averaging ~100 ms corresponds to ~10 Hz, making 4 Hz fundamentally insufficient.
    confidence: high
    relevance: medium
  - claim_id: length_balanced_training_data_covering_the_target_synthesis_duration_is
    role: supports
    claim: Length-balanced training data covering the target synthesis duration is necessary for autoregressive
      models to generalize to long-form speech.
    source: §7.3, Table 4
    evidence: VALL-E trained only on segments up to 28 seconds achieves WER 39.77% on test-90s, and training on
      longer data decreases but does not eliminate the gap; HALL-E's 8 Hz formulation resolves instability that
      remains even with longer VALL-E training data.
    confidence: high
    relevance: medium
  limitations:
  - 'Speaker similarity consistently degrades with frame rate reduction: HALL-E''s SIM scores are 0.025-0.042 lower
    than VALL-E across test conditions. This reflects a fundamental trade-off between temporal compression and acoustic
    fidelity preservation that the paper does not resolve.'
  - The current NAR model processes audio at full 48 Hz resolution, which limits NAR input length to around 28-54
    seconds during training and requires a sliding window at inference. Integrating cross-attention text conditioning
    addresses alignment but does not eliminate the NAR length bottleneck. The paper also does not compare against
    streaming or chunked autoregressive synthesis approaches, which represent a practical alternative. MinutesSpeech
    training data consists entirely of English podcast speech, and generalization to other languages, domains, or
    reading styles is untested.
  caveats: []
- id: iclr-2025-cuFzE8Jlvb
  published_date: "2025-04-24"
  entry_date: '2026-07-28'
  year: 2025
  venue: ICLR
  task:
  - TTS
  architecture:
  - autoregressive-LM
  - VAE
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: active_evidence
  method_family:
  - autoregressive_codec_language_models
  - vae_vector_quantized_codecs
  claims:
  - claim_id: continuous_latent_representations_can_replace_discrete_vector_quantization_in_autoregressive
    role: supports
    claim: Continuous latent representations can replace discrete vector quantization in autoregressive TTS without
      sacrificing generation quality.
    source: §5.1, Table 1, Table 2
    evidence: GMM-LM trained on continuous GMM-VAE encoder features outperforms VALL-E (RVQ-based) on WER, speaker
      similarity, Q-MOS, and S-MOS on LibriSpeech test-clean across all prompt lengths, while using 10.3% of VALL-E's
      parameter count.
    confidence: high
    relevance: medium
  - claim_id: longer_audio_prompts_do_not_uniformly_improve_zero_shot_speaker
    role: complicates
    claim: Longer audio prompts do not uniformly improve zero-shot speaker cloning across AR architectures.
    source: §5.1, Table 2
    evidence: VALL-E's WER increases monotonically with prompt length (6.04% at 3s, 7.54% at 8s, 9.68% at 15s),
      suggesting that simple cross-attention cannot leverage extended speaker context in AR decoding; the proposed
      GMM-LM shows the opposite trend, consistently benefiting from longer prompts.
    confidence: high
    relevance: medium
  - claim_id: strict_monotonic_alignment_substantially_reduces_word_error_rate_in_autoregressive
    role: supports
    claim: Strict monotonic alignment substantially reduces word error rate in autoregressive TTS compared to standard
      cross-attention and soft monotonic variants.
    source: Appendix A.1, Table 6
    evidence: Among alignment strategies tested on the same GMM-LM architecture, stochastic monotonic alignment
      with ST-Gumbel achieves WER 2.72% vs. 6.6% for cross-attention alone; even monotonic attention with Gumbel
      (without the stochastic binary forward pass) scores 3.34%.
    confidence: high
    relevance: medium
  - claim_id: continuous_speech_representations_improve_downstream_autoregressive_model_performance_relative_to
    role: supports
    claim: Continuous speech representations improve downstream autoregressive model performance relative to discrete
      counterparts, independent of the alignment mechanism.
    source: Appendix A.6, Table 10
    evidence: A head-to-head ablation comparing GMM-LM (continuous) against discrete AR models (VQ-VAE single codebook
      and DAC multi-codebook with delayed prediction), all using the proposed monotonic alignment, shows GMM-LM
      achieves WER 2.72% vs. 5.35% and 5.87% for the discrete variants.
    confidence: high
    relevance: medium
  - claim_id: increasing_the_number_of_gaussian_components_in_continuous_ar_modeling
    role: complicates
    claim: Increasing the number of Gaussian components in continuous AR modeling yields diminishing returns and
      can reduce quality through overfitting.
    source: §5.4, Table 4, Table 5, Appendix A.3
    evidence: GMM-LM with 6 diagonal-covariance Gaussians (WER 2.72%, SIM 0.91) outperforms 3-Gaussian (WER 2.89%,
      SIM 0.85), but 10-Gaussian degrades to WER 5.21%, SIM 0.71; the 6-mixture GMM-VAE also shows worse evaluation-set
      reconstruction than the 3-mixture model despite better training-set fit.
    confidence: high
    relevance: medium
  limitations:
  - The evaluation compares against VALL-E (2023), a dated AR baseline. Stronger AR systems published by the time
    of this paper's submission are not included, limiting the strength of the superiority claim for continuous AR
    over discrete AR in general.
  - The model is evaluated exclusively on English audiobook speech (LibriLight training, LibriSpeech evaluation).
    Generalisation to conversational speech, noisy in-the-wild data, or other languages is not demonstrated, though
    noise robustness experiments (Appendix A.4) show the method degrades gracefully under additive noise in prompts.
  - The GMM-VAE introduces an additional 76.5M-parameter component, partially offsetting the parameter savings claimed
    relative to VALL-E's RVQ codec (which uses 16.7M codec-related params). The total system size (GMM-VAE + GMM-LM-Mini)
    is 76.5M + 51.5M = 128M, larger than the headline 51.5M figure.
  - The stochastic monotonic alignment requires sequential per-step alignment computation (Algorithm 1), which may
    limit training throughput compared to fully parallelisable attention. The paper does not report training wall-clock
    times or throughput comparisons.
  - Code and pre-trained models are planned for release but were not available at submission time.
  caveats: []
- id: iclr-2025-dGSOn7sdWg
  published_date: "2025-04-24"
  entry_date: '2026-07-28'
  year: 2025
  venue: ICLR
  task:
  - SCA
  architecture:
  - autoregressive-LM
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: influential
  method_family:
  - autoregressive_codec_language_models
  claims:
  - claim_id: reducing_the_token_rate_of_speech_representations_below_10_hz
    role: supports
    claim: Reducing the token rate of speech representations below 10 Hz is sufficient to preserve semantic content
      adequate for spoken language modelling, while substantially improving training and inference efficiency.
    source: §5.6, Table 6
    evidence: SyllableLM at 6.25 Hz, 90M parameters, matches or exceeds TWIST models up to 13B parameters on sBLIMP
      semantic understanding benchmarks, with 30x less training compute and 4.5x faster inference than an equal-sized
      TWIST baseline.
    confidence: high
    relevance: medium
  - claim_id: the_loss_surface_of_a_masked_prediction_ssl_model_encodes
    role: supports
    claim: The loss surface of a masked-prediction SSL model encodes latent syllabic segmentation boundaries discoverable
      without additional supervised signal or cross-modal supervision.
    source: §3.1, Table 1
    evidence: LossPred, applied to a frozen HuBERT student-teacher pair without any training, achieves F1-50 of
      59.6 on syllabic boundary detection, outperforming the feature-similarity-based baseline (47.3) while requiring
      no fine-tuning.
    confidence: high
    relevance: medium
  - claim_id: iterative_student_teacher_distillation_over_pseudo_syllabic_boundaries_progressively_sharpens
    role: supports
    claim: Iterative student-teacher distillation over pseudo-syllabic boundaries progressively sharpens SSL encoder
      representations toward syllable-level organisation.
    source: §5.3, Table 2
    evidence: SylBoost applied to HuBERT improves boundary detection F1 from 60.1 (LossPred initialisation) to 70.2
      after two iterations; applying it to Data2Vec2 reaches 73.2, each iteration producing a measurable gain over
      the previous.
    confidence: high
    relevance: medium
  - claim_id: low_frequency_speech_units_that_improve_semantic_modelling_efficiency_may
    role: complicates
    claim: Low-frequency speech units that improve semantic modelling efficiency may sacrifice robustness to speaker
      rate variation.
    source: Appendix A.5, Table 11
    evidence: SylBoost unit counts collapse under audio speedups of 0.5x and 0.6x relative to original length, performing
      comparably to SD-HuBERT only at mild speedups (0.8x–0.9x range), while showing greater robustness to slowdowns.
    confidence: high
    relevance: medium
  - claim_id: units_optimised_for_semantic_modelling_in_audiobook_speech_may_lose
    role: complicates
    claim: Units optimised for semantic modelling in audiobook speech may lose paralinguistic information, limiting
      applicability to domains requiring prosodic or tonal fidelity.
    source: §6
    evidence: The authors note that low-frequency SylBoost units may lose paralinguistic features such as tone,
      and the entire evaluation is conducted on audiobook data (LibriSpeech / LibriLight); performance on spontaneous
      or multi-speaker speech is not reported.
    confidence: high
    relevance: medium
  limitations:
  - All training and evaluation uses English audiobook data (LibriSpeech, LibriLight). Generalisation to spontaneous
    conversational speech, other languages, or multi-speaker settings is not demonstrated.
  - The interleaved vocoder decoding pipeline introduces a dependency on the TWIST tokeniser and vocoder for resynthesis,
    meaning the system is not fully end-to-end and unit bitrate for the final waveform is partially bounded by TWIST's
    own quality ceiling (WER 6.3%). The efficiency gains in the SpeechLM do not fully apply to the decoding pipeline,
    which involves an additional language model. Scaling beyond 300M parameters was not attempted due to compute
    constraints, leaving open whether the efficiency advantage persists at very large model scales. The base encoder
    quality (Data2Vec2 vs newer models like w2v-BERT 2.0) is acknowledged as a confounding factor in cross-model
    comparisons.
  caveats: []
- id: iclr-2025-hQvX9MBowC
  published_date: "2025-04-24"
  entry_date: '2026-07-28'
  year: 2025
  venue: ICLR
  task:
  - TTS
  architecture:
  - diffusion
  - transformer-enc-dec
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: influential
  method_family:
  - diffusion_latent_codec_generators
  - transformer_encoder_decoder_tokenizers
  claims:
  - claim_id: diffusion_transformer_backbones_are_better_suited_to_tts_than_u
    role: supports
    claim: Diffusion Transformer backbones are better suited to TTS than U-Net backbones once domain-specific conditioning
      factors (phonemes, durations) are removed.
    source: §5.2, Table 4
    evidence: Under matched training conditions, replacing the DiT backbone with a U-Net (and a U-Net variant without
      down/up-sampling) increases WER from 2.93 to 3.7 and drops SIM-r from 0.588 to 0.389 on the English cross-sentence
      task.
    confidence: high
    relevance: medium
  - claim_id: predicting_total_target_length_and_generating_variable_length_sequences_outperforms
    role: supports
    claim: Predicting total target length and generating variable-length sequences outperforms fixed-length generation
      with padding in diffusion-based TTS.
    source: §5.2, Table 5
    evidence: Fixed-length modeling with padding reaches WER 6.81-8.89, while a learned speech length predictor
      with variable-length generation reaches WER 5.36-5.58 under otherwise identical settings.
    confidence: high
    relevance: medium
  - claim_id: aligning_text_and_speech_latent_representations_improves_cross_attention_conditioned
    role: supports
    claim: Aligning text and speech latent representations improves cross-attention-conditioned generation quality,
      independent of model or training-data scale.
    source: §5.2, Tables 6-7
    evidence: A speech codec fine-tuned with an auxiliary language-modeling loss against a frozen text encoder (Mel-VAE++)
      improves WER/SIM over the unaligned codec regardless of which text encoder (ByT5 or SpeechT5) is paired with
      it, and a jointly text-speech-trained text encoder (SpeechT5, 85M params) outperforms a larger text-only encoder
      (ByT5-base, 415M params) trained on more data.
    confidence: high
    relevance: medium
  - claim_id: removing_domain_specific_alignment_factors_from_ldm_based_tts_narrows
    role: complicates
    claim: Removing domain-specific alignment factors from LDM-based TTS narrows but does not eliminate the gap
      to phoneme-duration-based systems in speaker similarity.
    source: §5.1, Table 2
    evidence: DiTTo-en-XL reaches SIM-r 0.6554 on the cross-sentence task, below Voicebox's reported 0.681 (a phoneme/duration-based
      non-autoregressive model), even though DiTTo-en-XL is faster and matches or exceeds Voicebox on WER.
    confidence: high
    relevance: low
  - claim_id: codec_compression_ratio_not_codec_reconstruction_quality_alone_determines_suitability
    role: complicates
    claim: Codec compression ratio, not codec-reconstruction quality alone, determines suitability as a diffusion
      target for variable-length TTS.
    source: §5.2, Table 7
    evidence: DAC achieves higher PESQ and ViSQOL codec-reconstruction scores than Mel-VAE, but its 7-8x longer
      latent sequences make training and inference substantially less efficient and degrade end-to-end WER/SIM relative
      to the more compressed but lower-fidelity Mel-VAE.
    confidence: high
    relevance: high
  limitations:
  - Code and pretrained weights are not released at publication (only demo samples), and most baseline comparisons
    (Voicebox, VALL-E, NaturalSpeech 2/3) use numbers copied from the original papers' reported results rather than
    reproduced under DiTTo's own evaluation pipeline, limiting independent verification of head-to-head rankings.
  - The speaker-similarity gap to phoneme-duration-based systems like Voicebox persists even at the largest model
    scale tested, suggesting that removing domain-specific factors trades off some speaker-identity preservation
    for simplicity and speed. The paper also does not explore natural-language instruction following or fine-grained
    prosody control, both noted as future directions rather than addressed in this work. Multilingual results are
    reported on only 100 held-out examples per language, which is a narrow evaluation slice given nine languages
    with very different phonological properties.
  caveats: []
- id: iclr-2025-tQ1PmLfPBL
  published_date: "2025-04-24"
  entry_date: '2026-07-28'
  year: 2025
  venue: ICLR
  task:
  - TTS
  architecture:
  - flow-matching
  relevance: medium
  evidence_role:
  - architecture_variant
  - acceleration_evidence
  current_role: influential
  method_family:
  - flow_matching_codec_decoders
  claims:
  - claim_id: flow_matching_enables_higher_quality_waveform_generation_than_diffusion_with
    role: supports
    claim: Flow matching enables higher-quality waveform generation than diffusion with fewer inference steps.
    source: §4.6, Table 8
    evidence: PeriodWave with CFM and 6 steps achieves UTMOS 3.628 versus PeriodWave with DDPM at 50 steps achieving
      UTMOS 3.377 on LibriTTS; the CFM model also reaches comparable quality to 50-step DDPM at only 16 steps.
    confidence: high
    relevance: medium
  - claim_id: explicit_multi_period_decomposition_in_the_generator_architecture_improves_pitch
    role: supports
    claim: Explicit multi-period decomposition in the generator architecture improves pitch accuracy and periodicity
      over GAN and diffusion vocoders.
    source: §4.2, §4.6, Table 1, Table 7
    evidence: PeriodWave achieves pitch error of 15.04 cents and periodicity 0.0744 on LJSpeech, substantially below
      BigVGAN (19.02 cents, 0.0782) and all diffusion baselines; ablation shows monotonic improvement as more distinct
      prime-number periods are added.
    confidence: high
    relevance: medium
  - claim_id: single_step_gan_vocoders_achieve_significantly_faster_inference_than_iterative
    role: complicates
    claim: Single-step GAN vocoders achieve significantly faster inference than iterative flow-matching vocoders.
    source: §E, Table 15, Table 16
    evidence: PeriodWave at 16 steps runs at 7.48× real-time; HiFi-GAN runs at 166.70× real-time; even PeriodWave
      at 2 steps (56.36×) is slower than HiFi-GAN, though it already outperforms HiFi-GAN on all quality metrics.
    confidence: high
    relevance: medium
  - claim_id: iterative_waveform_generation_reduces_train_inference_mismatch_artefacts_in_two
    role: supports
    claim: Iterative waveform generation reduces train-inference mismatch artefacts in two-stage TTS relative to
      one-step GAN vocoders.
    source: §4.7, Table 9, §G
    evidence: In zero-shot TTS with ARDiT-TTS acoustic features, PeriodWave+FreeU achieves MOS 4.07 versus BigVGAN's
      4.03 and BigVSAN's 3.99; the iterative refinement allows the vocoder to correct imperfections in generated
      Mel-spectrograms rather than propagating them.
    confidence: high
    relevance: medium
  - claim_id: flow_matching_vocoders_can_decode_neural_codec_tokens_with_streaming
    role: supports
    claim: Flow-matching vocoders can decode neural codec tokens with streaming generation and minimal quality degradation.
    source: §5, Table 10, Table 11
    evidence: PeriodWave trained for parallel generation from Mimi (Q=8) tokens achieves CER 2.5% versus Mimi decoder's
      3.07%; streaming with single-token delay and 2-step sampling maintains comparable quality (CER 2.45%, UTMOS
      3.85 versus parallel 3.93).
    confidence: high
    relevance: high
  limitations:
  - 'Synthesis speed is the principal limitation: at 16 steps, PeriodWave runs at 7.48× real-time, substantially
    slower than one-step GAN vocoders (HiFi-GAN: 166×, BigVGAN: 38×). For latency-sensitive applications, the 2-step
    variant (56×) is practical but still 3× slower than HiFi-GAN.'
  - High-frequency reproduction remains challenging even with multi-band modeling and FreeU. The single-loss (CFM
    objective only) training means the model lacks the spectral feedback that GAN discriminators provide, and M-STFT
    metrics are generally worse than GAN baselines despite better perceptual scores. The authors plan to incorporate
    short-time Fourier convolution blocks or modified spectral objectives to address this.
  - The codec streaming mode uses a non-causal architecture with a one-token look-ahead, which introduces a small
    latency penalty. In-context streaming generation for longer sequences is identified as future work. Evaluation
    is confined to English speech and music; generalisation to other languages, accents, and audio domains remains
    untested in this paper.
  caveats: []
- id: iclr-2025-uxDFlPGRLX
  published_date: "2025-04-24"
  entry_date: '2026-07-28'
  year: 2025
  venue: ICLR
  task:
  - codec
  architecture:
  - flow-matching
  - GAN
  relevance: medium
  evidence_role:
  - architecture_variant
  - evaluation_caution
  current_role: active_evidence
  method_family:
  - flow_matching_codec_decoders
  - gan_neural_codecs_and_decoders
  claims:
  - claim_id: a_flow_matching_based_stochastic_postfilter_conditioned_on_a_deterministic
    role: supports
    claim: A flow matching-based stochastic postfilter conditioned on a deterministic codec's decoder output can
      replace adversarial training while achieving comparable subjective quality to a GAN-based codec.
    source: §5.2, Figure 6
    evidence: MUSHRA listening tests (Test A, 11 expert listeners) show no significant difference between FlowDec-75m
      and DAC-75 score distributions at matched bitrates of 4.5 and 7.5 kbit/s.
    confidence: high
    relevance: high
  - claim_id: coupling_the_flow_matching_source_distribution_to_the_conditioning_signal
    role: supports
    claim: Coupling the flow matching source distribution to the conditioning signal (rather than sampling it independently)
      removes the need for minibatch optimal-transport solvers and improves postfilter sample quality at low inference
      budgets.
    source: §5.1, Table 4
    evidence: At NFE=6, the proposed coupled formulation achieves FAD×100 of 1.62 versus 145.3 for the diffusion-based
      ScoreDec baseline and approximately 29 for an alternative constant-σ flow matching formulation, on the same
      underlying codec and test set.
    confidence: high
    relevance: medium
  - claim_id: improving_perceptual_distance_metrics_fad_via_generative_postfiltering_trades_off
    role: complicates
    claim: Improving perceptual distance metrics (FAD) via generative postfiltering trades off against intrusive
      distortion metrics relative to discriminator-trained codecs.
    source: §5.1, Figure 4, Figure 5
    evidence: Retrained non-adversarial DAC (NDAC) generally outperforms FlowDec on SI-SDR and fwSSNR even though
      FlowDec achieves better FAD, consistent with the perception-distortion tradeoff; the gap is small in the perceptually
      weighted fwSSNR but clear in SI-SDR.
    confidence: high
    relevance: medium
  - claim_id: generative_postfilters_trained_with_vanilla_score_or_flow_matching_formulations
    role: complicates
    claim: Generative postfilters trained with vanilla score- or flow-matching formulations using a fixed-variance
      or independent prior can fail to converge to the target signal or require expensive multi-step inference to
      reach acceptable quality.
    source: §5.1, Table 4
    evidence: ScoreDec (diffusion-based postfilter) produces unusable audio quality at a reduced inference budget
      of 6 function evaluations (FAD×100 = 145.3, SI-SDR = -27.23), only becoming competitive at roughly 50 evaluations;
      a constant-σ flow matching variant also underperforms the proposed coupled formulation at NFE=6.
    confidence: high
    relevance: medium
  limitations:
  - The proposed codec, like the DAC baseline it builds on, uses a noncausal architecture and is not streaming-capable,
    limiting applicability to real-time communication settings where low-latency, causal processing is required.
  - The authors note this could be addressed with a causal DNN design in future work, but no such variant is evaluated
    in this paper. Additionally, the postfilter and underlying codec are trained in two separate stages; joint training
    is identified as a possible quality improvement but is left unexplored due to potential training instability.
    The NCSN++ backbone used for the postfilter was originally designed for images, and the authors suggest audio-specific
    architectures could further improve quality, indicating the current network design is not optimized end-to-end
    for the audio domain. Listening test sample sizes (11 and 10 raters) and clip counts (21 ten-second clips) are
    modest for the audio-type breakdown analysis reported in the appendix.
  caveats: []
- id: 2025.findings-naacl.184
  published_date: "2025-04-29"
  entry_date: '2026-07-28'
  year: 2025
  venue: NAACL
  task:
  - TTS
  - codec
  architecture:
  - autoregressive-LM
  - VAE
  - flow-matching
  relevance: high
  evidence_role:
  - core_evidence
  - control_evidence
  current_role: active_evidence
  method_family:
  - autoregressive_codec_language_models
  - vae_vector_quantized_codecs
  - flow_matching_codec_decoders
  claims:
  - claim_id: continuous_speech_token_representations_preserve_high_frequency_information_more_effectively
    role: supports
    claim: Continuous speech token representations preserve high-frequency information more effectively than discrete
      RVQ tokens, with the gap largest in the 8kHz band.
    source: §3.3, Table 3
    evidence: Continuous speech token representations preserve high-frequency information more effectively than
      discrete RVQ tokens, with the gap largest in the 8kHz band.
    confidence: high
    relevance: high
  - claim_id: replacing_rvq_with_a_continuous_tokenizer_in_an_autoregressive_tts
    role: supports
    claim: Replacing RVQ with a continuous tokenizer in an autoregressive TTS framework reduces word error rate
      and improves speaker similarity compared to a discrete-token baseline.
    source: §3.2, Table 1
    evidence: Replacing RVQ with a continuous tokenizer in an autoregressive TTS framework reduces word error rate
      and improves speaker similarity compared to a discrete-token baseline.
    confidence: high
    relevance: high
  - claim_id: continuous_speech_tokenizers_exhibit_greater_robustness_to_sampling_rate_variation
    role: supports
    claim: Continuous speech tokenizers exhibit greater robustness to sampling rate variation than discrete tokenizers,
      particularly when window length ratios are modified.
    source: §3.4
    evidence: Continuous speech tokenizers exhibit greater robustness to sampling rate variation than discrete tokenizers,
      particularly when window length ratios are modified.
    confidence: high
    relevance: medium
  - claim_id: two_stage_training_tokenizer_pre_training_followed_by_joint_lm
    role: supports
    claim: Two-stage training (tokenizer pre-training followed by joint LM-tokenizer training) is necessary for
      continuous-token TTS; removing either stage degrades both intelligibility and speaker similarity.
    source: §2.4.3, Appendix A.1, Table 4
    evidence: Two-stage training (tokenizer pre-training followed by joint LM-tokenizer training) is necessary for
      continuous-token TTS; removing either stage degrades both intelligibility and speaker similarity.
    confidence: high
    relevance: low
  - claim_id: autoregressive_generation_over_continuous_speech_tokens_reduces_temporal_discontinuity_artefacts
    role: supports
    claim: Autoregressive generation over continuous speech tokens reduces temporal discontinuity artefacts relative
      to RVQ-based generation, as measured by NISQA continuity scores.
    source: §3.2, Table 2
    evidence: Autoregressive generation over continuous speech tokens reduces temporal discontinuity artefacts relative
      to RVQ-based generation, as measured by NISQA continuity scores.
    confidence: high
    relevance: high
  limitations:
  - '- Evaluation is English-only on LibriSpeech test-clean, a relatively clean/read-speech benchmark; generalization
    to spontaneous or multilingual speech is untested. - Model size is not reported, making cost comparisons with
    VALL-E/MELLE impossible. - The continuous LM output requires MSE regression rather than cross-entropy classification;
    training stability and mode averaging in the continuous space are not thoroughly discussed. - No subjective
    MOS evaluation from human raters; EMoS (automated MOS) is used instead. - The relationship to MELLE (which also
    drops discrete tokens, via mel-spectrogram prediction) deserves more analysis — they represent two distinct
    approaches to the same problem. - How the continuous tokenizer interacts with multi-speaker generalization is
    not tested.'
  caveats: []
- id: 2025.naacl-demo.12
  published_date: "2025-04-29"
  entry_date: '2026-07-28'
  year: 2025
  venue: NAACL
  task:
  - TTS
  - SCA
  architecture:
  - autoregressive-LM
  relevance: high
  evidence_role:
  - core_evidence
  - control_evidence
  - infrastructure
  current_role: active_evidence
  method_family:
  - autoregressive_codec_language_models
  claims:
  - claim_id: a_speech_language_model_can_be_initialized_from_a_pre
    role: supports
    claim: A speech language model can be initialized from a pre-trained text LLM and jointly trained on speech
      recognition, speech synthesis, text continuation, and audio continuation without substantially degrading the
      text-only capability of the base model.
    source: §4.3, Table 5
    evidence: The 1.7B multi-task model trained on ASR, TTS, TextLM, and AudioLM objectives scores MMLU 30.5, ARC-C
      41.3, and HellaSwag 50.4, close to the text-only LLaMA-3.2-1B baseline (32.2, 32.8, 41.2) despite carrying
      three additional speech tasks.
    confidence: high
    relevance: medium
  - claim_id: concatenating_neural_codec_tokens_with_self_supervised_speech_representations_frame
    role: supports
    claim: Concatenating neural codec tokens with self-supervised speech representations frame-by-frame is a viable
      tokenization strategy for both speech understanding and generation tasks within a single sequential model.
    source: §3.3
    evidence: The "Codec_SSL" scheme (ESPnet-Codec combined with XEUS SSL tokens) is used for the headline ASR and
      TTS experiments and the paper reports it "behaves well in both speech understanding and generation" *(§3.3)*.
    confidence: high
    relevance: high
  - claim_id: a_decoder_only_autoregressive_speech_language_model_can_match_or
    role: supports
    claim: A decoder-only autoregressive speech language model can match or exceed dedicated, larger ASR-only systems
      on English benchmarks while using substantially fewer parameters.
    source: §4.2, Table 3
    evidence: A 442M-parameter ESPnet-SpeechLM ASR model reaches average WER 5.4% across six English test sets,
      matching OWSM v3.1-medium (1.02B, 5.4%) and beating Whisper-small (244M, 6.4%) and Whisper-medium (769M, 5.7%).
    confidence: high
    relevance: medium
  - claim_id: cross_system_comparisons_of_speech_language_models_reported_in_the
    role: complicates
    claim: Cross-system comparisons of speech language models reported in the literature are frequently not run
      under matched conditions, limiting how much can be concluded from any single performance table.
    source: §4.3, Table 5
    evidence: In the multi-task comparison (Table 5), competitor numbers for Moshi, VITA, GLM-4-Voice, and others
      are taken directly from their own published reports rather than reproduced by the authors, and the paper explicitly
      flags this with footnote markers.
    confidence: high
    relevance: medium
  - claim_id: combining_a_codec_tokenizer_with_a_self_supervised_tokenizer_frame
    role: complicates
    claim: Combining a codec tokenizer with a self-supervised tokenizer frame-by-frame is reported as an effective
      design choice but is not validated against single-tokenizer ablations in the same controlled setting.
    source: §3.3
    evidence: The claim that Codec_SSL tokenization "behaves well" rests on a single line of justification without
      a paired ablation against codec-only or SSL-only tokenization on the same task and dataset.
    confidence: high
    relevance: high
  limitations:
  - The multi-task model's TTS quality (Proxy MOS 3.99, WER 6.0%) is noticeably weaker than the single-task TTS
    model trained on the same architecture (Proxy MOS 4.03, WER 3.1%), indicating a capacity or interference cost
    to joint multi-task training that the paper reports but does not analyze further. Most training and evaluation
    data is English-only (the multilingual text corpus is used only for the TextLM objective, not for speech tasks),
    so the demonstrated speech capabilities are not evidence of multilingual generalization. Competitor numbers
    in the multi-task comparison table are drawn from third-party reports under unmatched training data and conditions
    rather than reproduced by the authors, which the paper itself notes. As a system/demo paper, the contribution
    is the toolkit and its reference recipes rather than a novel architecture or training method; the headline numbers
    serve to validate functionality rather than push state of the art on any individual benchmark.
  caveats: []
- id: 2025.naacl-long.591
  published_date: "2025-04-29"
  entry_date: '2026-07-28'
  year: 2025
  venue: NAACL
  task:
  - TTS
  architecture:
  - transformer-enc-dec
  relevance: medium
  evidence_role:
  - architecture_variant
  - evaluation_caution
  current_role: active_evidence
  method_family:
  - transformer_encoder_decoder_tokenizers
  claims:
  - claim_id: monotonic_alignment_can_be_learned_end_to_end_as_a
    role: supports
    claim: Monotonic alignment can be learned end-to-end as a latent property of an encoder-decoder TTS model via
      backpropagation, without requiring forced alignments or dynamic programming during training.
    source: §3.3, §3.4
    evidence: VAT's alignment layer learns continuous alignment positions through interpolated relative position
      biases (IRPBs); alignment trajectories emerge from joint training with no external supervision, and generalise
      to utterances far longer than the training distribution.
    confidence: high
    relevance: medium
  - claim_id: augmenting_cross_attention_with_a_learned_monotonic_alignment_position_enables
    role: supports
    claim: Augmenting cross-attention with a learned monotonic alignment position enables unbounded length generalisation
      in encoder-decoder TTS without degrading naturalness relative to an unmodified Transformer baseline.
    source: §5.1, §5.2, §5.3, Table 1
    evidence: VAT achieves near-zero CER on inputs up to 1500 characters (~90 seconds) despite training only on
      utterances up to 9.6 seconds, while matching the T5 baseline in side-by-side naturalness evaluations (SxS
      -0.06 ± 0.14 on Lessac, 0.01 ± 0.14 on LibriTTS).
    confidence: high
    relevance: medium
  - claim_id: standard_mos_evaluations_are_insufficient_to_surface_robustness_failures_in
    role: complicates
    claim: Standard MOS evaluations are insufficient to surface robustness failures in autoregressive TTS because
      raters cannot detect dropped or repeated words without access to target transcripts.
    source: §5.1, §5.2, Table 1
    evidence: The T5 baseline achieves overlapping MOS with VAT (3.75 vs. 3.68 on Lessac) while producing a CER
      of 10.2 versus VAT's 3.3; the perceptual quality rating is statistically indistinguishable despite systematic
      robustness failures.
    confidence: high
    relevance: medium
  - claim_id: autoregressive_transformer_tts_without_explicit_alignment_guidance_fails_on_repeated
    role: complicates
    claim: Autoregressive Transformer TTS without explicit alignment guidance fails on repeated words even within
      training sequence length limits.
    source: §5.4
    evidence: The T5 baseline makes errors on 14 of 27 (52%) repeated-word test phrases, including phrases with
      as few as 2 repetitions of a single word, while VAT makes zero errors across all 27 templates.
    confidence: high
    relevance: medium
  - claim_id: duration_based_tts_achieves_better_asr_measured_character_error_rates
    role: refines
    claim: Duration-based TTS achieves better ASR-measured character error rates than expressive autoregressive
      models, but the gain is attributable to hyper-intelligible, monotone synthesis rather than superior text coverage.
    source: §5.3, Table 1
    evidence: NAT achieves CER 3.3 on LibriTTS, below both VAT (4.6) and ground truth (3.6), yet VAT is preferred
      over NAT in naturalness side-by-sides because NAT's unsupervised duration predictor produces robotic, monotonous
      prosody.
    confidence: high
    relevance: medium
  limitations:
  - Training speed is affected by the need to compute alignment positions serially during training, imposing a 12-20%
    slowdown relative to the T5 baseline depending on model scale. All experiments use English and a speaker-conditioned
    (non-zero-shot) setting; generalisation to other languages and to audio-prompted zero-shot scenarios is untested.
    Evaluation compares against T5, Tacotron-GMMA, and NAT; no direct comparison with codec LM systems (VALL-E,
    SPEAR-TTS, MQTTS) is provided, which the paper attributes to incompatible dataset scales and evaluation protocols.
    Hyper-parameter choices for the alignment layer, IRPB initialization, and maximum distance penalty are reported
    but not systematically ablated.
  caveats: []
- id: 2025.naacl-srw.6
  published_date: "2025-04-29"
  entry_date: '2026-07-28'
  year: 2025
  venue: NAACL
  task:
  - TTS
  - codec
  architecture:
  - autoregressive-LM
  - hybrid
  relevance: medium
  evidence_role:
  - architecture_variant
  - acceleration_evidence
  - control_evidence
  - evaluation_caution
  current_role: active_evidence
  method_family:
  - autoregressive_codec_language_models
  - hybrid_semantic_acoustic_tokenizers
  claims:
  - claim_id: non_overlapping_encoder_receptive_fields_in_rvq_codecs_improve_downstream
    role: supports
    claim: Non-overlapping encoder receptive fields in RVQ codecs improve downstream language model likelihood and
      end-to-end TTS metrics relative to the standard causal overlapping setup.
    source: §3.1, Table 1
    evidence: Replacing the overlapping causal encoder with a framewise encoder on a DAC-based codec reduces LM
      NLL by more than 8% and improves WER, NISQA, and speaker similarity on LibriTTS-R test-clean, despite slightly
      worsening Mel-L1 reconstruction.
    confidence: high
    relevance: high
  - claim_id: better_codec_audio_reconstruction_quality_does_not_reliably_predict_better
    role: complicates
    claim: Better codec audio reconstruction quality does not reliably predict better end-to-end speech generation
      quality in codec-LM systems.
    source: §4, Table 1
    evidence: The framewise encoder achieves higher NLL and better TTS metrics than the causal baseline while scoring
      slightly worse on Mel-spectral L1 reconstruction distance, demonstrating that reconstruction-optimised codecs
      can be suboptimal for downstream LM training.
    confidence: high
    relevance: high
  - claim_id: increasing_rvq_codec_frame_duration_can_substantially_reduce_codec_lm
    role: supports
    claim: Increasing RVQ codec frame duration can substantially reduce codec-LM inference latency with little or
      no degradation in TTS intelligibility and speaker similarity, provided the bitrate is held approximately constant
      by adjusting codebook depth.
    source: §3.3, §4, Table 2
    evidence: Doubling frame duration from 11ms to 22ms yields a 1.94x inference speedup with WER 4.21%, NISQA 4.42,
      and speaker similarity 81.0%, matching or improving on the 11ms framewise baseline. Quadrupling to 44ms further
      accelerates inference (3.2-3.8x) but substantially degrades WER and speaker similarity.
    confidence: high
    relevance: high
  - claim_id: a_single_lm_trained_with_codebook_level_dropout_can_efficiently
    role: supports
    claim: A single LM trained with codebook level dropout can efficiently approximate the performance profile of
      training one LM per candidate RVQ level count.
    source: §3.2, §4, Figure 2
    evidence: Training with 90%-full CL drop on a 12-level codec produces per-level performance curves for WER,
      NISQA, and speaker similarity that closely track those of 12 independently trained LMs across all Q' values
      1-12.
    confidence: high
    relevance: high
  - claim_id: the_optimal_number_of_rvq_codebook_levels_for_end_to
    role: refines
    claim: The optimal number of RVQ codebook levels for end-to-end codec-LM TTS differs across evaluation dimensions,
      and more levels are not universally better for end-to-end performance even when they monotonically improve
      codec reconstruction.
    source: Appendix B, Figure 3
    evidence: End-to-end FAD reaches a global minimum at 9 levels before degrading, while WER reaches its best at
      3-4 levels and NISQA and speaker similarity peak at approximately 9 levels, in contrast to codec Mel-L1 which
      improves monotonically.
    confidence: high
    relevance: high
  limitations:
  - The evaluation is conducted on a single English TTS corpus (LibriTTS-R test-clean) using automatic metrics only
    (WER via Whisper, NISQA, cosine speaker similarity). No human listening test is reported, so perceptual quality
    gains are estimated rather than directly validated.
  - The codec is trained on proprietary in-house podcast data (1.7K hours), which limits reproducibility for the
    codec training stage specifically. The LM training does use the public LibriTTS-R dataset. The codebook size
    hyperparameter (|V|) remains outside the scope of CL drop, requiring separate trial-and-error search. The paper
    does not investigate multilingual or noisy speech settings. The optimal frame duration finding (22ms being the
    sweet spot) is specific to this codec architecture and training data and may not generalise to other codec families.
  caveats: []
- id: '2505.02625'
  published_date: "2025-05-05"
  entry_date: '2026-07-28'
  year: 2025
  venue: arXiv
  task:
  - SCA
  architecture:
  - autoregressive-LM
  - flow-matching
  - hybrid
  relevance: medium
  evidence_role:
  - architecture_variant
  - acceleration_evidence
  - control_evidence
  current_role: active_evidence
  method_family:
  - autoregressive_codec_language_models
  - flow_matching_codec_decoders
  - hybrid_semantic_acoustic_tokenizers
  claims:
  - claim_id: autoregressive_streaming_tts_decoders_in_modular_spoken_conversational_agents_produce
    role: supports
    claim: Autoregressive streaming TTS decoders in modular spoken conversational agents produce higher-quality
      speech and smaller S2T-to-S2S accuracy gaps than non-autoregressive streaming decoders, at a modest latency
      cost.
    source: §5.1, Table 1
    evidence: Autoregressive streaming TTS decoders in modular spoken conversational agents produce higher-quality
      speech and smaller S2T-to-S2S accuracy gaps than non-autoregressive streaming decoders, at a modest latency
      cost.
    confidence: high
    relevance: low
  - claim_id: modular_spoken_conversational_agents_trained_on_tens_of_thousands_of
    role: supports
    claim: Modular spoken conversational agents trained on tens of thousands of hours of synthesized speech-to-speech
      dialogue can match the performance of native SpeechLMs trained on millions of hours of unsupervised speech
      data.
    source: §5.1, Table 1
    evidence: Modular spoken conversational agents trained on tens of thousands of hours of synthesized speech-to-speech
      dialogue can match the performance of native SpeechLMs trained on millions of hours of unsupervised speech
      data.
    confidence: high
    relevance: low
  - claim_id: jointly_conditioning_a_tts_language_model_on_llm_hidden_states
    role: supports
    claim: Jointly conditioning a TTS language model on LLM hidden states and text token embeddings via a learned
      gate fusion improves both semantic consistency (WER) and instruction-following quality over hidden-state-only
      conditioning.
    source: §5.2, Table 2
    evidence: Jointly conditioning a TTS language model on LLM hidden states and text token embeddings via a learned
      gate fusion improves both semantic consistency (WER) and instruction-following quality over hidden-state-only
      conditioning.
    confidence: high
    relevance: medium
  - claim_id: tts_language_model_pretraining_on_text_speech_pairs_is_a
    role: supports
    claim: TTS language model pretraining on text-speech pairs is a critical prerequisite for stable convergence
      in modular SpeechLMs; initializing from a language model alone is insufficient.
    source: §5.2, Table 3
    evidence: TTS language model pretraining on text-speech pairs is a critical prerequisite for stable convergence
      in modular SpeechLMs; initializing from a language model alone is insufficient.
    confidence: high
    relevance: medium
  - claim_id: in_streaming_speech_generation_the_write_chunk_size_w_primarily
    role: supports
    claim: In streaming speech generation, the write chunk size (W) primarily determines speech naturalness while
      the read chunk size (R) primarily determines text-speech alignment, with latency jointly determined by both.
    source: §5.2, Table 4
    evidence: In streaming speech generation, the write chunk size (W) primarily determines speech naturalness while
      the read chunk size (R) primarily determines text-speech alignment, with latency jointly determined by both.
    confidence: high
    relevance: low
  limitations:
  - All evaluations are conducted in English only. The model is trained on a single fixed output voice, and no multilingual
    or voice-diversity experiments are reported. Generalization to other languages or to emotionally expressive
    speech is untested.
  - The model cannot modulate speech style (emotion, speaking rate, dialect) in response to paralinguistic cues
    in the input, because training data contains only conventional speech-to-speech dialogue. The authors note this
    as planned future work. The benchmarks used (SpokenQA accuracy, ChatGPT score) are narrow and do not cover naturalness
    in unconstrained conversational settings or robustness to noisy or accented input. Comparisons to Minmo (a concurrent
    work using 1.4M hours) are mentioned in related work but not included in the main experimental table, leaving
    the data efficiency claim partially unverified against the most directly comparable system.
  caveats: []
- id: '2505.13000'
  published_date: "2025-05-19"
  entry_date: '2026-07-28'
  year: 2025
  venue: arXiv
  task:
  - codec
  - TTS
  architecture:
  - autoregressive-LM
  relevance: medium
  evidence_role:
  - architecture_variant
  - acceleration_evidence
  - control_evidence
  current_role: influential
  method_family:
  - autoregressive_codec_language_models
  claims:
  - claim_id: directly_encoding_ssl_features_into_the_first_rvq_layer_preserves
    role: supports
    claim: Directly encoding SSL features into the first RVQ layer preserves significantly more semantic content
      than distilling SSL representations into codec tokens, particularly for tonal languages where pitch information
      is phonemically critical.
    source: §4.2, Table 2
    evidence: Directly encoding SSL features into the first RVQ layer preserves significantly more semantic content
      than distilling SSL representations into codec tokens, particularly for tonal languages where pitch information
      is phonemically critical.
    confidence: high
    relevance: high
  - claim_id: neural_audio_codecs_operating_at_lower_frame_rates_with_more
    role: supports
    claim: Neural audio codecs operating at lower frame rates with more quantization layers achieve superior audio
      quality at equivalent bitrates compared to higher frame rate codecs with fewer quantization layers.
    source: §4.3, Table 3
    evidence: Neural audio codecs operating at lower frame rates with more quantization layers achieve superior
      audio quality at equivalent bitrates compared to higher frame rate codecs with fewer quantization layers.
    confidence: high
    relevance: medium
  - claim_id: semantic_enhancement_of_the_first_rvq_layer_improves_downstream_tts
    role: supports
    claim: Semantic enhancement of the first RVQ layer improves downstream TTS speaker similarity as well as intelligibility,
      because higher semantic fidelity in RVQ-1 enables the waveform stream's quantizers to focus on acoustic detail
      rather than recovering content information.
    source: §4.4, Table 4
    evidence: Semantic enhancement of the first RVQ layer improves downstream TTS speaker similarity as well as
      intelligibility, because higher semantic fidelity in RVQ-1 enables the waveform stream's quantizers to focus
      on acoustic detail rather than recovering content information.
    confidence: high
    relevance: high
  - claim_id: the_quality_gap_between_distillation_based_and_direct_encoding_semantic
    role: supports
    claim: The quality gap between distillation-based and direct-encoding semantic codecs is substantially larger
      in Mandarin than in English, revealing a systematic limitation of distillation approaches for tonal languages.
    source: §4.2, Table 2
    evidence: The quality gap between distillation-based and direct-encoding semantic codecs is substantially larger
      in Mandarin than in English, revealing a systematic limitation of distillation approaches for tonal languages.
    confidence: high
    relevance: medium
  limitations:
  - The 12.5Hz DualCodec-based TTS systems consistently underperform their 25Hz counterparts on both WER and speaker
    similarity (Table 4, Table 6), indicating that the quality upper bound of the 12.5Hz variant is not yet competitive
    with the best open-source systems at 50Hz despite the frame rate reduction improving inference speed.
  - The paper evaluates TTS only on Seed-TTS-Eval; no subjective TTS listening tests are reported, so the MUSHRA
    gains in codec reconstruction may not fully translate to perceived TTS naturalness. Speaker similarity scores
    with DualCodec-VALLE remain below those of MaskGCT baselines that use separate semantic and acoustic tokenizers,
    suggesting the unified approach has not yet matched the best-performing two-stage pipeline design. The DualCodec
    encoder is substantially heavier than baselines (628M vs 38M for Mimi) due to the frozen w2v-BERT-2.0 model,
    increasing training-time compute, though the decoder remains lightweight for inference.
  caveats: []
- id: '2506.10274'
  published_date: "2025-06-12"
  entry_date: '2026-07-28'
  year: 2025
  venue: arXiv
  task:
  - codec
  - TTS
  - evaluation
  architecture: []
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  - evaluation_caution
  - infrastructure
  current_role: influential
  method_family: []
  claims:
  - claim_id: no_discrete_audio_tokenizer_consistently_outperforms_others_across_reconstruction_downstream
    role: supports
    claim: No discrete audio tokenizer consistently outperforms others across reconstruction, downstream discriminative
      tasks, and acoustic language modeling; the optimal tokenizer is task- and domain-dependent.
    source: §3.4, Figure 4
    evidence: No discrete audio tokenizer consistently outperforms others across reconstruction, downstream discriminative
      tasks, and acoustic language modeling; the optimal tokenizer is task- and domain-dependent.
    confidence: high
    relevance: high
  - claim_id: increasing_the_number_of_codebooks_improves_signal_reconstruction_quality_but
    role: supports
    claim: Increasing the number of codebooks improves signal reconstruction quality but often degrades downstream
      task performance by adding redundancy that burdens representation-level learning.
    source: §3.2, §3.2 "Impact of Codebook Size"
    evidence: Increasing the number of codebooks improves signal reconstruction quality but often degrades downstream
      task performance by adding redundancy that burdens representation-level learning.
    confidence: high
    relevance: high
  - claim_id: semantic_distillation_aligning_early_rvq_layers_with_ssl_features_improves
    role: supports
    claim: Semantic distillation (aligning early RVQ layers with SSL features) improves phonetic content preservation
      in acoustic tokenizers but may reduce cross-domain generalization when the distillation source is speech-specific.
    source: §4.2 "Distillation Effect", Table 16
    evidence: Semantic distillation (aligning early RVQ layers with SSL features) improves phonetic content preservation
      in acoustic tokenizers but may reduce cross-domain generalization when the distillation source is speech-specific.
    confidence: high
    relevance: high
  - claim_id: domain_alignment_between_tokenizer_training_data_and_evaluation_domain_is
    role: supports
    claim: Domain alignment between tokenizer training data and evaluation domain is the dominant factor in discrete
      audio codec performance, outweighing quantization method or bitrate choices.
    source: §4.2 "Data Domains"
    evidence: Domain alignment between tokenizer training data and evaluation domain is the dominant factor in discrete
      audio codec performance, outweighing quantization method or bitrate choices.
    confidence: high
    relevance: high
  - claim_id: continuous_speech_representations_e_g_wavlm_large_consistently_outperform_all
    role: supports
    claim: Continuous speech representations (e.g., WavLM-large) consistently outperform all discrete tokenizers
      on discriminative tasks, with the gap widening in low-resource conditions.
    source: §3.2, Table 7
    evidence: Continuous speech representations (e.g., WavLM-large) consistently outperform all discrete tokenizers
      on discriminative tasks, with the gap widening in low-resource conditions.
    confidence: high
    relevance: medium
  limitations:
  - The TTS and audio LM evaluations are conducted with constrained training budgets (VALL-E on LibriTTS only; 300M
    audio LM at half the original training compute), which limits the practical transferability of findings on TTS
    tokenizer ranking to large-scale production settings.
  - The acoustic LM evaluation uses a single architecture (Qwen-2.5 based, 357M) and a single dataset (LibriHeavy),
    so findings on which tokenizers support better SLMs may not generalise to other LM architectures or training
    scales. The ablation study (Section 4) is restricted to a DAC-backbone framework; FSQ and SVQ conclusions may
    not transfer to other encoder-decoder designs. Evaluation metrics for audio generation (FAD, KLD, CLAP) conflate
    vocoder quality with language model quality, making it difficult to attribute performance differences to the
    tokenizer's representational properties versus its decoder quality. The paper does not evaluate streaming tokenizers
    under actual latency constraints, limiting guidance for real-time deployment. Trustworthiness considerations
    (voice deepfakes, bias) are raised as open concerns but not empirically evaluated.
  caveats: []
- id: '2507.02380'
  published_date: "2025-07-03"
  entry_date: '2026-07-28'
  year: 2025
  venue: arXiv
  task:
  - SCA
  - TTS
  architecture:
  - autoregressive-LM
  - flow-matching
  relevance: medium
  evidence_role:
  - architecture_variant
  - acceleration_evidence
  - control_evidence
  current_role: minor
  method_family:
  - autoregressive_codec_language_models
  - flow_matching_codec_decoders
  claims:
  - claim_id: routing_llm_hidden_states_into_the_tts_module_s_embedding
    role: supports
    claim: Routing LLM hidden states into the TTS module's embedding space can enable voice cloning in an end-to-end
      spoken chatbot without a dedicated speaker encoder.
    source: §1.2
    evidence: JoyTTS projects Qwen-7B hidden states via an MLP into the CosyVoice2-based LLM-TTS embedding, achieving
      SS of 0.73 on seed-tts-zh using prompt audio as the only speaker reference.
    confidence: high
    relevance: medium
  - claim_id: jointly_training_an_llm_chat_module_with_a_tts_module
    role: complicates
    claim: Jointly training an LLM-Chat module with a TTS module in a chatbot pipeline can degrade intelligibility
      relative to running the TTS component standalone, even when speaker similarity improves.
    source: §4, Table 1
    evidence: JoyTTS achieves WER of 5.09 compared to 1.45 for standalone CosyVoice2 on seed-tts-zh, despite closing
      the speaker similarity gap (JoyTTS SS 0.73 vs. CosyVoice2 SS 0.748).
    confidence: high
    relevance: low
  - claim_id: end_to_end_spoken_chatbot_systems_pairing_a_7b_parameter
    role: supports
    claim: End-to-end spoken chatbot systems pairing a 7B-parameter LLM with an autoregressive TTS module can achieve
      sub-2-second response latency on a single consumer GPU without specialised inference optimisations.
    source: §4
    evidence: JoyTTS reports 1.8-second end-to-end latency on a single NVIDIA 4090D with no engineering optimisation
      applied.
    confidence: high
    relevance: low
  limitations:
  - The evaluation is limited to a single Chinese benchmark (seed-tts-zh), leaving performance on English, multilingual,
    or spontaneous conversational speech uncharacterised. The WER gap between JoyTTS (5.09) and standalone CosyVoice2
    (1.45) is large and unexplained; the paper does not ablate whether the regression originates from the joint
    training procedure, the conversational data distribution, or the MLP projection coupling. No subjective listening
    tests are reported, making it impossible to assess naturalness or perceived quality beyond intelligibility and
    speaker similarity metrics. Emotion control, identified as a target for future work, is not implemented in the
    current system.
  caveats: []
- id: '2507.01348'
  published_date: "2025-07-08"
  entry_date: '2026-07-28'
  year: 2025
  venue: arXiv
  task:
  - VC
  - TTS
  architecture:
  - autoregressive-LM
  - VAE
  - GAN
  relevance: medium
  evidence_role:
  - architecture_variant
  current_role: active_evidence
  method_family:
  - autoregressive_codec_language_models
  - vae_vector_quantized_codecs
  - gan_neural_codecs_and_decoders
  claims:
  - claim_id: ctc_regularization_applied_before_vector_quantization_improves_the_temporal_locality
    role: supports
    claim: CTC regularization applied before vector quantization improves the temporal locality and temporal robustness
      of discrete speech content tokens.
    source: §5.2, Table 2
    evidence: SpeechCodeVAE achieves 59% higher De-duplication Efficiency and approximately 9 times better Speed
      Robustness than CosyVoice-50Hz; ablation without CTC loss collapses Speed Robustness from 0.219 to 0.009,
      identifying CTC as the critical factor.
    confidence: high
    relevance: high
  - claim_id: multi_task_learning_with_tts_as_an_auxiliary_objective_compensates
    role: supports
    claim: Multi-task learning with TTS as an auxiliary objective compensates for data scarcity in foreign accent
      conversion, improving both convergence and output quality.
    source: §3, §5.1, Table 1
    evidence: Joint FAC+TTS training on 370 hours of TTS data alongside 9.3 hours of FAC data yields 25% accentedness
      reduction and WER improvement from 14.4% to 9.1% relative to a standalone FAC baseline.
    confidence: high
    relevance: medium
  - claim_id: bert_style_masked_token_restoration_can_correct_stochastic_local_substitution
    role: supports
    claim: BERT-style masked token restoration can correct stochastic local substitution errors introduced by autoregressive
      speech token decoding, improving acoustic continuity.
    source: §5.3, Table 4
    evidence: Removing SpeechRestorer decreases TTS CMOS from 3.850 to 3.629, with the paper attributing the gain
      to error-correction of spurious token substitutions that cause acoustic discontinuities.
    confidence: high
    relevance: medium
  - claim_id: token_level_post_processing_for_llm_speech_generation_can_correct
    role: complicates
    claim: Token-level post-processing for LLM speech generation can correct local substitution errors but fails
      to address higher-level failure modes such as word skipping and repetition.
    source: §6
    evidence: The paper explicitly states that SpeechRestorer cannot fix skipped words or repetitions, as these
      require sequence-level rather than token-level correction.
    confidence: high
    relevance: medium
  - claim_id: training_data_scale_rather_than_architectural_design_is_the_primary
    role: refines
    claim: Training data scale, rather than architectural design, is the primary driver of quality gaps between
      LLM-based TTS systems at different performance levels.
    source: §5.3, Table 4
    evidence: SpeechAccentLLM trails NaturalSpeech2 in TTS naturalness (CMOS 3.850 vs. 3.944) despite comparable
      architecture; the gap is attributed to NS2 training on approximately two orders of magnitude more data.
    confidence: high
    relevance: medium
  limitations:
  - The FAC evaluation is restricted to four L1 backgrounds from L2-ARCTIC and uses a single TTS model (LJSpeech-trained
    VITS) to generate native accent counterparts. Generalisation to other accents, speaking styles, or higher-quality
    native reference speech is untested.
  - Prosody modelling is acknowledged as incomplete; the Variance Adapter models f0 only and does not capture rhythm,
    duration patterns, or prosodic phrasing beyond pitch. Timbre reconstruction quality is bounded by the frozen
    ECAPA-TDNN speaker encoder, which was not trained for the L2/accented-speech domain. SpeechRestorer cannot resolve
    sequence-level decoding failures (word skipping, repetition), leaving a category of LLM-generated errors unaddressed.
    The paper does not report total parameter counts for any module, which complicates direct comparison with other
    systems.
  caveats: []
- id: '2506.23325'
  published_date: "2025-07-09"
  entry_date: '2026-07-28'
  year: 2025
  venue: arXiv
  task:
  - codec
  architecture:
  - GAN
  - hybrid
  relevance: high
  evidence_role:
  - core_evidence
  current_role: active_evidence
  method_family:
  - gan_neural_codecs_and_decoders
  - hybrid_semantic_acoustic_tokenizers
  claims:
  - claim_id: minimizing_shared_parameters_between_semantic_and_acoustic_encoding_pathways_in
    role: supports
    claim: Minimizing shared parameters between semantic and acoustic encoding pathways in a speech codec reduces
      task conflict and enables simultaneous strong text alignment and high acoustic fidelity at low bitrates.
    source: §3.1, §4.4, Table 4
    evidence: XY-Tokenizer's dual-tower architecture (parameters shared only in the RVQ module) achieves WER 0.13
      and SPK-SIM 0.83 at 1 kbps; a matched single-channel model sharing encoder and RVQ across both tasks achieves
      SPK-SIM 0.77 with the same WER, confirming that parameter sharing degrades reconstruction.
    confidence: high
    relevance: high
  - claim_id: llm_based_asr_supervision_during_codec_pre_training_provides_stronger
    role: supports
    claim: LLM-based ASR supervision during codec pre-training provides stronger text alignment than SSL representation
      distillation for low-bitrate codecs.
    source: §4.3, Table 3
    evidence: At approximately 1 kbps, XY-Tokenizer (LLM-based ASR, WER 0.13) substantially outperforms SpeechTokenizer
      variants (distillation-based, WER 0.18-0.34) and Mimi-8 (distillation-based, WER 0.28) on the ASR probing
      task, while maintaining comparable or better reconstruction quality than those baselines.
    confidence: high
    relevance: high
  - claim_id: representation_distillation_from_ssl_models_for_semantic_alignment_introduces_reconstruction
    role: complicates
    claim: Representation distillation from SSL models for semantic alignment introduces reconstruction-quality
      penalties that worsen with distillation strength at low bitrates.
    source: §4.3, Table 3
    evidence: SpeechTokenizer-x3 (5x stronger distillation than the official setting) achieves WER 0.18 but SPK-SIM
      only 0.48 at 1.5 kbps, compared to SpeechTokenizer-x1 (WER 0.34, SPK-SIM 0.65), demonstrating that stronger
      distillation systematically degrades acoustic fidelity.
    confidence: high
    relevance: high
  - claim_id: allowing_the_llm_decoder_to_train_freely_during_multi_task
    role: complicates
    claim: Allowing the LLM decoder to train freely during multi-task codec pre-training improves autoregressive
      text generation but degrades the encoder's transferable text-alignment representations.
    source: §4.4, Table 5
    evidence: A trainable-LLM variant achieves lower LLM decoder WER (0.03 vs. 0.06 at 200K steps) but higher ASR
      probing WER (0.18 vs. 0.13), with the probing WER worsening progressively over 800K training steps, suggesting
      the encoder's text alignment capacity migrates into the flexible decoder.
    confidence: high
    relevance: high
  - claim_id: whisper_s_supervised_asr_pre_training_enables_better_paralinguistic_information
    role: supports
    claim: Whisper's supervised ASR pre-training enables better paralinguistic information preservation in codec
      encoder architectures than self-supervised alternatives.
    source: §2, Table 1
    evidence: Preliminary auto-encoder experiments with frozen encoders show Whisper achieves SPK-SIM 0.68, STOI
      0.88, and PESQ-NB 2.03 on LibriSpeech test-clean, compared to WavLM (SPK-SIM 0.53) and HuBERT (SPK-SIM 0.42),
      motivating Whisper as the codec encoder initialization.
    confidence: high
    relevance: high
  limitations:
  - Evaluation uses only objective metrics (WER, SPK-SIM, STOI, PESQ). No subjective listening tests or MOS scores
    are reported, making it impossible to assess perceptual audio quality relative to competing systems.
  - Achieving further reductions in bitrate below 1 kbps without performance degradation remains an open challenge
    noted by the authors. The scaling behavior of the two-stage training approach with respect to parameter count
    and dataset size is not characterized, leaving open questions about how to optimize training efficiency for
    larger variants. The LLM-based ASR decoder (Qwen2.5-0.5B) adds substantial parameter overhead during pre-training
    that is absent during inference, but the computational cost and training complexity of this stage relative to
    distillation-based approaches is not discussed.
  caveats: []
- id: '2412.18603'
  published_date: "2025-07-10"
  entry_date: '2026-07-28'
  year: 2025
  venue: ICML
  task:
  - SCA
  architecture:
  - autoregressive-LM
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  - evaluation_caution
  current_role: influential
  method_family:
  - autoregressive_codec_language_models
  claims:
  - claim_id: linear_time_state_space_sequence_models_enable_stable_long_form
    role: supports
    claim: Linear-time state-space sequence models enable stable long-form speech generation that Transformer spoken
      LMs cannot sustain at comparable scale.
    source: §7.1, §7.2, Tables 3-4
    evidence: SpeechSSM-2B maintains N-MOS of 4.12-4.16 across all time strata up to 4+ minutes and stable semantic
      coherence in 16-minute generation, while all Transformer-based baselines and windowed extensions of prior
      spoken LMs degrade rapidly beyond their training lengths.
    confidence: high
    relevance: medium
  - claim_id: separating_speaker_conditioning_into_a_dedicated_acoustic_synthesis_stage_preserves
    role: supports
    claim: Separating speaker conditioning into a dedicated acoustic synthesis stage preserves voice identity more
      effectively in long-form speech generation than encoding speaker information in semantic tokens.
    source: §6.1, §7.1, Tables 2-3
    evidence: SpeechSSM achieves SpkrSim of 0.85 in long-form evaluation (vs. 0.33-0.41 for GSLM and TWIST), attributed
      to SoundStorm's speaker-prompted acoustic stage handling identity independently of the semantic LM.
    confidence: high
    relevance: medium
  - claim_id: standard_evaluation_metrics_for_spoken_language_models_are_unreliable_or
    role: complicates
    claim: Standard evaluation metrics for spoken language models are unreliable or saturating for modern systems,
      particularly in long-form settings.
    source: §6.1, §4
    evidence: sWUGGY scores negatively correlate with generation quality at large token vocabularies (32k vs. hundreds);
      short-form holistic MOS approaches ground truth at 7 seconds while qualitative failures persist; transcript
      perplexity favors repetitive outputs at default temperatures. Reference-based embedding metrics and LLM-as-judge
      side-by-sides prove more discriminative.
    confidence: high
    relevance: medium
  - claim_id: long_form_speech_generation_remains_far_below_human_quality_even
    role: complicates
    claim: Long-form speech generation remains far below human quality even for state-of-the-art spoken language
      models.
    source: §7.1, Table 3
    evidence: No model achieves any wins against LibriSpeech-Long ground truth in 4-minute side-by-side comparisons
      judged by an LLM on fluency, coherence, logicality, and interestingness, demonstrating a large remaining quality
      gap despite strong short-form performance.
    confidence: high
    relevance: medium
  limitations:
  - Model weights are not released and inference requires TPU infrastructure; independent reproducibility of the
    long-form results is not currently possible.
  - 'Comparison fairness across baselines is imperfect: systems differ in training data scale, semantic tokenizer
    choice, token vocabulary size, and whether text was seen during training. The authors partially address this
    with the matched SpeechTransformer-2B baseline, but cannot control all confounds across the full set of comparisons.'
  - The semantic coherence advantage may partly reflect USM-v2's speaker invariance (reducing token entropy from
    vocal variation) rather than the SSM architecture per se. SpeechSSM-X, the extemporaneous variant trained on
    216k hours of monologue speech, is evaluated qualitatively only without quantitative benchmarks. LibriSpeech-Long
    covers read audiobook speech; generalization to spontaneous, multi-speaker, or conversational long-form speech
    is not evaluated.
  caveats: []
- id: '2507.07799'
  published_date: "2025-07-10"
  entry_date: '2026-07-28'
  year: 2025
  venue: arXiv
  task:
  - TTS
  architecture:
  - autoregressive-LM
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  - evaluation_caution
  current_role: minor
  method_family:
  - autoregressive_codec_language_models
  claims:
  - claim_id: instruction_conditioned_tts_enables_speaker_anonymization_without_relying_on_source
    role: supports
    claim: Instruction-conditioned TTS enables speaker anonymization without relying on source speaker audio or
      speaker embeddings.
    source: §3.3, §3.4.1, Table 1
    evidence: Parler-TTS with natural language speaker descriptions achieves FAR=0% on SLUE-VoxPopuli, matching
      zero-shot TTS baselines (VALL-E, XTTS-v2) that require reference audio, while offering explicit control over
      gender, pitch, accent, speaking rate, and channel conditions.
    confidence: high
    relevance: medium
  - claim_id: false_acceptance_rate_is_an_unreliable_discriminator_of_speaker_anonymization
    role: complicates
    claim: False acceptance rate is an unreliable discriminator of speaker anonymization quality when evaluation
      design guarantees disjoint training and test speakers.
    source: §3.4.1
    evidence: FAR=0% is achieved by every TTS-based system regardless of architecture or conditioning, because TTS
      models are trained on speakers absent from the evaluation set; the metric cannot distinguish the proposed
      method from any TTS-based pipeline in this setting.
    confidence: high
    relevance: medium
  - claim_id: asr_ner_pipeline_based_content_privacy_is_limited_by_recognition
    role: complicates
    claim: ASR + NER pipeline-based content privacy is limited by recognition and detection errors that allow a
      meaningful share of sensitive content to pass through.
    source: §3.4.1, §4
    evidence: The ASR component achieves 19.00% WER on original speech and DeBERTa-L NER achieves only 71.80% F1
      on predicted transcriptions, meaning many named entities are missed at the detection stage despite 99.95%
      replacement accuracy among detected entities.
    confidence: high
    relevance: medium
  - claim_id: generative_variability_in_prompt_based_tts_limits_consistent_voice_identity
    role: complicates
    claim: Generative variability in prompt-based TTS limits consistent voice identity reproduction across utterances
      sharing the same speaker description.
    source: §4
    evidence: Parler-TTS outputs may sound like different persons across generations even with an identical speaker
      description, undermining multi-utterance anonymization coherence and making deterministic identity control
      difficult.
    confidence: high
    relevance: medium
  - claim_id: speaker_attribute_choice_in_instruction_conditioned_tts_introduces_measurable_variation
    role: supports
    claim: Speaker attribute choice in instruction-conditioned TTS introduces measurable variation in downstream
      speech intelligibility that practitioners must account for.
    source: §3.4.2, Tables 2 and 3
    evidence: WER varies from 12.07% (Slovak accent) to 23.76% (Italian accent) across 38 accent descriptions, and
      from 12.39% (normal rate) to 18.56% (very fast) across speaking rate settings, while speaker privacy (FAR=0%)
      remains constant throughout.
    confidence: high
    relevance: medium
  limitations:
  - FAR=0% is observed for all TTS-based systems including both baselines, because the evaluation speakers are disjoint
    from TTS training speakers. The privacy metric cannot distinguish the proposed approach from any TTS pipeline
    in this experimental design, which materially limits the privacy claims.
  - The NER component (71.80% F1) misses a substantial share of named entities, so content privacy is incomplete
    even when replacement accuracy is near-perfect among detected entities. WER-based content privacy evaluation
    does not capture semantic re-identification risks from context, paraphrase, or implicit references.
  - The pipeline is irreversible by design, making it unsuitable for forensic or audit contexts where original content
    must be recoverable. Prosody cannot be preserved without risking speaker re-identification, since prosodic patterns
    carry speaker-discriminative information. The study covers six speaker attribute dimensions but uses a single
    dataset, limiting generalisability.
  caveats: []
- id: '2507.09070'
  published_date: "2025-07-11"
  entry_date: '2026-07-28'
  year: 2025
  venue: arXiv
  task:
  - VC
  architecture:
  - autoregressive-LM
  - flow-matching
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: active_evidence
  method_family:
  - autoregressive_codec_language_models
  - flow_matching_codec_decoders
  claims:
  - claim_id: audio_codec_and_self_supervised_speech_representations_inherently_encode_speaker
    role: supports
    claim: Audio codec and self-supervised speech representations inherently encode speaker identity, making timbre
      leakage a structural challenge in codec-based voice conversion.
    source: §4.1, Table 1
    evidence: Speaker classification accuracy on LibriHeavy using EnCodec token IDs reaches 96.7%, HuBERT layer
      9 discrete tokens 71.7%, and the authors' BEST-RQ tokenizer 82.05%, all substantially above chance.
    confidence: high
    relevance: high
  - claim_id: aligning_audio_encoder_outputs_to_speaker_independent_text_embeddings_via
    role: supports
    claim: Aligning audio encoder outputs to speaker-independent text embeddings via monotonic alignment search
      produces representations with negligible residual speaker identity.
    source: §4.1, §5, Table 1
    evidence: The SemAlign-trained semantic encoder Q_ϕ achieves 2.84% speaker classification accuracy on LibriHeavy,
      versus 82.05% for the same encoder's raw tokenizer, confirming that the alignment objective removes speaker
      information that CTC loss alone cannot.
    confidence: high
    relevance: medium
  - claim_id: zero_shot_voice_conversion_without_explicit_speaker_verification_embeddings_can
    role: supports
    claim: Zero-shot voice conversion without explicit speaker verification embeddings can achieve higher speaker
      similarity than systems that rely on them.
    source: §4.2, Tables 2-3
    evidence: SemAlignVC achieves the highest SMOS (3.29), WavLM speaker similarity (0.95), ECAPA-TDNN (0.82), and
      Resemblyzer (0.89) among KNNVC, HierSpeech++, and UniAudio on VCTK and LibriHeavy evaluations, despite using
      only a reference mel spectrogram excerpt rather than a speaker embedding.
    confidence: high
    relevance: low
  - claim_id: aggressive_timbre_disentanglement_via_text_embedding_alignment_introduces_intelligibility_trade
    role: complicates
    claim: Aggressive timbre disentanglement via text-embedding alignment introduces intelligibility trade-offs
      when BERT-derived representations replace phoneme-based encodings.
    source: §5, Table 3
    evidence: SemAlignVC achieves 12.31% WER on LibriHeavy, compared to HierSpeech++'s 8.24%, with the authors attributing
      the gap to occasional word substitutions caused by the semantic ambiguity of BERT token representations during
      generation.
    confidence: high
    relevance: medium
  limitations:
  - The audio tokenizer is trained on a proprietary internal dataset and is described as interchangeable, but its
    interaction with SemAlign has not been tested with public codecs. Results may not transfer directly to systems
    using EnCodec, SpeechTokenizer, or other publicly available tokenizers.
  - The evaluation is English-only and uses VCTK and LibriHeavy, which are relatively clean audiobook and read-speech
    corpora. Performance on spontaneous speech, accented speech, or noisy conditions is not assessed. The baseline
    set is modest (three systems), and no ablation isolates the contribution of the flow matching acoustic model
    relative to SemAlign itself. The WER gap relative to HierSpeech++ remains unexplained beyond the synonym-substitution
    hypothesis, which is not directly tested.
  caveats: []
- id: '2507.12197'
  published_date: "2025-07-16"
  entry_date: '2026-07-28'
  year: 2025
  venue: arXiv
  task:
  - TTS
  - singing
  architecture:
  - autoregressive-LM
  - GAN
  relevance: medium
  evidence_role:
  - architecture_variant
  - acceleration_evidence
  - control_evidence
  - evaluation_caution
  current_role: active_evidence
  method_family:
  - autoregressive_codec_language_models
  - gan_neural_codecs_and_decoders
  claims:
  - claim_id: scaling_the_number_of_rvq_codebooks_in_a_discrete_speech
    role: supports
    claim: Scaling the number of RVQ codebooks in a discrete speech codec reduces information loss and improves
      reconstruction quality for expressive and challenging vocal content.
    source: §3.2, Table 5
    evidence: QDAC reconstruction improves monotonically from 1 to 16 codebooks across PESQ, STOI, and SI-SDR; 16-codebook
      QDAC at 50Hz achieves PESQ 3.83 versus PESQ 2.98 for 8 codebooks, with similar gains on Mel distance and speaker
      similarity.
    confidence: high
    relevance: high
  - claim_id: higher_multi_codebook_reconstruction_fidelity_in_a_codec_does_not
    role: complicates
    claim: Higher multi-codebook reconstruction fidelity in a codec does not necessarily translate into better speaker
      identity preservation in downstream zero-shot TTS.
    source: §3.2, Table 6
    evidence: QTTS achieves Spk Sim 0.75 on SeedTTS-Easy compared to 0.82 and 0.81 for single-codebook CosyVoice
      v1 and v2, despite QDAC's superior reconstruction metrics in Table 5.
    confidence: high
    relevance: high
  - claim_id: semantic_disentanglement_in_audio_codecs_can_be_achieved_by_backpropagating
    role: supports
    claim: Semantic disentanglement in audio codecs can be achieved by backpropagating ASR loss exclusively through
      the first RVQ codebook, enforcing content isolation without relying on general-purpose self-supervised representations.
    source: §2.1.2, Table 5
    evidence: QDAC trains an AR-ASR module conditioned only on first-codebook tokens; the resulting WER at reconstruction
      (6.42 for 8cb/25Hz) is close to ground truth (6.01), indicating the first codebook encodes phoneme-level content
      while residual codebooks capture acoustic detail.
    confidence: high
    relevance: high
  - claim_id: multi_codebook_autoregressive_tts_admits_a_principled_speed_quality_trade
    role: refines
    claim: Multi-codebook autoregressive TTS admits a principled speed-quality trade-off by choosing between strict
      hierarchical inter-codebook conditioning and a delayed multi-head parallel prediction scheme.
    source: §2.2, §2.3, Tables 1, 3, 4
    evidence: Hierarchy Parallel (200M, dual-AR) and Multihead Delay (120M, parallel with fixed delay) achieve comparable
      TTFT at 512 tokens (26ms vs 24ms) but differ substantially in decode throughput; the Multihead variant reaches
      over 196K codebook tokens/s at short output lengths versus 105K for Hierarchy.
    confidence: high
    relevance: high
  - claim_id: mos_evaluations_in_zero_shot_tts_can_yield_above_reference
    role: complicates
    claim: MOS evaluations in zero-shot TTS can yield above-reference scores for synthesised speech, undermining
      direct absolute comparisons across studies.
    source: §3.2, Table 6
    evidence: Ground truth speech achieves MOS 2.7 while QTTS, CosyVoice, and CosyVoice2 all score between 3.01
      and 3.03 on the same test set, producing a ranking inconsistent with naturalness expectations.
    confidence: high
    relevance: medium
  limitations:
  - Training data is not disclosed, the PGC-hard benchmark is proprietary, and no code or demo is available, making
    results difficult to reproduce or build upon.
  - The evaluation compares only against the CosyVoice v1/v2 family, leaving open how QTTS performs relative to
    flow-matching systems, other multi-codebook approaches, or stronger autoregressive baselines. Speaker similarity
    is lower than both single-codebook baselines (0.75 vs 0.81-0.82), suggesting the multi-codebook generation pipeline
    needs further work to fully leverage the improved codec for speaker transfer. The paper positions singing and
    music synthesis as motivating use cases but does not evaluate on these tasks. Only the 8-codebook QTTS model
    is evaluated for TTS synthesis, leaving open whether 16-codebook generation would further improve or introduce
    new training challenges.
  caveats: []
- id: '2507.16632'
  published_date: "2025-07-22"
  entry_date: '2026-07-28'
  year: 2025
  venue: arXiv
  task:
  - SCA
  - TTS
  architecture:
  - autoregressive-LM
  - flow-matching
  - hybrid
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: influential
  method_family:
  - autoregressive_codec_language_models
  - flow_matching_codec_decoders
  - hybrid_semantic_acoustic_tokenizers
  claims:
  - claim_id: interleaving_discrete_audio_tokens_with_text_tokens_in_a_shared
    role: supports
    claim: Interleaving discrete audio tokens with text tokens in a shared language model vocabulary enables end-to-end
      spoken dialogue systems that respond coherently to paralinguistic cues without requiring separate speech synthesis
      pipelines.
    source: §3.1, §4.6
    evidence: Interleaving discrete audio tokens with text tokens in a shared language model vocabulary enables
      end-to-end spoken dialogue systems that respond coherently to paralinguistic cues without requiring separate
      speech synthesis pipelines.
    confidence: high
    relevance: high
  - claim_id: reinforcement_learning_applied_to_audio_language_models_can_improve_reasoning
    role: supports
    claim: Reinforcement learning applied to audio language models can improve reasoning efficiency in complex acoustic
      scenarios while preserving generation quality.
    source: §3.4
    evidence: Reinforcement learning applied to audio language models can improve reasoning efficiency in complex
      acoustic scenarios while preserving generation quality.
    confidence: high
    relevance: medium
  - claim_id: retrieval_augmented_generation_and_external_tool_calling_can_substantially_reduce
    role: supports
    claim: Retrieval-augmented generation and external tool calling can substantially reduce hallucination and expand
      capability (timbre switching, web-grounded responses) in large audio language models without architectural
      redesign.
    source: §3.1, §4.5
    evidence: Retrieval-augmented generation and external tool calling can substantially reduce hallucination and
      expand capability (timbre switching, web-grounded responses) in large audio language models without architectural
      redesign.
    confidence: high
    relevance: medium
  - claim_id: comprehensive_multi_task_pre_training_across_asr_tts_translation_and
    role: supports
    claim: Comprehensive multi-task pre-training across ASR, TTS, translation, and conversation significantly improves
      spoken dialogue performance in low-resource languages and accented speech.
    source: §3.2, §4.1
    evidence: Comprehensive multi-task pre-training across ASR, TTS, translation, and conversation significantly
      improves spoken dialogue performance in low-resource languages and accented speech.
    confidence: high
    relevance: low
  - claim_id: existing_audio_language_model_benchmarks_fail_to_capture_fine_grained
    role: complicates
    claim: Existing audio language model benchmarks fail to capture fine-grained paralinguistic comprehension and
      voice-triggered tool invocation, leaving important capability dimensions systematically unmeasured.
    source: §4.2, §4.5
    evidence: Existing audio language model benchmarks fail to capture fine-grained paralinguistic comprehension
      and voice-triggered tool invocation, leaving important capability dimensions systematically unmeasured.
    confidence: high
    relevance: medium
  limitations:
  - Two of the main evaluation benchmarks (StepEval-Audio-Paralinguistic, StepEval-Audio-Toolcall) are introduced
    by the authors themselves and have not been validated independently. Results on these benchmarks may overstate
    absolute capability levels even if relative comparisons are informative.
  - Model size is not reported, which makes parameter-count comparisons with other open-source systems (Kimi-Audio,
    Qwen2.5-Omni) difficult to interpret fairly. The open-source mini variant uses Qwen2.5-7B as its backbone, but
    the full model's parameter count remains undisclosed.
  - The audio search tool relies on a proprietary library of hundreds of thousands of speech samples, limiting reproducibility
    of this capability. Latency characteristics are not reported; real-time performance is claimed via VAD and deployment
    infrastructure inherited from Step-Audio but not benchmarked here.
  - Evaluation is primarily Chinese-English bilingual. Performance on other language families, especially those
    with minimal pre-training coverage, is largely uncharacterized. The tool-calling evaluation uses synthesized
    speech conversations generated from text scripts rather than natural speech, which may not reflect real-world
    tool invocation patterns.
  caveats: []
- id: '2507.21138'
  published_date: "2025-07-22"
  entry_date: '2026-07-28'
  year: 2025
  venue: arXiv
  task:
  - TTS
  architecture:
  - autoregressive-LM
  relevance: high
  evidence_role:
  - core_evidence
  - control_evidence
  current_role: active_evidence
  method_family:
  - autoregressive_codec_language_models
  claims:
  - claim_id: rl_alignment_with_composite_perceptual_rewards_improves_multiple_speech_quality
    role: supports
    claim: RL alignment with composite perceptual rewards improves multiple speech quality dimensions simultaneously
      over SFT-only baselines in autoregressive codec TTS.
    source: §3.5, Table 8
    evidence: GRPO with combined WER + speaker similarity + DNSMOS rewards achieves gains in all three metrics over
      the SFT baseline, while single-reward models improve only their target dimension; TTS-1 WER drops from 7.9%
      (SFT) to 6.3% (RL-aligned).
    confidence: high
    relevance: high
  - claim_id: audio_pre_training_on_large_scale_raw_speech_substantially_improves
    role: supports
    claim: Audio pre-training on large-scale raw speech substantially improves subsequent SFT quality in LLM-based
      TTS beyond what instruction-tuned LLM initialisation provides.
    source: §3.4, Figure 5
    evidence: Initialising SFT from an audio pre-trained LLaMA-3.2-1B checkpoint yields lower SFT loss, approximately
      15% lower WER, and approximately 3% higher speaker similarity than initialising from the base LLaMA-3.2-1B-Instruct
      checkpoint.
    confidence: high
    relevance: medium
  - claim_id: scaling_speechlm_parameter_count_in_autoregressive_codec_tts_consistently_improves
    role: supports
    claim: Scaling SpeechLM parameter count in autoregressive codec TTS consistently improves intelligibility and
      speaker fidelity across languages.
    source: §4, Figure 8, Table 8
    evidence: TTS-1-Max (8.8B) achieves 5.1% overall WER and higher SIM across all 11 evaluated languages compared
      to TTS-1 (1.6B) at 6.3% WER, with performance correlating with pre-training loss differences.
    confidence: high
    relevance: high
  - claim_id: style_conditioning_via_discrete_text_tags_conflicts_with_speaker_identity
    role: complicates
    claim: Style conditioning via discrete text tags conflicts with speaker identity preservation in single-codebook
      codec TTS architectures.
    source: §3.6
    evidence: Direct prepending of style markup tags during SFT produced no effect on synthesized speech because
      the single-codebook design entangles acoustic and semantic information; successful style control required
      constructing paired neutral/stylized utterances from the same speaker and applying LoRA fine-tuning.
    confidence: high
    relevance: high
  - claim_id: streaming_audio_delivery_in_autoregressive_tts_introduces_audible_artifacts_and
    role: complicates
    claim: Streaming audio delivery in autoregressive TTS introduces audible artifacts and volume inconsistencies
      at segment boundaries that require specific engineering mitigations independent of the generative model's
      quality.
    source: §5.1
    evidence: Without concatenation restricted to non-voicing regions and context-aware decoder decoding with extended
      audio prompt context, segment boundaries introduce clicks and volume drops; these are engineering-layer problems
      independent of SpeechLM quality.
    confidence: high
    relevance: low
  limitations:
  - Model weights are not publicly released, making independent benchmarking and replication impossible. All evaluations
    use proprietary or internal test sets; the internal TTS arena covers only English and uses approximately 20
    annotators with a modest vote count per pair.
  - The training data is drawn from a proprietary mixture of public and licensed sources whose exact composition
    is not disclosed, limiting reproducibility. The evaluation framework does not include standard public TTS benchmarks
    (e.g., LibriTTS or VCTK test sets), making direct numerical comparison with published systems that do report
    on these benchmarks difficult.
  - The paper notes that speaker similarity metrics fluctuate with emotionally expressive speech, suggesting that
    current automated evaluation protocols may not fully capture perceptual quality in dynamic scenarios. The audio
    markup system, while effective for English, shows "reduced fidelity" when generalizing style control to non-English
    languages. Prompt audio caching can cause prosodic bleed from the reference audio into generated speech, and
    longer sequences generated from short prompts may degrade in quality.
  caveats: []
- id: '2505.15670'
  published_date: "2025-07-25"
  entry_date: '2026-07-28'
  year: 2025
  venue: arXiv
  task:
  - SCA
  architecture:
  - autoregressive-LM
  relevance: medium
  evidence_role:
  - architecture_variant
  - acceleration_evidence
  - control_evidence
  current_role: active_evidence
  method_family:
  - autoregressive_codec_language_models
  claims:
  - claim_id: full_duplex_spoken_dialogue_systems_can_be_built_without_speech
    role: supports
    claim: Full-duplex spoken dialogue systems can be built without speech-text pretraining by routing user audio
      through a pretrained streaming encoder rather than requiring the LLM to learn audio representations end-to-end.
    source: §3, §6.1, §6.2, Tables 2–3
    evidence: SALM-Duplex skips speech pretraining entirely and instead uses a 100M streaming CTC encoder for user
      input; it still outperforms Moshi on barge-in success rate (94.5% vs. 55.1%) and reasoning GPT scores across
      all five evaluation sets.
    confidence: high
    relevance: low
  - claim_id: asymmetric_duplex_architectures_that_separate_user_and_agent_speech_pathways
    role: supports
    claim: Asymmetric duplex architectures that separate user and agent speech pathways enable independent specialization,
      including speaker-specific codec fine-tuning without affecting user comprehension.
    source: §3.2, §6.3, Table 4
    evidence: Personalized 0.6 kbps NanoCodec (fine-tuned on 21k target-speaker utterances) outperforms Moshi's
      Mimi at 1.1 kbps and untuned NanoCodec at 1.2 kbps on MOS, CER, and SECS, while operating at roughly half
      the bitrate.
    confidence: high
    relevance: high
  - claim_id: end_to_end_speech_to_speech_models_do_not_consistently
    role: complicates
    claim: End-to-end speech-to-speech models do not consistently match optimal cascaded systems in reasoning quality,
      even when the cascaded oracle has access to ground-truth ASR transcriptions of user speech.
    source: §6.2, Table 3
    evidence: SALM-Duplex outperforms GT+LLM on Roleplay and ASR-QA but underperforms on UltraChat (3.5 vs. 6.4),
      Topic, and Alpaca; the gap reflects compounding ASR error and limited backbone reasoning capacity at 1.1B
      parameters.
    confidence: high
    relevance: medium
  - claim_id: barge_in_success_rate_and_latency_together_constitute_more_discriminative
    role: supports
    claim: Barge-in success rate and latency together constitute more discriminative signals for evaluating full-duplex
      systems than speech quality metrics such as UTMOS.
    source: §5.2, §6.1, Table 2
    evidence: On the Impatient set, SALM-Duplex and Moshi have zero false alarms each and UTMOS within 0.2 points
      (4.0 vs. 3.8), yet differ by 39.4 percentage points in barge-in success rate, making success rate the dominant
      differentiating metric.
    confidence: high
    relevance: low
  limitations:
  - All agent speech in training data is synthesized using a TTS model with a fixed speaker. Generalization to diverse
    agent voices or real conversational speech has not been demonstrated.
  - The 1.1B TinyLlama backbone limits reasoning ceiling; the gap to the GT+LLM oracle on complex dialogue tasks
    suggests ASR error compounds with limited LLM capacity. The quantitative comparison is restricted to Moshi;
    other contemporaneous duplex systems (OmniFlatten, SALMONN-Omni, MinMo) are discussed in related work but not
    benchmarked. The hardcoded 0.64s post-user-turn silence used to suppress unexpected agent barge-in may not transfer
    to naturally paced conversation. Evaluation datasets are largely synthetic, so performance on real recorded
    conversational speech remains an open question.
  caveats: []
- id: '2507.18897'
  published_date: "2025-07-25"
  entry_date: '2026-07-28'
  year: 2025
  venue: arXiv
  task:
  - codec
  architecture:
  - GAN
  relevance: high
  evidence_role:
  - core_evidence
  current_role: active_evidence
  method_family:
  - gan_neural_codecs_and_decoders
  claims:
  - claim_id: single_quantizer_neural_codecs_can_match_multi_quantizer_systems_in
    role: supports
    claim: Single-quantizer neural codecs can match multi-quantizer systems in perceived speech quality at far lower
      bitrates when combined with stabilized VQ spaces and asymmetric decoder architectures.
    source: §4.3, Table 1
    evidence: HH-Codec achieves UTMOS 3.61 on LibriTTS test-clean at 0.3 kbps with a single quantizer, competitive
      with DAC's 4-quantizer configuration at 4 kbps (UTMOS 3.41) and markedly above SpeechTokenizer's single-quantizer
      variant at 0.75 kbps (UTMOS 1.26).
    confidence: high
    relevance: medium
  - claim_id: multi_layer_vq_training_with_single_layer_inference_where_auxiliary
    role: supports
    claim: Multi-layer VQ training with single-layer inference, where auxiliary quantizer layers act as regularizers,
      substantially improves codebook utilization in high-compression single-quantizer settings.
    source: §4.4, Tables 2-3
    evidence: SLM-VQ achieves 98% codebook utilization at 8192 entries versus 56% for Classic VQ and 92% for single-layer
      SLM-VQ, while improving UTMOS from 2.76 (Classic VQ) to 3.07 on LibriTTS test-other.
    confidence: high
    relevance: high
  - claim_id: dual_domain_supervision_combining_intermediate_mel_spectrogram_and_final_audio
    role: supports
    claim: Dual-domain supervision combining intermediate mel-spectrogram and final audio reconstruction objectives
      is critical for stable high-compression neural codec training.
    source: §4.4, Table 2
    evidence: Reducing to single audio-domain supervision drops UTMOS from 3.07 to 1.85 and SPK-SIM from 0.64 to
      0.33 on LibriTTS test-other, the largest degradation across all ablation variants.
    confidence: high
    relevance: high
  - claim_id: standard_adversarial_codec_training_recipes_break_down_at_extreme_compression
    role: complicates
    claim: Standard adversarial codec training recipes break down at extreme compression ratios, requiring architectural
      and procedural modifications to avoid collapse.
    source: §1
    evidence: Below 0.3 kbps with existing methods, the paper documents adversarial training collapse, a 63% UTMOS
      drop below 30 tokens/s, 43% codebook utilization at 8192 entries, and minimal benefit from expanding training
      data, all addressed in HH-Codec through SLM-VQ and progressive training.
    confidence: high
    relevance: high
  limitations:
  - HuBERT-based semantic distillation is trained on English, so multilingual performance of SLM-VQ is unknown.
    The downstream spoken language modeling experiment measures only training loss reduction rather than end-to-end
    TTS quality, leaving it unclear how the 24-token-per-second compression affects downstream synthesis intelligibility
    and naturalness. Training data conditions differ across compared baselines (WavTokenizer and DAC use larger
    or different datasets), which limits direct attribution of gains to architecture versus data. The ablation study
    is conducted on a subset of training data only (LibriTTS train-100/360, not the full training set including
    Emilia), which may underestimate some component contributions.
  caveats: []
- id: 2025.acl-long.1043
  published_date: "2025-07-27"
  entry_date: '2026-07-28'
  year: 2025
  venue: ACL
  task:
  - TTS
  architecture:
  - flow-matching
  - hybrid
  relevance: medium
  evidence_role:
  - architecture_variant
  current_role: influential
  method_family:
  - flow_matching_codec_decoders
  - hybrid_semantic_acoustic_tokenizers
  claims:
  - claim_id: replacing_gaussian_noise_with_a_learned_prior_as_the_starting
    role: supports
    claim: Replacing Gaussian noise with a learned prior as the starting point for flow matching reduces the number
      of required inference steps to one without needing a separate distillation stage.
    source: §3.1, §3.3
    evidence: Replacing Gaussian noise with a learned prior as the starting point for flow matching reduces the
      number of required inference steps to one without needing a separate distillation stage.
    confidence: high
    relevance: medium
  - claim_id: flow_matching_tts_systems_trained_on_traditional_ot_cfm_are
    role: complicates
    claim: Flow-matching TTS systems trained on traditional OT-CFM are data-hungry and fail to generalise when retrained
      on limited data, while neural codec-based systems remain effective with as few as 500 hours.
    source: §4.2, Table 1
    evidence: Flow-matching TTS systems trained on traditional OT-CFM are data-hungry and fail to generalise when
      retrained on limited data, while neural codec-based systems remain effective with as few as 500 hours.
    confidence: high
    relevance: high
  - claim_id: zero_shot_tts_systems_trained_exclusively_on_clean_prompts_degrade
    role: supports
    claim: Zero-shot TTS systems trained exclusively on clean prompts degrade substantially in intelligibility when
      given noisy reference audio, with autoregressive codec models being especially vulnerable.
    source: §4.4, Table 4
    evidence: Zero-shot TTS systems trained exclusively on clean prompts degrade substantially in intelligibility
      when given noisy reference audio, with autoregressive codec models being especially vulnerable.
    confidence: high
    relevance: high
  - claim_id: factorised_codec_representations_that_balance_acoustic_and_semantic_attributes_trade
    role: supports
    claim: Factorised codec representations that balance acoustic and semantic attributes trade perceptual naturalness
      (UTMOS) for intelligibility (WER) relative to codecs that prioritise acoustic fidelity.
    source: §4.2
    evidence: Factorised codec representations that balance acoustic and semantic attributes trade perceptual naturalness
      (UTMOS) for intelligibility (WER) relative to codecs that prioritise acoustic fidelity.
    confidence: high
    relevance: high
  - claim_id: fine_tuning_a_zero_shot_tts_model_on_noise_augmented
    role: supports
    claim: Fine-tuning a zero-shot TTS model on noise-augmented prompts preserves intelligibility and substantially
      recovers acoustic quality metrics under low-SNR conditions.
    source: §4.4, Table 4
    evidence: Fine-tuning a zero-shot TTS model on noise-augmented prompts preserves intelligibility and substantially
      recovers acoustic quality metrics under low-SNR conditions.
    confidence: high
    relevance: medium
  limitations:
  - '- UTMOS and speaker similarity lag behind F5-TTS and VoiceCraft; the prior-based approach trades acoustic naturalness
    for intelligibility and speed. - Duration predictor rounding errors (integer quantization of phoneme durations)
    introduce temporal domain artifacts. - FACodec dependency means reproduction requires NaturalSpeech 3''s codec
    infrastructure. - Noise robustness fine-tuning improves non-WER metrics but was validated only on the QUT-NOISE
    database. - Future work: multilingual extension, adaptive noise filtering integration.'
  caveats: []
- id: 2025.acl-long.1498
  published_date: "2025-07-27"
  entry_date: '2026-07-28'
  year: 2025
  venue: ACL
  task:
  - TTS
  - codec
  architecture:
  - autoregressive-LM
  - GAN
  relevance: high
  evidence_role:
  - core_evidence
  - control_evidence
  current_role: active_evidence
  method_family:
  - autoregressive_codec_language_models
  - gan_neural_codecs_and_decoders
  claims:
  - claim_id: context_dependent_encoding_in_rvq_based_neural_audio_codecs_causes
    role: supports
    claim: Context-dependent encoding in RVQ-based neural audio codecs causes perceptually equivalent audio segments
      to produce divergent discrete token sequences, increasing prediction uncertainty in downstream codec language
      models.
    source: §2, §2.3
    evidence: Context-dependent encoding in RVQ-based neural audio codecs causes perceptually equivalent audio segments
      to produce divergent discrete token sequences, increasing prediction uncertainty in downstream codec language
      models.
    confidence: high
    relevance: high
  - claim_id: codec_token_consistency_and_downstream_autoregressive_tts_quality_are_monotonically
    role: supports
    claim: 'Codec token consistency and downstream autoregressive TTS quality are monotonically correlated: improvements
      in consistency accuracy reliably reduce word error rate and increase speaker similarity.'
    source: §5.2, Figure 4
    evidence: 'Codec token consistency and downstream autoregressive TTS quality are monotonically correlated: improvements
      in consistency accuracy reliably reduce word error rate and increase speaker similarity.'
    confidence: high
    relevance: high
  - claim_id: auxiliary_consistency_losses_applied_during_codec_training_can_substantially_increase
    role: supports
    claim: Auxiliary consistency losses applied during codec training can substantially increase token consistency
      with negligible impact on reconstruction quality.
    source: §5.1, Table 1
    evidence: Auxiliary consistency losses applied during codec training can substantially increase token consistency
      with negligible impact on reconstruction quality.
    confidence: high
    relevance: high
  - claim_id: consistency_constraint_methods_applied_to_codec_training_generalize_across_neural
    role: supports
    claim: Consistency constraint methods applied to codec training generalize across neural codec architectures
      and autoregressive LM backbones, as demonstrated by cross-system experiments on both EnCodec-VALL-E and FunCodec-UniAudio
      pipelines.
    source: §5.2, Table 3
    evidence: Consistency constraint methods applied to codec training generalize across neural codec architectures
      and autoregressive LM backbones, as demonstrated by cross-system experiments on both EnCodec-VALL-E and FunCodec-UniAudio
      pipelines.
    confidence: high
    relevance: high
  - claim_id: in_rvq_codecs_deeper_codebook_layers_suffer_disproportionately_from_context
    role: supports
    claim: In RVQ codecs, deeper codebook layers suffer disproportionately from context-induced inconsistency, because
      they encode fine-grained acoustic detail that is more sensitive to contextual perturbation than the semantic
      information stored in shallow layers.
    source: §2.3, Appendix A.3, Table 6
    evidence: In RVQ codecs, deeper codebook layers suffer disproportionately from context-induced inconsistency,
      because they encode fine-grained acoustic detail that is more sensitive to contextual perturbation than the
      semantic information stored in shallow layers.
    confidence: high
    relevance: high
  limitations:
  - The DRI fix is applied at codec training time; it does not address the underlying structural cause (large receptive
    fields in the convolutional encoder). A follow-up could explore whether directly reducing receptive field size
    with compensating distillation achieves similar or better consistency.
  - The method is validated on English speech (LibriTTS and MLS, which is multilingual but dominated by English).
    Whether DRI and its mitigation generalize to tonal languages (where fine-grained acoustic context is phonemically
    contrastive) is unstudied.
  - Evaluation is restricted to the VALL-E architecture for downstream TTS. The generalizability claim to "any neural
    codec language model" rests on two additional cross-codec experiments (Table 3), which is suggestive but not
    comprehensive. In particular, the paper does not evaluate on flow-matching TTS (e.g. CosyVoice or E2-TTS), where
    consistency constraints on the codec may have different downstream effects.
  - The MOS and SMOS evaluations use 50 audio samples — a small panel by modern standards. Confidence intervals
    are not reported for subjective metrics.
  - DRI may interact with audio compression artifacts differently at different bitrates; the paper evaluates 4 kbps
    and 8 kbps but does not explore very low bitrates where context-dependence may be even stronger.
  caveats: []
- id: 2025.acl-long.346
  published_date: "2025-07-27"
  entry_date: '2026-07-28'
  year: 2025
  venue: ACL
  task:
  - TTS
  architecture:
  - transformer-enc-dec
  - hybrid
  relevance: medium
  evidence_role:
  - architecture_variant
  - acceleration_evidence
  - control_evidence
  current_role: active_evidence
  method_family:
  - transformer_encoder_decoder_tokenizers
  - hybrid_semantic_acoustic_tokenizers
  claims:
  - claim_id: disentangling_speaker_timbre_and_speaking_style_into_separate_codec_representations
    role: supports
    claim: Disentangling speaker timbre and speaking style into separate codec representations is a necessary condition
      for simultaneous zero-shot speaker cloning and style control; without explicit decoupling, the two conditioning
      signals interfere and controllability collapses.
    source: §4.3, Table 4
    evidence: Disentangling speaker timbre and speaking style into separate codec representations is a necessary
      condition for simultaneous zero-shot speaker cloning and style control; without explicit decoupling, the two
      conditioning signals interfere and controllability collapses.
    confidence: high
    relevance: high
  - claim_id: natural_language_style_descriptions_have_an_inherent_many_to_many
    role: supports
    claim: Natural language style descriptions have an inherent many-to-many relationship with audio that cannot
      be resolved by timbre conditioning alone, requiring a probabilistic model of style variation such as a mixture
      density network.
    source: §3.3
    evidence: Natural language style descriptions have an inherent many-to-many relationship with audio that cannot
      be resolved by timbre conditioning alone, requiring a probabilistic model of style variation such as a mixture
      density network.
    confidence: high
    relevance: medium
  - claim_id: zero_shot_speaker_cloning_capability_in_style_controllable_tts_can
    role: supports
    claim: Zero-shot speaker cloning capability in style-controllable TTS can be achieved by building on a large-scale
      pre-trained disentangled codec without sacrificing audio quality relative to dedicated zero-shot TTS systems.
    source: §4.2, Table 2
    evidence: Zero-shot speaker cloning capability in style-controllable TTS can be achieved by building on a large-scale
      pre-trained disentangled codec without sacrificing audio quality relative to dedicated zero-shot TTS systems.
    confidence: high
    relevance: high
  - claim_id: probabilistic_sampling_from_a_mixture_density_model_of_style_representations
    role: supports
    claim: Probabilistic sampling from a mixture density model of style representations improves both style diversity
      and generalization to out-of-domain style descriptions compared to deterministic style encoding.
    source: §4.2, §4.3, Table 3
    evidence: Probabilistic sampling from a mixture density model of style representations improves both style diversity
      and generalization to out-of-domain style descriptions compared to deterministic style encoding.
    confidence: high
    relevance: medium
  - claim_id: pitch_control_is_measurably_harder_to_preserve_when_timbre_and
    role: supports
    claim: Pitch control is measurably harder to preserve when timbre and style are controlled simultaneously, suggesting
      that pitch conditioning interacts with speaker identity in ways that speed, volume, and emotion do not.
    source: §4.2, Table 1
    evidence: Pitch control is measurably harder to preserve when timbre and style are controlled simultaneously,
      suggesting that pitch conditioning interacts with speaker identity in ways that speed, volume, and emotion
      do not.
    confidence: high
    relevance: low
  limitations:
  - 'The paper explicitly notes two limitations: (1) The training dataset is still limited in scale for style-controllable
    TTS; significantly larger datasets (tens of thousands of hours with style annotations) may be needed to achieve
    more advanced controllability. (2) The exploration of generative model architectures is narrow — only non-autoregressive
    parallel decoding is tried. Diffusion or flow-matching generators operating in the disentangled codec space
    might offer better quality or diversity.'
  - 'Beyond the paper''s self-assessment: pitch accuracy is the only metric where ControlSpeech underperforms style-only
    baselines on both in-domain and out-of-domain test sets. The authors attribute this to simultaneous timbre-style
    control, but the mechanism is unexplained and unresolved. The paper does not evaluate on standard TTS benchmarks
    (LibriSpeech, VCTK), relying entirely on VccmDataset, which makes external comparison difficult. The demo availability
    is not confirmed in the paper. Ethical risks from zero-shot voice cloning are acknowledged but only partially
    addressed (watermarking is proposed as future work).'
  caveats: []
- id: 2025.acl-long.65
  published_date: "2025-07-27"
  entry_date: '2026-07-28'
  year: 2025
  venue: ACL
  task:
  - TTS
  architecture:
  - autoregressive-LM
  - VAE
  - hybrid
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: influential
  method_family:
  - autoregressive_codec_language_models
  - vae_vector_quantized_codecs
  - hybrid_semantic_acoustic_tokenizers
  claims:
  - claim_id: discrete_codec_representations_introduce_a_quantifiable_fidelity_loss_relative_to
    role: supports
    claim: Discrete codec representations introduce a quantifiable fidelity loss relative to continuous mel-spectrogram
      representations even at high codebook counts, measurable in both WER and speaker similarity.
    source: §5, Table 1
    evidence: Discrete codec representations introduce a quantifiable fidelity loss relative to continuous mel-spectrogram
      representations even at high codebook counts, measurable in both WER and speaker similarity.
    confidence: high
    relevance: high
  - claim_id: continuous_valued_autoregressive_speech_synthesis_can_achieve_robustness_and_naturalness
    role: supports
    claim: Continuous-valued autoregressive speech synthesis can achieve robustness and naturalness on par with
      codec-based two-stage systems when paired with appropriate regularization objectives.
    source: §5.1, §5.2, Table 1, Table 3
    evidence: Continuous-valued autoregressive speech synthesis can achieve robustness and naturalness on par with
      codec-based two-stage systems when paired with appropriate regularization objectives.
    confidence: high
    relevance: high
  - claim_id: a_variational_latent_sampling_module_applied_to_continuous_spectrogram_prediction
    role: supports
    claim: A variational latent sampling module applied to continuous spectrogram prediction provides diversity
      and robustness benefits analogous to top-p sampling for discrete tokens, without the instability caused by
      the high similarity of consecutive acoustic codes.
    source: §3.2.2, §5.3, Table 4
    evidence: A variational latent sampling module applied to continuous spectrogram prediction provides diversity
      and robustness benefits analogous to top-p sampling for discrete tokens, without the instability caused by
      the high similarity of consecutive acoustic codes.
    confidence: high
    relevance: medium
  - claim_id: bypassing_the_non_autoregressive_second_stage_in_codec_language_model
    role: supports
    claim: Bypassing the non-autoregressive second stage in codec language model pipelines reduces inference time
      while maintaining competitive output quality.
    source: §5.4, Table 5
    evidence: Bypassing the non-autoregressive second stage in codec language model pipelines reduces inference
      time while maintaining competitive output quality.
    confidence: high
    relevance: high
  - claim_id: prediction_quality_in_continuous_valued_autoregressive_tts_degrades_gracefully_with
    role: complicates
    claim: Prediction quality in continuous-valued autoregressive TTS degrades gracefully with reduction factor
      increases, enabling a controllable quality-efficiency trade-off unavailable in discrete-token systems.
    source: §5.1, Table 1, Table 2
    evidence: Prediction quality in continuous-valued autoregressive TTS degrades gracefully with reduction factor
      increases, enabling a controllable quality-efficiency trade-off unavailable in discrete-token systems.
    confidence: high
    relevance: medium
  limitations:
  - '- English-only evaluation; multilingual extension not attempted. - Vocoder quality bottleneck: uses open-source
    HiFi-GAN trained on 585h LibriTTS; Voicebox''s proprietary vocoder trained on 60Kh provides higher quality ceiling.
    - Mel-spectrogram as the only continuous representation explored; VAE latent states suggested as future work.
    - SMOS exceeding ground truth may partly reflect the test setup''s limitation (inter-speaker/inter-session variation
    in the reference set rather than genuine quality superiority). - No streaming or low-latency inference analysis.'
  caveats: []
- id: 2025.acl-long.654
  published_date: "2025-07-27"
  entry_date: '2026-07-28'
  year: 2025
  venue: ACL
  task:
  - TTS
  - codec
  architecture:
  - GAN
  - VAE
  relevance: high
  evidence_role:
  - core_evidence
  - control_evidence
  current_role: active_evidence
  method_family:
  - gan_neural_codecs_and_decoders
  - vae_vector_quantized_codecs
  claims:
  - claim_id: codec_token_distributions_in_the_first_rvq_channel_are_a
    role: supports
    claim: Codec token distributions in the first RVQ channel are a meaningful bottleneck for autoregressive generation
      from text, independent of reconstruction quality.
    source: §1, §3.3
    evidence: Codec token distributions in the first RVQ channel are a meaningful bottleneck for autoregressive
      generation from text, independent of reconstruction quality.
    confidence: high
    relevance: high
  - claim_id: redistributing_information_load_uniformly_across_the_first_few_rvq_codebook
    role: supports
    claim: Redistributing information load uniformly across the first few RVQ codebook channels via parallel masked
      quantization consistently improves speaker similarity in downstream autoregressive TTS.
    source: §4.4, Table 3
    evidence: Redistributing information load uniformly across the first few RVQ codebook channels via parallel
      masked quantization consistently improves speaker similarity in downstream autoregressive TTS.
    confidence: high
    relevance: high
  - claim_id: a_fourier_based_decoder_with_a_self_attention_module_achieves
    role: supports
    claim: A Fourier-based decoder with a self-attention module achieves better codec reconstruction quality than
      a transposed-convolution upsampler, without length extrapolation issues.
    source: §3.2, Appendix G, Table 9
    evidence: A Fourier-based decoder with a self-attention module achieves better codec reconstruction quality
      than a transposed-convolution upsampler, without length extrapolation issues.
    confidence: high
    relevance: high
  - claim_id: codec_reconstruction_quality_does_not_scale_substantially_with_training_data
    role: supports
    claim: Codec reconstruction quality does not scale substantially with training data volume beyond a few hundred
      hours, while domain generalization does benefit from larger and more diverse datasets.
    source: Appendix A, Table 5
    evidence: Codec reconstruction quality does not scale substantially with training data volume beyond a few hundred
      hours, while domain generalization does benefit from larger and more diverse datasets.
    confidence: high
    relevance: high
  - claim_id: the_choice_of_underlying_codec_has_a_larger_impact_on
    role: supports
    claim: The choice of underlying codec has a larger impact on zero-shot TTS speaker similarity than on intelligibility,
      with codec swaps producing 10–15% SPK-SIM gains while WER differences remain within noise.
    source: §4.3, Table 2
    evidence: The choice of underlying codec has a larger impact on zero-shot TTS speaker similarity than on intelligibility,
      with codec swaps producing 10–15% SPK-SIM gains while WER differences remain within noise.
    confidence: high
    relevance: high
  limitations:
  - '- Language-Codec is trained and evaluated exclusively on speech; audio, music, and environmental sound domains
    are explicitly left as future work. The codec''s suitability for general audio language models is therefore
    unvalidated. - MCRVQ prediction accuracy drops when more than 4 codebook channels are used in downstream models,
    suggesting the information-redistribution benefit weakens at higher bitrates. The mechanism for this degradation
    is not fully explained. - The paper evaluates downstream quality only with VALL-E and MobileSpeech; it is unclear
    whether the SPK-SIM gains extend to flow-matching or diffusion-based TTS backends. - No demo page is linked
    in the paper, making it difficult to subjectively verify the quality claims beyond the crowd-sourced MOS. -
    The internal 20,000-hour Chinese dataset is not publicly available, limiting full reproducibility of the 50k-hour
    training run.'
  caveats: []
- id: 2025.acl-long.682
  published_date: "2025-07-27"
  entry_date: '2026-07-28'
  year: 2025
  venue: ACL
  task:
  - TTS
  - SCA
  - evaluation
  architecture:
  - autoregressive-LM
  - GAN
  - hybrid
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  - infrastructure
  current_role: active_evidence
  method_family:
  - autoregressive_codec_language_models
  - gan_neural_codecs_and_decoders
  - hybrid_semantic_acoustic_tokenizers
  claims:
  - claim_id: gan_based_vocoders_occupy_the_dominant_position_in_production_speech
    role: supports
    claim: GAN-based vocoders occupy the dominant position in production speech generation systems because their
      computational efficiency advantage over autoregressive and diffusion alternatives is orders of magnitude,
      at acceptable perceptual quality cost.
    source: §3.3.1, §F.3, Table 8
    evidence: GAN-based vocoders occupy the dominant position in production speech generation systems because their
      computational efficiency advantage over autoregressive and diffusion alternatives is orders of magnitude,
      at acceptable perceptual quality cost.
    confidence: high
    relevance: medium
  - claim_id: continued_pretraining_from_a_text_language_model_checkpoint_improves_speechlm
    role: supports
    claim: Continued pretraining from a text language model checkpoint improves SpeechLM convergence and downstream
      task performance compared to random initialization.
    source: §4.2.1
    evidence: Continued pretraining from a text language model checkpoint improves SpeechLM convergence and downstream
      task performance compared to random initialization.
    confidence: high
    relevance: medium
  - claim_id: semantic_tokenizers_and_acoustic_tokenizers_represent_complementary_capability_profiles_strong
    role: supports
    claim: Semantic tokenizers and acoustic tokenizers represent complementary capability profiles — strong semantic
      content fidelity versus strong acoustic reconstruction fidelity — and no single tokenizer type dominates both
      dimensions.
    source: §3.1, §F.2, Table 6
    evidence: Semantic tokenizers and acoustic tokenizers represent complementary capability profiles — strong semantic
      content fidelity versus strong acoustic reconstruction fidelity — and no single tokenizer type dominates both
      dimensions.
    confidence: high
    relevance: medium
  - claim_id: post_alignment_of_speechlms_via_preference_optimization_addresses_qualitatively_different
    role: supports
    claim: Post-alignment of SpeechLMs via preference optimization addresses qualitatively different failure modes
      (semantic inconsistency, token distribution mismatch) than post-alignment of text LLMs.
    source: §4.2.3
    evidence: Post-alignment of SpeechLMs via preference optimization addresses qualitatively different failure
      modes (semantic inconsistency, token distribution mismatch) than post-alignment of text LLMs.
    confidence: high
    relevance: medium
  - claim_id: full_duplex_spoken_interaction_simultaneous_bidirectional_speech_with_interruption_support
    role: supports
    claim: Full-duplex spoken interaction — simultaneous bidirectional speech with interruption support — requires
      joint modeling of both speaker streams and remains an open research challenge.
    source: §4.3
    evidence: Full-duplex spoken interaction — simultaneous bidirectional speech with interruption support — requires
      joint modeling of both speaker streams and remains an open research challenge.
    confidence: high
    relevance: medium
  limitations:
  - 'Coverage necessarily lags the field: systems published after mid-2024 receive limited treatment, and the overall
    corpus skews heavily toward English and Mandarin. The safety section identifies toxicity and speaker privacy
    risks but does not analyze them quantitatively. End-to-end training that backpropagates gradients from vocoder
    output to tokenizer input is flagged as a potentially high-value research direction but remains unexplored.
    The question of whether incorporating text modality fundamentally benefits or constrains speech intelligence
    is left open.'
  caveats: []
- id: 2025.acl-long.817
  published_date: "2025-07-27"
  entry_date: '2026-07-28'
  year: 2025
  venue: ACL
  task:
  - SCA
  - TTS
  architecture:
  - autoregressive-LM
  relevance: medium
  evidence_role:
  - architecture_variant
  - acceleration_evidence
  - control_evidence
  current_role: active_evidence
  method_family:
  - autoregressive_codec_language_models
  claims:
  - claim_id: boundary_aware_speech_representations_are_critical_for_enabling_offline_trained
    role: supports
    claim: Boundary-aware speech representations are critical for enabling offline-trained speech LLMs to perform
      simultaneous inference via wait-k strategies.
    source: §5.1, §5.2, Tables 4, 11
    evidence: CIF-based boundary-aware prompts outperform fixed downsampling by approximately 4 ASR-BLEU points
      at equivalent latency on Es-En, Fr-En, and De-En CVSS-C test sets, with a parallel ~4 BLEU gap on text output
      confirming the bottleneck is LLM prediction rather than speech synthesis.
    confidence: high
    relevance: medium
  - claim_id: offline_training_combined_with_test_time_simultaneous_inference_policies_can
    role: supports
    claim: Offline training combined with test-time simultaneous inference policies can match or outperform systems
      trained specifically for streaming in speech-to-speech translation.
    source: §5.1, Fig. 4, Table 1
    evidence: SimulS2S-LLM, trained offline, consistently outperforms the streaming-trained StreamSpeech model across
      all three language pairs at comparable latency, achieving up to 4 ASR-BLEU improvement on Es-En.
    confidence: high
    relevance: low
  - claim_id: aggregating_llm_hidden_states_across_multiple_layers_improves_discrete_speech
    role: supports
    claim: Aggregating LLM hidden states across multiple layers improves discrete speech token prediction compared
      to using only the final layer.
    source: §5.3, Fig. 6
    evidence: Multi-layer hidden state weighting yields approximately 1 ASR-BLEU improvement over last-layer-only
      decoding on CVSS-C Es-En Simul-S2ST, attributed to the final layer's focus on semantic text information at
      the expense of acoustic richness needed for speech token generation.
    confidence: high
    relevance: high
  - claim_id: llm_based_approaches_to_simultaneous_speech_generation_face_a_latency
    role: complicates
    claim: LLM-based approaches to simultaneous speech generation face a latency penalty from LLM inference overhead
      that narrows the practical quality-latency advantage over non-LLM methods.
    source: §D, Tables 8-10, Limitations
    evidence: Computation-aware ATD for SimulS2S-LLM is substantially higher than standard ATD (e.g., 4239ms vs.
      3440ms at k=8 on Es-En), and the system is not evaluated at very low latency regimes (AL < 1s) where reordering
      requirements make offline-trained models unsuitable.
    confidence: high
    relevance: low
  - claim_id: shallow_fusion_of_n_gram_language_models_with_ctc_decoding
    role: supports
    claim: Shallow fusion of n-gram language models with CTC decoding of discrete speech tokens improves simultaneous
      speech translation quality without increasing latency.
    source: §5.4, Table 2
    evidence: n-gram LM fusion over greedy CTC search improves ASR-BLEU from 24.7 to 26.3 on CVSS-C Es-En at identical
      ATD of 3439ms.
    confidence: high
    relevance: high
  limitations:
  - The system is not evaluated at very low latency (AL < 1s), a regime the authors identify as unsuitable for offline-trained
    models due to reordering requirements. This excludes SimulS2S-LLM from the most latency-critical applications.
    Computation-aware latency is substantially higher than the reported ATD, and all experiments use 7B/8B open-source
    LLMs on small datasets (70-174 hours per language pair), leaving scalability to larger models and data unverified.
  - The evaluation is limited to three European language pairs in a single translation direction each. Language
    pairs with greater structural divergence or more extensive reordering would stress the wait-k assumption more
    severely. The system has not been evaluated on offline inference or tasks other than S2ST and S2TT, despite
    the claim that offline training preserves such capabilities. Long-form simultaneous speech translation is also
    untested due to lack of suitable data.
  caveats: []
- id: 2025.acl-long.87
  published_date: "2025-07-27"
  entry_date: '2026-07-28'
  year: 2025
  venue: ACL
  task:
  - VC
  architecture:
  - flow-matching
  relevance: medium
  evidence_role:
  - architecture_variant
  - acceleration_evidence
  - control_evidence
  current_role: active_evidence
  method_family:
  - flow_matching_codec_decoders
  claims:
  - claim_id: combining_asr_derived_phonetic_features_and_quantized_self_supervised_representations
    role: supports
    claim: Combining ASR-derived phonetic features and quantized self-supervised representations via adaptive fusion
      reduces timbre leakage while preserving paralinguistic content in zero-shot voice conversion.
    source: §5.3, Table 3
    evidence: Removing the PPG branch (WavLM-only) causes SMOS to drop from 4.11 to 3.07 and SECS from 0.71 to 0.45
      on LibriTTS, indicating that SSL features alone carry substantial timbre leakage; removing the SSL branch
      degrades NMOS and WER, confirming PPGs alone lose paralinguistic richness.
    confidence: high
    relevance: medium
  - claim_id: flow_matching_provides_faster_inference_than_diffusion_based_voice_conversion
    role: supports
    claim: Flow matching provides faster inference than diffusion-based voice conversion systems without sacrificing
      speaker similarity or naturalness.
    source: §5.1, Table 1
    evidence: Takin-VC achieves RTF 0.154, lower than DiffVC (0.294), NS2VC (0.347), and SeedVC (0.341), while simultaneously
      outperforming these baselines on NMOS, SMOS, and SECS.
    confidence: high
    relevance: low
  - claim_id: global_time_invariant_speaker_embeddings_are_insufficient_for_robust_timbre
    role: complicates
    claim: Global, time-invariant speaker embeddings are insufficient for robust timbre modeling in expressive zero-shot
      voice conversion.
    source: §5.3, Table 4
    evidence: Removing the context-aware cross-attention module (which aligns source content with target timbre
      dynamically) drops SMOS from 4.11 to 3.61 and SECS from 0.71 to 0.58, while the memory-augmented module removal
      drops SECS to 0.52. Both modules provide content-sensitive timbre conditioning beyond a static speaker embedding
      alone.
    confidence: high
    relevance: medium
  - claim_id: cross_gender_voice_conversion_consistently_yields_lower_speaker_similarity_than
    role: complicates
    claim: Cross-gender voice conversion consistently yields lower speaker similarity than same-gender conversion
      even in well-trained systems.
    source: §5.2, Table 2
    evidence: 'On the large-scale multilingual dataset, same-gender pairs (F2F: SECS 0.74; M2M: 0.73) outperform
      cross-gender pairs (F2M: 0.71; M2F: 0.70) in speaker embedding cosine similarity, a gap that persists across
      all conversion directions.'
    confidence: high
    relevance: low
  - claim_id: quantizing_self_supervised_speech_features_before_content_encoding_reduces_timbre
    role: refines
    claim: Quantizing self-supervised speech features before content encoding reduces timbre leakage more effectively
      than using continuous SSL representations directly.
    source: §3.2, §5.3, Table 3
    evidence: The RVQ quantizer (codebook size 8,200) applied to WavLM features is the key mechanism for timbre
      suppression in the hybrid encoder; ablation with WavLM-only (continuous features without adaptive fusion)
      shows SECS drops to 0.45 compared to 0.71 for the full model, consistent with timbre leakage from unquantized
      SSL features.
    confidence: high
    relevance: medium
  limitations:
  - All large-scale training data (500k hours) and the 100-speaker evaluation set are proprietary and not publicly
    available. The large-scale results cannot be reproduced by external researchers, and it is unclear how much
    of the gain over competitive baselines is attributable to data scale rather than the proposed modules.
  - The paper does not include targeted evaluation of paralinguistic preservation (breathing, crying, emotion transfer),
    despite listing this as a primary contribution. NMOS and SMOS measure general naturalness and speaker similarity
    but are not designed to capture expressive fidelity specifically. Speech editing under zero-shot conditions
    is acknowledged as out of scope and a direction for future work. Ethical risks from voice impersonation are
    noted but no technical mitigations are proposed.
  caveats: []
- id: 2025.acl-long.912
  published_date: "2025-07-27"
  entry_date: '2026-07-28'
  year: 2025
  venue: ACL
  task:
  - SCA
  - TTS
  architecture:
  - autoregressive-LM
  - flow-matching
  - hybrid
  relevance: medium
  evidence_role:
  - architecture_variant
  - acceleration_evidence
  - control_evidence
  current_role: active_evidence
  method_family:
  - autoregressive_codec_language_models
  - flow_matching_codec_decoders
  - hybrid_semantic_acoustic_tokenizers
  claims:
  - claim_id: autoregressive_streaming_speech_decoders_in_modular_spoken_conversational_agents_substantially
    role: supports
    claim: Autoregressive streaming speech decoders in modular spoken conversational agents substantially improve
      naturalness over non-autoregressive alternatives with comparable latency.
    source: §5.1, Table 1
    evidence: Autoregressive streaming speech decoders in modular spoken conversational agents substantially improve
      naturalness over non-autoregressive alternatives with comparable latency.
    confidence: high
    relevance: low
  - claim_id: a_gate_fusion_mechanism_that_adaptively_blends_llm_hidden_states
    role: supports
    claim: A gate fusion mechanism that adaptively blends LLM hidden states with text token embeddings as input
      to the TTS language model improves both instruction following quality and text-speech consistency over simpler
      additive fusion.
    source: §5.2, Table 2
    evidence: A gate fusion mechanism that adaptively blends LLM hidden states with text token embeddings as input
      to the TTS language model improves both instruction following quality and text-speech consistency over simpler
      additive fusion.
    confidence: high
    relevance: low
  - claim_id: streaming_tts_pretraining_on_speech_dialogue_data_is_critical_for
    role: complicates
    claim: 'Streaming TTS pretraining on speech dialogue data is critical for quality: initializing from a text-only
      pretrained model degrades performance substantially, and training from scratch fails to converge.'
    source: §5.2, Table 3
    evidence: 'Streaming TTS pretraining on speech dialogue data is critical for quality: initializing from a text-only
      pretrained model degrades performance substantially, and training from scratch fails to converge.'
    confidence: high
    relevance: low
  - claim_id: multi_turn_dialogue_training_data_consistently_outperforms_single_turn_data
    role: supports
    claim: Multi-turn dialogue training data consistently outperforms single-turn data of the same total size for
      modular speech language models across spoken QA and instruction-following benchmarks.
    source: §5.3, Table 5
    evidence: Multi-turn dialogue training data consistently outperforms single-turn data of the same total size
      for modular speech language models across spoken QA and instruction-following benchmarks.
    confidence: high
    relevance: low
  - claim_id: the_s2t_to_s2s_accuracy_gap_in_spoken_question_answering
    role: supports
    claim: The S2T-to-S2S accuracy gap in spoken question answering is a meaningful indicator of speech generation
      quality, and autoregressive TTS decoders reduce this gap relative to non-autoregressive alternatives.
    source: §5.1, Table 1
    evidence: The S2T-to-S2S accuracy gap in spoken question answering is a meaningful indicator of speech generation
      quality, and autoregressive TTS decoders reduce this gap relative to non-autoregressive alternatives.
    confidence: high
    relevance: medium
  limitations:
  - The model generates speech in a single fixed output style; it cannot modulate emotion, speaking rate, or dialect
    in response to the content or paralinguistic cues of the input speech. All evaluation is in English. The output
    voice is fixed during training (a single uniform voice for all responses), which limits expressiveness and speaker
    diversity. The system is inherently a response-after-input architecture and does not support full-duplex conversation.
    Latency, while adequate for real-time interaction, leaves room for further reduction via engineering optimization.
    Whether the gate fusion approach generalizes to multilingual settings or to more expressive speech styles is
    not explored.
  caveats: []
- id: 2025.acl-long.937
  published_date: "2025-07-27"
  entry_date: '2026-07-28'
  year: 2025
  venue: ACL
  task:
  - codec
  architecture:
  - VAE
  relevance: medium
  evidence_role:
  - architecture_variant
  current_role: active_evidence
  method_family:
  - vae_vector_quantized_codecs
  claims:
  - claim_id: domain_adaptive_partitioning_within_a_shared_codebook_resolves_the_performance
    role: supports
    claim: Domain-adaptive partitioning within a shared codebook resolves the performance degradation that plagues
      unified single-codebook codecs on mixed multi-domain audio.
    source: §3.2, §5.3, Tables 2, 4, 6
    evidence: Removing the partitioned codebook raises Mel Distance from 0.344 to 0.487 on speech and from 0.396
      to 0.506 on music in ablation; the full UniCodec outperforms WavTokenizer (unified) on all three domains in
      both objective and MUSHRA evaluation.
    confidence: high
    relevance: high
  - claim_id: self_supervised_mask_prediction_objectives_integrated_directly_into_codec_training
    role: supports
    claim: Self-supervised mask prediction objectives integrated directly into codec training improve semantic representation
      without requiring external pretrained SSL encoders.
    source: §3.4, §5.2, Table 5
    evidence: The semantic training stage improves RAVDESS classification accuracy from 36.81 to 40.28 and Audio-MNIST
      from 69.84 to 70.94 compared to the codec without this stage; no auxiliary SSL model is used.
    confidence: high
    relevance: high
  - claim_id: joint_training_of_reconstruction_and_self_supervised_semantic_objectives_from
    role: complicates
    claim: Joint training of reconstruction and self-supervised semantic objectives from scratch degrades single-codebook
      codec performance; sequential staging is required.
    source: §3.4
    evidence: Preliminary experiments show that training reconstruction, mask prediction, and contrastive loss simultaneously
      is infeasible; the two-stage approach (acoustic training first, then semantic stage) is necessary for stable
      convergence.
    confidence: high
    relevance: high
  - claim_id: large_scale_diverse_audio_training_data_introduces_noise_that_degrades
    role: complicates
    claim: Large-scale diverse audio training data introduces noise that degrades codec reconstruction quality,
      requiring targeted fine-tuning on curated high-quality subsets.
    source: §4, §5.3, Appendix C, Table 6
    evidence: Training on 80K hours of mixed data without the fine-tuning stage raises Mel Distance by 0.103 on
      speech (0.448 vs 0.344) relative to the fine-tuned model; high-quality fine-tuning on LibriTTS clean, VCTK,
      and LJSpeech recovers most of the degradation.
    confidence: high
    relevance: high
  - claim_id: single_codebook_codecs_at_75_tokens_per_second_can_achieve
    role: supports
    claim: Single-codebook codecs at 75 tokens per second can achieve acoustic reconstruction quality competitive
      with multi-layer RVQ codecs operating at 600 tokens per second when combined with domain-adaptive design.
    source: §5.1, Table 3
    evidence: UniCodec PESQ 3.03 and STOI 0.949 on LibriTTS test-clean exceeds Encodec (2.72, 0.939) and SpeechTokenizer
      (2.61, 0.917) at 600 TPS, with eight times fewer tokens per second.
    confidence: high
    relevance: high
  limitations:
  - 'Training is sensitive to noisy or low-quality input: large-scale noisy data alone degrades reconstruction quality,
    and the fine-tuning stage is essential. This dependency on curated data limits scalability.'
  - The single-codebook codec struggles to balance both acoustic reconstruction fidelity and semantic density across
    diverse domains simultaneously, as these objectives conflict. Streaming performance degrades relative to non-streaming
    inference. The paper does not demonstrate integration with an actual audio language model, leaving downstream
    generation quality unvalidated.
  caveats: []
- id: 2025.acl-short.81
  published_date: "2025-07-27"
  entry_date: '2026-07-28'
  year: 2025
  venue: ACL
  task:
  - TTS
  - evaluation
  architecture:
  - autoregressive-LM
  - transformer-enc-dec
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: active_evidence
  method_family:
  - autoregressive_codec_language_models
  - transformer_encoder_decoder_tokenizers
  claims:
  - claim_id: long_form_training_audio_10_20_seconds_per_segment_with
    role: supports
    claim: Long-form training audio (10–20 seconds per segment) with explicit speaker identities improves zero-shot
      TTS quality for low-resource tonal languages compared to training on short-segment corpora.
    source: §3.1, §4
    evidence: Long-form training audio (10–20 seconds per segment) with explicit speaker identities improves zero-shot
      TTS quality for low-resource tonal languages compared to training on short-segment corpora.
    confidence: high
    relevance: medium
  - claim_id: multilingual_voice_cloning_models_such_as_xtts_v2_exhibit_architectural
    role: supports
    claim: Multilingual voice-cloning models such as XTTS-v2 exhibit architectural failure modes on short input
      sequences that are not corrected by data augmentation with short clips.
    source: §4
    evidence: Multilingual voice-cloning models such as XTTS-v2 exhibit architectural failure modes on short input
      sequences that are not corrected by data augmentation with short clips.
    confidence: high
    relevance: medium
  - claim_id: autoregressive_codec_language_models_vall_e_voicecraft_generalize_better_than
    role: supports
    claim: Autoregressive codec language models (VALL-E, VoiceCraft) generalize better than Tortoise-based models
      to short-sentence inputs in low-resource language fine-tuning.
    source: §4, Table 2
    evidence: Autoregressive codec language models (VALL-E, VoiceCraft) generalize better than Tortoise-based models
      to short-sentence inputs in low-resource language fine-tuning.
    confidence: high
    relevance: high
  - claim_id: a_dataset_curation_pipeline_based_on_dual_asr_agreement_filtering
    role: supports
    claim: A dataset curation pipeline based on dual-ASR agreement filtering produces higher-quality transcriptions
      for audiobook audio than single-model transcription alone, enabling more reliable TTS training.
    source: §2.1
    evidence: A dataset curation pipeline based on dual-ASR agreement filtering produces higher-quality transcriptions
      for audiobook audio than single-model transcription alone, enabling more reliable TTS training.
    confidence: high
    relevance: medium
  limitations:
  - The paper does not evaluate code-switching scenarios (mixed Vietnamese-English input), which is relevant in
    practice. The dataset is audiobook domain only, so speaking style coverage is narrower than general-purpose
    datasets. All models are fine-tuned rather than trained from scratch, which means performance is bounded by
    the pre-trained model's multilingual capacity. The architecture issue observed with XTTS-v2 on short sentences
    is identified but not resolved. The dataset is released for non-commercial use only, which limits industrial
    adoption. It is also unclear how the system handles tonal phonology beyond phonemizer outputs, and no ablation
    on augmented data proportion is presented.
  caveats: []
- id: 2025.findings-acl.1051
  published_date: "2025-07-27"
  entry_date: '2026-07-28'
  year: 2025
  venue: ACL
  task:
  - TTS
  - SCA
  architecture:
  - autoregressive-LM
  relevance: high
  evidence_role:
  - core_evidence
  - acceleration_evidence
  - control_evidence
  current_role: active_evidence
  method_family:
  - autoregressive_codec_language_models
  claims:
  - claim_id: decoupling_speech_synthesis_from_llm_text_generation_via_a_lightweight
    role: supports
    claim: Decoupling speech synthesis from LLM text generation via a lightweight autoregressive module can preserve
      the base LLM's reasoning quality while achieving competitive streaming latency.
    source: §6.4, Table 1
    evidence: Decoupling speech synthesis from LLM text generation via a lightweight autoregressive module can preserve
      the base LLM's reasoning quality while achieving competitive streaming latency.
    confidence: high
    relevance: low
  - claim_id: end_to_end_speech_enabled_llms_that_fine_tune_or
    role: supports
    claim: End-to-end speech-enabled LLMs that fine-tune or condition the base LLM on speech data consistently show
      degraded language understanding compared to systems that keep the LLM frozen.
    source: §6.4, Table 1
    evidence: End-to-end speech-enabled LLMs that fine-tune or condition the base LLM on speech data consistently
      show degraded language understanding compared to systems that keep the LLM frozen.
    confidence: high
    relevance: medium
  - claim_id: a_single_layer_rvq_codec_is_sufficient_for_low_latency
    role: supports
    claim: A single-layer RVQ codec is sufficient for low-latency autoregressive TTS generation when paired with
      a compact decoder transformer, avoiding the complexity of multi-codebook prediction.
    source: §3.1
    evidence: A single-layer RVQ codec is sufficient for low-latency autoregressive TTS generation when paired with
      a compact decoder transformer, avoiding the complexity of multi-codebook prediction.
    confidence: high
    relevance: high
  - claim_id: streaming_tts_quality_improves_with_larger_decode_chunk_sizes_with
    role: supports
    claim: Streaming TTS quality improves with larger decode chunk sizes, with WER and UTMOS gains achievable without
      substantially increasing end-to-end latency.
    source: §6.4, Figure 6
    evidence: Streaming TTS quality improves with larger decode chunk sizes, with WER and UTMOS gains achievable
      without substantially increasing end-to-end latency.
    confidence: high
    relevance: low
  - claim_id: language_adaptation_of_a_codec_based_tts_module_can_be
    role: supports
    claim: Language adaptation of a codec-based TTS module can be achieved by replacing training data alone, without
      architectural changes or explicit G2P conversion for the new language.
    source: §6.5, Table 3
    evidence: Language adaptation of a codec-based TTS module can be achieved by replacing training data alone,
      without architectural changes or explicit G2P conversion for the new language.
    confidence: high
    relevance: high
  limitations:
  - LLMVoX is single-speaker — no voice cloning or speaker reference support. The Arabic model was trained on XTTS-synthesized
    data, so XTTS acts as an upper bound (CER 1.7% vs. LLMVoX 8.2%). The streaming pipeline does not yet extend
    to the ASR front-end. Latency with 70B LLMs exceeds 1.9s, making real-time use marginal. The quality improvement
    from larger chunk sizes (UTMOS 3.75→4.41) suggests that the 475ms latency figure is somewhat optimistic for
    maximum quality operation.
  caveats: []
- id: 2025.findings-acl.115
  published_date: "2025-07-27"
  entry_date: '2026-07-28'
  year: 2025
  venue: ACL
  task:
  - SCA
  - TTS
  architecture:
  - autoregressive-LM
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: active_evidence
  method_family:
  - autoregressive_codec_language_models
  claims:
  - claim_id: grouping_discrete_semantic_tokens_during_autoregressive_generation_reduces_training_and
    role: supports
    claim: Grouping discrete semantic tokens during autoregressive generation reduces training and inference costs
      by alleviating the frequency mismatch between text and audio token streams.
    source: §5.4.1, Table 6
    evidence: Semantic Group Modeling with G=3 reduces SLAM-Omni training from 126 to 60 GPU hours and ASR-WER from
      18.23% (G=1) to 4.54%, with G=3 providing the best quality-efficiency trade-off across five group sizes tested.
    confidence: high
    relevance: high
  - claim_id: multi_stage_pre_training_on_modality_specific_tasks_for_spoken
    role: refines
    claim: Multi-stage pre-training on modality-specific tasks for spoken dialogue systems can degrade instruction-following
      ability despite improving modality alignment metrics.
    source: §5.4.2, Table 7
    evidence: ASR pre-training reduces ChatGPT Score from 39.32 to 34.02 and TTS pre-training to 27.22, while ASR-WER
      improves only marginally (4.38% / 4.53% vs. 4.54%), indicating that modality-specific pre-training shifts
      the model away from general instruction-following capabilities.
    confidence: high
    relevance: low
  - claim_id: zero_shot_timbre_control_in_spoken_dialogue_systems_achieves_competitive
    role: complicates
    claim: Zero-shot timbre control in spoken dialogue systems achieves competitive speaker similarity but remains
      constrained by training data volume relative to dedicated TTS systems.
    source: §5.2, Table 5
    evidence: SLAM-Omni reaches SIM-o of 0.517, comparable to FireRedTTS (0.486) but below CosyVoice2 (0.684), with
      the gap attributed to approximately 50x less training data; limited and less diverse data cause ambiguities
      in timbre rendering during vocoder synthesis.
    confidence: high
    relevance: low
  - claim_id: compressing_multi_turn_dialogue_history_to_text_representation_sacrifices_non
    role: complicates
    claim: Compressing multi-turn dialogue history to text representation sacrifices non-verbal paralinguistic context
      that may be important for maintaining dialogue coherence across turns.
    source: §3.5, §6 Limitations
    evidence: Historical Text Prompting stores only text history, explicitly trading away emotional and prosodic
      signals from previous dialogue turns for computational efficiency; the paper identifies this as a primary
      limitation affecting scenarios requiring sustained dialogue coherence.
    confidence: high
    relevance: high
  - claim_id: semantic_token_based_speech_generation_in_spoken_dialogue_systems_provides
    role: supports
    claim: Semantic token-based speech generation in spoken dialogue systems provides tighter speech-text alignment
      than acoustic codec-based approaches, as measured by word error rate between generated speech and corresponding
      text.
    source: §5.1, Table 3
    evidence: SLAM-Omni achieves the lowest ASR-WER (4.54%) among all evaluated spoken dialogue models, outperforming
      larger models including Moshi (7.18%), GLM-4-Voice (12.71%), and Freeze-Omni (16.32%), using single-layer
      semantic tokens rather than multi-codebook acoustic tokens.
    confidence: high
    relevance: high
  limitations:
  - The finding that single-stage training outperforms multi-stage pre-training is demonstrated only at 0.5B scale
    with limited data (400K utterances). The authors explicitly note that extending to larger LLMs would require
    substantially more training data, and whether the result holds at larger scales or with more diverse corpora
    is untested.
  - 'Historical text prompting provides computational efficiency at the cost of discarding non-verbal paralinguistic
    signals (emotion, prosody) from previous dialogue turns. This limits the system''s ability to maintain affective
    continuity across multi-turn conversations. The evaluation is also narrow: the custom 8-task benchmark measures
    general spoken interaction but not domain-specific or emotion-aware dialogue quality.'
  caveats: []
- id: 2025.findings-ijcnlp.49
  published_date: "2025-07-27"
  entry_date: '2026-07-28'
  year: 2025
  venue: ACL
  task:
  - SCA
  architecture:
  - autoregressive-LM
  relevance: medium
  evidence_role:
  - architecture_variant
  - acceleration_evidence
  - control_evidence
  current_role: active_evidence
  method_family:
  - autoregressive_codec_language_models
  claims:
  - claim_id: a_full_duplex_speech_language_model_can_perform_structured_dialogue
    role: supports
    claim: A full-duplex speech language model can perform structured dialogue state tracking when DST is executed
      entirely through the model's text stream using explicit delimiter tokens.
    source: §3, §5.3, Table 1
    evidence: J-Moshi-ext fine-tuned on synthesized JMultiWOZ data achieves JGA 65.69 and Slot F1 97.78 on the JMultiWOZ
      test set, with both metrics improving monotonically as training data volume increases from 25% to 100%.
    confidence: high
    relevance: low
  - claim_id: extending_full_duplex_speech_models_to_task_oriented_dialogue_does
    role: complicates
    claim: Extending full-duplex speech models to task-oriented dialogue does not close the gap with text-based
      systems on response generation quality.
    source: §5.3, Table 1
    evidence: The proposed method reaches BLEU 0.129 and BERTScore 0.689, against text-based upper bounds of BLEU
      0.364 and BERTScore 0.830 (T5-large), with the gap attributed to insufficient training data and to the difficulty
      of aligning text and audio modalities during response generation.
    confidence: high
    relevance: low
  - claim_id: tts_data_quality_limitations_compound_into_downstream_training_noise_when
    role: complicates
    claim: TTS data quality limitations compound into downstream training noise when synthesized speech is used
      to construct full-duplex dialogue training corpora.
    source: §5.3
    evidence: The multilingual OuteTTS model used for Japanese speech synthesis generated speech with reduced clarity,
      causing recognition errors during time-annotated tokenization that hindered the model's learning of coherent
      text-audio token sequences for response generation.
    confidence: high
    relevance: low
  - claim_id: dst_timing_estimation_is_feasible_within_a_full_duplex_spoken
    role: supports
    claim: DST timing estimation is feasible within a full-duplex spoken dialogue model but introduces substantial
      prediction error tied to the model's speech perception accuracy.
    source: §5.1, §5.3
    evidence: DST initiation timing (defined as the time of `<bs>` token generation) achieved a MAE of 4.9 seconds,
      with discrepancies attributed to the difficulty of accurately predicting user utterance boundaries from the
      model's audio token representations alone.
    confidence: high
    relevance: low
  limitations:
  - All experiments are conducted in Japanese using J-Moshi, which was trained on less Japanese data than the original
    English Moshi. The authors explicitly note that performance on English equivalents (MultiWOZ, SpokenWOZ) could
    be substantially higher, and cross-lingual generalisability is untested.
  - Evaluation is fully automatic (JGA, Slot F1, BLEU, BERTScore); no subjective listening tests or user studies
    are reported, leaving response naturalness, latency under real deployment conditions, and the effect of DST
    pauses on conversational flow unassessed. The experiment covers only the domains included in JMultiWOZ (travel
    planning), and extension to diverse or open-domain settings is planned but not yet evaluated. Future work includes
    experiments with the original Moshi on SpokenWOZ and DSTC11 datasets and application to additional dialogue
    domains.
  caveats: []
- id: 2025.ccl-1.80
  published_date: "2025-08-01"
  entry_date: '2026-07-28'
  year: 2025
  venue: workshop
  task:
  - TTS
  architecture:
  - autoregressive-LM
  - hybrid
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: active_evidence
  method_family:
  - autoregressive_codec_language_models
  - hybrid_semantic_acoustic_tokenizers
  claims:
  - claim_id: projecting_phoneme_representations_from_two_languages_into_a_shared_latent
    role: supports
    claim: Projecting phoneme representations from two languages into a shared latent space reduces cross-lingual
      phoneme confusion and improves naturalness in code-switched speech synthesis.
    source: §3.2, Table 2
    evidence: Projecting phoneme representations from two languages into a shared latent space reduces cross-lingual
      phoneme confusion and improves naturalness in code-switched speech synthesis.
    confidence: high
    relevance: medium
  - claim_id: per_token_language_id_conditioning_helps_a_multilingual_codec_lm
    role: supports
    claim: Per-token language ID conditioning helps a multilingual codec LM distinguish phonetic characteristics
      across languages in code-switched synthesis, though its impact is smaller than that of shared phoneme representations.
    source: §3.3, Table 4
    evidence: Per-token language ID conditioning helps a multilingual codec LM distinguish phonetic characteristics
      across languages in code-switched synthesis, though its impact is smaller than that of shared phoneme representations.
    confidence: high
    relevance: high
  - claim_id: codec_language_model_architectures_outperform_vae_based_seq2seq_systems_for
    role: supports
    claim: Codec language model architectures outperform VAE-based seq2seq systems for code-switched TTS when trained
      on monolingual data only.
    source: §4.3.2, Table 3
    evidence: Codec language model architectures outperform VAE-based seq2seq systems for code-switched TTS when
      trained on monolingual data only.
    confidence: high
    relevance: high
  - claim_id: code_switched_tts_systems_can_be_trained_effectively_from_monolingual
    role: supports
    claim: Code-switched TTS systems can be trained effectively from monolingual corpora alone, without requiring
      real bilingual or code-switched training audio.
    source: §3.1, §4.1
    evidence: Code-switched TTS systems can be trained effectively from monolingual corpora alone, without requiring
      real bilingual or code-switched training audio.
    confidence: high
    relevance: medium
  limitations:
  - '- Proprietary Lao dataset is not publicly available, limiting reproducibility. - Evaluation uses a small number
    of listeners (10 Lao, 10 English), raising statistical concerns. - RMSE is used as the sole objective metric;
    no WER, CER, or SPK-SIM reported. - The approach does not generalize beyond Lao-English without new proprietary
    data per language. - No streaming or real-time inference analysis. - How the model handles intra-word code-switching
    (vs. inter-sentence) is not evaluated.'
  caveats: []
- id: '2507.22746'
  published_date: "2025-08-01"
  entry_date: '2026-07-28'
  year: 2025
  venue: arXiv
  task:
  - TTS
  architecture:
  - hybrid
  relevance: medium
  evidence_role:
  - architecture_variant
  - acceleration_evidence
  - control_evidence
  current_role: active_evidence
  method_family:
  - hybrid_semantic_acoustic_tokenizers
  claims:
  - claim_id: composing_autoregressive_generation_across_fixed_length_token_blocks_with_parallel
    role: supports
    claim: Composing autoregressive generation across fixed-length token blocks with parallel flow-matching denoising
      within each block can simultaneously provide KV-cache efficiency and bidirectional contextual refinement.
    source: §3.1, Table 3
    evidence: Composing autoregressive generation across fixed-length token blocks with parallel flow-matching denoising
      within each block can simultaneously provide KV-cache efficiency and bidirectional contextual refinement.
    confidence: high
    relevance: medium
  - claim_id: neural_codecs_using_finite_scalar_quantisation_can_preserve_speaker_similarity
    role: supports
    claim: Neural codecs using finite scalar quantisation can preserve speaker similarity and intelligibility at
      frame rates (12.5 Hz) where STFT-based vocoders suffer significant quality degradation.
    source: §4.3.3, Table 4
    evidence: Neural codecs using finite scalar quantisation can preserve speaker similarity and intelligibility
      at frame rates (12.5 Hz) where STFT-based vocoders suffer significant quality degradation.
    confidence: high
    relevance: low
  - claim_id: continuous_denoising_models_can_implicitly_classify_discrete_token_targets_through
    role: supports
    claim: Continuous denoising models can implicitly classify discrete token targets through appropriate embedding
      design, without requiring a separate discrete language model head.
    source: §3.1
    evidence: Continuous denoising models can implicitly classify discrete token targets through appropriate embedding
      design, without requiring a separate discrete language model head.
    confidence: high
    relevance: medium
  - claim_id: reducing_the_token_frame_rate_is_a_more_tractable_path
    role: supports
    claim: Reducing the token frame rate is a more tractable path to low-latency hybrid AR-diffusion TTS than increasing
      diffusion step efficiency alone, given the quadratic scaling of self-attention with sequence length.
    source: §3.1, §4.3.2
    evidence: Reducing the token frame rate is a more tractable path to low-latency hybrid AR-diffusion TTS than
      increasing diffusion step efficiency alone, given the quadratic scaling of self-attention with sequence length.
    confidence: high
    relevance: low
  limitations:
  - No MOS or SMOS listening test results are reported. All quality comparisons use SPK-SIM, WER, and FAD on an
    internal podcast dataset. The absence of subjective evaluation and fair comparison against published baselines
    (VALL-E 2, E2 TTS, NaturalSpeech 3) on a public benchmark makes it impossible to independently verify naturalness
    claims.
  - The model is trained and evaluated on English podcast data only. Generalisation to other languages, controlled
    studio-quality TTS, and expressive speech domains is untested. The podcast use-case naturally emphasises diversity
    and disfluency tolerance over precise prosody control, so the evaluation protocol may not transfer to production
    TTS settings.
  - The codec and acoustic model are proprietary (Microsoft internal), with no public code or demo reported. Reproducibility
    relies entirely on the architectural description in the paper.
  - Mean flow optimisation for step reduction is borrowed from concurrent work (Geng et al. 2025); the sensitivity
    of Dragon-FM quality to NFE count at scale is not fully characterised — ablations cover only 2, 4, 6, 12, and
    24 steps.
  caveats: []
- id: '2508.02849'
  published_date: "2025-08-04"
  entry_date: '2026-07-28'
  year: 2025
  venue: arXiv
  task:
  - codec
  architecture:
  - VAE
  relevance: medium
  evidence_role:
  - architecture_variant
  current_role: active_evidence
  method_family:
  - vae_vector_quantized_codecs
  claims:
  - claim_id: distillation_from_self_supervised_models_such_as_hubert_or_wavlm
    role: supports
    claim: Distillation from self-supervised models such as HuBERT or WavLM does not achieve true semantic disentanglement
      in speech codecs, as these representations inherently retain paralinguistic content.
    source: §II.A, §V.B
    evidence: Distillation from self-supervised models such as HuBERT or WavLM does not achieve true semantic disentanglement
      in speech codecs, as these representations inherently retain paralinguistic content.
    confidence: high
    relevance: medium
  - claim_id: frame_level_cross_modal_contrastive_learning_between_phoneme_and_speech
    role: supports
    claim: Frame-level cross-modal contrastive learning between phoneme and speech representations produces cleaner
      semantic-paralinguistic separation than phoneme classification loss in neural codecs.
    source: §V.D, Table IV
    evidence: Frame-level cross-modal contrastive learning between phoneme and speech representations produces cleaner
      semantic-paralinguistic separation than phoneme classification loss in neural codecs.
    confidence: high
    relevance: medium
  - claim_id: vae_augmented_finite_scalar_quantization_vae_fsq_achieves_substantially_higher
    role: supports
    claim: VAE-augmented finite scalar quantization (VAE+FSQ) achieves substantially higher codebook utilisation
      and reduced long-tail token distribution compared to VQ-VAE in single-codebook speech codecs.
    source: §V.C, Figure 4, Table II
    evidence: VAE-augmented finite scalar quantization (VAE+FSQ) achieves substantially higher codebook utilisation
      and reduced long-tail token distribution compared to VQ-VAE in single-codebook speech codecs.
    confidence: high
    relevance: high
  - claim_id: explicit_modeling_of_paralinguistic_information_as_a_reconstruction_bridge_between
    role: supports
    claim: Explicit modeling of paralinguistic information as a reconstruction bridge between semantic and acoustic
      encodings improves both semantic completeness and reconstruction fidelity in low-bitrate streaming codecs.
    source: §III.B, §V.B
    evidence: Explicit modeling of paralinguistic information as a reconstruction bridge between semantic and acoustic
      encodings improves both semantic completeness and reconstruction fidelity in low-bitrate streaming codecs.
    confidence: high
    relevance: high
  - claim_id: staged_training_that_freezes_acoustic_modules_before_introducing_semantic_and
    role: supports
    claim: Staged training that freezes acoustic modules before introducing semantic and KL losses is necessary
      for stable convergence in multi-objective codec training.
    source: §III.E, §V.C, Table II
    evidence: Staged training that freezes acoustic modules before introducing semantic and KL losses is necessary
      for stable convergence in multi-objective codec training.
    confidence: high
    relevance: high
  limitations:
  - The contrastive learning objective requires duration-aligned phoneme-level text labels during training. The
    authors acknowledge this as a key limitation and flag unsupervised disentanglement as the primary future direction.
  - Evaluation is conducted only on English (LibriTTS) and Mandarin (AISHELL-3) data; generalisation to other languages,
    especially low-resource or tonal languages with different phoneme structures, is unverified. The model uses
    HiFi-GAN as its vocoder, which introduces separate dependencies, vocoder artefacts, and the bulk of the decoding
    RTF — excluding vocoder time, the decoder RTF drops to 0.001. Codebook utilisation of 98.06% is reported for
    the VAE-FSQ variant but without comparison to the final SecoustiCodec configuration in the ablation table directly;
    the connection between codebook saturation and downstream LM training quality is asserted but not demonstrated
    empirically in TTS or dialogue tasks.
  caveats: []
- id: '2504.10352'
  published_date: "2025-08-05"
  entry_date: '2026-07-28'
  year: 2025
  venue: arXiv
  task:
  - TTS
  architecture:
  - autoregressive-LM
  - hybrid
  relevance: medium
  evidence_role:
  - architecture_variant
  - evaluation_caution
  current_role: influential
  method_family:
  - autoregressive_codec_language_models
  - hybrid_semantic_acoustic_tokenizers
  claims:
  - claim_id: pseudo_autoregressive_generation_which_commits_spans_left_to_right_within
    role: supports
    claim: Pseudo-autoregressive generation, which commits spans left-to-right within a bidirectional masked transformer,
      achieves constant inference steps regardless of target speech duration while maintaining temporal coherence.
    source: §3, §5.4
    evidence: Pseudo-autoregressive generation, which commits spans left-to-right within a bidirectional masked
      transformer, achieves constant inference steps regardless of target speech duration while maintaining temporal
      coherence.
    confidence: high
    relevance: medium
  - claim_id: a_model_trained_on_580_hours_of_english_speech_can
    role: supports
    claim: A model trained on 580 hours of English speech can match or exceed the intelligibility of NAR flow-matching
      systems trained on 100,000+ hours when temporal ordering is explicitly enforced during generation.
    source: §5.3, Table 1
    evidence: A model trained on 580 hours of English speech can match or exceed the intelligibility of NAR flow-matching
      systems trained on 100,000+ hours when temporal ordering is explicitly enforced during generation.
    confidence: high
    relevance: medium
  - claim_id: confidence_guided_iterative_nar_refinement_of_an_initial_par_generation
    role: supports
    claim: Confidence-guided iterative NAR refinement of an initial PAR generation substantially reduces word error
      rate with only a small number of additional inference steps.
    source: §5.5, Figure 4
    evidence: Confidence-guided iterative NAR refinement of an initial PAR generation substantially reduces word
      error rate with only a small number of additional inference steps.
    confidence: high
    relevance: medium
  - claim_id: temporally_unordered_nar_generation_produces_higher_alignment_errors_than_span
    role: supports
    claim: Temporally unordered NAR generation produces higher alignment errors than span-level causal generation
      across both continuation and cross-sentence evaluation tasks.
    source: §5.4, Table 3
    evidence: Temporally unordered NAR generation produces higher alignment errors than span-level causal generation
      across both continuation and cross-sentence evaluation tasks.
    confidence: high
    relevance: medium
  - claim_id: separate_model_capacity_for_each_generation_stage_is_necessary_unifying
    role: supports
    claim: Separate model capacity for each generation stage is necessary; unifying PAR and NAR refinement into
      a single multitask model degrades cross-sentence intelligibility by approximately 20%.
    source: §5.5
    evidence: Separate model capacity for each generation stage is necessary; unifying PAR and NAR refinement into
      a single multitask model degrades cross-sentence intelligibility by approximately 20%.
    confidence: high
    relevance: medium
  limitations:
  - PALLE is evaluated only on English (LibriTTS). The 100-step inference (with 7 refinement steps) may still be
    too slow for the most latency-sensitive streaming applications despite the 10x speedup. Duration estimation
    for the cross-sentence task relies on a simple linear heuristic; errors in duration estimation lead to modest
    quality degradation (WER-H 2.83 vs. 2.62 with GT duration). The shared architecture between stage one and stage
    two (joint multitask fine-tuning) causes stage two loss to degrade stage one performance, suggesting that separate
    model capacity is required.
  caveats: []
- id: '2508.04141'
  published_date: "2025-08-06"
  entry_date: '2026-07-28'
  year: 2025
  venue: arXiv
  task:
  - TTS
  architecture:
  - autoregressive-LM
  - hybrid
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: active_evidence
  method_family:
  - autoregressive_codec_language_models
  - hybrid_semantic_acoustic_tokenizers
  claims:
  - claim_id: generating_semantic_and_acoustic_tokens_simultaneously_in_a_single_autoregressive
    role: supports
    claim: Generating semantic and acoustic tokens simultaneously in a single autoregressive forward pass, rather
      than cascading semantic prediction before acoustic prediction, reduces word error rate and improves naturalness
      in zero-shot TTS.
    source: §V.A, Table I
    evidence: Generating semantic and acoustic tokens simultaneously in a single autoregressive forward pass, rather
      than cascading semantic prediction before acoustic prediction, reduces word error rate and improves naturalness
      in zero-shot TTS.
    confidence: high
    relevance: medium
  - claim_id: combining_specialist_ssl_models_for_distinct_speech_attributes_semantic_content
    role: supports
    claim: Combining specialist SSL models for distinct speech attributes (semantic content, acoustic texture, speaker
      identity) as frozen feature extractors enables more effective token-level disentanglement than using a single
      encoder for all attributes.
    source: §III.A, Tables III–IV
    evidence: Combining specialist SSL models for distinct speech attributes (semantic content, acoustic texture,
      speaker identity) as frozen feature extractors enables more effective token-level disentanglement than using
      a single encoder for all attributes.
    confidence: high
    relevance: medium
  - claim_id: a_hybrid_ar_nar_design_that_enforces_independence_at_the
    role: supports
    claim: A hybrid AR+NAR design that enforces independence at the coarse token level and interdependence at the
      fine-grained level distributes modeling complexity more effectively than an all-AR or all-NAR approach.
    source: §V.B, Tables III–IV
    evidence: A hybrid AR+NAR design that enforces independence at the coarse token level and interdependence at
      the fine-grained level distributes modeling complexity more effectively than an all-AR or all-NAR approach.
    confidence: high
    relevance: medium
  - claim_id: parallel_semantic_acoustic_modeling_improves_naturalness_and_intelligibility_without_fully
    role: supports
    claim: Parallel semantic-acoustic modeling improves naturalness and intelligibility without fully closing the
      speaker similarity gap relative to systems with dedicated speaker embedding refinement.
    source: §V.A, Table I
    evidence: Parallel semantic-acoustic modeling improves naturalness and intelligibility without fully closing
      the speaker similarity gap relative to systems with dedicated speaker embedding refinement.
    confidence: high
    relevance: low
  limitations:
  - Speaker similarity lags behind CosyVoice (SMOS gap ~0.15–0.2 on English), suggesting the parallel architecture
    does not yet fully leverage speaker conditioning. UTMOS scores, while competitive, do not reach ground-truth
    levels. Model size is not reported, making compute comparisons difficult. The internal Chinese dataset and preprocessing
    pipeline (Emilia + NCSSD) are not publicly released, limiting reproducibility on that front. Extending the framework
    to prosody control, emotion conditioning, or cross-lingual voice conversion is not explored. The subjective
    decoupling evaluation (Section V.C) relies on 90% evaluator agreement rather than a standardized metric, leaving
    quantitative disentanglement assessment as an open question.
  caveats: []
- id: '2508.04585'
  published_date: "2025-08-06"
  entry_date: '2026-07-28'
  year: 2025
  venue: ACM MM
  task:
  - TTS
  - SCA
  architecture:
  - autoregressive-LM
  - flow-matching
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: active_evidence
  method_family:
  - autoregressive_codec_language_models
  - flow_matching_codec_decoders
  claims:
  - claim_id: matching_the_token_rates_of_speech_and_facial_landmark_codecs
    role: supports
    claim: Matching the token rates of speech and facial landmark codecs enables frame-level synchronisation between
      synthesised speech and talking-face animations without post-hoc alignment.
    source: §4.1.3, §4.2
    evidence: Matching the token rates of speech and facial landmark codecs enables frame-level synchronisation
      between synthesised speech and talking-face animations without post-hoc alignment.
    confidence: high
    relevance: medium
  - claim_id: llm_based_joint_prediction_of_interleaved_speech_and_visual_tokens
    role: supports
    claim: LLM-based joint prediction of interleaved speech and visual tokens in dialogue context outperforms cascaded
      speech-then-video generation on both emotional accuracy and lip synchronisation.
    source: §6.2, §6.3, Table 2, Table 3
    evidence: LLM-based joint prediction of interleaved speech and visual tokens in dialogue context outperforms
      cascaded speech-then-video generation on both emotional accuracy and lip synchronisation.
    confidence: high
    relevance: low
  - claim_id: including_visual_dialogue_history_talking_face_animations_of_prior_turns
    role: supports
    claim: Including visual dialogue history (talking-face animations of prior turns) improves emotion accuracy
      and speaker consistency in conversational speech synthesis relative to audio-text-only context.
    source: §6.4, Table 4
    evidence: Including visual dialogue history (talking-face animations of prior turns) improves emotion accuracy
      and speaker consistency in conversational speech synthesis relative to audio-text-only context.
    confidence: high
    relevance: low
  - claim_id: low_rate_discrete_tokenisation_of_facial_landmarks_1_token_per
    role: supports
    claim: Low-rate discrete tokenisation of facial landmarks (1 token per frame) is more effective for LLM contextual
      modelling than higher-rate representations, even at a cost in geometric reconstruction fidelity.
    source: §6.1, §6.2, Table 1, Table 2
    evidence: Low-rate discrete tokenisation of facial landmarks (1 token per frame) is more effective for LLM contextual
      modelling than higher-rate representations, even at a cost in geometric reconstruction fidelity.
    confidence: high
    relevance: medium
  - claim_id: emotion_guided_conditioning_of_the_speech_renderer_including_predicted_emotion
    role: supports
    claim: Emotion-guided conditioning of the speech renderer, including predicted emotion labels as explicit conditioning,
      improves measured emotional expressiveness over systems that rely on implicit contextual inference alone.
    source: §6.4, Table 4
    evidence: Emotion-guided conditioning of the speech renderer, including predicted emotion labels as explicit
      conditioning, improves measured emotional expressiveness over systems that rely on implicit contextual inference
      alone.
    confidence: high
    relevance: low
  limitations:
  - The Talking-face Animations Renderer (EchoMimic) is a pre-trained third-party module that receives no additional
    fine-tuning in this pipeline. Its outputs are constrained by the quality ceiling and biases of the EchoMimic
    base model, limiting the paper's ability to attribute animation quality gains to UniTalker specifically vs.
    the renderer.
  - 'Rendering latency is notable: speech generation takes approximately 2 seconds and animation rendering takes
    approximately 5 seconds per 25 frames on an RTX 4080 with 32 GB RAM, making the system unsuitable for real-time
    interaction in its current form. The paper acknowledges this and lists streaming optimisation as future work.'
  - The training data for dialogue context is limited to 113 hours of audio-only and 307 hours of visual dialogue,
    which is modest compared to the 170,000-hour pretraining corpus used for the speech renderer. Generalisation
    to diverse speaking styles and languages beyond the training distribution is untested. All subjective evaluations
    used 30 non-native English speakers recruited locally, raising questions about evaluation representativeness
    for naturalness judgements.
  - The CSVS task definition assumes that the target utterance's text is known at inference time. This is a constrained
    setting (closer to expressive TTS than to open-ended dialogue response generation) and does not address the
    harder problem of jointly generating text, speech, and visual responses.
  caveats: []
- id: '2508.05207'
  published_date: "2025-08-07"
  entry_date: '2026-07-28'
  year: 2025
  venue: arXiv
  task:
  - codec
  architecture:
  - GAN
  relevance: medium
  evidence_role:
  - architecture_variant
  - acceleration_evidence
  current_role: active_evidence
  method_family:
  - gan_neural_codecs_and_decoders
  claims:
  - claim_id: operating_in_the_time_frequency_domain_enables_neural_codecs_to
    role: supports
    claim: Operating in the time-frequency domain enables neural codecs to achieve higher perceptual quality for
      full-band audio than equivalent waveform-domain architectures, especially at low bit rates.
    source: §2, §4, Table 1
    evidence: Operating in the time-frequency domain enables neural codecs to achieve higher perceptual quality
      for full-band audio than equivalent waveform-domain architectures, especially at low bit rates.
    confidence: high
    relevance: medium
  - claim_id: cross_channel_phase_coherence_in_multi_channel_neural_codecs_requires
    role: supports
    claim: Cross-channel phase coherence in multi-channel neural codecs requires joint processing of audio channels
      in at least some encoder layers, and neither fully independent nor fully joint encoding is optimal.
    source: §2
    evidence: Cross-channel phase coherence in multi-channel neural codecs requires joint processing of audio channels
      in at least some encoder layers, and neither fully independent nor fully joint encoding is optimal.
    confidence: high
    relevance: medium
  - claim_id: multi_scale_spectral_discriminators_are_effective_for_training_high_quality
    role: supports
    claim: Multi-scale spectral discriminators are effective for training high-quality neural codecs without requiring
      waveform-domain discriminators.
    source: §3
    evidence: Multi-scale spectral discriminators are effective for training high-quality neural codecs without
      requiring waveform-domain discriminators.
    confidence: high
    relevance: medium
  - claim_id: biased_quantizer_dropout_towards_low_codebook_counts_during_training_improves
    role: supports
    claim: Biased quantizer dropout towards low codebook counts during training improves codec quality at the low
      bit rate end without sacrificing high bit rate performance.
    source: §3.1.1
    evidence: Biased quantizer dropout towards low codebook counts during training improves codec quality at the
      low bit rate end without sacrificing high bit rate performance.
    confidence: high
    relevance: high
  - claim_id: real_time_streaming_neural_codec_inference_at_48_khz_stereo
    role: supports
    claim: Real-time streaming neural codec inference at 48 kHz stereo is achievable on a desktop CPU with an 80
      ms architectural latency when using causal convolutions and a minimal look-ahead.
    source: §1, §2
    evidence: Real-time streaming neural codec inference at 48 kHz stereo is achievable on a desktop CPU with an
      80 ms architectural latency when using causal convolutions and a minimal look-ahead.
    confidence: high
    relevance: high
  limitations:
  - Training data is proprietary and the only baseline is DAC; results cannot be reproduced and the comparison does
    not include SoundStream, EnCodec, or Mimi, leaving SpectroStream's position in the broader codec landscape unclear.
  - Evaluation is restricted to music (MUSDB18) despite the paper's "general audio" framing. Speech quality at 48
    kHz stereo is not reported. The A/B preference protocol does not include a MUSHRA-style anchor, making absolute
    quality judgements difficult. The latency of 80 ms is described as suitable for streaming but is not benchmarked
    against real-time constraints in actual deployment. The delayed-fusion fusion point is treated as a design choice
    found empirically — no ablation is provided to quantify the quality/coherence trade-off as a function of fusion
    layer depth.
  caveats: []
- id: '2508.06262'
  published_date: "2025-08-08"
  entry_date: '2026-07-28'
  year: 2025
  venue: arXiv
  task:
  - TTS
  architecture:
  - autoregressive-LM
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: active_evidence
  method_family:
  - autoregressive_codec_language_models
  claims:
  - claim_id: token_verification_is_necessary_for_multi_token_prediction_to_be
    role: supports
    claim: 'Token verification is necessary for multi-token prediction to be effective in autoregressive TTS: without
      it, WER increases from 3.07% to 14.37% and speaker similarity drops from 0.570 to 0.463.'
    source: §V, Table III
    evidence: 'Token verification is necessary for multi-token prediction to be effective in autoregressive TTS:
      without it, WER increases from 3.07% to 14.37% and speaker similarity drops from 0.570 to 0.463.'
    confidence: high
    relevance: low
  - claim_id: plug_and_play_mtp_modules_trained_on_a_modest_dataset
    role: supports
    claim: Plug-and-play MTP modules trained on a modest dataset can accelerate a frozen autoregressive TTS backbone
      by up to 1.48x without sacrificing generation quality on standard benchmarks.
    source: §IV.B, Table I
    evidence: Plug-and-play MTP modules trained on a modest dataset can accelerate a frozen autoregressive TTS backbone
      by up to 1.48x without sacrificing generation quality on standard benchmarks.
    confidence: high
    relevance: medium
  - claim_id: under_quality_maximizing_inference_settings_mtp_with_verification_can_improve
    role: supports
    claim: Under quality-maximizing inference settings, MTP with verification can improve intelligibility beyond
      the backbone baseline, likely due to extended look-ahead context from the cascaded hidden states.
    source: §IV.A, Table I
    evidence: Under quality-maximizing inference settings, MTP with verification can improve intelligibility beyond
      the backbone baseline, likely due to extended look-ahead context from the cascaded hidden states.
    confidence: high
    relevance: medium
  - claim_id: converting_a_non_causal_codec_decoder_to_a_causal_streaming
    role: supports
    claim: Converting a non-causal codec decoder to a causal streaming architecture via lightweight fine-tuning
      preserves approximately 95% of reconstruction quality, making streaming reconstruction viable without full
      retraining.
    source: §IV.C, Table II
    evidence: Converting a non-causal codec decoder to a causal streaming architecture via lightweight fine-tuning
      preserves approximately 95% of reconstruction quality, making streaming reconstruction viable without full
      retraining.
    confidence: high
    relevance: high
  - claim_id: attention_based_mtp_modules_substantially_outperform_mlp_based_equivalents_of
    role: supports
    claim: Attention-based MTP modules substantially outperform MLP-based equivalents of similar parameter count
      in both intelligibility and speaker similarity for TTS acceleration.
    source: §IV.B, Table I
    evidence: Attention-based MTP modules substantially outperform MLP-based equivalents of similar parameter count
      in both intelligibility and speaker similarity for TTS acceleration.
    confidence: high
    relevance: low
  limitations:
  - Evaluation is English-only (LibriTTS training, Seed-TTS-eval-en test), and speaker generalization to out-of-distribution
    languages or accents is untested. The verification overhead (one additional LM forward pass for verification
    at each step) partially offsets the MTP speedup, especially at strict topk values. The 1.48x figure assumes
    topk=500, which allows some quality degradation; the fully lossless speedup (topk=100) is closer to 1.42x. Scaling
    MTP to larger models (Llasa-3B, Llasa-8B) is not investigated.
  caveats: []
- id: '2508.07302'
  published_date: "2025-08-10"
  entry_date: '2026-07-28'
  year: 2025
  venue: arXiv
  task:
  - TTS
  architecture:
  - autoregressive-LM
  - flow-matching
  relevance: medium
  evidence_role:
  - architecture_variant
  - acceleration_evidence
  - control_evidence
  current_role: active_evidence
  method_family:
  - autoregressive_codec_language_models
  - flow_matching_codec_decoders
  claims:
  - claim_id: language_agnostic_emotional_embeddings_from_pre_trained_models_can_serve
    role: supports
    claim: Language-agnostic emotional embeddings from pre-trained models can serve as a reliable cross-lingual
      retrieval signal for zero-shot emotion transfer in TTS.
    source: §III.C, §IV.C
    evidence: Language-agnostic emotional embeddings from pre-trained models can serve as a reliable cross-lingual
      retrieval signal for zero-shot emotion transfer in TTS.
    confidence: high
    relevance: low
  - claim_id: retrieval_augmented_prompting_reduces_foreign_accent_artefacts_in_cross_lingual
    role: supports
    claim: Retrieval-augmented prompting reduces foreign-accent artefacts in cross-lingual emotional speech synthesis
      more effectively than direct prosody transfer between typologically distant languages.
    source: §III.C, §IV.E
    evidence: Retrieval-augmented prompting reduces foreign-accent artefacts in cross-lingual emotional speech synthesis
      more effectively than direct prosody transfer between typologically distant languages.
    confidence: high
    relevance: low
  - claim_id: flow_matching_alignment_between_discrete_codec_tokens_and_mel_spectrograms
    role: supports
    claim: Flow-matching alignment between discrete codec tokens and mel-spectrograms improves speaker identity
      preservation as well as prosodic naturalness in multilingual synthesis.
    source: §III.B, §IV.E
    evidence: Flow-matching alignment between discrete codec tokens and mel-spectrograms improves speaker identity
      preservation as well as prosodic naturalness in multilingual synthesis.
    confidence: high
    relevance: high
  - claim_id: clustering_based_retrieval_strategies_over_large_emotional_speech_pools_maintain
    role: supports
    claim: Clustering-based retrieval strategies over large emotional speech pools maintain higher accuracy and
      lower latency than exhaustive cosine similarity search as pool size grows.
    source: §IV.D, Table II
    evidence: Clustering-based retrieval strategies over large emotional speech pools maintain higher accuracy and
      lower latency than exhaustive cosine similarity search as pool size grows.
    confidence: high
    relevance: low
  - claim_id: two_stage_fine_tuning_first_on_phonetics_then_on_expressiveness
    role: supports
    claim: Two-stage fine-tuning — first on phonetics, then on expressiveness — enables effective emotion adaptation
      in low-resource target languages from a strong multilingual foundation model.
    source: §III.D, §IV.A
    evidence: Two-stage fine-tuning — first on phonetics, then on expressiveness — enables effective emotion adaptation
      in low-resource target languages from a strong multilingual foundation model.
    confidence: high
    relevance: low
  limitations:
  - All evaluation is conducted on internal, non-public datasets with a narrow test configuration (one Chinese speaker,
    proprietary Thai data). There is no standard benchmark, no released code, and no cross-lab reproducibility path.
    Claims about emotion transfer quality are difficult to verify independently.
  - The evaluation covers only the Chinese-to-Thai direction. The paper claims the framework is language-agnostic,
    but this is stated as future work rather than demonstrated. The listener panel is small (15 raters) and the
    Thai subset is particularly small (5 raters), raising questions about statistical reliability. The EMOS metric
    used here is a custom MOS variant not directly comparable to published results elsewhere. Comparisons with Typhoon2-Audio
    are limited to CER only, leaving the emotional quality comparison against a strong 8B-parameter baseline unanswered.
  caveats: []
- id: '2508.08399'
  published_date: "2025-08-11"
  entry_date: '2026-07-28'
  year: 2025
  venue: arXiv
  task:
  - codec
  architecture:
  - VAE
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: active_evidence
  method_family:
  - vae_vector_quantized_codecs
  claims:
  - claim_id: fully_discrete_disentanglement_of_phonetic_prosodic_and_speaker_information_in
    role: complicates
    claim: Fully discrete disentanglement of phonetic, prosodic, and speaker information in a speech codec is achievable
      without phoneme labels or F0 supervision, at the cost of a small reconstruction quality degradation.
    source: §III, §IV.B, Table II
    evidence: Fully discrete disentanglement of phonetic, prosodic, and speaker information in a speech codec is
      achievable without phoneme labels or F0 supervision, at the cost of a small reconstruction quality degradation.
    confidence: high
    relevance: high
  - claim_id: quantizing_speaker_vectors_into_discrete_codes_reduces_speaker_identity_fidelity
    role: complicates
    claim: Quantizing speaker vectors into discrete codes reduces speaker identity fidelity compared to continuous
      speaker representations, posing a fundamental trade-off between LLM compatibility and speaker retention.
    source: §IV.B, Table III
    evidence: Quantizing speaker vectors into discrete codes reduces speaker identity fidelity compared to continuous
      speaker representations, posing a fundamental trade-off between LLM compatibility and speaker retention.
    confidence: high
    relevance: medium
  - claim_id: instance_normalization_of_ssl_residual_features_provides_a_label_free
    role: supports
    claim: Instance normalization of SSL residual features provides a label-free mechanism to separate time-invariant
      speaker statistics from time-variant prosodic content.
    source: §III.B
    evidence: Instance normalization of SSL residual features provides a label-free mechanism to separate time-invariant
      speaker statistics from time-variant prosodic content.
    confidence: high
    relevance: medium
  - claim_id: fully_discrete_speech_codecs_can_match_conventional_voice_conversion_methods
    role: supports
    claim: Fully discrete speech codecs can match conventional voice conversion methods on intelligibility and naturalness
      while enabling attribute manipulation through codebook-level operations.
    source: §IV.B, Table III
    evidence: Fully discrete speech codecs can match conventional voice conversion methods on intelligibility and
      naturalness while enabling attribute manipulation through codebook-level operations.
    confidence: high
    relevance: high
  limitations:
  - All experiments use LibriSpeech clean speech (16 kHz, studio conditions); performance on noisy, spontaneous,
    or out-of-domain speech is untested. The one-shot VC evaluation uses only two reference speakers (one male,
    one female), limiting statistical confidence in the speaker similarity results.
  - The model is not tested on any downstream application (TTS, ASR, speech LM), despite this being the stated motivation.
    Whether the disentangled discrete tokens actually improve over non-disentangled tokens on downstream tasks remains
    an open question — the paper acknowledges this as future work. The GRVQ codebook dimensionality analysis shows
    a clear trade-off between bitrate and speaker identity, but optimal bitrate allocation across the three streams
    is not systematically explored. Prosody quantization codebook interpretability beyond F0 correlation (e.g.,
    energy, duration, speaking rate) is not investigated.
  caveats: []
- id: '2508.08961'
  published_date: "2025-08-12"
  entry_date: '2026-07-28'
  year: 2025
  venue: arXiv
  task:
  - TTS
  - SCA
  - VC
  architecture:
  - autoregressive-LM
  relevance: medium
  evidence_role:
  - architecture_variant
  current_role: active_evidence
  method_family:
  - autoregressive_codec_language_models
  claims:
  - claim_id: separating_the_token_used_for_llm_input_from_the_token
    role: supports
    claim: Separating the token used for LLM input from the token used for generation output can resolve the information-level
      conflict that makes joint optimisation of understanding and generation tasks difficult in a shared-token speech
      LLM.
    source: §DualSpeechLM, §Results and Analyses
    evidence: Separating the token used for LLM input from the token used for generation output can resolve the
      information-level conflict that makes joint optimisation of understanding and generation tasks difficult in
      a shared-token speech LLM.
    confidence: high
    relevance: medium
  - claim_id: training_a_speech_tokenizer_directly_against_a_text_llm_s
    role: supports
    claim: Training a speech tokenizer directly against a text LLM's next-token prediction objective is more effective
      at reducing the speech-text modality gap than optimising against ASR or SSL reconstruction losses alone.
    source: §USTokenizer, Table 1
    evidence: Training a speech tokenizer directly against a text LLM's next-token prediction objective is more
      effective at reducing the speech-text modality gap than optimising against ASR or SSL reconstruction losses
      alone.
    confidence: high
    relevance: medium
  - claim_id: understanding_task_supervision_produces_representations_that_transfer_to_generation_quality
    role: supports
    claim: Understanding-task supervision produces representations that transfer to generation quality improvements,
      but the converse — generation-task supervision improving understanding — is weaker and less consistent.
    source: §Ablation Study, §H. Discussion
    evidence: Understanding-task supervision produces representations that transfer to generation quality improvements,
      but the converse — generation-task supervision improving understanding — is weaker and less consistent.
    confidence: high
    relevance: medium
  - claim_id: stochastic_conditioning_during_training_exposing_a_generation_module_to_varied
    role: supports
    claim: Stochastic conditioning during training (exposing a generation module to varied subsets of its conditioning
      signals) improves robustness to imperfect upstream predictions at inference time.
    source: §Ablation Study, Table 6
    evidence: Stochastic conditioning during training (exposing a generation module to varied subsets of its conditioning
      signals) improves robustness to imperfect upstream predictions at inference time.
    confidence: high
    relevance: medium
  limitations:
  - The entire evaluation is conducted at 4.5K hours of training data with parameter-efficient LoRA fine-tuning.
    The claim that USTokens reduce data requirements is plausible but untested at the scale (70K–570K hours) where
    competing systems are evaluated. Whether the dual-token architecture and the understanding-driven tokeniser
    remain advantageous at scale is an open question.
  - No code or demo is linked in the paper, limiting reproducibility. The 4.5K-hour training regime excludes noisy,
    in-the-wild, and multilingual data, so generalisation to these conditions is untested despite the paper's stated
    future direction of expanding to multilingual and cross-domain data. The USTokenizer's understanding-driven
    loss requires a frozen LLM during tokeniser training, adding a significant computational overhead at the tokenisation
    stage (288% memory increase) even if this overhead disappears at DualSpeechLM inference. The model size of the
    full system (Phi-3.5-3B + AcousticGPT) is not explicitly stated in aggregate, and the AcousticGPT's token generation
    speed relative to real-time is not reported.
  caveats: []
- id: '2504.12867'
  published_date: "2025-08-13"
  entry_date: '2026-07-28'
  year: 2025
  venue: arXiv
  task:
  - TTS
  architecture:
  - autoregressive-LM
  - flow-matching
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: active_evidence
  method_family:
  - autoregressive_codec_language_models
  - flow_matching_codec_decoders
  claims:
  - claim_id: fine_grained_natural_language_emotion_descriptions_provide_richer_control_over
    role: complicates
    claim: Fine-grained natural language emotion descriptions provide richer control over expressive speech synthesis
      than coarse categorical labels, at the cost of requiring emotion-specific training data.
    source: §2.1, §6.1
    evidence: Fine-grained natural language emotion descriptions provide richer control over expressive speech synthesis
      than coarse categorical labels, at the cost of requiring emotion-specific training data.
    confidence: high
    relevance: low
  - claim_id: parallel_phoneme_token_prediction_as_a_secondary_output_head_reduces
    role: supports
    claim: Parallel phoneme token prediction as a secondary output head reduces intelligibility errors in LLM-based
      TTS, particularly on challenging inputs such as rare words and tongue twisters.
    source: §6.2.1, §6.2.2, Table 5, Table 6
    evidence: Parallel phoneme token prediction as a secondary output head reduces intelligibility errors in LLM-based
      TTS, particularly on challenging inputs such as rare words and tongue twisters.
    confidence: high
    relevance: medium
  - claim_id: llm_pretraining_initialisation_meaningfully_benefits_emotion_controllable_tts_models_without
    role: supports
    claim: 'LLM pretraining initialisation meaningfully benefits emotion-controllable TTS: models without it show
      substantially higher word error rates and weaker emotion transfer.'
    source: §6.2.4, Table 8
    evidence: 'LLM pretraining initialisation meaningfully benefits emotion-controllable TTS: models without it
      show substantially higher word error rates and weaker emotion transfer.'
    confidence: high
    relevance: low
  - claim_id: automatic_emotion_similarity_metrics_e_g_emotion2vec_cosine_similarity_correlate
    role: supports
    claim: Automatic emotion similarity metrics (e.g. emotion2vec cosine similarity) correlate reasonably at the
      system level but poorly at the utterance level with human perceptual judgments, limiting their utility for
      fine-grained model comparison.
    source: §7, Table 10
    evidence: Automatic emotion similarity metrics (e.g. emotion2vec cosine similarity) correlate reasonably at
      the system level but poorly at the utterance level with human perceptual judgments, limiting their utility
      for fine-grained model comparison.
    confidence: high
    relevance: low
  - claim_id: multimodal_llms_are_not_yet_reliable_judges_of_emotional_speech
    role: supports
    claim: Multimodal LLMs are not yet reliable judges of emotional speech quality, exhibiting both low correlation
      with human ratings and inter-run instability.
    source: §7, Table 10
    evidence: Multimodal LLMs are not yet reliable judges of emotional speech quality, exhibiting both low correlation
      with human ratings and inter-run instability.
    confidence: high
    relevance: medium
  limitations:
  - 'The English model is trained and evaluated entirely on synthetic data generated by GPT-4o-audio. Both EmoVoice-DB
    (training) and the test set are GPT-4o-audio outputs, creating circularity: the model learns to mimic GPT-4o-audio''s
    synthesis style rather than natural human emotional speech. Generalisation to real human emotional recordings
    or to out-of-distribution TTS systems is undemonstrated.'
  - The emotion recall evaluation omits three of seven emotion categories (disgusted, fearful, surprised) due to
    low recognition accuracy from emotion2vec, which limits the scope of the emotional expressiveness claims.
  - 'EmoVoice does not support zero-shot speaker generalisation: speaker identity is controlled via a reference
    speech prompt passed to the CosyVoice flow-matching module, but the system is not evaluated in a cross-speaker
    zero-shot scenario.'
  - The phoneme-boost variant (EmoVoice-PP) requires phoneme sequences at training time, adding a preprocessing
    dependency (Phonemizer) and making extension to languages with complex or poorly supported phonemisers non-trivial.
  - No efficiency or latency analysis is provided; the streaming characteristics of the grouped decoding approach
    relative to real-time deployment requirements are not reported.
  caveats: []
- id: '2508.11224'
  published_date: "2025-08-15"
  entry_date: '2026-07-28'
  year: 2025
  venue: ASRU
  task:
  - evaluation
  - TTS
  architecture: []
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: active_evidence
  method_family: []
  claims:
  - claim_id: ssl_models_using_frame_wise_masked_prediction_capture_relative_prosodic
    role: supports
    claim: SSL models using frame-wise masked prediction capture relative prosodic contours within an utterance
      rather than absolute acoustic magnitudes, making them insensitive to global intensity rescaling.
    source: §V-A, Fig. 1
    evidence: SSL models using frame-wise masked prediction capture relative prosodic contours within an utterance
      rather than absolute acoustic magnitudes, making them insensitive to global intensity rescaling.
    confidence: high
    relevance: medium
  - claim_id: models_pretrained_to_predict_discrete_targets_encode_phoneme_like_structure
    role: supports
    claim: Models pretrained to predict discrete targets encode phoneme-like structure effectively at small cluster
      sizes, while models pretrained on continuous targets require larger cluster sizes to approach the same phonemic
      alignment.
    source: §V-A, Fig. 3
    evidence: Models pretrained to predict discrete targets encode phoneme-like structure effectively at small cluster
      sizes, while models pretrained on continuous targets require larger cluster sizes to approach the same phonemic
      alignment.
    confidence: high
    relevance: medium
  - claim_id: training_k_means_clustering_on_emotionally_expressive_speech_increases_the
    role: supports
    claim: Training k-means clustering on emotionally expressive speech increases the prosodic sensitivity of resulting
      tokens for most SSL model and layer combinations.
    source: §V-B, Table I
    evidence: Training k-means clustering on emotionally expressive speech increases the prosodic sensitivity of
      resulting tokens for most SSL model and layer combinations.
    confidence: high
    relevance: medium
  - claim_id: applying_a_temporal_moving_average_to_ssl_features_before_k
    role: complicates
    claim: Applying a temporal moving average to SSL features before k-means clustering provides an adjustable trade-off
      between prosodic sensitivity and speaker invariance, with intermediate window sizes improving both simultaneously.
    source: §V-C, Fig. 5
    evidence: Applying a temporal moving average to SSL features before k-means clustering provides an adjustable
      trade-off between prosodic sensitivity and speaker invariance, with intermediate window sizes improving both
      simultaneously.
    confidence: high
    relevance: medium
  - claim_id: differences_between_ssl_pretraining_objectives_in_their_token_level_linguistic
    role: supports
    claim: Differences between SSL pretraining objectives in their token-level linguistic and prosodic encoding
      are concentrated in the final transformer layers, while intermediate layers exhibit largely similar behaviour
      across model families.
    source: §V-A, §V-B
    evidence: Differences between SSL pretraining objectives in their token-level linguistic and prosodic encoding
      are concentrated in the final transformer layers, while intermediate layers exhibit largely similar behaviour
      across model families.
    confidence: high
    relevance: medium
  limitations:
  - The analysis measures sensitivity via TER — a proxy for how much token sequences change in response to acoustic
    manipulation — rather than directly probing what information is decodable from the tokens. Whether the observed
    sensitivity differences translate to actual gains in downstream prosody-related tasks (prosody-conditional TTS,
    emphasis transfer, emotion recognition from discrete tokens) remains untested. The evaluation uses a single
    corpus (TIMIT) of read speech by native English speakers, which may limit generalisability to spontaneous, conversational,
    or multilingual speech. Duration was excluded from prosody analysis because deduplication collapses durational
    information; this omission means the benchmark does not cover the full prosody space. The study does not test
    acoustic tokens (neural codec outputs), focusing exclusively on semantic tokens.
  caveats: []
- id: '2508.11326'
  published_date: "2025-08-15"
  entry_date: '2026-07-28'
  year: 2025
  venue: arXiv
  task:
  - TTS
  architecture:
  - autoregressive-LM
  - diffusion
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: active_evidence
  method_family:
  - autoregressive_codec_language_models
  - diffusion_latent_codec_generators
  claims:
  - claim_id: freezing_the_backbone_llm_and_routing_modality_specific_tokens_to
    role: supports
    claim: Freezing the backbone LLM and routing modality-specific tokens to separate expert sets can preserve pre-trained
      text understanding capabilities during speech generation fine-tuning.
    source: §3.2
    evidence: Freezing the backbone LLM and routing modality-specific tokens to separate expert sets can preserve
      pre-trained text understanding capabilities during speech generation fine-tuning.
    confidence: high
    relevance: medium
  - claim_id: instruction_conditioned_tts_systems_trained_on_tag_derived_description_datasets
    role: supports
    claim: Instruction-conditioned TTS systems trained on tag-derived description datasets exhibit significant performance
      degradation when faced with figurative or metaphorical natural language at inference time.
    source: §1, §4.1
    evidence: Instruction-conditioned TTS systems trained on tag-derived description datasets exhibit significant
      performance degradation when faced with figurative or metaphorical natural language at inference time.
    confidence: high
    relevance: medium
  - claim_id: commercial_speech_synthesis_products_are_not_immune_to_the_out
    role: supports
    claim: Commercial speech synthesis products are not immune to the out-of-domain description challenge, suggesting
      that this generalisation gap is not solved by scale alone.
    source: §4.2, Table 2
    evidence: Commercial speech synthesis products are not immune to the out-of-domain description challenge, suggesting
      that this generalisation gap is not solved by scale alone.
    confidence: high
    relevance: medium
  - claim_id: modality_separation_techniques_from_multimodal_vision_language_research_transfer_meaningfully
    role: supports
    claim: Modality separation techniques from multimodal vision-language research transfer meaningfully to the
      speech domain, reducing catastrophic forgetting without requiring multi-modal data mixing during pre-training.
    source: §2.2, §3.2
    evidence: Modality separation techniques from multimodal vision-language research transfer meaningfully to the
      speech domain, reducing catastrophic forgetting without requiring multi-modal data mixing during pre-training.
    confidence: high
    relevance: medium
  limitations:
  - The evaluation is based on 20 in-domain and 40 out-of-domain test samples, annotated by 21 evaluators. These
    are very small test sets for drawing strong comparative conclusions. Both the in-domain and out-of-domain test
    sets were constructed by the MoE-TTS authors, introducing potential design bias toward cases where the proposed
    approach excels.
  - The system currently supports only English text descriptions, due to the limited scope of available open-source
    description-based TTS datasets. The LLM architecture sensitivity is unexplored — all experiments use Qwen3-4B,
    and the impact of model scale (smaller or larger LLM backbones) on the MoE approach is left for future work.
    The diffusion and VAEGAN components are adapted from Stable Audio without fine-tuning on description-based data,
    and their contribution to description alignment is not ablated.
  caveats: []
- id: interspeech-2025-0115
  published_date: "2025-08-17"
  entry_date: '2026-07-28'
  year: 2025
  venue: Interspeech
  task:
  - codec
  architecture:
  - transformer-enc-dec
  relevance: medium
  evidence_role:
  - architecture_variant
  current_role: active_evidence
  method_family:
  - transformer_encoder_decoder_tokenizers
  claims:
  - claim_id: even_codecs_explicitly_trained_with_disentanglement_objectives_fail_to_cleanly
    role: complicates
    claim: Even codecs explicitly trained with disentanglement objectives fail to cleanly separate pitch from other
      speech attributes in their token embeddings.
    source: §2.3, §2.4
    evidence: Even codecs explicitly trained with disentanglement objectives fail to cleanly separate pitch from
      other speech attributes in their token embeddings.
    confidence: high
    relevance: medium
  - claim_id: linguistic_content_in_neural_audio_codec_representations_concentrates_in_the
    role: supports
    claim: Linguistic content in neural audio codec representations concentrates in the lowest RVQ scales regardless
      of whether distillation was used, but leaks into higher scales when the frame rate is very low.
    source: §2.1
    evidence: Linguistic content in neural audio codec representations concentrates in the lowest RVQ scales regardless
      of whether distillation was used, but leaks into higher scales when the frame rate is very low.
    confidence: high
    relevance: high
  - claim_id: a_masked_autoencoder_framework_can_bridge_codec_tokens_and_perceptual
    role: supports
    claim: A masked-autoencoder framework can bridge codec tokens and perceptual speech attributes bidirectionally,
      enabling voice conversion at dramatically lower bitrates than spectrogram-based equivalents.
    source: §3.1, §3.2
    evidence: A masked-autoencoder framework can bridge codec tokens and perceptual speech attributes bidirectionally,
      enabling voice conversion at dramatically lower bitrates than spectrogram-based equivalents.
    confidence: high
    relevance: high
  - claim_id: post_hoc_interpretability_tools_reveal_systematic_trade_offs_between_content
    role: supports
    claim: Post-hoc interpretability tools reveal systematic trade-offs between content accuracy and synthesis quality
      that differ by codec design, complicating the choice of codec for controllable speech generation.
    source: §3.2, Table 1
    evidence: Post-hoc interpretability tools reveal systematic trade-offs between content accuracy and synthesis
      quality that differ by codec design, complicating the choice of codec for controllable speech generation.
    confidence: high
    relevance: high
  limitations:
  - All experiments use LibriSpeech, a clean read-speech corpus with limited acoustic diversity. Whether the observed
    encoding patterns hold for spontaneous speech, expressive data, or noise-conditioned codecs is untested.
  - The study covers four specific codecs; the broader generalisation across the growing landscape of codec designs
    (including future multi-scale or end-to-end codec-LM systems) is an open question. Pitch estimation from codec
    tokens remains poor, and no remedy is proposed — it is unclear whether this is a fundamental limitation of RVQ-based
    representations or an artefact of the particular codecs studied. The bidirectional AnCoGen-Codec is compared
    only against the Melspectrogram baseline and not against dedicated disentanglement-oriented codec frameworks
    such as FreeCodec or SpeechFlow, which would provide stronger context for the synthesis results.
  caveats: []
- id: interspeech-2025-0196
  published_date: "2025-08-17"
  entry_date: '2026-07-28'
  year: 2025
  venue: Interspeech
  task:
  - codec
  architecture:
  - GAN
  - VAE
  relevance: medium
  evidence_role:
  - architecture_variant
  current_role: active_evidence
  method_family:
  - gan_neural_codecs_and_decoders
  - vae_vector_quantized_codecs
  claims:
  - claim_id: imposing_spectral_structure_on_codec_latent_embeddings_through_supervised_disentanglement
    role: supports
    claim: Imposing spectral structure on codec latent embeddings through supervised disentanglement improves compression
      efficiency over unstructured group quantization.
    source: §2.2, §3.3, Table 3
    evidence: Imposing spectral structure on codec latent embeddings through supervised disentanglement improves
      compression efficiency over unstructured group quantization.
    confidence: high
    relevance: high
  - claim_id: inter_group_prediction_in_neural_codecs_where_high_frequency_representations
    role: supports
    claim: Inter-group prediction in neural codecs, where high-frequency representations are predicted from quantized
      low-frequency components, reduces bitstream redundancy without proportional increases in model size.
    source: §2.2, §3.1
    evidence: Inter-group prediction in neural codecs, where high-frequency representations are predicted from quantized
      low-frequency components, reduces bitstream redundancy without proportional increases in model size.
    confidence: high
    relevance: medium
  - claim_id: the_performance_advantage_of_latent_space_spectral_decomposition_in_neural
    role: supports
    claim: The performance advantage of latent-space spectral decomposition in neural codecs scales with bitrate,
      with larger gains at higher bitrates where low-frequency quality is sufficient to anchor the prediction.
    source: §3.3, Table 3
    evidence: The performance advantage of latent-space spectral decomposition in neural codecs scales with bitrate,
      with larger gains at higher bitrates where low-frequency quality is sufficient to anchor the prediction.
    confidence: high
    relevance: high
  - claim_id: causal_time_domain_neural_speech_codecs_can_match_or_exceed
    role: supports
    claim: Causal, time-domain neural speech codecs can match or exceed traditional codecs (EVS, Opus) on perceptual
      quality metrics at equivalent bitrates.
    source: §3.2, Table 2
    evidence: Causal, time-domain neural speech codecs can match or exceed traditional codecs (EVS, Opus) on perceptual
      quality metrics at equivalent bitrates.
    confidence: high
    relevance: medium
  limitations:
  - The subjective evaluation uses only 7 listeners; this falls below the minimum typically required to draw statistically
    robust MOS conclusions. The test set of 30 samples, while drawn from standardised ITU-T and EVS material, is
    small for a codec evaluation claiming state-of-the-art status.
  - The evaluation covers only speech; generalisability of the SP approach to music or general audio is not addressed.
    The cross-codec comparisons in Table 2 mix sampling rates (16 kHz vs. 24 vs. 32 kHz), so direct MOS comparison
    between SPCODEC and Encodec requires care. The ablation holds the encoder-decoder fixed to DAC's architecture,
    which means the gains may differ when applied to other backbone designs. No streaming or latency analysis is
    provided beyond the real-time factor on one CPU configuration.
  caveats: []
- id: interspeech-2025-0246
  published_date: "2025-08-17"
  entry_date: '2026-07-28'
  year: 2025
  venue: Interspeech
  task:
  - codec
  architecture:
  - transformer-enc-dec
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: active_evidence
  method_family:
  - transformer_encoder_decoder_tokenizers
  claims:
  - claim_id: speaker_invariant_discrete_speech_tokens_trained_with_a_dual_codebook
    role: supports
    claim: Speaker-invariant discrete speech tokens, trained with a dual-codebook objective, can improve both spoken
      language model performance and speech resynthesis intelligibility simultaneously.
    source: §3.2, §3.3, Tables 1–3
    evidence: Speaker-invariant discrete speech tokens, trained with a dual-codebook objective, can improve both
      spoken language model performance and speech resynthesis intelligibility simultaneously.
    confidence: high
    relevance: high
  - claim_id: n_gram_predictability_and_phoneme_character_mutual_information_are_stronger
    role: supports
    claim: N-gram predictability and phoneme/character mutual information are stronger proxies for SLM downstream
      performance than the widely-used ABX error rate.
    source: §3.5, Figure 3b
    evidence: N-gram predictability and phoneme/character mutual information are stronger proxies for SLM downstream
      performance than the widely-used ABX error rate.
    confidence: high
    relevance: medium
  - claim_id: ssl_based_tokenizers_can_match_or_exceed_neural_codec_tokenizers
    role: supports
    claim: SSL-based tokenizers can match or exceed neural codec tokenizers on speech intelligibility (WER) at substantially
      lower bitrates when the encoder is fine-tuned to suppress speaker variation.
    source: §3.3, Table 3
    evidence: SSL-based tokenizers can match or exceed neural codec tokenizers on speech intelligibility (WER) at
      substantially lower bitrates when the encoder is fine-tuned to suppress speaker variation.
    confidence: high
    relevance: high
  - claim_id: supervised_fine_tuning_with_phoneme_recognition_targets_provides_consistent_gains
    role: supports
    claim: Supervised fine-tuning with phoneme recognition targets provides consistent gains over ASR targets for
      speech resynthesis intelligibility, suggesting phoneme alignment is more directly beneficial than word-level
      transcription for unit-based vocoders.
    source: §3.3, Table 3
    evidence: Supervised fine-tuning with phoneme recognition targets provides consistent gains over ASR targets
      for speech resynthesis intelligibility, suggesting phoneme alignment is more directly beneficial than word-level
      transcription for unit-based vocoders.
    confidence: high
    relevance: medium
  - claim_id: scaling_slm_model_size_has_diminishing_returns_on_tasks_that
    role: supports
    claim: Scaling SLM model size has diminishing returns on tasks that require sentence-level semantic coherence
      when the tokenizer quality is the primary bottleneck.
    source: §3.2, Table 2
    evidence: Scaling SLM model size has diminishing returns on tasks that require sentence-level semantic coherence
      when the tokenizer quality is the primary bottleneck.
    confidence: high
    relevance: medium
  limitations:
  - The unconstrained comparison in Table 2 pits a 150M SLM trained on 6k hours against models using 13B parameters
    and 150k+ hours. While the framing as "limited-resource" is accurate, direct comparisons against high-resource
    systems risk being misleading — the TSC gap (70.7 vs. 82.9 for SPIRIT LM) is large and likely attributable to
    scale rather than tokenizer quality.
  - The evaluation is English-only, and the authors acknowledge multilinguality as future work. The HiFi-GAN vocoder
    receives speaker and style IDs externally, so the speaker-invariance of the tokens is not tested in an open-vocabulary,
    zero-shot resynthesis scenario. The Expresso WER figures (>10%) are elevated by the expressive speaking styles
    (laughter, whisper), making absolute comparisons with non-expressive benchmarks impractical. The paper also
    does not report wall-clock training times or resource costs in sufficient detail for full reproducibility assessment.
  caveats: []
- id: interspeech-2025-0253
  published_date: "2025-08-17"
  entry_date: '2026-07-28'
  year: 2025
  venue: Interspeech
  task:
  - TTS
  architecture:
  - autoregressive-LM
  - hybrid
  relevance: medium
  evidence_role:
  - architecture_variant
  current_role: active_evidence
  method_family:
  - autoregressive_codec_language_models
  - hybrid_semantic_acoustic_tokenizers
  claims:
  - claim_id: maintaining_dynamically_updated_compressed_memory_across_sentences_improves_naturalness_and
    role: supports
    claim: Maintaining dynamically updated compressed memory across sentences improves naturalness and coherence
      in paragraph-level TTS compared to methods that use fixed-window preceding sentences.
    source: §3.4.1, Table 1
    evidence: Maintaining dynamically updated compressed memory across sentences improves naturalness and coherence
      in paragraph-level TTS compared to methods that use fixed-window preceding sentences.
    confidence: high
    relevance: medium
  - claim_id: speech_context_representations_are_more_effective_than_text_context_representations
    role: supports
    claim: Speech context representations are more effective than text context representations for guiding prosodic
      coherence in autoregressive LM-based TTS, due to the one-to-many relationship between text and speech.
    source: §3.4.2, Table 1
    evidence: Speech context representations are more effective than text context representations for guiding prosodic
      coherence in autoregressive LM-based TTS, due to the one-to-many relationship between text and speech.
    confidence: high
    relevance: medium
  - claim_id: applying_bidirectional_attention_to_prefix_tokens_via_a_prefix_mask
    role: supports
    claim: Applying bidirectional attention to prefix tokens via a prefix mask enhances in-context learning in decoder-only
      TTS LMs without compromising autoregressive generation consistency.
    source: §2.3, §3.4.2, Table 1
    evidence: Applying bidirectional attention to prefix tokens via a prefix mask enhances in-context learning in
      decoder-only TTS LMs without compromising autoregressive generation consistency.
    confidence: high
    relevance: medium
  - claim_id: excessively_long_variable_length_inference_prompts_increase_hallucination_and_content
    role: supports
    claim: Excessively long variable-length inference prompts increase hallucination and content errors in autoregressive
      TTS, while fixed-length context representations mitigate this instability.
    source: §3.4.1, Table 1
    evidence: Excessively long variable-length inference prompts increase hallucination and content errors in autoregressive
      TTS, while fixed-length context representations mitigate this instability.
    confidence: high
    relevance: medium
  limitations:
  - '- Evaluation is monolingual (Chinese Mandarin only); generalizability to other languages is unconfirmed. -
    The gap to Ground Truth (MOS 4.406 vs 3.796) remains substantial; paragraph-level naturalness is still an open
    problem. - No code release; replication requires re-implementing CAM on a CosyVoice backbone. - The model is
    evaluated only on single-speaker audiobooks; how it handles multi-speaker paragraphs is unknown. - The fixed
    32-embedding Perceiver Resampler bottleneck may lose fine-grained phonetic detail.'
  caveats: []
- id: interspeech-2025-0310
  published_date: "2025-08-17"
  entry_date: '2026-07-28'
  year: 2025
  venue: Interspeech
  task:
  - SCA
  - codec
  architecture:
  - autoregressive-LM
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: active_evidence
  method_family:
  - autoregressive_codec_language_models
  claims:
  - claim_id: moderately_coarse_fixed_width_segmentation_around_80_ms_combined_with
    role: supports
    claim: Moderately coarse fixed-width segmentation (around 80 ms) combined with a large K-means vocabulary outperforms
      fine-grained original-resolution tokenization on zero-shot spoken language understanding tasks without sacrificing
      accuracy.
    source: §4, Table 3
    evidence: Moderately coarse fixed-width segmentation (around 80 ms) combined with a large K-means vocabulary
      outperforms fine-grained original-resolution tokenization on zero-shot spoken language understanding tasks
      without sacrificing accuracy.
    confidence: high
    relevance: medium
  - claim_id: larger_segmentation_widths_require_proportionally_larger_vocabularies_to_preserve_phonetic
    role: supports
    claim: Larger segmentation widths require proportionally larger vocabularies to preserve phonetic discriminability,
      analogous to the phoneme-morpheme relationship in linguistics.
    source: §5.1
    evidence: Larger segmentation widths require proportionally larger vocabularies to preserve phonetic discriminability,
      analogous to the phoneme-morpheme relationship in linguistics.
    confidence: high
    relevance: medium
  - claim_id: variable_width_segmentation_based_on_linguistic_units_phoneme_syllable_word
    role: supports
    claim: Variable-width segmentation based on linguistic units (phoneme, syllable, word boundaries) does not consistently
      outperform fixed-width segmentation of matched median duration, and incurs additional computational cost.
    source: §5.3
    evidence: Variable-width segmentation based on linguistic units (phoneme, syllable, word boundaries) does not
      consistently outperform fixed-width segmentation of matched median duration, and incurs additional computational
      cost.
    confidence: high
    relevance: medium
  - claim_id: optimal_tokenization_settings_vary_across_spoken_language_understanding_benchmarks_suggesting
    role: supports
    claim: Optimal tokenization settings vary across spoken language understanding benchmarks, suggesting that ensembling
      multiple tokenization schemes may be necessary for broad SLU capability.
    source: §5.2
    evidence: Optimal tokenization settings vary across spoken language understanding benchmarks, suggesting that
      ensembling multiple tokenization schemes may be necessary for broad SLU capability.
    confidence: high
    relevance: medium
  limitations:
  - '- Evaluation is exclusively on SLU (understanding) tasks; no speech generation quality assessment. - Training
    data (LibriSpeech 960h) is small by current SLM standards; findings may differ at larger scale (LibriLight 60k).
    - Only HuBERT layer-9 as the SSL backbone; other models (WavLM, wav2vec 2.0) may yield different tradeoff curves.
    - Variable-width segmentation using predicted (rather than oracle) boundaries introduces inaccuracy; learned
    segment representations (as in Sylber) may be better than raw pooling. - Optimal tokenization varies per benchmark,
    motivating multi-token ensemble strategies not explored here.'
  caveats: []
- id: interspeech-2025-0319
  published_date: "2025-08-17"
  entry_date: '2026-07-28'
  year: 2025
  venue: Interspeech
  task:
  - TTS
  architecture:
  - autoregressive-LM
  relevance: high
  evidence_role:
  - core_evidence
  - control_evidence
  current_role: active_evidence
  method_family:
  - autoregressive_codec_language_models
  claims:
  - claim_id: llm_based_zero_shot_tts_systems_are_more_sensitive_to
    role: supports
    claim: LLM-based zero-shot TTS systems are more sensitive to noise in audio prompts than speaker-embedding-based
      approaches, because their in-context learning mechanism preserves the acoustic environment of the prompt.
    source: §1
    evidence: LLM-based zero-shot TTS systems are more sensitive to noise in audio prompts than speaker-embedding-based
      approaches, because their in-context learning mechanism preserves the acoustic environment of the prompt.
    confidence: high
    relevance: medium
  - claim_id: performing_speech_enhancement_in_the_discrete_acoustic_token_domain_outperforms
    role: supports
    claim: Performing speech enhancement in the discrete acoustic token domain outperforms waveform-domain SE methods
      in both speech quality and computational efficiency, achieving higher DNSMOS scores at roughly one-third the
      FLOPs.
    source: §4.1, Table 1
    evidence: Performing speech enhancement in the discrete acoustic token domain outperforms waveform-domain SE
      methods in both speech quality and computational efficiency, achieving higher DNSMOS scores at roughly one-third
      the FLOPs.
    confidence: high
    relevance: high
  - claim_id: waveform_domain_speech_enhancement_introduces_artifacts_that_degrade_speaker_identity
    role: supports
    claim: Waveform-domain speech enhancement introduces artifacts that degrade speaker identity in the enhanced
      prompt, resulting in lower speaker similarity in downstream zero-shot TTS compared to codec-domain denoising.
    source: §4.2, Table 3
    evidence: Waveform-domain speech enhancement introduces artifacts that degrade speaker identity in the enhanced
      prompt, resulting in lower speaker similarity in downstream zero-shot TTS compared to codec-domain denoising.
    confidence: high
    relevance: high
  - claim_id: predicting_only_the_first_two_rvq_groups_of_clean_acoustic
    role: supports
    claim: Predicting only the first two RVQ groups of clean acoustic tokens is sufficient for effective token-domain
      speech enhancement; predicting more groups increases complexity without improving quality.
    source: §4.1, Table 2
    evidence: Predicting only the first two RVQ groups of clean acoustic tokens is sufficient for effective token-domain
      speech enhancement; predicting more groups increases complexity without improving quality.
    confidence: high
    relevance: high
  - claim_id: the_vq_bottleneck_of_neural_codecs_acts_as_an_implicit
    role: supports
    claim: The VQ bottleneck of neural codecs acts as an implicit noise filter during quantization, providing a
      structural advantage for denoising in the token domain relative to signal-domain methods.
    source: §4.1
    evidence: The VQ bottleneck of neural codecs acts as an implicit noise filter during quantization, providing
      a structural advantage for denoising in the token domain relative to signal-domain methods.
    confidence: high
    relevance: medium
  limitations:
  - '- Specific to LauraTTS + Encodec (FunCodeec); adaptation to other codecs (DAC, SoundStream, EnCodec original)
    requires retraining. - Training uses synthetic noise (DNS Challenge 2022 + WHAM!); real-world noise types (music,
    babble, room impulse responses) may not be fully covered. - No ablation on LauraTTS model size or the effect
    of codec VQ count K on denoiser difficulty. - Intelligibility (WER 2.44%) is slightly worse than clean-prompt
    LauraTTS (2.33%), though close. - No evaluation of the codec denoiser on other TTS or SE tasks beyond the paired
    experiment.'
  caveats: []
- id: interspeech-2025-0347
  published_date: "2025-08-17"
  entry_date: '2026-07-28'
  year: 2025
  venue: Interspeech
  task:
  - codec
  - singing
  architecture:
  - GAN
  - VAE
  relevance: medium
  evidence_role:
  - architecture_variant
  - acceleration_evidence
  - control_evidence
  current_role: active_evidence
  method_family:
  - gan_neural_codecs_and_decoders
  - vae_vector_quantized_codecs
  claims:
  - claim_id: injecting_explicit_periodic_signals_into_a_neural_codec_decoder_enables
    role: supports
    claim: Injecting explicit periodic signals into a neural codec decoder enables independent F0 control during
      waveform reconstruction, decoupling pitch from the discrete token stream.
    source: §4.2, §4.3, Table 1, Figure 2
    evidence: Period variants achieve substantially lower F0-RMSE than Base (which embeds pitch implicitly in tokens)
      across all pitch shift conditions, and MOS improves from 2.37 to 3.28 at no shift.
    confidence: high
    relevance: high
  - claim_id: including_singing_voice_data_in_codec_training_improves_f0_accuracy
    role: supports
    claim: Including singing voice data in codec training improves F0 accuracy at high pitch ranges that speech-only
      corpora do not cover.
    source: §4.2, Figure 2, Figure 3
    evidence: Period+GT reduces F0-RMSE relative to Period when log F0 is shifted upward by 6-12 semitones, corresponding
      to the extended high-pitch coverage of GTSinger vs. LibriTTS.
    confidence: high
    relevance: high
  - claim_id: gradient_reversal_based_pitch_disentanglement_in_neural_codecs_does_not
    role: complicates
    claim: Gradient reversal-based pitch disentanglement in neural codecs does not reliably improve perceptual quality
      and may degrade it.
    source: §4.3, Table 1
    evidence: Period-GRL scores 2.44 MOS vs. Period at 3.28 MOS at no shift; Period-GRL+GT scores 3.33 vs. Period+GT
      at 3.55. The authors note the effect varies with training data domain, leaving the mechanism unclear.
    confidence: high
    relevance: medium
  - claim_id: codec_training_on_data_with_a_wider_pitch_range_can
    role: complicates
    claim: Codec training on data with a wider pitch range can improve quality at high pitches but reduces quality
      at lower pitches that are underrepresented in the new data.
    source: §4.3, Table 1
    evidence: At -6 semitone shift, Period+GT (3.29 MOS) scores lower than Period (3.42 MOS), attributed to the
      model allocating capacity to the wider pitch range of GTSinger at the cost of fidelity in the lower range.
    confidence: high
    relevance: high
  limitations:
  - All subjective evaluation is conducted on a proprietary Japanese children's song dataset recorded by two singers
    unseen during training. Results may not generalise to other singing styles, languages, or recording conditions.
  - 'The GRL disentanglement module did not improve naturalness in subjective evaluation and in some conditions
    worsened it; the authors acknowledge this interaction with training data domain requires clarification. The
    evaluation is a codec reconstruction task only: no downstream discrete-token-based singing synthesis system
    is presented, so the utility of PeriodCodec in a full SVS pipeline remains undemonstrated. Model size is not
    reported.'
  caveats: []
- id: interspeech-2025-0355
  published_date: "2025-08-17"
  entry_date: '2026-07-28'
  year: 2025
  venue: Interspeech
  task:
  - codec
  - evaluation
  architecture: []
  relevance: medium
  evidence_role:
  - architecture_variant
  - evaluation_caution
  current_role: active_evidence
  method_family: []
  claims:
  - claim_id: codec_noise_robustness_does_not_reliably_predict_across_bitrate_regimes
    role: supports
    claim: Codec noise robustness does not reliably predict across bitrate regimes, and clean-speech quality metrics
      are a poor proxy for noisy-condition performance.
    source: §4, Figure 1 in paper
    evidence: At the highest bitrate, DAC outperforms EnCodec in noise robustness; at 3 kbps, EnCodec surpasses
      DAC on most downstream metrics under noise. HiFi-Codec matches EnCodec in clean-speech perceptual quality
      but degrades far more under noise.
    confidence: high
    relevance: high
  - claim_id: neural_speech_codecs_exhibit_non_linear_input_output_behavior_violations
    role: supports
    claim: Neural speech codecs exhibit non-linear input-output behavior (violations of additivity and homogeneity)
      that partially explains their robustness failures under noise and signal overlap.
    source: §5, Figure 2 in paper
    evidence: Additivity improves with bitrate across all evaluated codecs, and codecs with better additivity (DAC)
      show stronger noise robustness. HiFi-Codec and SpeechTokenizer exhibit uneven homogeneity at extreme gains,
      aligning with their higher degradation in speaker and emotion metrics.
    confidence: high
    relevance: medium
  - claim_id: codec_evaluation_using_only_clean_speech_reconstruction_metrics_provides_an
    role: complicates
    claim: Codec evaluation using only clean-speech reconstruction metrics provides an incomplete characterization
      of codec behavior for real-world speech processing pipelines.
    source: §4
    evidence: Perceptual quality (PESQ) converges to similar values across codecs under severe noise, masking large
      differences in phonetic content (WER), speaker identity (EER), and emotional fidelity (SER-ACC).
    confidence: high
    relevance: high
  - claim_id: low_frequency_spectral_emphasis_introduced_by_time_domain_training_losses
    role: supports
    claim: Low-frequency spectral emphasis introduced by time-domain training losses is a consistent artifact across
      neural speech codecs and can cause unintended spectral coloration.
    source: §6, Figure 3 in paper
    evidence: All five evaluated codecs exhibit low-frequency boosting below 100 Hz under sine sweep analysis, most
      pronounced in EnCodec; the effect is attributed to waveform-level L1/L2 losses used in training.
    confidence: high
    relevance: medium
  limitations:
  - The evaluation is limited to five codecs, all of which use residual vector quantization. More recent codecs
    with different quantization strategies (e.g., fully transformer-based or finite scalar quantization approaches)
    are not included. The noise conditions, while diverse, use fixed SNR ranges and do not cover all real-world
    degradation types (e.g., codec artifacts, packet loss, or non-stationary noise sources). All test data are English;
    whether these robustness rankings generalize across languages and speaking styles is unexamined. The emotion
    preservation evaluation relies on a single emotion recognizer (emotion2vec) trained on limited data, and speaker
    identity assessment uses a single ECAPA-TDNN model, which may introduce evaluator-specific biases. No human
    perceptual evaluation is included.
  caveats: []
- id: interspeech-2025-0464
  published_date: "2025-08-17"
  entry_date: '2026-07-28'
  year: 2025
  venue: Interspeech
  task:
  - VC
  - codec
  architecture:
  - autoregressive-LM
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: active_evidence
  method_family:
  - autoregressive_codec_language_models
  claims:
  - claim_id: explicit_mutual_information_minimisation_at_the_codec_embedding_level_provides
    role: supports
    claim: Explicit mutual information minimisation at the codec-embedding level provides effective prosody-timbre
      disentanglement for voice conversion.
    source: §3.5, Table 3
    evidence: Removing the MI loss (L_MI) from the full system leads to a notably higher normalised F0 distance
      in the prosody-from-source scenario (3.28 vs. 2.82), while quality and timbre metrics change only modestly,
      isolating prosody control as the primary benefit of the MI objective.
    confidence: high
    relevance: high
  - claim_id: in_context_learning_codec_lms_can_serve_as_controllable_vc
    role: supports
    claim: In-context learning codec LMs can serve as controllable VC backbones when augmented with prosody-disentangling
      encoder modules.
    source: §3.4, §3.5, Table 2, Table 3
    evidence: The proposed system builds on VALL-E X's ICL mechanism and outperforms VALL-E X in speaker similarity
      (ASV 0.91 vs. 0.84), intelligibility (WER 0.101 vs. 0.115), naturalness (MOS 4.36 vs. 4.19), and prosody alignment
      (F0 distance 2.70 vs. 3.10) in the prompt-based scenario.
    confidence: high
    relevance: high
  - claim_id: prosody_disentanglement_at_the_codec_level_introduces_a_small_trade
    role: complicates
    claim: Prosody disentanglement at the codec level introduces a small trade-off in absolute codec reconstruction
      fidelity compared to the unmodified encoder.
    source: §3.3, Table 1
    evidence: PACE's ASV score (0.662) and NISQA score (3.98) are lower than the baseline EnCodec encoder (0.681,
      4.17), though the gap does not substantially affect system-level VC performance.
    confidence: high
    relevance: high
  - claim_id: prosody_from_source_and_prosody_from_prompt_are_distinct_capability
    role: refines
    claim: Prosody-from-source and prosody-from-prompt are distinct capability axes in voice conversion; systems
      strong at one do not automatically handle the other.
    source: §3.5, Table 3
    evidence: VALL-E X supports only prosody-from-prompt and is excluded from the source-prosody evaluation; TriAAN-VC
      and ProsoVC support only source-prosody and are excluded from the prompt-prosody evaluation. Only the proposed
      system is evaluated in both modes.
    confidence: high
    relevance: low
  limitations:
  - All evaluation is conducted on LibriTTS-clean-100 and test-clean, a relatively clean single-domain corpus with
    247 speakers. Generalisation to noisy environments, expressive or emotional speech, or cross-lingual settings
    is not tested.
  - The 54-hour training dataset is modest for a codec language model approach; it is unclear whether the disentanglement
    quality degrades with longer or more expressive source utterances. No code or demo is reported, limiting reproducibility.
    The paper does not ablate the number of RVQ codebooks or the sensitivity of the MI-minimisation trade-off weight
    (lambda_MI), leaving the robustness of the disentanglement objective undercharacterised. Prosody is operationalised
    solely through f0 and UV binary flags; richer prosodic dimensions such as energy, speaking rate, and phrase-level
    structure are not captured.
  caveats: []
- id: interspeech-2025-0468
  published_date: "2025-08-17"
  entry_date: '2026-07-28'
  year: 2025
  venue: Interspeech
  task:
  - codec
  architecture:
  - VAE
  - GAN
  relevance: medium
  evidence_role:
  - architecture_variant
  - acceleration_evidence
  current_role: active_evidence
  method_family:
  - vae_vector_quantized_codecs
  - gan_neural_codecs_and_decoders
  claims:
  - claim_id: directly_encoding_ssl_features_as_a_first_class_codec_stream
    role: supports
    claim: Directly encoding SSL features as a first-class codec stream produces stronger semantic preservation
      in RVQ-1 tokens than distillation from an SSL model, particularly for tonal languages where pitch fidelity
      is critical.
    source: §4.2, Table 2
    evidence: Directly encoding SSL features as a first-class codec stream produces stronger semantic preservation
      in RVQ-1 tokens than distillation from an SSL model, particularly for tonal languages where pitch fidelity
      is critical.
    confidence: high
    relevance: high
  - claim_id: operating_a_neural_codec_at_lower_frame_rates_with_more
    role: supports
    claim: Operating a neural codec at lower frame rates with more RVQ layers at fixed token rate improves audio
      quality over higher-frame-rate codecs with fewer layers at the same bitrate.
    source: §4.3, Table 3
    evidence: Operating a neural codec at lower frame rates with more RVQ layers at fixed token rate improves audio
      quality over higher-frame-rate codecs with fewer layers at the same bitrate.
    confidence: high
    relevance: high
  - claim_id: semantic_quality_of_rvq_1_tokens_is_a_primary_determinant
    role: supports
    claim: Semantic quality of RVQ-1 tokens is a primary determinant of downstream TTS intelligibility in autoregressive
      codec-based systems, independent of codec audio reconstruction quality.
    source: §4.4, Table 4
    evidence: Semantic quality of RVQ-1 tokens is a primary determinant of downstream TTS intelligibility in autoregressive
      codec-based systems, independent of codec audio reconstruction quality.
    confidence: high
    relevance: high
  - claim_id: an_ssl_based_semantic_stream_in_a_codec_encoder_can
    role: supports
    claim: An SSL-based semantic stream in a codec encoder can improve perceptual audio quality beyond what waveform-only
      codecs achieve, even when using the same decoder architecture.
    source: §4.3, Table 3
    evidence: An SSL-based semantic stream in a codec encoder can improve perceptual audio quality beyond what waveform-only
      codecs achieve, even when using the same decoder architecture.
    confidence: high
    relevance: high
  limitations:
  - The 12.5 Hz DualCodec-based TTS lags behind the 25 Hz variant in both WER and speaker similarity, indicating
    that the more aggressive downsampling introduces a ceiling on semantic accuracy that affects TTS quality. The
    paper acknowledges this gap as the primary remaining challenge.
  - The SSL model (w2v-BERT-2.0, 600M parameters, frozen) is required at TTS training time but not inference. This
    makes the training pipeline heavier than pure waveform codec approaches. It is also unclear whether the approach
    generalises to SSL models other than w2v-BERT-2.0, or whether the chosen 16th layer feature is optimal across
    languages beyond English and Mandarin.
  - Codec evaluation uses a single controlled bitrate band (~0.75 kbps); performance at higher bitrates typical
    of studio-quality TTS is not evaluated. The subjective test has only 8 participants, limiting statistical confidence
    for MUSHRA comparisons between closely-scoring systems.
  caveats: []
- id: interspeech-2025-0506
  published_date: "2025-08-17"
  entry_date: '2026-07-28'
  year: 2025
  venue: Interspeech
  task:
  - codec
  architecture:
  - transformer-enc-dec
  relevance: medium
  evidence_role:
  - architecture_variant
  current_role: active_evidence
  method_family:
  - transformer_encoder_decoder_tokenizers
  claims:
  - claim_id: neural_codec_token_spaces_capture_semantically_useful_information_beyond_audio
    role: supports
    claim: Neural codec token spaces capture semantically useful information beyond audio reconstruction, enabling
      effective use as pretraining targets for general audio representation models.
    source: §4, Table 1
    evidence: EnCodecMAE, which predicts frozen EnCodec RVQ tokens, outperforms BEATs (which uses random quantizer
      targets) and AudioMAE on global HEAREval score, with large+ST models reaching 99.2 on Audioset-only pretraining
      versus 96.0 for BEATs Iter 3.
    confidence: high
    relevance: high
  - claim_id: frame_level_temporal_representations_outperform_patch_based_representations_for_speech
    role: supports
    claim: Frame-level temporal representations outperform patch-based representations for speech tasks in universal
      audio models, while both approaches perform comparably on environmental sound classification.
    source: §4, Table 1
    evidence: EnCodecMAE (frame-level) achieves 96.3% speech command accuracy and 75.5% emotion recognition versus
      patch-based BEATs Iter 1's 91.0% and 68.7%, while matching or being slightly behind on ESC-50 (79.8 vs. 82.3).
    confidence: high
    relevance: medium
  - claim_id: self_supervised_universal_audio_representations_do_not_close_the_gap
    role: complicates
    claim: Self-supervised universal audio representations do not close the gap with dedicated speech SSL models
      on phoneme-level tasks such as ASR.
    source: §4, Table 2
    evidence: EnCodecMAE Large+ST achieves 8.59% WER on LibriSpeech with a language model, versus 2.94% for HuBERT
      Large trained solely on speech data. The paper attributes this to differences in target definition (MFCC clusters
      correlating with phonemes) and input representation (raw waveform).
    confidence: high
    relevance: medium
  - claim_id: the_optimal_input_representation_for_self_supervised_audio_models_is
    role: refines
    claim: The optimal input representation for self-supervised audio models is task-dependent rather than universally
      preferable across audio domains.
    source: §4, Table 1
    evidence: Melspectrograms yield higher global HEAREval scores than EnCodec encoder outputs (95.9 vs. 83.4 base
      model), but EnCodec input outperforms melspectrograms on pitch prediction (NSynth), where codec features capture
      finer harmonic structure.
    confidence: high
    relevance: high
  - claim_id: pretraining_data_diversity_is_essential_for_universal_audio_representation_models
    role: supports
    claim: Pretraining data diversity is essential for universal audio representation; models pretrained on a single
      domain generalise poorly outside that domain.
    source: §4, Table 1
    evidence: Speech-only pretraining (LL6K only) gives a global score of 75.3, music-only (FMA only) gives 89.1,
      while the full mixture (AS+FMA+LL6K) reaches 97.1 global, with gains concentrated in the non-primary domain
      for each restricted setting.
    confidence: high
    relevance: high
  limitations:
  - The ASR evaluation uses the SUPERB protocol (frozen encoder, BiLSTM probe, 100h fine-tune), which is designed
    to measure representation quality, not end-to-end ASR. The resulting WERs cannot be directly compared against
    fine-tuned ASR systems trained end-to-end and should be treated as a proxy for phoneme-level representation
    quality only.
  - Evaluation on HEAREval uses instance-level embeddings derived by averaging frame-level features, which discards
    temporal structure useful for tasks requiring sequence-level reasoning. The self-training stage cluster targets
    (k-means over 10k samples with k=1024) may be too coarse for high-resolution phonetic discrimination, which
    the authors identify as a direction for future work alongside exploring alternative training targets to improve
    ASR without sacrificing music and environmental performance. All models are evaluated with a frozen upstream
    encoder; fine-tuning experiments are not reported, leaving open whether the representations would improve further
    under task-specific adaptation.
  caveats: []
- id: interspeech-2025-0575
  published_date: "2025-08-17"
  entry_date: '2026-07-28'
  year: 2025
  venue: Interspeech
  task:
  - VC
  - TTS
  architecture:
  - VAE
  relevance: high
  evidence_role:
  - core_evidence
  - control_evidence
  current_role: active_evidence
  method_family:
  - vae_vector_quantized_codecs
  claims:
  - claim_id: watermarks_embedded_in_the_speaker_specific_latent_space_of_a
    role: supports
    claim: Watermarks embedded in the speaker-specific latent space of a neural codec survive zero-shot voice cloning
      synthesis, whereas waveform-level watermarks do not.
    source: §1, §3.5, Table 1
    evidence: Watermarks embedded in the speaker-specific latent space of a neural codec survive zero-shot voice
      cloning synthesis, whereas waveform-level watermarks do not.
    confidence: high
    relevance: high
  - claim_id: the_effectiveness_of_latent_space_watermarking_in_zero_shot_vc
    role: supports
    claim: The effectiveness of latent-space watermarking in zero-shot VC scenarios depends on the VC model preserving
      speaker-specific latents to achieve high speaker similarity.
    source: §1, §2.1
    evidence: The effectiveness of latent-space watermarking in zero-shot VC scenarios depends on the VC model preserving
      speaker-specific latents to achieve high speaker similarity.
    confidence: high
    relevance: low
  - claim_id: vc_simulated_augmentation_during_training_without_exposure_to_actual_vc
    role: supports
    claim: VC-simulated augmentation during training — without exposure to actual VC model outputs — is sufficient
      to achieve robust watermark recovery from synthesized audio.
    source: §2.3, §3.6, Table 2
    evidence: VC-simulated augmentation during training — without exposure to actual VC model outputs — is sufficient
      to achieve robust watermark recovery from synthesized audio.
    confidence: high
    relevance: medium
  - claim_id: codec_based_watermarking_pipelines_introduce_perceptible_audio_quality_degradation_compared
    role: complicates
    claim: Codec-based watermarking pipelines introduce perceptible audio quality degradation compared to waveform-level
      methods, representing a trade-off between VC resistance and transparency.
    source: §3.7, Table 3
    evidence: Codec-based watermarking pipelines introduce perceptible audio quality degradation compared to waveform-level
      methods, representing a trade-off between VC resistance and transparency.
    confidence: high
    relevance: high
  limitations:
  - The model is trained only on VCTK, a small clean dataset, which limits robustness to out-of-distribution audio
    editing (the paper notes lower-than-1.0 ACC on traditional editing, likely due to this). The audio quality impact
    (PESQ 2.2 vs. 4.32 for AudioSeal) is significant and may be prohibitive for some use cases. The approach does
    not address adversarial attacks specifically designed to remove latent-space watermarks. Coverage is limited
    to English; multilingual generalization is untested. Finally, the approach assumes the zero-shot VC model must
    preserve speaker-specific latents for high similarity — models that do not operate this way could evade detection.
  caveats: []
- id: interspeech-2025-0669
  published_date: "2025-08-17"
  entry_date: '2026-07-28'
  year: 2025
  venue: Interspeech
  task:
  - codec
  - TTS
  architecture:
  - hybrid
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: active_evidence
  method_family:
  - hybrid_semantic_acoustic_tokenizers
  claims:
  - claim_id: supervised_phonetic_data_ctc_and_phoneme_classification_can_replace_ssl
    role: supports
    claim: Supervised phonetic data (CTC and phoneme classification) can replace SSL pseudo-label distillation as
      the phonetic supervision signal in hybrid speech tokenizers, achieving superior phonetic representation without
      requiring a pretrained SSL model.
    source: §3.3, §5.1, Table 1
    evidence: Supervised phonetic data (CTC and phoneme classification) can replace SSL pseudo-label distillation
      as the phonetic supervision signal in hybrid speech tokenizers, achieving superior phonetic representation
      without requiring a pretrained SSL model.
    confidence: high
    relevance: high
  - claim_id: ctc_character_match_loss_is_the_dominant_driver_of_phonetic
    role: supports
    claim: CTC character-match loss is the dominant driver of phonetic encoding quality in RVQ-based tokenizers,
      contributing more than phoneme classification alone.
    source: §5.3, Table 4
    evidence: CTC character-match loss is the dominant driver of phonetic encoding quality in RVQ-based tokenizers,
      contributing more than phoneme classification alone.
    confidence: high
    relevance: high
  - claim_id: a_transformer_encoder_inserted_before_the_rvq_quantizer_improves_phonetic
    role: supports
    claim: A transformer encoder inserted before the RVQ quantizer improves phonetic representation, but requires
      stochastic skip-connection dropout during training to prevent the network from bypassing it.
    source: §3.2, §5.3, Table 5
    evidence: A transformer encoder inserted before the RVQ quantizer improves phonetic representation, but requires
      stochastic skip-connection dropout during training to prevent the network from bypassing it.
    confidence: high
    relevance: high
  - claim_id: hybrid_tokenizers_that_optimize_phonetic_encoding_via_direct_supervision_can
    role: supports
    claim: Hybrid tokenizers that optimize phonetic encoding via direct supervision can approach pure acoustic codecs
      in reconstruction quality while substantially surpassing SSL-distilled baselines.
    source: §5.1, Table 2
    evidence: Hybrid tokenizers that optimize phonetic encoding via direct supervision can approach pure acoustic
      codecs in reconstruction quality while substantially surpassing SSL-distilled baselines.
    confidence: high
    relevance: medium
  - claim_id: speech_tokenizers_that_better_encode_phonetic_structure_yield_stronger_downstream
    role: supports
    claim: Speech tokenizers that better encode phonetic structure yield stronger downstream speech language model
      performance on lexical discrimination benchmarks.
    source: §5.2, Table 3
    evidence: Speech tokenizers that better encode phonetic structure yield stronger downstream speech language
      model performance on lexical discrimination benchmarks.
    confidence: high
    relevance: medium
  limitations:
  - PAST requires labeled phoneme/character data, which constrains multilingual scalability — the paper explicitly
    acknowledges this and targets it as future work. PAST's reconstruction quality (SISNR=4.84) falls below pure
    EnCodec (SISNR=7.49), reflecting the phonetic-acoustic trade-off. The model is 185M parameters, larger than
    some baseline tokenizers, partly due to the transformer encoder. Evaluation is English-only on clean speech;
    robustness to noise, accents, and spontaneous speech is not assessed.
  caveats: []
- id: interspeech-2025-0874
  published_date: "2025-08-17"
  entry_date: '2026-07-28'
  year: 2025
  venue: Interspeech
  task:
  - SCA
  architecture:
  - autoregressive-LM
  - hybrid
  relevance: medium
  evidence_role:
  - architecture_variant
  - acceleration_evidence
  - control_evidence
  - evaluation_caution
  current_role: influential
  method_family:
  - autoregressive_codec_language_models
  - hybrid_semantic_acoustic_tokenizers
  claims:
  - claim_id: separating_user_and_agent_stream_representations_using_a_pretrained_speech
    role: supports
    claim: Separating user and agent stream representations — using a pretrained speech encoder for input and a
      neural codec for generation — allows full-duplex S2S models to bypass LLM speech pretraining without sacrificing
      conversation quality.
    source: §3, §6.1, §6.2
    evidence: Separating user and agent stream representations — using a pretrained speech encoder for input and
      a neural codec for generation — allows full-duplex S2S models to bypass LLM speech pretraining without sacrificing
      conversation quality.
    confidence: high
    relevance: high
  - claim_id: codec_personalisation_through_fine_tuning_on_target_speaker_data_can
    role: supports
    claim: Codec personalisation through fine-tuning on target-speaker data can recover audio quality at half the
      bitrate of an untuned codec, as measured by MOS, CER, and speaker similarity.
    source: §6.3, Table 4
    evidence: Codec personalisation through fine-tuning on target-speaker data can recover audio quality at half
      the bitrate of an untuned codec, as measured by MOS, CER, and speaker similarity.
    confidence: high
    relevance: high
  - claim_id: turn_level_alignment_between_text_and_speech_tokens_in_duplex
    role: supports
    claim: Turn-level alignment between text and speech tokens in duplex training is sufficient to learn barge-in
      behaviour; word-level alignment provides no measurable improvement.
    source: §3.1
    evidence: Turn-level alignment between text and speech tokens in duplex training is sufficient to learn barge-in
      behaviour; word-level alignment provides no measurable improvement.
    confidence: high
    relevance: medium
  - claim_id: full_duplex_end_to_end_models_remain_at_a_reasoning
    role: supports
    claim: Full-duplex end-to-end models remain at a reasoning disadvantage compared to cascaded oracle systems,
      though the gap narrows as backbone LLM quality increases.
    source: §6.2, Table 3
    evidence: Full-duplex end-to-end models remain at a reasoning disadvantage compared to cascaded oracle systems,
      though the gap narrows as backbone LLM quality increases.
    confidence: high
    relevance: medium
  - claim_id: open_source_availability_of_training_code_and_model_weights_is
    role: supports
    claim: Open-source availability of training code and model weights is a critical bottleneck for research progress
      in full-duplex spoken dialogue, given the historical concentration of such systems in closed industrial labs.
    source: §1
    evidence: Open-source availability of training code and model weights is a critical bottleneck for research
      progress in full-duplex spoken dialogue, given the historical concentration of such systems in closed industrial
      labs.
    confidence: high
    relevance: low
  limitations:
  - The backbone is TinyLlama-1.1B, a relatively small LLM. The reasoning gap between the end-to-end model and the
    GT+LLM cascaded oracle is real and acknowledged, particularly on QA tasks. Scaling to a larger backbone remains
    untested and its interaction with the duplex architecture is an open question.
  - Training data is entirely synthetic (TTS-generated user and agent speech) except for the ASR-QA portion, which
    introduces a domain mismatch with natural conversation. The fixed 0.64-second silence inserted before agent
    turns is a hard-coded heuristic that will affect latency in practice and may not generalise to more varied conversational
    pacing. The evaluation does not include a listening test (MOS via human raters) for duplex conversation quality
    — UTMOS and GPT score are proxies. The first-response latency metric is not comparable to Moshi because Moshi's
    proactive interruption behaviour makes the metric inapplicable, which limits direct system comparison.
  caveats: []
- id: interspeech-2025-0984
  published_date: "2025-08-17"
  entry_date: '2026-07-28'
  year: 2025
  venue: Interspeech
  task:
  - codec
  - evaluation
  architecture: []
  relevance: medium
  evidence_role:
  - architecture_variant
  - evaluation_caution
  current_role: active_evidence
  method_family: []
  claims:
  - claim_id: neural_speech_codecs_operating_below_approximately_1_100_bps_degrade
    role: supports
    claim: Neural speech codecs operating below approximately 1,100 bps degrade in intelligibility relative to uncoded
      speech, even when subjective quality metrics suggest acceptable performance.
    source: §4.1
    evidence: Neural speech codecs operating below approximately 1,100 bps degrade in intelligibility relative to
      uncoded speech, even when subjective quality metrics suggest acceptable performance.
    confidence: high
    relevance: medium
  - claim_id: wer_derived_from_large_vocabulary_asr_systems_is_not_a
    role: supports
    claim: WER derived from large-vocabulary ASR systems is not a valid proxy for subjective speech intelligibility
      in codec benchmarking on isolated word stimuli.
    source: §4.3
    evidence: WER derived from large-vocabulary ASR systems is not a valid proxy for subjective speech intelligibility
      in codec benchmarking on isolated word stimuli.
    confidence: high
    relevance: high
  - claim_id: stoi_and_estoi_correlate_well_with_subjective_drt_intelligibility_scores
    role: complicates
    claim: STOI and ESTOI correlate well with subjective DRT intelligibility scores when averaged across talker
      gender and wordlists, but fail to capture the finer-grained variation that subjective tests reveal.
    source: §4.3
    evidence: STOI and ESTOI correlate well with subjective DRT intelligibility scores when averaged across talker
      gender and wordlists, but fail to capture the finer-grained variation that subjective tests reveal.
    confidence: high
    relevance: medium
  - claim_id: talker_gender_significantly_affects_subjective_intelligibility_scores_independently_of_codec
    role: supports
    claim: Talker gender significantly affects subjective intelligibility scores independently of codec condition,
      and this variation is not reliably reflected in objective intelligibility metrics.
    source: §4.1, §4.2
    evidence: Talker gender significantly affects subjective intelligibility scores independently of codec condition,
      and this variation is not reliably reflected in objective intelligibility metrics.
    confidence: high
    relevance: high
  limitations:
  - The female reference signal shows consistently lower DRT scores than the male reference (M=85 vs. M=91.87),
    suggesting potential issues with the audio materials rather than codec effects. The authors flag this for further
    investigation, and it limits confidence in the codec-level conclusions for female talkers.
  - The study uses closed-set word tests (DRT/MRT), which are prone to ceiling effects at higher codec qualities
    — a limitation the authors acknowledge by suggesting future use of open-set SUS tests. The benchmark is limited
    to English and to 16 kHz stimuli (codecs designed for higher sampling rates were resampled), which may not reflect
    their intended operating conditions. The crowdsourced sample of 82 participants is distributed unevenly across
    eight test blocks, with some blocks approaching the recommended minimum of eight participants per condition.
    WER results were obtained via Whisper, which already yields 19–25% WER on clean single-word utterances, questioning
    whether the metric has sufficient headroom to distinguish codec conditions meaningfully.
  caveats: []
- id: interspeech-2025-0989
  published_date: "2025-08-17"
  entry_date: '2026-07-28'
  year: 2025
  venue: Interspeech
  task:
  - TTS
  - evaluation
  architecture:
  - autoregressive-LM
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  - infrastructure
  current_role: active_evidence
  method_family:
  - autoregressive_codec_language_models
  claims:
  - claim_id: speaker_diversity_in_training_data_is_a_stronger_driver_of
    role: complicates
    claim: Speaker diversity in training data is a stronger driver of zero-shot TTS generalization than audio quality
      or dataset size alone, as a 10-speaker high-quality dataset fails catastrophically on unseen speakers despite
      controlled recording conditions.
    source: §4.3, Table 3
    evidence: Speaker diversity in training data is a stronger driver of zero-shot TTS generalization than audio
      quality or dataset size alone, as a 10-speaker high-quality dataset fails catastrophically on unseen speakers
      despite controlled recording conditions.
    confidence: high
    relevance: medium
  - claim_id: mixed_bandwidth_audio_in_large_scale_speech_corpora_degrades_codec
    role: supports
    claim: Mixed-bandwidth audio in large-scale speech corpora degrades codec and vocoder training, making bandwidth
      estimation and filtering an essential step in high-bandwidth TTS data preparation.
    source: §1, §2.3
    evidence: Mixed-bandwidth audio in large-scale speech corpora degrades codec and vocoder training, making bandwidth
      estimation and filtering an essential step in high-bandwidth TTS data preparation.
    confidence: high
    relevance: high
  - claim_id: restoring_punctuation_and_capitalization_to_asr_derived_transcripts_is_feasible
    role: supports
    claim: Restoring punctuation and capitalization to ASR-derived transcripts is feasible at scale via text matching
      (87% coverage) with neural prediction for remaining cases, and meaningfully improves transcript quality for
      TTS prosody modeling.
    source: §2.1
    evidence: Restoring punctuation and capitalization to ASR-derived transcripts is feasible at scale via text
      matching (87% coverage) with neural prediction for remaining cases, and meaningfully improves transcript quality
      for TTS prosody modeling.
    confidence: high
    relevance: low
  - claim_id: providing_per_utterance_quality_metadata_wer_cer_bandwidth_speaker_count
    role: supports
    claim: Providing per-utterance quality metadata (WER, CER, bandwidth, speaker count) rather than applying fixed
      thresholds increases dataset utility by allowing downstream researchers to select quality-volume trade-offs
      appropriate to their application.
    source: §2.5, §2.6
    evidence: Providing per-utterance quality metadata (WER, CER, bandwidth, speaker count) rather than applying
      fixed thresholds increases dataset utility by allowing downstream researchers to select quality-volume trade-offs
      appropriate to their application.
    confidence: high
    relevance: medium
  limitations:
  - The dataset is English-only. Audio quality is not filtered by SNR, accepting noise present in LibriVox recordings;
    this is intentional (modern TTS can handle noise) but may affect some applications. The 44.1 kHz subset's bandwidth
    ranges from 13–22 kHz (mixed-bandwidth within the subset), which may complicate training for systems requiring
    uniform bandwidth. No listening-test based MOS evaluation is reported; evaluation relies on automatic metrics
    (SQUIM-MOS, SSIM). The Koel-TTS codec details are not specified in this paper. Future work on non-English and
    non-audiobook sources is mentioned but not pursued.
  caveats: []
- id: interspeech-2025-1020
  published_date: "2025-08-17"
  entry_date: '2026-07-28'
  year: 2025
  venue: Interspeech
  task:
  - TTS
  - evaluation
  architecture:
  - VAE
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: minor
  method_family:
  - vae_vector_quantized_codecs
  claims:
  - claim_id: larger_codebook_sizes_are_the_primary_lever_for_improving_discrete
    role: supports
    claim: Larger codebook sizes are the primary lever for improving discrete prosody representation quality, with
      embedding vector dimensionality having minimal independent effect.
    source: §5
    evidence: Across 150+ experiments, FFE improves monotonically with codebook bin count (20 to 320), while embedding
      size had "surprisingly insignificant" effect with smaller sizes often outperforming larger ones.
    confidence: high
    relevance: high
  - claim_id: joint_discrete_encoding_of_f0_and_energy_is_achievable_with
    role: supports
    claim: Joint discrete encoding of F0 and Energy is achievable with minimal reconstruction trade-off relative
      to F0-only encoding, when codebook capacity is matched.
    source: §5, Table 5
    evidence: The FI strategy reaches 0.49% FFE 20 % while simultaneously embedding Energy at 6.33 dB RMSE, outperforming
      the F0-only reference (1.60% FFE 20 %) from prior work.
    confidence: high
    relevance: high
  - claim_id: speaker_normalisation_of_f0_improves_energy_reconstruction_but_degrades_f0
    role: complicates
    claim: Speaker normalisation of F0 improves energy reconstruction but degrades F0 reconstruction accuracy when
      unvoiced regions must be preserved.
    source: §5, Tables 5-6
    evidence: FN+EN+VM achieves energy RMSE of 3.00 dB vs 6.33 dB for FI, but FFE 20 % rises to 3.30% vs 0.49% for
      FI, indicating a direct quality trade-off between unvoiced-region handling and energy embedding fidelity.
    confidence: high
    relevance: medium
  - claim_id: downstream_utility_of_prosody_embeddings_in_tts_or_speech_lm
    role: complicates
    claim: Downstream utility of prosody embeddings in TTS or speech LM systems is unverified; embedding quality
      is measured only through reconstruction accuracy rather than impact on synthesis naturalness.
    source: §4, §6
    evidence: Evaluation is limited to VDE and FFE reconstruction metrics on LibriTTS; no TTS or downstream speech
      generation experiments are reported.
    confidence: high
    relevance: low
  limitations:
  - No downstream TTS evaluation is provided, so the practical benefit of these embeddings for synthesis quality
    is undemonstrated. The evaluation is limited to English multispeaker speech (LibriTTS), and it is unclear how
    well the learned codebooks would transfer to other languages or acoustic conditions. The paper excludes audio
    files shorter than one second from training, which may bias the codebook away from short utterances common in
    conversational speech. Future work proposed by the authors involves prosody-conditioned text generation, which
    would serve as the first downstream validation of the released embeddings.
  caveats: []
- id: interspeech-2025-1084
  published_date: "2025-08-17"
  entry_date: '2026-07-28'
  year: 2025
  venue: Interspeech
  task:
  - TTS
  architecture:
  - autoregressive-LM
  relevance: high
  evidence_role:
  - core_evidence
  - acceleration_evidence
  - control_evidence
  - evaluation_caution
  current_role: active_evidence
  method_family:
  - autoregressive_codec_language_models
  claims:
  - claim_id: mamba_based_sequence_models_can_match_or_exceed_the_quality
    role: supports
    claim: Mamba-based sequence models can match or exceed the quality of larger Transformer-based TTS systems while
      enabling real-time streaming inference on CPU hardware.
    source: §4.5, §4.6, Table 1
    evidence: SMAM+MLM (26M params) achieves MOS 4.02 and CER 2.73%, matching Lee et al. (2024) at 263M params (MOS
      4.00, CER 4.01%) while reducing first-token latency from 26.5s to 0.065s on a single-threaded CPU.
    confidence: high
    relevance: low
  - claim_id: iterative_depthwise_refinement_of_rvq_tokens_substantially_improves_codec_tts
    role: supports
    claim: Iterative depthwise refinement of RVQ tokens substantially improves codec TTS quality over single-pass
      parallel depth prediction.
    source: §4.7, Table 1
    evidence: Replacing MLM depthwise decoding with a single-pass no-masking baseline (SMAM+noMLM) causes a significant
      drop in all quality metrics (MOS from 4.02 to 3.89, CER from 2.73% to 4.12%, UTMOS from 4.13 to 3.83) with
      negligible change in RTF and latency.
    confidence: high
    relevance: high
  - claim_id: objective_speaker_similarity_metrics_based_on_embedding_cosine_distance_do
    role: complicates
    claim: Objective speaker similarity metrics based on embedding cosine distance do not reliably predict subjective
      speaker similarity as judged by human listeners.
    source: §4.6, Table 1
    evidence: SMAM+MLM scores SECS 0.816 (below Lee et al.'s 0.863) but achieves higher SMOS of 3.36 vs. 3.27, indicating
      a divergence between embedding-space distance and perceptual similarity that has practical implications for
      zero-shot TTS evaluation.
    confidence: high
    relevance: low
  - claim_id: depthwise_decoding_strategies_for_rvq_present_an_explicit_quality_speed
    role: supports
    claim: Depthwise decoding strategies for RVQ present an explicit quality-speed trade-off that system designers
      can exploit based on deployment constraints.
    source: §3.3, §4.5, §4.6, Table 1
    evidence: SMAM+MLM (iterative, 3 passes) achieves MOS 4.02 and RTF 0.701, while SMAM+INR (single forward pass)
      achieves MOS 3.97 and RTF 0.568, demonstrating a consistent quality-speed trade-off across both objective
      and subjective evaluations.
    confidence: high
    relevance: high
  limitations:
  - Evaluation is limited to LibriTTS test-clean (English, read speech), leaving performance on spontaneous speech,
    noisy environments, and non-English languages uncharacterized. The SECS speaker similarity scores for the proposed
    models fall below the strongest baseline (Lee et al. 2024), indicating room for improvement in speaker faithfulness
    despite strong subjective SMOS scores. The paper does not release code, limiting reproducibility and adoption.
    RTF comparisons are not fully apples-to-apples since baselines generate complete utterances in batch mode while
    SMAM operates incrementally. Future directions mentioned include a fully streaming pipeline covering codec processing
    and applying depthwise decoding strategies to decoder-only speech language models.
  caveats: []
- id: interspeech-2025-1106
  published_date: "2025-08-17"
  entry_date: '2026-07-28'
  year: 2025
  venue: Interspeech
  task:
  - codec
  architecture:
  - VAE
  - GAN
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: active_evidence
  method_family:
  - vae_vector_quantized_codecs
  - gan_neural_codecs_and_decoders
  claims:
  - claim_id: explicit_speaker_perturbation_during_codec_training_is_more_effective_for
    role: supports
    claim: Explicit speaker perturbation during codec training is more effective for speaker disentanglement than
      relying on implicit information bottleneck alone.
    source: §3.3, Table 2; §3.4
    evidence: LSCodec achieves higher target-speaker SECS (0.852 at 50Hz) and lower speaker probing accuracy than
      TiCodec 1VQ (SECS 0.714), which uses an implicit VQ bottleneck without any perturbation, even though LSCodec
      operates at lower bitrate (0.45 kbps vs 0.75 kbps).
    confidence: high
    relevance: high
  - claim_id: single_codebook_discrete_speech_codecs_can_match_or_exceed_multi
    role: supports
    claim: Single-codebook discrete speech codecs can match or exceed multi-codebook acoustic codec reconstruction
      quality at ultra-low bitrates when content and speaker information are explicitly decoupled.
    source: §3.2, Table 1
    evidence: LSCodec-50Hz (V=300, 0.45 kbps) achieves WER 3.33% and MOS 4.49 on LibriTTS test-clean, outperforming
      all single-codebook baselines including WavTokenizer-small (WER 7.86, MOS 4.14 at 0.48 kbps) and multi-codebook
      SemantiCodec (WER 4.16 at 0.63 kbps).
    confidence: high
    relevance: high
  - claim_id: reducing_codec_frame_rate_through_temporal_downsampling_degrades_content_intelligibility
    role: complicates
    claim: Reducing codec frame rate through temporal downsampling degrades content intelligibility without proportional
      improvement in speaker disentanglement.
    source: §3.2, Table 1; §3.3, Table 2
    evidence: Halving the frame rate from 50Hz to 25Hz reduces bitrate from 0.45 to 0.25 kbps but increases reconstruction
      WER from 3.33% to 5.46% and VC WER from 4.04% to 6.32%, while reconstruction SECS changes only marginally
      (0.954 to 0.945).
    confidence: high
    relevance: high
  - claim_id: an_auxiliary_ssl_token_prediction_objective_is_necessary_for_maintaining
    role: refines
    claim: An auxiliary SSL token prediction objective is necessary for maintaining content intelligibility when
      an information bottleneck is used to remove speaker timbre.
    source: §3.5, Table 3
    evidence: Ablating the SSL token prediction loss increases VAE-stage WER from 4.96% to 11.22% while SECS remains
      essentially unchanged (0.811 to 0.811), confirming that the SSL prediction task guides content encoding independently
      of the speaker removal objective.
    confidence: high
    relevance: medium
  - claim_id: multi_stage_codec_training_establishing_a_continuous_disentangled_space_before
    role: supports
    claim: Multi-stage codec training, establishing a continuous disentangled space before quantization, improves
      both content preservation and speaker disentanglement compared to direct VQ training.
    source: §3.5, Table 3
    evidence: Skipping stage 1 (VAE pre-training) and training VQ-VAE directly degrades WER from 3.39% to 3.84%
      and SECS from 0.817 to 0.800 in the VQ-VAE stage, confirming that continuous-space initialization benefits
      discrete representation quality.
    confidence: high
    relevance: high
  limitations:
  - The model is evaluated exclusively on English LibriTTS data, leaving multilingual and cross-lingual generalization
    untested. The vocoder (CTX-vec2wav alpha) is trained on a fixed 24 kHz corpus, so quality at other sampling
    rates or in noisy conditions is unclear. Speaker probing uses a single X-vector classifier on LibriTTS speakers,
    which may not detect all forms of residual speaker information. The paper notes that stronger perturbation methods,
    better content preservation at 25Hz, and scaling to larger data are open directions. No code or pre-trained
    models are released.
  caveats: []
- id: interspeech-2025-1289
  published_date: "2025-08-17"
  entry_date: '2026-07-28'
  year: 2025
  venue: Interspeech
  task:
  - codec
  architecture:
  - hybrid
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: active_evidence
  method_family:
  - hybrid_semantic_acoustic_tokenizers
  claims:
  - claim_id: constant_frame_rate_coding_introduces_temporal_redundancy_in_neural_speech
    role: supports
    claim: Constant-frame-rate coding introduces temporal redundancy in neural speech codecs by allocating equal
      resolution to silence and phonetically dense regions alike.
    source: §1, §2.1
    evidence: Constant-frame-rate coding introduces temporal redundancy in neural speech codecs by allocating equal
      resolution to silence and phonetically dense regions alike.
    confidence: high
    relevance: medium
  - claim_id: dynamically_allocating_coarser_temporal_frames_to_low_entropy_speech_regions
    role: supports
    claim: Dynamically allocating coarser temporal frames to low-entropy speech regions reduces token sequence length
      without proportional degradation in reconstruction quality.
    source: §4.2, Table 1
    evidence: Dynamically allocating coarser temporal frames to low-entropy speech regions reduces token sequence
      length without proportional degradation in reconstruction quality.
    confidence: high
    relevance: medium
  - claim_id: reducing_the_number_of_encoded_frames_at_equivalent_bitrate_can
    role: supports
    claim: Reducing the number of encoded frames at equivalent bitrate can improve intelligibility, suggesting that
      sequence compactness benefits autoregressive downstream models independently of bitrate.
    source: §4.2, Table 1
    evidence: Reducing the number of encoded frames at equivalent bitrate can improve intelligibility, suggesting
      that sequence compactness benefits autoregressive downstream models independently of bitrate.
    confidence: high
    relevance: high
  - claim_id: variable_frame_rate_allocation_and_variable_bitrate_control_are_orthogonal
    role: supports
    claim: Variable frame rate allocation and variable bitrate control are orthogonal axes in neural codec design
      and can be combined additively.
    source: §3, §5
    evidence: Variable frame rate allocation and variable bitrate control are orthogonal axes in neural codec design
      and can be combined additively.
    confidence: high
    relevance: high
  limitations:
  - Evaluation is restricted to a single codec backbone (DAC) and a single dataset (LibriTTS). No human listening
    study is reported; all quality judgements rest on UTMOS, STOI, WER, and spectral distances. Generalisability
    to other architectures or acoustic conditions is untested.
  - The paper does not report downstream TTS or speech LM experiments, so the claimed benefit of shorter token sequences
    for generation quality and latency is prospective rather than demonstrated. The entropy-based frame allocation
    heuristic uses fixed hyperparameters (bin count N, smoothing σ) without ablation; sensitivity to these choices
    is unknown. Training with mixed granularity ratios uses a fixed heuristic allocation that may not be optimal.
    The granularity ratios at inference must be set by the user; no automatic target-rate optimisation is described.
  caveats: []
- id: interspeech-2025-1440
  published_date: "2025-08-17"
  entry_date: '2026-07-28'
  year: 2025
  venue: Interspeech
  task:
  - codec
  architecture:
  - VAE
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  - evaluation_caution
  current_role: active_evidence
  method_family:
  - vae_vector_quantized_codecs
  claims:
  - claim_id: self_supervised_disentanglement_of_speech_into_content_speaker_and_prosody
    role: supports
    claim: Self-supervised disentanglement of speech into content, speaker, and prosody streams can match or exceed
      supervised codec quality at significantly lower bitrate.
    source: §4.1, Table 1, Table 2
    evidence: Self-supervised disentanglement of speech into content, speaker, and prosody streams can match or
      exceed supervised codec quality at significantly lower bitrate.
    confidence: high
    relevance: high
  - claim_id: codec_coding_efficiency_is_more_sensitive_to_information_factorisation_than
    role: supports
    claim: Codec coding efficiency is more sensitive to information factorisation than to raw model capacity or
      bitrate allocation.
    source: §4.1, Table 1
    evidence: Codec coding efficiency is more sensitive to information factorisation than to raw model capacity
      or bitrate allocation.
    confidence: high
    relevance: high
  - claim_id: routing_wavlm_supervision_to_the_decoder_rather_than_the_encoder
    role: supports
    claim: Routing WavLM supervision to the decoder rather than the encoder during training improves speaker-content
      disentanglement for voice conversion.
    source: §2.5, §4.2
    evidence: Routing WavLM supervision to the decoder rather than the encoder during training improves speaker-content
      disentanglement for voice conversion.
    confidence: high
    relevance: medium
  - claim_id: ultra_low_bitrate_codecs_below_0_5_kbps_can_achieve
    role: supports
    claim: Ultra-low-bitrate codecs (below 0.5 kbps) can achieve MUSHRA scores competitive with codecs operating
      at 2–3 kbps when disentanglement is applied to reduce frame-level redundancy.
    source: §4.1, Table 2
    evidence: Ultra-low-bitrate codecs (below 0.5 kbps) can achieve MUSHRA scores competitive with codecs operating
      at 2–3 kbps when disentanglement is applied to reduce frame-level redundancy.
    confidence: high
    relevance: high
  limitations:
  - The demo and code availability are not confirmed in the paper or metadata. Reproducibility relies on external
    checkpoints for baselines — FACodec and SpeechTokenizer results are inferred from official checkpoints under
    potentially different conditions than the re-trained TiCodec and DAC baselines.
  - Evaluation is restricted to English (LibriSpeech and VCTK). Generalisation to other languages, accents, or spontaneous-speech
    domains is untested. The prosody encoder's low-mel-bin design is validated empirically via t-SNE visualisation
    but without a formal mutual information analysis. It is unclear how much prosody actually remains once the speaker
    and content encoders are also active during decoding — partial speaker clustering in Fig. 2 suggests the separation
    is not complete.
  caveats: []
- id: interspeech-2025-1538
  published_date: "2025-08-17"
  entry_date: '2026-07-28'
  year: 2025
  venue: Interspeech
  task:
  - VC
  architecture:
  - autoregressive-LM
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: active_evidence
  method_family:
  - autoregressive_codec_language_models
  claims:
  - claim_id: conditioning_acoustic_generation_on_explicitly_predicted_text_tokens_reduces_intelligibility
    role: supports
    claim: Conditioning acoustic generation on explicitly predicted text tokens reduces intelligibility errors in
      autoregressive voice conversion relative to purely acoustic-domain approaches.
    source: §3.3.3, Table 1
    evidence: Removing text token generation from StarVC raises WER from 6.27% to 7.30% and SECS-WavLM drops from
      0.472 to 0.382; StarVC achieves the lowest WER and CER among all compared systems including diffusion-based
      CosyVoice (8.24%/4.27%).
    confidence: high
    relevance: medium
  - claim_id: multi_stage_training_that_initializes_voice_conversion_with_asr_pretraining
    role: supports
    claim: Multi-stage training that initializes voice conversion with ASR pretraining improves both content preservation
      and speaker similarity relative to single-stage training.
    source: §3.3.3, Table 1
    evidence: Removing multi-stage training degrades SECS-Res from 0.835 to 0.812 and raises WER from 6.27% to 7.24%;
      multi-stage training is the single largest contributor in the ablation study.
    confidence: high
    relevance: low
  - claim_id: objective_speaker_embedding_metrics_and_perceptual_speaker_similarity_ratings_can
    role: complicates
    claim: Objective speaker embedding metrics and perceptual speaker similarity ratings can diverge for codec-based
      voice conversion systems trained with strong linguistic objectives.
    source: §3.3.1, §3.3.2, Tables 1-2
    evidence: StarVC scores marginally below CosyVoice on SECS-Res (0.835 vs. 0.839) and SECS-WavLM (0.472 vs. 0.478),
      yet exceeds CosyVoice on subjective SMOS (3.98 vs. 3.94), suggesting embedding-based metrics underestimate
      perceived similarity for this system class.
    confidence: high
    relevance: high
  - claim_id: autoregressive_voice_conversion_systems_can_produce_explicit_transcription_output_alongside
    role: refines
    claim: Autoregressive voice conversion systems can produce explicit transcription output alongside converted
      audio at negligible additional cost, enabling inline content verification without separate ASR inference.
    source: §3.3.1, Table 1
    evidence: StarVC generates text tokens with WER-Text of 4.95% and CER-Text of 1.51% as a byproduct of the VC
      decoding process, providing word-level content verification as part of the conversion pipeline.
    confidence: high
    relevance: medium
  limitations:
  - Subjective MOS evaluation involves only 20 listeners and 20 source-target pairs, making the reported SMOS and
    NMOS advantages over CosyVoice and OpenVoice V2 (all within overlapping confidence intervals) difficult to interpret
    as significant.
  - The evaluation covers English only on a single clean corpus (LibriTTS test-clean). Generalization to cross-lingual
    conversion, noisy conditions, or longer conversational utterances is untested. The three-stage training pipeline
    requires 180 GPU-hours on 8 H100s, representing a substantial compute cost that may limit practical adoption.
    Data augmentation relies on OpenVoice V2-synthesized speech, which could propagate artifacts from that system
    into StarVC's training distribution. Whether the text-before-speech decoding constraint generalizes to expressive
    or emotional speech conversion, where prosody is not captured by a pure transcription, remains an open question.
  caveats: []
- id: interspeech-2025-1639
  published_date: "2025-08-17"
  entry_date: '2026-07-28'
  year: 2025
  venue: Interspeech
  task:
  - codec
  - VC
  architecture:
  - GAN
  relevance: high
  evidence_role:
  - core_evidence
  - control_evidence
  current_role: active_evidence
  method_family:
  - gan_neural_codecs_and_decoders
  claims:
  - claim_id: targeted_distillation_into_specific_rvq_layers_can_isolate_fine_grained
    role: supports
    claim: Targeted distillation into specific RVQ layers can isolate fine-grained speaking style attributes (such
      as vocal effort) independently of semantic content in neural speech codecs.
    source: §2.3, §3.1, Table 2
    evidence: LombardTokenizer conditions the second RVQ layer via cosine distillation from Lombard speech encoders
      while keeping RVQ layer 1 semantically focused via mHuBERT; the resulting system achieves WER 10.97% and EER
      6.67% on vocal effort conversion on AVID, substantially outperforming FreeVC (WER 20.35%, EER 16.67%) while
      retaining comparable synthesis quality (PESQ 3.07 vs. SpeechTokenizer 3.01).
    confidence: high
    relevance: high
  - claim_id: disentanglement_constraints_in_neural_codecs_reduce_unconstrained_reconstruction_quality_relative
    role: complicates
    claim: Disentanglement constraints in neural codecs reduce unconstrained reconstruction quality relative to
      codecs without regularisation.
    source: §3.1, Table 1
    evidence: EnCodec (no disentanglement) achieves PESQ 3.32 and STOI 0.94 on LibriSpeech, while LombardTokenizer's
      dual distillation (semantic and Lombard) yields PESQ 3.07 and STOI 0.93, consistent with SpeechTokenizer's
      3.01; the quality gap widens on in-distribution neutral speech but narrows on expressive zero-shot data.
    confidence: high
    relevance: medium
  - claim_id: providing_a_specialized_style_encoder_to_a_voice_conversion_model
    role: complicates
    claim: Providing a specialized style encoder to a voice conversion model architecture does not guarantee that
      the encoder's information will be effectively exploited for style control.
    source: §3.2, Table 2, Figure 2
    evidence: FreeVC modified with the same Lombard encoder (FVClmb) fails to produce statistically distinguishable
      vocal effort levels in perceptual evaluation, and degrades WER on FLombard (32.63%) relative to the standard
      speaker-encoder variant (24.04%), while LT1 using the same encoder via RVQ distillation achieves WER 17.74%
      and accurate perceptual effort transfer.
    confidence: high
    relevance: medium
  - claim_id: zero_shot_generalization_of_speaking_style_transfer_across_languages_is
    role: supports
    claim: Zero-shot generalization of speaking style transfer across languages is achievable when codec disentanglement
      is guided by multilingual self-supervised representations.
    source: §2.1, §3.2, Table 2
    evidence: LombardTokenizer trained on English AVID achieves WER 15.77% on the unseen French FLombard dataset
      (zero-shot), compared to FreeVC's 24.04–32.63%; the paper attributes part of this advantage to replacing HuBERT
      with multilingual mHuBERT in the semantic RVQ layer.
    confidence: high
    relevance: high
  - claim_id: codec_level_disentanglement_enables_robust_cross_speaker_style_transfer_with
    role: supports
    claim: Codec-level disentanglement enables robust cross-speaker style transfer with low speaker identity leakage
      under both intra-speaker and inter-speaker conditions.
    source: §3.2, Table 2, Figure 2
    evidence: LombardTokenizer inter-speaker vocal effort conversion produces perceptual distributions not significantly
      different from intra-speaker conversion (Dunn's test), with EER remaining below 8% on AVID for both LT1 (6.67%)
      and LT2 (7.83%), indicating preserved speaker identity.
    confidence: high
    relevance: high
  limitations:
  - The zero-shot evaluation is from English training to French test data, and both languages are Indo-European
    — the claim of cross-lingual generalisation has not been tested on typologically distant languages.
  - The model is evaluated on controlled intensity datasets (AVID, FLombard) and trained on studio-quality recordings;
    performance in real-world noisy conditions or spontaneous Lombard speech is unknown. The AVID dataset uses instructed
    intensity levels rather than naturally elicited Lombard speech, which may reduce ecological validity. Synthesis
    quality is measured with PESQ and STOI, which emphasise intelligibility and signal fidelity but may not capture
    naturalness of expressive speech styles. The study does not address how many simultaneous disentanglement targets
    an RVQ structure can support before quality degrades substantially.
  caveats: []
- id: interspeech-2025-1641
  published_date: "2025-08-17"
  entry_date: '2026-07-28'
  year: 2025
  venue: Interspeech
  task:
  - TTS
  architecture:
  - autoregressive-LM
  relevance: high
  evidence_role:
  - core_evidence
  - control_evidence
  current_role: active_evidence
  method_family:
  - autoregressive_codec_language_models
  claims:
  - claim_id: explicit_phoneme_position_supervision_during_autoregressive_codec_training_eliminates_alignment
    role: supports
    claim: Explicit phoneme position supervision during autoregressive codec training eliminates alignment errors
      more effectively than phoneme identity prediction or monotonic decoding constraints.
    source: §4.2.1, Table 2; §4.3, Table 4
    evidence: Explicit phoneme position supervision during autoregressive codec training eliminates alignment errors
      more effectively than phoneme identity prediction or monotonic decoding constraints.
    confidence: high
    relevance: high
  - claim_id: alignment_failures_in_codec_language_model_tts_including_phoneme_skipping
    role: supports
    claim: Alignment failures in codec language model TTS — including phoneme skipping, repetition, and one-to-many
      correspondence — are fundamentally a training-objective problem rather than an inference-time problem.
    source: §4.2.1, §4.3, Table 4
    evidence: Alignment failures in codec language model TTS — including phoneme skipping, repetition, and one-to-many
      correspondence — are fundamentally a training-objective problem rather than an inference-time problem.
    confidence: high
    relevance: high
  - claim_id: jointly_predicting_phoneme_identity_and_position_introduces_conflicting_signals_that
    role: supports
    claim: Jointly predicting phoneme identity and position introduces conflicting signals that degrade pronunciation
      accuracy compared to position-only prediction.
    source: §4.2.1, Table 2
    evidence: Jointly predicting phoneme identity and position introduces conflicting signals that degrade pronunciation
      accuracy compared to position-only prediction.
    confidence: high
    relevance: medium
  - claim_id: robustness_improvements_in_autoregressive_codec_tts_can_be_achieved_without
    role: supports
    claim: Robustness improvements in autoregressive codec TTS can be achieved without changes to inference-time
      decoding strategy or additional duration prediction stages.
    source: §3.3, §4.2.1, Table 1
    evidence: Robustness improvements in autoregressive codec TTS can be achieved without changes to inference-time
      decoding strategy or additional duration prediction stages.
    confidence: high
    relevance: high
  limitations:
  - The model is trained and evaluated exclusively on Mandarin using a proprietary G2P toolkit and character-level
    duration annotations from the WenetSpeech4TTS dataset. Generalisation to languages without character-aligned
    duration labels, or to datasets where forced-alignment quality is lower, is untested.
  - The approach requires phoneme duration annotations at training time to construct the position sequence, which
    constrains its applicability to datasets with reliable forced alignments. At inference, the enrollment speech
    prompt must include a duration estimate — the paper does not discuss what happens when this estimate is imprecise.
  - The subjective evaluation is limited (15 listeners, 30 samples per system, Mandarin only), so the CMOS and SMOS
    figures should not be treated as strong evidence of naturalness gains beyond what the objective CER improvements
    already imply. Model size is not reported, making direct reproducibility assessment difficult.
  caveats: []
- id: interspeech-2025-1776
  published_date: "2025-08-17"
  entry_date: '2026-07-28'
  year: 2025
  venue: Interspeech
  task:
  - TTS
  architecture:
  - hybrid
  relevance: high
  evidence_role:
  - core_evidence
  current_role: active_evidence
  method_family:
  - hybrid_semantic_acoustic_tokenizers
  claims:
  - claim_id: multi_task_joint_training_on_synthesis_editing_and_continuation_tasks
    role: supports
    claim: Multi-task joint training on synthesis, editing, and continuation tasks improves speech synthesis quality
      over single-task training in non-autoregressive codec models.
    source: §3.2, §3.3, Tables 1, 3
    evidence: In all three tokenizer configurations (ST, STDAC, HuDAC), multi-task SpeechSEC consistently outperforms
      the corresponding single-task baseline on MOS, voice preservation, WER, and CER, with gains confirmed by ablation.
    confidence: high
    relevance: high
  - claim_id: in_multi_task_speech_generation_training_editing_tasks_primarily_contribute
    role: refines
    claim: In multi-task speech generation training, editing tasks primarily contribute intelligibility improvements
      while continuation tasks primarily contribute acoustic quality and voice preservation.
    source: §3.3, Table 3
    evidence: Ablation removing the editing task increases WER by up to 4.4 points with minimal audio quality change;
      removing continuation degrades MOS by up to 0.18 and voice preservation by up to 0.06 with smaller intelligibility
      effects.
    confidence: high
    relevance: medium
  - claim_id: non_autoregressive_masked_token_prediction_frameworks_can_unify_speech_synthesis
    role: supports
    claim: Non-autoregressive masked token prediction frameworks can unify speech synthesis, editing, and continuation
      tasks through task-specific input conditioning within a single model.
    source: §2, §3.2, Table 2
    evidence: SpeechSEC handles all three tasks with a shared Conformer backbone, differentiating tasks via a Task
      Register embedding and per-task masking strategies, achieving competitive quality on editing (MOS 3.93) and
      continuation (MOS 3.63) alongside synthesis.
    confidence: high
    relevance: medium
  - claim_id: the_choice_of_semantic_and_acoustic_token_extractor_significantly_affects
    role: complicates
    claim: The choice of semantic and acoustic token extractor significantly affects absolute synthesis quality
      in codec-based TTS, even when model architecture and training are held constant.
    source: §3.2, Table 1
    evidence: With the same SpeechSEC architecture and training scheme, MOS ranges from 3.65 (STDAC) to 4.20 (ST)
      across the three tokenizer configurations, and voice preservation from 0.61 to 0.72, indicating that tokenizer
      quality is a dominant factor.
    confidence: high
    relevance: high
  limitations:
  - The evaluation is restricted to LibriTTS-R, a clean studio-quality English corpus, leaving generalization to
    noisy, spontaneous, or multilingual speech untested. Cross-system comparisons with SoundStorm use independently
    reported numbers from separate evaluations, weakening the claim of surpassing prior state of the art. Model
    parameter count is not reported, preventing meaningful comparisons of capacity-normalised performance. Speech
    continuation lacks intelligibility metrics (WER, CER) by design, limiting interpretability of those results.
    The paper does not evaluate the editing task on real-world editing scenarios beyond random masking.
  caveats: []
- id: interspeech-2025-1993
  published_date: "2025-08-17"
  entry_date: '2026-07-28'
  year: 2025
  venue: Interspeech
  task:
  - TTS
  - VC
  architecture:
  - autoregressive-LM
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: active_evidence
  method_family:
  - autoregressive_codec_language_models
  claims:
  - claim_id: embedding_watermark_detection_directly_into_codec_encoder_training_is_a
    role: supports
    claim: Embedding watermark detection directly into codec encoder training is a viable alternative to post-hoc
      or hard-coded watermark gates for protecting open-source zero-shot TTS models.
    source: §2.2, §3.3.1
    evidence: Embedding watermark detection directly into codec encoder training is a viable alternative to post-hoc
      or hard-coded watermark gates for protecting open-source zero-shot TTS models.
    confidence: high
    relevance: high
  - claim_id: neural_codec_architectures_are_a_natural_intervention_point_for_access
    role: supports
    claim: Neural codec architectures are a natural intervention point for access-control in speaker-conditioned
      TTS because they mediate all speaker information transfer from prompt to synthesis.
    source: §1, §2.3
    evidence: Neural codec architectures are a natural intervention point for access-control in speaker-conditioned
      TTS because they mediate all speaker information transfer from prompt to synthesis.
    confidence: high
    relevance: high
  - claim_id: training_time_augmentation_with_common_audio_distortions_substantially_improves_a
    role: supports
    claim: Training-time augmentation with common audio distortions substantially improves a codec's robustness
      to watermark removal attacks without degrading reconstruction quality on clean audio.
    source: §2.2, Table 1, Table 2
    evidence: Training-time augmentation with common audio distortions substantially improves a codec's robustness
      to watermark removal attacks without degrading reconstruction quality on clean audio.
    confidence: high
    relevance: high
  - claim_id: codec_level_defenses_for_voice_cloning_create_a_structural_barrier
    role: supports
    claim: Codec-level defenses for voice cloning create a structural barrier to adaptation attacks because TTS
      models trained on modified codec distributions cannot be trivially swapped to unprotected codecs without retraining.
    source: §2.3, §3.3.2
    evidence: Codec-level defenses for voice cloning create a structural barrier to adaptation attacks because TTS
      models trained on modified codec distributions cannot be trivially swapped to unprotected codecs without retraining.
    confidence: high
    relevance: high
  limitations:
  - The defense is effective only against speech watermarked with the specific watermarking model (AudioSeal) used
    during codec training. A copyrighted voice that is unwatermarked — or that is protected with a different, unseen
    watermarking system — receives no protection. The attacker simply needs to avoid using an AudioSeal-watermarked
    prompt.
  - The evaluation is conducted only on the VALL-E architecture and EnCodec backbone; whether the approach generalizes
    to other zero-shot TTS architectures (e.g., flow-matching or diffusion-based systems) is not tested. All data
    is clean studio speech (LibriSpeech/LibriTTS-R); robustness in noisy or spontaneous speech conditions is unknown.
    The speed adjustment and low-pass filter attacks are reported as failures for the codec detector, though the
    authors argue these attacks also degrade clean prompts, partially neutralising them as practical bypass routes.
    No listening test compares watermark-rejected output against human expectations of what a "protection failure"
    looks and sounds like.
  caveats: []
- id: interspeech-2025-2075
  published_date: "2025-08-17"
  entry_date: '2026-07-28'
  year: 2025
  venue: Interspeech
  task:
  - codec
  architecture: []
  relevance: medium
  evidence_role:
  - architecture_variant
  current_role: active_evidence
  method_family: []
  claims:
  - claim_id: multi_granularity_quantization_codebooks_encode_paralinguistic_and_prosodic_features_more
    role: supports
    claim: Multi-granularity quantization codebooks encode paralinguistic and prosodic features more efficiently
      than single-resolution codebooks at equivalent or higher bitrates.
    source: §6.2, Table 3
    evidence: SVCs (k=500 per codebook, ~544 bits/s on Expresso) outperform frame-level k=2000 baselines (~548 bits/s)
      on all emotion sub-categories and prominence classification, with angry speech F1 rising from 0.298 to 0.614.
    confidence: high
    relevance: medium
  - claim_id: pooling_continuous_speech_representations_before_discretization_retains_more_paralinguistic_and
    role: supports
    claim: Pooling continuous speech representations before discretization retains more paralinguistic and prosodic
      information than pooling after discretization.
    source: §6.1, Table 2
    evidence: Pre-pooling consistently outperforms post-pooling in both utterance-level SER (0.5074 vs 0.2834 accuracy)
      and word-level prominence classification (0.3423 vs 0.1210 F-score) across single-level and multi-level codebook
      conditions.
    confidence: high
    relevance: medium
  - claim_id: dsu_based_expressive_speech_resynthesis_retains_a_large_gap_in
    role: complicates
    claim: DSU-based expressive speech resynthesis retains a large gap in style fidelity relative to continuous-feature
      resynthesis, even with improved codebook designs.
    source: §6.2, §6.3, Table 4
    evidence: SVCs achieve 41.22% style classification accuracy on Expresso versus 74.72% for continuous features
      and 88.42% for ground truth, a 33-point gap despite SVCs outperforming all other DSU baselines tested.
    confidence: high
    relevance: high
  - claim_id: forced_alignment_requirements_constrain_multi_granularity_dsu_methods_to_text
    role: complicates
    claim: Forced-alignment requirements constrain multi-granularity DSU methods to text-paired speech settings,
      limiting their applicability to spontaneous or unlabelled data.
    source: §4.1, §7
    evidence: Phone, word, and utterance boundaries for SVC segmentation are derived from forced alignments using
      the Montreal Forced Aligner and HTK; no unsupervised segmentation is evaluated, and the authors list automatic
      segmentation as future work.
    confidence: high
    relevance: medium
  limitations:
  - The method depends on forced alignments at three linguistic levels (phone, word, utterance), which requires
    text transcriptions and limits use in low-resource or spontaneous speech domains. The authors acknowledge this
    and identify unsupervised or automatic segmentation as the most promising avenue for removing this constraint.
  - Evaluation uses automated quality proxies only (UTMOS, style classifier accuracy, WER). The authors note that
    human listening tests and qualitative error analysis are needed and explicitly defer them to future work.
  - The merged frame-level DSU representation discards part of the factorized structure SVCs provide. Whether architectures
    that natively consume multi-stream DSUs would yield further gains is unanswered.
  - Finally, all codebook vocabulary sizes are fixed at k=500. The paper suggests that per-level vocabulary tuning
    could improve task adaptation, but this is not explored empirically.
  caveats: []
- id: interspeech-2025-2447
  published_date: "2025-08-17"
  entry_date: '2026-07-28'
  year: 2025
  venue: Interspeech
  task:
  - TTS
  architecture:
  - autoregressive-LM
  relevance: medium
  evidence_role:
  - architecture_variant
  - acceleration_evidence
  - control_evidence
  - evaluation_caution
  current_role: active_evidence
  method_family:
  - autoregressive_codec_language_models
  claims:
  - claim_id: speculative_decoding_adapted_for_speech_can_reduce_autoregressive_inference_latency
    role: supports
    claim: Speculative decoding adapted for speech can reduce autoregressive inference latency without measurable
      degradation in subjective naturalness or speaker similarity.
    source: §4.1, §4.2
    evidence: Speculative decoding adapted for speech can reduce autoregressive inference latency without measurable
      degradation in subjective naturalness or speaker similarity.
    confidence: high
    relevance: low
  - claim_id: speech_token_sequences_exhibit_many_to_one_mappings_to_perceived
    role: supports
    claim: Speech token sequences exhibit many-to-one mappings to perceived quality, enabling relaxed acceptance
      criteria that improve decoding throughput over strict token-distribution matching.
    source: §2.2, §4.3
    evidence: Speech token sequences exhibit many-to-one mappings to perceived quality, enabling relaxed acceptance
      criteria that improve decoding throughput over strict token-distribution matching.
    confidence: high
    relevance: medium
  - claim_id: initialising_a_lightweight_draft_model_from_the_upper_layers_of
    role: supports
    claim: Initialising a lightweight draft model from the upper layers of the target model provides immediate vocabulary
      alignment and reduces the data requirements for draft model training.
    source: §2.3
    evidence: Initialising a lightweight draft model from the upper layers of the target model provides immediate
      vocabulary alignment and reduces the data requirements for draft model training.
    confidence: high
    relevance: medium
  - claim_id: inference_stage_acceleration_of_autoregressive_tts_is_achievable_without_fine
    role: supports
    claim: Inference-stage acceleration of autoregressive TTS is achievable without fine-tuning the target model,
      preserving deployment flexibility for frozen production systems.
    source: §2, §4.1
    evidence: Inference-stage acceleration of autoregressive TTS is achievable without fine-tuning the target model,
      preserving deployment flexibility for frozen production systems.
    confidence: high
    relevance: medium
  limitations:
  - The WER increase (3.67% → 5.70%) is unexplained beyond a data-scale hypothesis. The draft model's limited training
    data (LibriTTS, ~580h vs. CosyVoice 2's proprietary corpus) is identified as the likely cause, but this is not
    verified experimentally — e.g., by scaling draft training data or by ablating with matched data.
  - Evaluation is restricted to a single target model (CosyVoice 2) on a single English benchmark. Generalisation
    of SSD to multilingual systems, streaming inference contexts, or multi-codebook AR models (e.g., VALL-E-style
    RVQ decoding) is not explored. The tolerance factor β is treated as a fixed hyperparameter tuned on objective
    metrics; its interaction with speaker diversity and domain shift is unexamined. The reported 1.4× speedup measures
    LM-RTF only and does not account for the draft model's own compute overhead in the total pipeline time.
  caveats: []
- id: interspeech-2025-2564
  published_date: "2025-08-17"
  entry_date: '2026-07-28'
  year: 2025
  venue: Interspeech
  task:
  - SCA
  architecture:
  - autoregressive-LM
  relevance: medium
  evidence_role:
  - architecture_variant
  - acceleration_evidence
  - control_evidence
  current_role: active_evidence
  method_family:
  - autoregressive_codec_language_models
  claims:
  - claim_id: large_scale_monolingual_pre_training_followed_by_stereo_dialogue_fine
    role: supports
    claim: Large-scale monolingual pre-training followed by stereo dialogue fine-tuning enables spoken dialogue
      models to acquire language-specific conversational behaviors.
    source: §5, Table 3
    evidence: J-Moshi, trained on J-CHAT and stereo Japanese dialogue, exhibits more speech overlaps (5.0s/min)
      and more IPUs (53.2/min) than English Moshi (1.2s overlap, 35.1 IPUs), consistent with Japanese conversational
      norms.
    confidence: high
    relevance: low
  - claim_id: synthetic_spoken_dialogue_generated_by_multi_stream_tts_improves_language
    role: supports
    claim: Synthetic spoken dialogue generated by multi-stream TTS improves language capability in full-duplex dialogue
      models when added to fine-tuning data.
    source: §4.3, Table 2
    evidence: J-Moshi-ext (trained with 602 hours of TTS-synthesized dialogue added) achieves meaningfulness 2.30
      versus J-Moshi's 2.19, a statistically distinguishable improvement, with no degradation in naturalness.
    confidence: high
    relevance: low
  - claim_id: neural_audio_codecs_pre_trained_on_one_language_can_transfer
    role: complicates
    claim: Neural audio codecs pre-trained on one language can transfer to another with minimal acoustic degradation,
      but the autoregressive language model component requires substantial retraining to achieve acceptable dialogue
      quality.
    source: §4.3, Table 2
    evidence: Mimi re-synthesis of Japanese speech degrades by approximately 0.5 MOS from ground truth, while J-Moshi
      (with RQ-Transformer adapted) degrades by more than 1 MOS, identifying the language model as the primary quality
      bottleneck.
    confidence: high
    relevance: low
  - claim_id: morphological_density_differences_across_languages_affect_full_duplex_dialogue_model
    role: complicates
    claim: 'Morphological density differences across languages affect full-duplex dialogue model training dynamics:
      languages with higher phoneme-to-token ratios produce sparser text-to-audio token alignments that may require
      adjusted training objectives.'
    source: §5
    evidence: Japanese data preprocessing results in 88% PAD tokens in text sequences versus 65% for English in
      Moshi, reflecting that kanji characters encode more phonemes per token, and the authors flag this as a design
      consideration for future Japanese-specific training objectives.
    confidence: high
    relevance: low
  limitations:
  - Both J-Moshi and J-Moshi-ext score above 1 MOS below the Mimi re-synthesis ceiling, indicating the autoregressive
    RQ-Transformer component is a major quality bottleneck in Japanese. Mimi itself degrades by approximately 0.5
    MOS from ground truth when applied to Japanese without adaptation, suggesting codec fine-tuning for Japanese
    will be necessary for production-quality systems.
  - The comparison with English Moshi in Table 3 is not conducted under identical experimental conditions (different
    test sets, possibly different prompt lengths), so turn-taking statistics should be interpreted as indicative
    rather than rigorously controlled. The 24.6% overall WER of the TTS-synthesized augmentation data introduces
    noise, and the effect of this noise on specific error categories is not analyzed. The paper does not evaluate
    spoken dialogue content quality beyond naturalness and meaningfulness, leaving turn-taking appropriateness and
    response coherence unmeasured.
  caveats: []
- id: interspeech-2025-2726
  published_date: "2025-08-17"
  entry_date: '2026-07-28'
  year: 2025
  venue: Interspeech
  task:
  - codec
  architecture:
  - GAN
  - hybrid
  relevance: high
  evidence_role:
  - core_evidence
  current_role: active_evidence
  method_family:
  - gan_neural_codecs_and_decoders
  - hybrid_semantic_acoustic_tokenizers
  claims:
  - claim_id: staged_codec_training_that_separates_quantizer_optimization_from_decoder_optimization
    role: supports
    claim: Staged codec training that separates quantizer optimization from decoder optimization can improve single-codebook
      reconstruction quality beyond joint training.
    source: §3.6, Table 3
    evidence: 'DS-Codec''s two-stage framework (mirror Stage 1 to train the quantizer, non-mirror Stage 2 to specialize
      the decoder) outperforms APCodec+''s single-stage joint training at both stages: Stage 1 UTMOS 4.123 vs. 4.113,
      PESQ 2.768 vs. 2.632; final model UTMOS 4.214 vs. 4.186 on LibriSpeech.'
    confidence: high
    relevance: high
  - claim_id: a_mirrored_encoder_decoder_constraint_during_quantizer_training_reduces_the
    role: supports
    claim: A mirrored encoder-decoder constraint during quantizer training reduces the input-output MSE of the quantization
      module, producing more robust codebooks.
    source: §3.5, Figure 2
    evidence: VQ loss curves during Stage 1 show the mirrored structure achieves lower quantization MSE than the
      non-mirrored structure across training epochs, even though the non-mirrored structure has lower VQ loss; the
      paper interprets smaller MSE as higher codebook fidelity.
    confidence: high
    relevance: medium
  - claim_id: product_quantization_over_multiple_small_sub_codebooks_enables_large_effective
    role: supports
    claim: Product quantization over multiple small sub-codebooks enables large effective codebook sizes while preserving
      the single-token-per-frame interface required by LLM-based TTS.
    source: §2.2.2, Table 1
    evidence: DS-Codec-PQ combines four 16-code VQ modules to produce a 65,536-code effective codebook, indexed
      as a single integer, achieving UTMOS 4.214 and PESQ 2.882 at 1.28kbps with 80 tokens/second.
    confidence: high
    relevance: high
  - claim_id: a_stronger_decoder_does_not_straightforwardly_compensate_for_a_weaker
    role: complicates
    claim: A stronger decoder does not straightforwardly compensate for a weaker quantizer when both are trained
      jointly.
    source: §3.6, Table 3
    evidence: APCodec+'s joint training with a non-mirrored (stronger) decoder in Stage 1 yields lower Stage 1 PESQ
      (2.632) than DS-Codec's mirror Stage 1 with a weaker mirrored decoder (PESQ 2.768), suggesting quantizer quality
      dominates reconstruction fidelity early in training.
    confidence: high
    relevance: medium
  limitations:
  - Evaluation is limited to English read speech (LibriSpeech) and one supplementary in-domain set (LJSpeech). Performance
    on noisy, spontaneous, or multilingual speech is untested. Model size is not reported, making it impossible
    to assess parameter efficiency relative to BigCodec (159M) or DAC (74M). The UTMOS and PESQ metrics used are
    objective proxies for perceptual quality; no formal subjective listening study is reported.
  - The Stage 2 improvement from retaining decoder weights vs. reinitializing them is stated but not ablated directly
    — the comparison to APCodec+ involves multiple differences (stage design, weight retention, architecture), so
    the isolated effect of weight retention is unclear.
  caveats: []
- id: '2508.15827'
  published_date: "2025-08-18"
  entry_date: '2026-07-28'
  year: 2025
  venue: arXiv
  task:
  - SCA
  architecture:
  - autoregressive-LM
  relevance: medium
  evidence_role:
  - architecture_variant
  - acceleration_evidence
  current_role: active_evidence
  method_family:
  - autoregressive_codec_language_models
  claims:
  - claim_id: in_speech_models_reasoning_depth_and_response_latency_are_not
    role: supports
    claim: In speech models, reasoning depth and response latency are not fundamentally in conflict when the model's
      token generation rate substantially exceeds the real-time audio playback rate.
    source: §1, §3.1
    evidence: In speech models, reasoning depth and response latency are not fundamentally in conflict when the
      model's token generation rate substantially exceeds the real-time audio playback rate.
    confidence: high
    relevance: low
  - claim_id: interleaving_silent_reasoning_tokens_with_spoken_response_tokens_at_a
    role: supports
    claim: Interleaving silent reasoning tokens with spoken response tokens at a fixed ratio can improve accuracy
      on structured reasoning tasks while reducing audible output length.
    source: §2.3, §4.4, Table 2
    evidence: Interleaving silent reasoning tokens with spoken response tokens at a fixed ratio can improve accuracy
      on structured reasoning tasks while reducing audible output length.
    confidence: high
    relevance: medium
  - claim_id: the_thinking_before_speaking_paradigm_when_applied_directly_to_speech
    role: supports
    claim: The "thinking-before-speaking" paradigm, when applied directly to speech, produces user-facing latency
      or verbosity that impairs conversational quality independently of reasoning correctness.
    source: §1, §2.2
    evidence: The "thinking-before-speaking" paradigm, when applied directly to speech, produces user-facing latency
      or verbosity that impairs conversational quality independently of reasoning correctness.
    confidence: high
    relevance: low
  - claim_id: multi_stage_training_separating_modality_alignment_reasoning_transfer_and_acoustic
    role: supports
    claim: Multi-stage training — separating modality alignment, reasoning transfer, and acoustic synthesis — is
      an effective strategy for progressively adapting an existing speech LLM to a new generation paradigm.
    source: §3.3
    evidence: Multi-stage training — separating modality alignment, reasoning transfer, and acoustic synthesis —
      is an effective strategy for progressively adapting an existing speech LLM to a new generation paradigm.
    confidence: high
    relevance: medium
  - claim_id: synthetic_speech_based_mathematical_reasoning_datasets_constructed_from_text_corpora
    role: supports
    claim: Synthetic speech-based mathematical reasoning datasets constructed from text corpora via TTS can provide
      sufficient training signal for spoken reasoning capabilities.
    source: §3.2, §4.4
    evidence: Synthetic speech-based mathematical reasoning datasets constructed from text corpora via TTS can provide
      sufficient training signal for spoken reasoning capabilities.
    confidence: high
    relevance: medium
  limitations:
  - The entire evaluation uses a single benchmark (Spoken-MQA) focused on mathematics. There is no assessment of
    speech naturalness, intelligibility, or reasoning accuracy on open-domain conversational tasks. Reported latency
    claims refer to the absence of a pre-speech reasoning phase rather than to measured real-time performance metrics.
  - The fixed 2:8 interleaving ratio is derived from a throughput estimate for a specific GPU configuration; it
    is not adaptive and may be suboptimal for different deployment environments or model sizes. The training data
    is entirely synthetic — both the audio (produced by CosyVoice2-0.5B) and the reasoning traces (constructed algorithmically
    from text datasets). Whether the model generalises to naturalistic spoken queries beyond maths problems is untested.
    The GPT-based verification stage for dataset quality introduces a dependency on a proprietary closed model that
    is not reproducible.
  caveats: []
- id: '2508.15931'
  published_date: "2025-08-21"
  entry_date: '2026-07-28'
  year: 2025
  venue: arXiv
  task:
  - evaluation
  architecture:
  - transformer-enc-dec
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: active_evidence
  method_family:
  - transformer_encoder_decoder_tokenizers
  claims:
  - claim_id: pairwise_comparison_framing_reduces_annotation_subjectivity_and_label_sparsity_in
    role: supports
    claim: Pairwise comparison framing reduces annotation subjectivity and label sparsity in perceptual attribute
      modelling more effectively than scalar labelling alone.
    source: §2.2, §3.1
    evidence: Pairwise comparison framing reduces annotation subjectivity and label sparsity in perceptual attribute
      modelling more effectively than scalar labelling alone.
    confidence: high
    relevance: medium
  - claim_id: differential_attention_that_subtracts_shared_information_between_two_representations_improves
    role: supports
    claim: Differential attention that subtracts shared information between two representations improves cross-speaker
      generalisation in timbre attribute discrimination.
    source: §3.2, Table 1
    evidence: Differential attention that subtracts shared information between two representations improves cross-speaker
      generalisation in timbre attribute discrimination.
    confidence: high
    relevance: medium
  - claim_id: transitivity_based_data_augmentation_over_sparse_pairwise_labels_significantly_expands
    role: supports
    claim: Transitivity-based data augmentation over sparse pairwise labels significantly expands training coverage
      and improves model robustness without requiring additional human annotation.
    source: §3.1, Table 2
    evidence: Transitivity-based data augmentation over sparse pairwise labels significantly expands training coverage
      and improves model robustness without requiring additional human annotation.
    confidence: high
    relevance: medium
  - claim_id: acoustic_signal_features_and_human_perceptual_judgements_of_timbre_attributes
    role: supports
    claim: Acoustic signal features and human perceptual judgements of timbre attributes are systematically inconsistent,
      limiting the effectiveness of handcrafted feature approaches.
    source: §2.2, Figure 1
    evidence: Acoustic signal features and human perceptual judgements of timbre attributes are systematically inconsistent,
      limiting the effectiveness of handcrafted feature approaches.
    confidence: high
    relevance: medium
  limitations:
  - Evaluation is confined to VCTK-RVA, a dataset of 110 clean studio-recorded English speakers. Generalisation
    to spontaneous speech, noisy conditions, or non-English timbre descriptors is entirely unvalidated.
  - The model targets classification accuracy but does not demonstrate downstream impact on any generation task.
    Whether improved timbre attribute detection actually produces better controllable TTS or voice editing remains
    an open question. The FACodec encoder is frozen throughout; it is unclear whether joint fine-tuning would further
    improve performance. The 34 timbre attributes in VCTK-RVA are English-centric and human-annotated by a small
    group, raising questions about attribute definition consistency and cultural transferability.
  caveats: []
- id: '2508.16790'
  published_date: "2025-08-22"
  entry_date: '2026-07-28'
  year: 2025
  venue: arXiv
  task:
  - codec
  - TTS
  architecture:
  - diffusion
  - transformer-enc-dec
  - autoregressive-LM
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: influential
  method_family:
  - diffusion_latent_codec_generators
  - transformer_encoder_decoder_tokenizers
  - autoregressive_codec_language_models
  claims:
  - claim_id: text_conditioning_in_the_codec_decoder_rather_than_in_the
    role: supports
    claim: Text conditioning in the codec decoder, rather than in the language model alone, is a viable lever for
      achieving extreme compression rates in speech tokenization without adversarial training.
    source: §3.1, Table 4
    evidence: Text conditioning in the codec decoder, rather than in the language model alone, is a viable lever
      for achieving extreme compression rates in speech tokenization without adversarial training.
    confidence: high
    relevance: high
  - claim_id: a_single_end_to_end_training_objective_flow_matching_loss
    role: supports
    claim: A single end-to-end training objective (flow-matching loss) is sufficient to jointly optimise quantization
      and reconstruction in a speech codec, eliminating the need for multi-stage pipelines.
    source: §3.1, §4.2.2
    evidence: A single end-to-end training objective (flow-matching loss) is sufficient to jointly optimise quantization
      and reconstruction in a speech codec, eliminating the need for multi-stage pipelines.
    confidence: high
    relevance: high
  - claim_id: the_reconstruction_generation_gap_the_degradation_in_intelligibility_when_tokens
    role: supports
    claim: The reconstruction-generation gap — the degradation in intelligibility when tokens are predicted by a
      language model rather than encoding reference speech — varies substantially across tokenizer architectures
      and is not captured by reconstruction metrics alone.
    source: §4.3, Figure 3
    evidence: The reconstruction-generation gap — the degradation in intelligibility when tokens are predicted by
      a language model rather than encoding reference speech — varies substantially across tokenizer architectures
      and is not captured by reconstruction metrics alone.
    confidence: high
    relevance: medium
  - claim_id: lower_token_rates_in_speech_tokenizers_can_improve_autoregressive_tts
    role: supports
    claim: Lower token rates in speech tokenizers can improve autoregressive TTS intelligibility by shortening prediction
      sequences and reducing error accumulation, particularly on linguistically challenging inputs.
    source: §4.3, Table 5
    evidence: Lower token rates in speech tokenizers can improve autoregressive TTS intelligibility by shortening
      prediction sequences and reducing error accumulation, particularly on linguistically challenging inputs.
    confidence: high
    relevance: medium
  - claim_id: binary_spherical_quantization_without_a_commitment_loss_achieves_stable_end
    role: supports
    claim: Binary Spherical Quantization without a commitment loss achieves stable end-to-end training of a speech
      codec and produces superior representations to standard VQ under equal codebook sizes.
    source: §4.2.2, Table 4
    evidence: Binary Spherical Quantization without a commitment loss achieves stable end-to-end training of a speech
      codec and produces superior representations to standard VQ under equal codebook sizes.
    confidence: high
    relevance: high
  limitations:
  - 'TaDiCodec''s text-aware decoder is not a general-purpose audio codec: it requires a transcript at both training
    and inference time. The strong reconstruction and TTS results are conditional on text availability; performance
    at 6.25 Hz without text conditioning is not competitive (WER exceeds 10% at 12.5 Hz without text, per Table
    4). This limits applicability to codec-transmission, speech enhancement, or any scenario where transcriptions
    are unavailable.'
  - The diffusion decoder introduces multi-step inference latency. At 32 steps, decoding speed is acceptable for
    generation but higher than GAN vocoders; reducing to 5 steps degrades quality noticeably. The authors propose
    distillation as future work but have not yet demonstrated single-step performance.
  - The system has been validated on TTS only; whether TaDiCodec tokens are useful for spoken dialogue modeling,
    speech understanding, or other downstream tasks remains untested. The prompt mechanism, while effective, means
    full reconstruction quality depends on a speech prefix being available — similar to zero-shot TTS conditioning,
    but potentially limiting in some SCA deployment scenarios.
  caveats: []
- id: '2508.19205'
  published_date: "2025-08-26"
  entry_date: '2026-07-28'
  year: 2025
  venue: arXiv
  task:
  - TTS
  architecture:
  - hybrid
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: influential
  method_family:
  - hybrid_semantic_acoustic_tokenizers
  claims:
  - claim_id: extreme_acoustic_codec_compression_single_codebook_vae_at_7_5
    role: supports
    claim: Extreme acoustic codec compression (single-codebook VAE at 7.5 Hz) can achieve superior perceptual quality
      over multi-codebook discrete codecs operating at much higher frame rates.
    source: §3.3, Table 3
    evidence: Extreme acoustic codec compression (single-codebook VAE at 7.5 Hz) can achieve superior perceptual
      quality over multi-codebook discrete codecs operating at much higher frame rates.
    confidence: high
    relevance: high
  - claim_id: long_form_multi_speaker_tts_benefits_from_separate_acoustic_and
    role: supports
    claim: Long-form multi-speaker TTS benefits from separate acoustic and semantic tokenizers trained with task-specific
      objectives rather than a single unified codec.
    source: §2.1
    evidence: Long-form multi-speaker TTS benefits from separate acoustic and semantic tokenizers trained with task-specific
      objectives rather than a single unified codec.
    confidence: high
    relevance: high
  - claim_id: scaling_the_llm_backbone_in_a_next_token_diffusion_speech
    role: supports
    claim: Scaling the LLM backbone in a next-token diffusion speech system yields consistent gains in perceptual
      quality, speaker similarity, and expressiveness.
    source: §3.1, Table 1
    evidence: Scaling the LLM backbone in a next-token diffusion speech system yields consistent gains in perceptual
      quality, speaker similarity, and expressiveness.
    confidence: high
    relevance: low
  - claim_id: token_level_diffusion_conditioned_on_llm_hidden_states_enables_streaming
    role: supports
    claim: Token-level diffusion conditioned on LLM hidden states enables streaming speech generation without the
      codebook constraints of discrete autoregressive systems.
    source: §2.2
    evidence: Token-level diffusion conditioned on LLM hidden states enables streaming speech generation without
      the codebook constraints of discrete autoregressive systems.
    confidence: high
    relevance: high
  - claim_id: tts_systems_optimised_for_long_form_conversational_content_retain_competitive
    role: supports
    claim: TTS systems optimised for long-form conversational content retain competitive performance on short-utterance
      benchmarks without dedicated fine-tuning.
    source: §3.2, Table 2
    evidence: TTS systems optimised for long-form conversational content retain competitive performance on short-utterance
      benchmarks without dedicated fine-tuning.
    confidence: high
    relevance: medium
  limitations:
  - Training data is not disclosed. The paper is from Microsoft Research but does not specify the data composition,
    size, or any cleaning procedures, making it impossible to assess whether the reported gains are attributable
    to architecture or data advantage.
  - The model is limited to English and Chinese; other languages produce unpredictable outputs. The system does
    not model overlapping speech — a significant gap for realistic conversational audio. The subjective evaluation
    used only 8 long-form test conversations, which is a narrow sample; standard benchmark evaluations (SEED) are
    short-utterance only and do not capture the long-form quality the paper targets. Speaker similarity at 7.5 Hz
    remains below the best short-utterance systems (e.g., Seed-TTS at 0.762 SIM for English), suggesting the compressed
    representation sacrifices some speaker identity fidelity. The maximum of 4 speakers is a hard constraint imposed
    by the context design.
  caveats: []
- id: '2508.20660'
  published_date: "2025-08-28"
  entry_date: '2026-07-28'
  year: 2025
  venue: arXiv
  task:
  - codec
  - evaluation
  architecture: []
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  - infrastructure
  current_role: active_evidence
  method_family: []
  claims:
  - claim_id: acoustic_reconstruction_quality_and_semantic_information_retention_are_largely_orthogonal
    role: supports
    claim: Acoustic reconstruction quality and semantic information retention are largely orthogonal objectives
      in neural audio codec design, and optimising for one does not reliably improve the other.
    source: §4.3, Figure 2
    evidence: Acoustic reconstruction quality and semantic information retention are largely orthogonal objectives
      in neural audio codec design, and optimising for one does not reliably improve the other.
    confidence: high
    relevance: high
  - claim_id: codecs_trained_without_explicit_semantic_objectives_achieve_state_of_the
    role: supports
    claim: Codecs trained without explicit semantic objectives achieve state-of-the-art signal reconstruction but
      exhibit substantially higher word error rates than semantically-trained codecs of comparable bitrate.
    source: §4.3.1
    evidence: Codecs trained without explicit semantic objectives achieve state-of-the-art signal reconstruction
      but exhibit substantially higher word error rates than semantically-trained codecs of comparable bitrate.
    confidence: high
    relevance: high
  - claim_id: single_codebook_low_bitrate_codecs_that_preserve_textual_semantic_content
    role: complicates
    claim: Single-codebook, low-bitrate codecs that preserve textual semantic content tend to sacrifice paralinguistic
      information such as speaker emotion and music characteristics.
    source: §4.3.2, Table 5
    evidence: Single-codebook, low-bitrate codecs that preserve textual semantic content tend to sacrifice paralinguistic
      information such as speaker emotion and music characteristics.
    confidence: high
    relevance: high
  - claim_id: codec_performance_on_clean_speech_benchmarks_does_not_generalise_to
    role: supports
    claim: Codec performance on clean speech benchmarks does not generalise to noisier, more expressive, or musically
      rich audio domains, with performance gaps widening at low bitrates.
    source: §4.2, Table 4
    evidence: Codec performance on clean speech benchmarks does not generalise to noisier, more expressive, or musically
      rich audio domains, with performance gaps widening at low bitrates.
    confidence: high
    relevance: high
  - claim_id: flow_matching_based_audio_codecs_achieve_competitive_perceptual_quality_on
    role: supports
    claim: Flow-matching-based audio codecs achieve competitive perceptual quality on non-speech audio relative
      to traditional RVQ-GAN-based codecs at equivalent bitrate and codebook size.
    source: §4.2
    evidence: Flow-matching-based audio codecs achieve competitive perceptual quality on non-speech audio relative
      to traditional RVQ-GAN-based codecs at equivalent bitrate and codebook size.
    confidence: high
    relevance: high
  limitations:
  - The benchmark's semantic evaluation relies on two embedding-based proxy tasks — ASR probing and classification
    — that do not directly measure how codec representations affect generation quality or dialogue coherence in
    a deployed SLM. The authors acknowledge this, noting that semantic information is "inherently broad and multifaceted"
    and their current methods are "relatively simplistic." The self-collected dataset of 400 entries is small and
    may not generalise beyond the specific Bilibili content sources. All audio is downsampled to 16 kHz for fair
    comparison, which disadvantages high-sample-rate codecs (DAC-44k, FlowDec-48k) and blurs the bitrate advantage
    those models provide at their native resolution. Token-based semantic evaluation was explored but proved impractical
    for large-codebook single-VQ models (Stable-Codec, X-Codec-2.0) given the embedding mapping difficulty on LibriSpeech
    alone.
  caveats: []
- id: '2509.00503'
  published_date: "2025-08-30"
  entry_date: '2026-07-28'
  year: 2025
  venue: arXiv
  task:
  - VC
  - codec
  architecture:
  - autoregressive-LM
  - transformer-enc-dec
  relevance: medium
  evidence_role:
  - architecture_variant
  current_role: active_evidence
  method_family:
  - autoregressive_codec_language_models
  - transformer_encoder_decoder_tokenizers
  claims:
  - claim_id: adaptive_entropy_based_segmentation_of_discrete_speech_tokens_preserves_more
    role: supports
    claim: Adaptive entropy-based segmentation of discrete speech tokens preserves more task-relevant linguistic
      information than fixed-length downsampling at equivalent compression ratios.
    source: §6.1, Table 3; §6.2, Table 4
    evidence: Adaptive entropy-based segmentation of discrete speech tokens preserves more task-relevant linguistic
      information than fixed-length downsampling at equivalent compression ratios.
    confidence: high
    relevance: high
  - claim_id: optimal_token_granularity_differs_systematically_between_understanding_and_generation_tasks
    role: supports
    claim: 'Optimal token granularity differs systematically between understanding and generation tasks: recognition-oriented
      tasks (ASR, ST) benefit from moderate compression near phoneme rate, while voice conversion requires finer
      token density to maintain acoustic fidelity.'
    source: §5.1, Table 1; §6.3, Table 5
    evidence: 'Optimal token granularity differs systematically between understanding and generation tasks: recognition-oriented
      tasks (ASR, ST) benefit from moderate compression near phoneme rate, while voice conversion requires finer
      token density to maintain acoustic fidelity.'
    confidence: high
    relevance: high
  - claim_id: ssl_derived_semantic_tokens_at_standard_rates_50_hz_contain
    role: supports
    claim: SSL-derived semantic tokens at standard rates (50 Hz) contain substantial redundancy that can be removed
      without degrading — and occasionally improving — downstream task performance.
    source: §6.1, Table 3
    evidence: SSL-derived semantic tokens at standard rates (50 Hz) contain substantial redundancy that can be removed
      without degrading — and occasionally improving — downstream task performance.
    confidence: high
    relevance: medium
  - claim_id: entropy_boundaries_in_compressed_token_sequences_align_with_linguistically_meaningful
    role: supports
    claim: 'Entropy boundaries in compressed token sequences align with linguistically meaningful units: 15 Hz compression
      achieves 83.2% phoneme boundary alignment, while 7 Hz aligns primarily with word boundaries (89.7%).'
    source: Appendix C.3, Table 11
    evidence: 'Entropy boundaries in compressed token sequences align with linguistically meaningful units: 15 Hz
      compression achieves 83.2% phoneme boundary alignment, while 7 Hz aligns primarily with word boundaries (89.7%).'
    confidence: high
    relevance: high
  limitations:
  - 'Voice conversion quality degrades noticeably with compression: entropy-guided 15 Hz already falls below HuBERT
    deduplicated (Q-MOS 3.85 vs 4.12), and the gap widens at higher compression. For generation tasks, this framework
    does not improve on simply using deduplicated tokens — only for understanding tasks does compression help.'
  - The framework is evaluated exclusively on HuBERT-derived tokens; whether the entropy-based approach transfers
    to other SSL features (WavLM, w2v-BERT) or supervised tokenizers (S3, FACodec) is untested. The entropy LLM
    requires pre-training on 20k hours of speech, adding a pipeline step beyond vanilla k-means clustering. The
    method is tested on English only, and its behaviour on morphologically complex or tonal languages — where token
    redundancy patterns may differ — remains unknown. All evaluations use the English MLS training corpus, so generalisation
    across domains and recording conditions is unexplored.
  caveats: []
- id: '2509.02020'
  published_date: "2025-09-02"
  entry_date: '2026-07-28'
  year: 2025
  venue: arXiv
  task:
  - TTS
  - SCA
  architecture:
  - autoregressive-LM
  relevance: medium
  evidence_role:
  - architecture_variant
  - acceleration_evidence
  - control_evidence
  current_role: influential
  method_family:
  - autoregressive_codec_language_models
  claims:
  - claim_id: reducing_speech_tokenizer_frame_rate_to_12_5hz_with_explicit
    role: supports
    claim: Reducing speech tokenizer frame rate to 12.5Hz with explicit semantic supervision produces tokens that
      enable more stable text-to-token modelling over long dialogue sequences than higher-rate tokenizers without
      semantic injection.
    source: §2.1, §4.1, Table 1
    evidence: Reducing speech tokenizer frame rate to 12.5Hz with explicit semantic supervision produces tokens
      that enable more stable text-to-token modelling over long dialogue sequences than higher-rate tokenizers without
      semantic injection.
    confidence: high
    relevance: low
  - claim_id: a_dual_transformer_architecture_for_multi_layer_rvq_prediction_achieves
    role: supports
    claim: A dual-transformer architecture for multi-layer RVQ prediction achieves lower first-packet latency than
      the delay-pattern while providing stronger contextual conditioning from prior turns.
    source: §2.2
    evidence: A dual-transformer architecture for multi-layer RVQ prediction achieves lower first-packet latency
      than the delay-pattern while providing stronger contextual conditioning from prior turns.
    confidence: high
    relevance: high
  - claim_id: autoregressive_tts_systems_trained_on_multi_speaker_dialogue_data_with
    role: supports
    claim: Autoregressive TTS systems trained on multi-speaker dialogue data with text-speech interleaved formatting
      can infer and adjust prosody and emotion from implicit conversational context without explicit emotion labels.
    source: §3.2, §4.3, Table 3
    evidence: Autoregressive TTS systems trained on multi-speaker dialogue data with text-speech interleaved formatting
      can infer and adjust prosody and emotion from implicit conversational context without explicit emotion labels.
    confidence: high
    relevance: low
  - claim_id: sentence_by_sentence_multi_speaker_dialogue_tts_systems_produce_more
    role: supports
    claim: Sentence-by-sentence multi-speaker dialogue TTS systems produce more coherent prosody across turns than
      approaches that concatenate monologue TTS outputs or model a mixed audio track.
    source: §4.4, Table 4
    evidence: Sentence-by-sentence multi-speaker dialogue TTS systems produce more coherent prosody across turns
      than approaches that concatenate monologue TTS outputs or model a mixed audio track.
    confidence: high
    relevance: low
  - claim_id: fine_tuning_a_post_trained_dialogue_tts_model_on_as
    role: supports
    claim: Fine-tuning a post-trained dialogue TTS model on as little as 50 hours of speaker-specific data is sufficient
      to produce synthesis that is perceptually indistinguishable from human recordings in a majority of trials.
    source: §4.4, Figure 4
    evidence: Fine-tuning a post-trained dialogue TTS model on as little as 50 hours of speaker-specific data is
      sufficient to produce synthesis that is perceptually indistinguishable from human recordings in a majority
      of trials.
    confidence: high
    relevance: low
  limitations:
  - '- Currently limited to 3-minute dialogues with up to 4 speakers; scaling requires extending training corpus.
    - English speaker similarity (SIM 0.665) lags Mandarin (0.736), attributed to limited English voice diversity
    in training data — a data rather than architectural limitation. - Trails Mimi on PESQ metrics, likely because
    Mimi was trained on a massive English-only corpus closely matching LibriSpeech. - Emotion fine-tuning is demonstrated
    for a single distinctive female voice; generalisation to arbitrary voices and more nuanced emotional transitions
    is not evaluated. - No ablation of the semantic supervision contribution vs. the lower frame rate independently.'
  caveats: []
- id: '2509.02244'
  published_date: "2025-09-02"
  entry_date: '2026-07-28'
  year: 2025
  venue: arXiv
  task:
  - codec
  architecture:
  - VAE
  - GAN
  relevance: medium
  evidence_role:
  - architecture_variant
  - acceleration_evidence
  current_role: minor
  method_family:
  - vae_vector_quantized_codecs
  - gan_neural_codecs_and_decoders
  claims:
  - claim_id: single_stage_vector_quantization_without_residual_refinement_can_achieve_intelligibility
    role: supports
    claim: Single-stage vector quantization without residual refinement can achieve intelligibility and perceptual
      quality competitive with multi-level RVQ codecs at comparable bitrates.
    source: §5.2, Table 1
    evidence: Single-stage vector quantization without residual refinement can achieve intelligibility and perceptual
      quality competitive with multi-level RVQ codecs at comparable bitrates.
    confidence: high
    relevance: high
  - claim_id: training_a_hifi_gan_vocoder_on_codec_reconstructed_spectrograms_rather
    role: supports
    claim: Training a HiFi-GAN vocoder on codec-reconstructed spectrograms rather than clean references improves
      robustness to codec artefacts.
    source: §4.2
    evidence: Training a HiFi-GAN vocoder on codec-reconstructed spectrograms rather than clean references improves
      robustness to codec artefacts.
    confidence: high
    relevance: high
  - claim_id: patchwise_quantization_of_mel_spectrograms_produces_a_2d_discrete_token
    role: supports
    claim: Patchwise quantization of mel spectrograms produces a 2D discrete token grid compatible with low-latency
      streaming at practical bitrates without architectural complexity.
    source: §3, §5.4
    evidence: Patchwise quantization of mel spectrograms produces a 2D discrete token grid compatible with low-latency
      streaming at practical bitrates without architectural complexity.
    confidence: high
    relevance: low
  - claim_id: objective_metrics_such_as_mcd_can_diverge_from_perceptual_quality
    role: supports
    claim: Objective metrics such as MCD can diverge from perceptual quality metrics like PESQ when comparing codecs
      across different sampling rates and architectures.
    source: §5.2, Table 1
    evidence: Objective metrics such as MCD can diverge from perceptual quality metrics like PESQ when comparing
      codecs across different sampling rates and architectures.
    confidence: high
    relevance: medium
  limitations:
  - The training dataset is described only as a "multilingual speech corpus" without further detail — speaker count,
    language distribution, total hours, and data sources are unreported. This makes the results impossible to reproduce
    and limits the generalisability of claims about performance on "diverse" speech.
  - No subjective evaluation (MOS or MUSHRA) is included; all comparisons are objective-only, which is particularly
    limiting for a perceptual task like codec quality assessment. The paper does not ablate the patch size (4×4
    is the only configuration tested) or codebook size. MCD comparisons between 16kHz and 24kHz systems may be confounded
    by the sampling rate difference. The codec's real-world performance at varying packet-loss rates or noisy input
    conditions is untested. Future directions flagged by the authors include reducing the frequency-axis dimensionality
    toward a 1D token sequence, joint codec-vocoder training, and downstream use with autoregressive language models.
  caveats: []
- id: '2509.04685'
  published_date: "2025-09-04"
  entry_date: '2026-07-28'
  year: 2025
  venue: arXiv
  task:
  - codec
  architecture:
  - GAN
  relevance: high
  evidence_role:
  - core_evidence
  - acceleration_evidence
  - control_evidence
  current_role: active_evidence
  method_family:
  - gan_neural_codecs_and_decoders
  claims:
  - claim_id: content_adaptive_token_allocation_in_acoustic_tokenisers_achieves_better_reconstruction
    role: supports
    claim: Content-adaptive token allocation in acoustic tokenisers achieves better reconstruction quality than
      fixed-rate designs at equal or lower token budgets.
    source: §4.2, Table 1
    evidence: VARSTok at 30.95 Hz achieves UTMOS 3.8949 on LibriTTS test-clean, surpassing the 40 Hz WavTokenizer
      (3.6107) while using 23% fewer tokens; the 36.81 Hz configuration (UTMOS 4.000) nearly matches the 75 Hz WavTokenizer
      (4.025) with fewer than half the tokens.
    confidence: high
    relevance: medium
  - claim_id: dynamically_segmented_speech_tokens_carry_more_semantically_discriminative_information_than
    role: supports
    claim: Dynamically segmented speech tokens carry more semantically discriminative information than uniformly
      sampled tokens at the same average rate.
    source: §4.3, Table 2
    evidence: All VARSTok configurations outperform the 40 Hz WavTokenizer on all four ARCH benchmark classification
      tasks (emotion, digit recognition, intent), despite operating at lower average frame rates.
    confidence: high
    relevance: medium
  - claim_id: encoding_token_duration_implicitly_in_the_vq_codebook_index_eliminates
    role: supports
    claim: Encoding token duration implicitly in the VQ codebook index eliminates the need for auxiliary duration
      predictors and preserves compatibility with autoregressive speech language models.
    source: §3.4, §4.4, Table 3
    evidence: The implicit duration coding scheme maps each cluster's content index k and duration d to a single
      token ID D = (d-1)*K + k, enabling a standard cross-entropy autoregressive model to generate variable-rate
      token sequences without modification; MOS and WER improve over the fixed-rate baseline in zero-shot TTS.
    confidence: high
    relevance: high
  - claim_id: more_aggressive_temporal_compression_in_variable_rate_tokenisers_trades_reconstruction
    role: complicates
    claim: More aggressive temporal compression in variable-rate tokenisers trades reconstruction quality for token
      efficiency beyond a practical compression threshold.
    source: §4.2, Table 1
    evidence: Increasing S_max from 2 to 8 reduces the average frame rate from 46.5 Hz to 22.38 Hz but degrades
      UTMOS from 4.038 to 3.647 and PESQ from 2.069 to 1.453 on LibriTTS test-clean; the optimal configuration (tau=0.7,
      S_max=4 at 30.95 Hz) sits at the knee of this trade-off curve.
    confidence: high
    relevance: high
  - claim_id: inference_speed_in_autoregressive_speech_lms_depends_primarily_on_sequence
    role: refines
    claim: Inference speed in autoregressive speech LMs depends primarily on sequence length rather than vocabulary
      size, so variable-rate tokenisers with expanded vocabularies still accelerate decoding.
    source: §J, Table 5
    evidence: VARSTok (tau=0.6) achieves RTF 0.487 versus 0.766 for the 40 Hz WavTokenizer baseline (36% speedup)
      despite expanding the token vocabulary from K to K*S_max = 16,384 entries, because shorter sequences reduce
      the dominant cost of attention computation over more function evaluations.
    confidence: high
    relevance: medium
  limitations:
  - Evaluation is restricted to English (LibriTTS). The clustering algorithm relies on cosine similarity in a WavTokenizer
    embedding space trained on English read speech; whether the density-peak boundaries remain meaningful for other
    languages, spontaneous speech, or emotionally expressive styles is untested.
  - Speaker similarity under more aggressive compression (tau=0.6, 26.29 Hz) does show a statistically modest decline
    in objective SIM (0.880 vs 0.918 for the baseline), and while subjective SMOS remains comparable, the long-tail
    impact on voices far from the training distribution is unknown. Codebook collapse becomes severe for K above
    4096 in the expanded index space, suggesting that very large vocabulary configurations require dedicated regularisation
    strategies not addressed here. The clustering algorithm is not differentiable, so joint end-to-end training
    with a downstream TTS model is not straightforward.
  caveats: []
- id: '2509.05863'
  published_date: "2025-09-06"
  entry_date: '2026-07-28'
  year: 2025
  venue: arXiv
  task:
  - TTS
  architecture:
  - autoregressive-LM
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  - evaluation_caution
  current_role: active_evidence
  method_family:
  - autoregressive_codec_language_models
  claims:
  - claim_id: preference_based_alignment_dpo_improves_intelligibility_in_multilingual_autoregressive_tts
    role: supports
    claim: Preference-based alignment (DPO) improves intelligibility in multilingual autoregressive TTS systems
      beyond what supervised fine-tuning achieves.
    source: §4.3, §5.2, Table 2, Table 3
    evidence: LatinX (DPO) reduces WER across nearly all 30 cross-lingual language pairs compared to the supervised
      fine-tuned baseline, and outperforms XTTSv2 in most pairs; Romanian-source conditions show particularly large
      gains (e.g., ro-to-es at 0.45% WER).
    confidence: high
    relevance: medium
  - claim_id: automated_speaker_similarity_metrics_based_on_speaker_encoder_embeddings_do
    role: complicates
    claim: Automated speaker similarity metrics based on speaker encoder embeddings do not reliably reflect human
      perceptual judgments of voice identity in zero-shot TTS.
    source: §5.2, §6, Table 4, Table 5
    evidence: XTTSv2 achieves higher Sim-O scores than both LatinX models, yet human evaluators strongly prefer
      LatinX speaker similarity (SMOS 3.63/3.54 vs. 3.24); the paper explicitly flags this as a divergence between
      objective and subjective evaluation.
    confidence: high
    relevance: low
  - claim_id: dpo_alignment_in_tts_involves_a_trade_off_optimizing_for
    role: complicates
    claim: 'DPO alignment in TTS involves a trade-off: optimizing for intelligibility and objective similarity can
      reduce naturalness MOS and, in some language conditions, perceptual similarity relative to the fine-tuned
      baseline.'
    source: §5.2, §6, Table 5, Table 6
    evidence: LatinX (DPO) improves WER and Sim-E over the fine-tuned model but achieves lower average MOS (3.35
      vs. 3.41) and lower SMOS in several cross-lingual conditions; the paper attributes this partly to the codec
      introducing artifacts that cap perceptual quality.
    confidence: high
    relevance: medium
  - claim_id: lossy_neural_audio_codecs_set_a_perceptual_quality_ceiling_in
    role: complicates
    claim: Lossy neural audio codecs set a perceptual quality ceiling in codec-based TTS that preference alignment
      cannot overcome, because the model learns to replicate codec artifacts introduced during reference encoding.
    source: §6
    evidence: The paper notes that the VQ-VAE codec is lossy and the model learns to reproduce its artifacts, limiting
      the maximum perceptual quality achievable regardless of post-training alignment method.
    confidence: high
    relevance: high
  limitations:
  - The evaluation uses an internal test set of unseen speakers with no publicly named benchmark, and the human
    rating pool is predominantly English and Portuguese native speakers. Conclusions about multilingual naturalness
    and similarity, especially for French, Italian, and Romanian, should be treated with caution.
  - The DPO preference signal is constructed solely from WER and speaker similarity; no prosody, naturalness, or
    rhythm metric is incorporated, which likely explains the MOS regression relative to the fine-tuned baseline.
    The preference labeling is fully automated with no human verification of winner/loser assignments. The real-time
    factor of 4.85 makes the system unsuitable for real-time applications, and the authors note that non-autoregressive
    architectures are a necessary direction. The Romanian evaluation suffers from very small rater counts and predominantly
    non-native listeners, undermining the interpretation of the unusually high SMOS scores that exceed real audio.
  caveats: []
- id: '2509.09174'
  published_date: "2025-09-11"
  entry_date: '2026-07-28'
  year: 2025
  venue: arXiv
  task:
  - SCA
  architecture:
  - autoregressive-LM
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: active_evidence
  method_family:
  - autoregressive_codec_language_models
  claims:
  - claim_id: decoupling_semantic_training_objectives_from_acoustic_token_prediction_substantially_reduces
    role: supports
    claim: Decoupling semantic training objectives from acoustic token prediction substantially reduces knowledge
      degradation in speech-to-speech LLMs.
    source: §5.1, Table 4
    evidence: EchoX's Echo training, which generates speech targets from the model's own semantic hidden states,
      raises average QA accuracy from 24.3 (T2C without Echo) to 37.1 (EchoX-3B) on the same data, compared to 12.8
      for direct interleaved training.
    confidence: high
    relevance: medium
  - claim_id: unit_based_speech_token_compression_via_language_model_segmentation_improves
    role: supports
    claim: Unit-based speech token compression via language-model segmentation improves downstream accuracy and
      reduces sequence length without sacrificing audio quality.
    source: §5.3, Table 5, Figure 7
    evidence: Unit language achieves 4.57 length ratio vs. 9.31 for raw units while improving accuracy on all three
      QA benchmarks and maintaining comparable audio quality in spectral comparison.
    confidence: high
    relevance: high
  - claim_id: streaming_inference_in_speech_llms_can_be_achieved_with_minimal
    role: supports
    claim: Streaming inference in speech LLMs can be achieved with minimal accuracy degradation when the segmentation
      boundary is determined by semantic similarity rather than fixed length.
    source: §5.4, Table 6
    evidence: EchoX's cosine-similarity trigger reduces first-token latency from 138 to 27 tokens at 3B scale with
      less than 1.5 percentage points of accuracy drop on any benchmark.
    confidence: high
    relevance: low
  - claim_id: training_data_efficiency_in_speech_llms_may_depend_more_on
    role: complicates
    claim: Training data efficiency in speech LLMs may depend more on the training paradigm than on data volume.
    source: §4.2, Table 2
    evidence: EchoX achieves competitive performance on spoken QA against models trained on millions of hours using
      only approximately 6,200 hours, but this result holds specifically for factual QA and has not been tested
      on broader spoken dialogue tasks.
    confidence: high
    relevance: medium
  - claim_id: speech_naturalness_and_response_helpfulness_are_not_jointly_optimised_by
    role: complicates
    claim: Speech naturalness and response helpfulness are not jointly optimised by the same training signal in
      speech-to-speech LLMs.
    source: §Appendix C, Figure 8
    evidence: Human evaluation shows EchoX wins clearly on helpfulness but performs only competitively on naturalness,
      reflecting a training objective focused on semantic correctness rather than prosodic quality.
    confidence: high
    relevance: medium
  limitations:
  - Speech quality is assessed only via brief spectral comparison (Figure 7) and a 5-rater human study. No perceptual
    quality metric (MOS, DNSMOS) or automatic speech recognition accuracy on generated audio is reported as a primary
    evaluation result, making it difficult to characterise the system's output quality independently of QA accuracy.
  - The evaluation benchmarks are limited to factual knowledge QA (Llama Questions, Web Questions, TriviaQA). It
    is unclear whether Echo training retains its advantage on open-ended dialogue, instruction following, or longer-form
    conversational tasks. The human evaluation was conducted with only five raters on one dataset (AlpacaEval),
    which limits statistical confidence in the naturalness comparison. The model trains on synthesised assistant
    audio from GPT-SoVITS, which may introduce a fixed timbre bias and limit voice diversity. The streaming threshold
    and window size are fixed hyperparameters with no ablation reported on their sensitivity.
  caveats: []
- id: '2509.09201'
  published_date: "2025-09-11"
  entry_date: '2026-07-28'
  year: 2025
  venue: arXiv
  task:
  - codec
  - VC
  architecture:
  - GAN
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: active_evidence
  method_family:
  - gan_neural_codecs_and_decoders
  claims:
  - claim_id: neural_audio_codecs_can_disentangle_speech_and_background_sound_in
    role: supports
    claim: Neural audio codecs can disentangle speech and background sound in the representation domain, enabling
      downstream tasks to selectively access either component without explicit front-end separation.
    source: §4.3, §4.5, §6.1, Table 4, Table 6
    evidence: DeCodec's SOP+RST design achieves orthogonal speech and background subspaces; recombining representations
      enables one-shot VC (SPK-SIM 0.83, WER 50.46) and speech enhancement (DNSMOS OVL 3.39) without denoising pre-processing,
      outperforming cascaded StoRM+SpeechTokenizer on ASR (WER* 26.7 vs. 34.5).
    confidence: high
    relevance: high
  - claim_id: semantic_guidance_from_self_supervised_features_in_the_first_quantiser
    role: supports
    claim: Semantic guidance from self-supervised features in the first quantiser layer improves the noise-robustness
      of codec representations used for downstream ASR.
    source: §6.1.4, Table 5
    evidence: Ablation-3 (SOP+RST without SG) achieves WER* 41.9 on noisy speech; adding HuBERT-L9 guidance reduces
      it to 25.8 (causal) and 23.6 (non-causal), confirming that semantic anchoring helps the quantiser concentrate
      linguistic content in a background-invariant first layer.
    confidence: high
    relevance: high
  - claim_id: explicit_disentanglement_constraints_in_neural_codecs_introduce_a_reconstruction_quality
    role: complicates
    claim: Explicit disentanglement constraints in neural codecs introduce a reconstruction quality trade-off relative
      to purely reconstruction-optimised designs.
    source: §6.1.1, Table 2
    evidence: DeCodec's mel distance on clean speech (0.89) is worse than DAC (0.65) and HiFi-Codec (0.75), which
      use no orthogonality or disentanglement losses, suggesting that forcing orthogonal subspace separation increases
      spectral distortion.
    confidence: high
    relevance: medium
  - claim_id: hierarchical_codec_disentanglement_enables_controllable_background_sound_handling_in_zero
    role: supports
    claim: Hierarchical codec disentanglement enables controllable background sound handling in zero-shot TTS without
      retraining the downstream generation model.
    source: §6.2.2, Table 7
    evidence: VALL-E trained on DeCodec tokens achieves MOS 4.09 with background preserved (BPMOS 4.19) vs. MOS
      3.96 without, controlled solely by whether the BRVQ-1:8 background tokens are included at inference; the TTS
      model itself was trained only on clean speech.
    confidence: high
    relevance: high
  - claim_id: voice_conversion_on_noisy_speech_via_representation_recombination_produces_high
    role: complicates
    claim: Voice conversion on noisy speech via representation recombination produces high WER even when speaker
      similarity is well-preserved, due to voiced/unvoiced segment boundary mismatches between source and reference
      utterances.
    source: §6.1.3, Table 4
    evidence: DeCodec one-shot VC achieves SPK-SIM 0.83 (matching the reference ceiling of 0.69 being clearly exceeded)
      but WER 50.46, which the authors attribute to structural mismatch in speech segment timing rather than semantic
      content corruption.
    confidence: high
    relevance: high
  limitations:
  - Subjective evaluations use panels of 10-12 volunteers on 300-clip test sets, which is small for MOS-based conclusions.
    The noisy speech test set is constructed by synthetically mixing clean LibriSpeech with DNS-Noise at controlled
    SNRs (-5 to 20 dB), which may not reflect the diversity of real-world background conditions.
  - The model operates at 16 kHz and is trained on a speech-dominant corpus; performance on music or non-speech
    audio types is not evaluated despite the universal codec framing. The high VC WER (50.46) under noisy conditions
    indicates that the current representation recombination approach for voice conversion degrades intelligibility
    significantly and would likely need additional design work before practical deployment. Model size is not reported,
    making direct efficiency comparisons with baselines difficult.
  caveats: []
- id: '2509.09550'
  published_date: "2025-09-11"
  entry_date: '2026-07-28'
  year: 2025
  venue: arXiv
  task:
  - codec
  architecture:
  - hybrid
  - GAN
  relevance: high
  evidence_role:
  - core_evidence
  current_role: active_evidence
  method_family:
  - hybrid_semantic_acoustic_tokenizers
  - gan_neural_codecs_and_decoders
  claims:
  - claim_id: fsq_based_neural_audio_codecs_develop_inherent_redundancy_in_their
    role: supports
    claim: FSQ-based neural audio codecs develop inherent redundancy in their code representations, enabling multiple
      independently trained encoders to produce radically different code sequences that nonetheless decode to perceptually
      equivalent audio.
    source: §4.2, Table 2
    evidence: Encoder distillation experiment on NeuCodec showing only 2% element-wise code match between original
      (635M params) and distilled (42M params) encoders, while cosine similarity between pre-quantization projections
      is 0.73 and reconstruction metrics (WER, CER, STOI, PESQ) remain within 0.5% WER of each other.
    confidence: high
    relevance: medium
  - claim_id: fsq_codecs_are_substantially_more_robust_to_bit_level_transmission
    role: supports
    claim: FSQ codecs are substantially more robust to bit-level transmission errors than RVQ codecs of comparable
      codebook size, maintaining intelligibility at bit-flip rates an order of magnitude higher than the RVQ failure
      threshold.
    source: §5, Figure 3
    evidence: Binary symmetric channel simulation on LibriSpeech test-clean showing FSQ codecs (NeuCodec, Distill-NeuCodec,
      StableCodec) maintain stable STOI and PESQ up to 10% bit-flip probability, whereas RVQ codecs (EnCodec, DAC)
      experience sharp quality collapse above 1% bit-flip rate.
    confidence: high
    relevance: high
  - claim_id: the_perturbation_robustness_advantage_of_fsq_over_rvq_is_structural
    role: refines
    claim: 'The perturbation robustness advantage of FSQ over RVQ is structural: FSQ''s fixed-grid quantization
      maps bit-flip perturbations to bounded, predictable steps in the embedding space, whereas RVQ code perturbations
      produce arbitrarily large embedding-space displacements.'
    source: §4.2, §6
    evidence: Implicit codebook confusion matrices show that 93% of level predictions between original and distilled
      encoders are either correct or off by exactly one neighbouring level, confirming local structure in the FSQ
      embedding space.
    confidence: high
    relevance: high
  - claim_id: fsq_based_codec_architectures_that_achieve_competitive_intelligibility_metrics_often
    role: complicates
    claim: FSQ-based codec architectures that achieve competitive intelligibility metrics often rely on very large
      pretrained semantic encoders, concentrating most of the parameter budget in a component that contributes indirectly
      to the quantization benefit.
    source: §3, §4.1, Table 2
    evidence: Wav2Vec2-BERT-large accounts for 600M of NeuCodec's 635M total parameters; replacing it with DistilHuBERT
      reduces the model to 42M parameters with only a 0.5% WER increase, suggesting the semantic representation
      is partially substitutable without sacrificing the FSQ robustness property.
    confidence: high
    relevance: high
  limitations:
  - Evaluation uses WER, CER, STOI, and PESQ but no perceptual naturalness metric (MOS or MUSHRA), making it impossible
    to assess how NeuCodec compares to state-of-the-art codecs on speech generation quality. The perturbation experiment
    uses a binary symmetric channel that does not reflect real-world transmission protocols (e.g., structured burst
    errors or packet loss), so robustness claims cannot be extrapolated directly to deployment scenarios.
  - Training includes 1,000 hours of proprietary data, partially limiting reproducibility. The model is evaluated
    only at 16kHz and 24kHz; the 24kHz upsampling decoder is trained on only 2,600 hours of data, and high-fidelity
    (48kHz) audio is not addressed. The paper does not evaluate whether FSQ robustness properties are preserved
    when the codec is used as a tokenizer for downstream TTS or language model tasks, which is the stated application
    motivation.
  caveats: []
- id: '2509.09631'
  published_date: "2025-09-11"
  entry_date: '2026-07-28'
  year: 2025
  venue: arXiv
  task:
  - TTS
  architecture:
  - flow-matching
  - transformer-enc-dec
  relevance: medium
  evidence_role:
  - architecture_variant
  - acceleration_evidence
  - control_evidence
  current_role: active_evidence
  method_family:
  - flow_matching_codec_decoders
  - transformer_encoder_decoder_tokenizers
  claims:
  - claim_id: discrete_flow_matching_defined_directly_over_factorized_speech_token_subspaces
    role: supports
    claim: Discrete flow matching defined directly over factorized speech token subspaces achieves competitive naturalness
      and superior prosody reconstruction compared to continuous-space flow and diffusion TTS baselines trained
      on comparable data.
    source: §4.2, Table 1, Table 2
    evidence: Discrete flow matching defined directly over factorized speech token subspaces achieves competitive
      naturalness and superior prosody reconstruction compared to continuous-space flow and diffusion TTS baselines
      trained on comparable data.
    confidence: high
    relevance: low
  - claim_id: factorized_multi_head_velocity_prediction_separate_prediction_heads_for_distinct
    role: supports
    claim: Factorized multi-head velocity prediction — separate prediction heads for distinct speech attribute subspaces
      within a single discrete flow model — improves prosody and acoustic fidelity over a single-head alternative.
    source: §4.3, Table 4
    evidence: Factorized multi-head velocity prediction — separate prediction heads for distinct speech attribute
      subspaces within a single discrete flow model — improves prosody and acoustic fidelity over a single-head
      alternative.
    confidence: high
    relevance: low
  - claim_id: global_speaker_embedding_conditioning_via_adaln_is_insufficient_for_reliable
    role: supports
    claim: Global speaker embedding conditioning via AdaLN is insufficient for reliable zero-shot speaker similarity
      in discrete token-space TTS; perceptual similarity judgements diverge from embedding-based automatic metrics
      (SIM-O) in this setting.
    source: §4.2, Table 1, Table 2
    evidence: Global speaker embedding conditioning via AdaLN is insufficient for reliable zero-shot speaker similarity
      in discrete token-space TTS; perceptual similarity judgements diverge from embedding-based automatic metrics
      (SIM-O) in this setting.
    confidence: high
    relevance: low
  - claim_id: non_autoregressive_discrete_flow_models_with_compact_dit_backbones_can
    role: supports
    claim: Non-autoregressive discrete flow models with compact DiT backbones can match the inference latency of
      single-step flow matching systems while operating at higher NFE, without compromising data efficiency.
    source: §4.2, Table 3
    evidence: Non-autoregressive discrete flow models with compact DiT backbones can match the inference latency
      of single-step flow matching systems while operating at higher NFE, without compromising data efficiency.
    confidence: high
    relevance: low
  limitations:
  - 'The model''s weakest dimension is automatic speaker similarity (SIM-O): simple global AdaLN speaker conditioning
    is insufficient for reliable timbre reproduction, especially in zero-shot settings. The authors acknowledge
    this and suggest cross-attention over local timbre embeddings as future work.'
  - Because FACodec separates speaker identity via a timbre embedding external to the discrete token streams, integrating
    speaker conditioning into a DFM framework is non-trivial; alternative codecs such as EnCodec implicitly embed
    speaker information within VQ codebooks, potentially enabling stronger speaker adaptation in future discrete-space
    models.
  - The evaluation uses a single language (English) and a moderately sized training set (470h), leaving open questions
    about multilingual scalability and behavior at larger data scales. The gap between SIM-O ranking (third among
    five) and perceptual similarity MOS ranking (best) also points to known deficiencies in embedding-based automatic
    speaker similarity metrics, and a better evaluation framework for zero-shot speaker identity remains an open
    research direction.
  caveats: []
- id: '2509.11425'
  published_date: "2025-09-14"
  entry_date: '2026-07-28'
  year: 2025
  venue: arXiv
  task:
  - codec
  - TTS
  architecture:
  - GAN
  - autoregressive-LM
  relevance: high
  evidence_role:
  - core_evidence
  - control_evidence
  current_role: active_evidence
  method_family:
  - gan_neural_codecs_and_decoders
  - autoregressive_codec_language_models
  claims:
  - claim_id: injecting_multimodal_guidance_directly_into_the_encoder_latent_space_of
    role: supports
    claim: Injecting multimodal guidance directly into the encoder latent space of a neural codec improves reconstruction
      quality beyond similarity-based supervision of the quantized layer.
    source: §2.3.1, §3.2.1, Table 2
    evidence: FuseCodec-Fusion, which fuses global semantic and contextual vectors into Z via additive fusion, achieves
      WER 3.99, ViSQOL 3.47, and PESQ 3.13 on LibriSpeech test-clean, outperforming FuseCodec-Distill (ViSQOL 3.43,
      PESQ 3.06) and FuseCodec-ContextAlign, both of which only supervise Q(1) without modifying the latent.
    confidence: high
    relevance: high
  - claim_id: neural_codecs_trained_with_joint_semantic_and_contextual_supervision_generalize
    role: supports
    claim: Neural codecs trained with joint semantic and contextual supervision generalize to unseen languages without
      multilingual training data.
    source: §3.3, Table 5
    evidence: FuseCodec, trained exclusively on English LibriSpeech train-clean-100, achieves the best WER and ViSQOL
      in the majority of 7 tested languages from Multilingual LibriSpeech and outperforms all baselines on PESQ
      by at least 0.3 in most languages.
    confidence: high
    relevance: medium
  - claim_id: codec_representations_enriched_with_semantic_and_contextual_signals_support_stronger
    role: supports
    claim: Codec representations enriched with semantic and contextual signals support stronger downstream task
      generalization (emotion recognition, audio event classification) than acoustic-only codecs.
    source: §3.2.2, Table 3
    evidence: On CodecSUPERB at 4 kbps, FuseCodec-Fusion achieves emotion recognition accuracy of 73.96% and audio
      signal quality 0.785, versus 66.18% and 0.697 for EnCodec at 6 kbps. All FuseCodec variants exceed SpeechTokenizer,
      EnCodec, and DAC on audio signal quality.
    confidence: high
    relevance: high
  - claim_id: fine_grained_temporal_alignment_between_text_tokens_and_acoustic_frames
    role: complicates
    claim: Fine-grained temporal alignment between text tokens and acoustic frames improves local interpretability
      but is constrained relative to global supervision strategies.
    source: §2.3.3, §3.2.1, Table 2
    evidence: FuseCodec-ContextAlign, which aligns contextual embeddings to RVQ tokens via a windowed similarity
      matching algorithm, achieves WER 4.15 and ViSQOL 3.18, lagging FuseCodec-Fusion (WER 3.99, ViSQOL 3.47) and
      FuseCodec-Distill (ViSQOL 3.43). The paper attributes this to constrained local alignment limiting global
      contextual guidance.
    confidence: high
    relevance: medium
  - claim_id: distilling_both_semantic_self_supervised_speech_model_and_contextual_language
    role: supports
    claim: Distilling both semantic (self-supervised speech model) and contextual (language model) signals into
      codec token supervision outperforms semantic-only distillation for perceptual naturalness.
    source: §3.2.1, Table 2
    evidence: FuseCodec-Distill achieves UTMOS 3.65 and Similarity 0.996 on LibriSpeech test-clean, while codecs
      using only semantic distillation (SpeechTokenizer, Mimi, X-Codec2) score below 3.55 UTMOS and fail to consistently
      match speaker similarity.
    confidence: high
    relevance: high
  limitations:
  - The codec is trained on LibriSpeech train-clean-100 (100 hours, English read speech), which limits conclusions
    about robustness to spontaneous speech, diverse accents, or noisy conditions. Multilingual generalization results
    are promising but the training data and BERT model are English-only, leaving the mechanism behind cross-lingual
    transfer unclear. The ASR transcription step (wav2vec 2.0) introduces an error-prone intermediate representation
    during training; the effect of ASR errors on contextual embedding quality is not quantified. Model size is not
    reported, making it difficult to assess computational cost relative to baselines. The TTS evaluation compares
    only against other codec-based systems trained on LibriTTS and does not benchmark against the strongest current
    TTS models.
  caveats: []
- id: '2509.13068'
  published_date: "2025-09-16"
  entry_date: '2026-07-28'
  year: 2025
  venue: arXiv
  task:
  - TTS
  - codec
  - VC
  architecture:
  - autoregressive-LM
  - VAE
  relevance: high
  evidence_role:
  - core_evidence
  - acceleration_evidence
  - control_evidence
  current_role: active_evidence
  method_family:
  - autoregressive_codec_language_models
  - vae_vector_quantized_codecs
  claims:
  - claim_id: cascaded_residual_codec_architectures_can_enforce_attribute_disentanglement_through_structure
    role: supports
    claim: Cascaded residual codec architectures can enforce attribute disentanglement through structure rather
      than through adversarial training objectives.
    source: §2.1, §3.3.3, Table 3
    evidence: MSR-Codec achieves clean separation of timbre, prosody, and semantic content by having each stream
      operate on residuals from the previous stage, without adversarial disentanglement loss; VC experiments confirm
      independent manipulation of each attribute.
    confidence: high
    relevance: high
  - claim_id: explicit_prosodic_supervision_in_a_dedicated_codec_stream_promotes_measurable
    role: supports
    claim: Explicit prosodic supervision in a dedicated codec stream promotes measurable disentanglement of pitch
      from speaker identity.
    source: §2.1.2, §3.3.3, Table 3
    evidence: VQ1 (prosody stream) is trained with MSE loss against ground-truth F0 and spectral energy; prosody-only
      VC achieves low ΔF0,tar (12.3-14.2 Hz) while maintaining high SIM-src (0.59-0.64), confirming that prosody
      and timbre are independently manipulable.
    confidence: high
    relevance: high
  - claim_id: disentangled_codec_designs_can_achieve_competitive_speaker_similarity_at_lower
    role: supports
    claim: Disentangled codec designs can achieve competitive speaker similarity at lower bitrates than undifferentiated
      RVQ codecs.
    source: §3.3.1, Table 1
    evidence: MSR-Codec-424 achieves SPK-SIM 0.80 at 424 bps, higher than WavTokenizer (0.67 at 900 bps) and X-Codec
      (0.72 at 1000 bps), attributed to the time-invariant timbre stream which preserves speaker identity without
      scaling with utterance length.
    confidence: high
    relevance: high
  - claim_id: data_efficient_tts_systems_built_on_factorized_codec_representations_can
    role: supports
    claim: Data-efficient TTS systems built on factorized codec representations can achieve competitive intelligibility
      relative to larger models trained on more data.
    source: §3.3.2, Table 2
    evidence: The 0.2B MSR-Codec-524 TTS model trained on 45k hours achieves WER 3.07% on Seed-TTS-eval English,
      outperforming Llasa-1B trained on 250k hours (WER 3.22%) and FireRedTTS-0.4B trained on 150k hours (WER 3.82%).
    confidence: high
    relevance: high
  - claim_id: signal_fidelity_codec_metrics_stoi_pesq_and_speaker_similarity_diverge
    role: complicates
    claim: Signal-fidelity codec metrics (STOI, PESQ) and speaker similarity diverge at low bitrates, making holistic
      quality assessment difficult.
    source: §3.3.1, Table 1
    evidence: MSR-Codec-424 achieves the highest SPK-SIM (0.80) among codecs at comparable bitrates but lower STOI
      (0.84) and PESQ-WB (1.82) than some baselines, indicating that speaker preservation and signal-level fidelity
      are optimized differently by the multi-stream design.
    confidence: high
    relevance: high
  limitations:
  - No subjective listening test (MOS/MUSHRA) is reported for any condition; all quality comparisons rely on automatic
    metrics (UTMOS, STOI, PESQ, WER, SPK-SIM). Conclusions about perceived naturalness cannot be confirmed from
    the available data.
  - 'The VC evaluation protocol is small in scope: 8 target speakers from VCTK and 100 source utterances from LibriTTS.
    Generalisation to more diverse speakers, accents, or noisy conditions is not assessed. The FreGAN vocoder operates
    at 16 kHz, and the Mel-spectrogram-based pipeline may impose a quality ceiling relative to waveform-domain codecs.
    The TTS model is evaluated only on English; the codec was trained on Mandarin and English data but multilingual
    TTS capability is not demonstrated. Model size figures for the codec itself are not reported; only the TTS model
    size (0.2B) is provided.'
  caveats: []
- id: '2412.16846'
  published_date: "2025-09-17"
  entry_date: '2026-07-28'
  year: 2025
  venue: arXiv
  task:
  - TTS
  architecture:
  - autoregressive-LM
  - VAE
  relevance: medium
  evidence_role:
  - architecture_variant
  - acceleration_evidence
  - control_evidence
  current_role: active_evidence
  method_family:
  - autoregressive_codec_language_models
  - vae_vector_quantized_codecs
  claims:
  - claim_id: distributional_training_objectives_for_continuous_ar_speech_modeling_achieve_higher
    role: supports
    claim: Distributional training objectives for continuous AR speech modeling achieve higher intelligibility than
      regression-based alternatives.
    source: §TTS Evaluation, Table 2; §Ablation Study, Table 5
    evidence: KALL-E with KL divergence loss achieves WER 1.94 / CER 0.96 on Seed-TTS test sets, below all discrete-token
      and regression-based baselines; ablation replacing Flow-VAE with Stable Audio VAE (near-zero KL weight, approaching
      a plain autoencoder) collapses CER from 2.79 to 40.09 at the same latent dimension.
    confidence: high
    relevance: medium
  - claim_id: low_frame_rate_continuous_representations_reduce_autoregressive_tts_inference_compute
    role: supports
    claim: Low frame-rate continuous representations reduce autoregressive TTS inference compute by over an order
      of magnitude without sacrificing synthesis quality.
    source: §TTS Evaluation, Table 3, Table 4
    evidence: KALL-E at 12.5 Hz requires 7,947 GFLOPs to synthesize 10 seconds vs. 122,170 for Llasa-1B at 50 Hz,
      while achieving higher MOS (4.17 vs. 3.92) and lower WER (1.94 vs. 3.6) on the same test set.
    confidence: high
    relevance: medium
  - claim_id: objective_speaker_similarity_metrics_for_zero_shot_tts_are_unreliable
    role: complicates
    claim: Objective speaker similarity metrics for zero-shot TTS are unreliable for cross-system comparisons when
      decoder architectures differ in their use of reference audio.
    source: §TTS Evaluation, Table 2, Table 3
    evidence: Discrete-token systems (Seed-TTS SIM 0.796, FireRedTTS SIM 0.635) score differently on objective SPK-SIM
      than KALL-E (SIM 0.646/0.568), but KALL-E receives higher listener naturalness ratings; the authors attribute
      the gap to those systems conditioning the waveform decoder on the reference utterance at decode time, which
      inflates the metric independent of perceived speaker fidelity.
    confidence: high
    relevance: low
  - claim_id: test_time_adaptation_from_a_single_reference_utterance_improves_speaker
    role: supports
    claim: Test-time adaptation from a single reference utterance improves speaker similarity in continuous-representation
      AR TTS without requiring full model retraining.
    source: §Test Time Training; §TTS Evaluation, Table 2
    evidence: KALL-E (TTT) improves SPK-SIM from 0.568 to 0.611 on test-en using N=200 latent sequences sampled
      from the reference utterance's Flow-VAE distribution, with WER remaining stable at 1.90.
    confidence: high
    relevance: high
  - claim_id: increasing_vae_kl_regularization_weight_trades_reconstruction_fidelity_for_a
    role: complicates
    claim: Increasing VAE KL regularization weight trades reconstruction fidelity for a latent space structure that
      is more suitable for downstream generative modeling.
    source: §VAE Evaluation, Table 1; §Ablation Study, Table 5
    evidence: Flow-VAE uses KL weight 32 and scores PESQ-WB 3.26 at 512-dim/12.5 Hz, below Stable Audio VAE (3.11)
      at the same frame rate with near-zero KL weight; however, Stable Audio VAE's latent space causes CER to collapse
      when used as the AR LM encoder, demonstrating that reconstruction quality and generation compatibility impose
      conflicting constraints on VAE training.
    confidence: high
    relevance: medium
  limitations:
  - 'Objective speaker similarity remains below discrete-token systems that condition their decoders on the reference
    audio, suggesting the Flow-VAE''s information bottleneck trades some speaker detail for a more LM-friendly latent
    space. The TTT procedure assumes the transcript of the reference utterance is available, which may not hold
    in all deployment settings. Overfitting risk in TTT is real: CER rises after N=200 in ablation, limiting the
    effective adaptation set size. Evaluations are conducted solely on the Seed-TTS test sets; generalization to
    other benchmarks, out-of-distribution speakers, or noisy acoustic conditions is not assessed. Training data
    composition differs from the most directly comparable system (Llasa-1B), making it difficult to fully isolate
    architecture from data quality as the source of WER gains.'
  caveats: []
- id: '2509.13670'
  published_date: "2025-09-17"
  entry_date: '2026-07-28'
  year: 2025
  venue: arXiv
  task:
  - codec
  architecture:
  - GAN
  relevance: high
  evidence_role:
  - core_evidence
  - acceleration_evidence
  current_role: active_evidence
  method_family:
  - gan_neural_codecs_and_decoders
  claims:
  - claim_id: knowledge_distillation_from_a_non_causal_high_complexity_teacher_can
    role: supports
    claim: Knowledge distillation from a non-causal, high-complexity teacher can effectively recover reconstruction
      quality degraded by model causalization and channel pruning in low-latency streamable neural codecs.
    source: §IV.A, Table I
    evidence: StreamCodec2 (NH→CL direct KD) achieves PESQ 2.744 and ViSQOL 4.313 vs. 2.650 and 4.290 for the undistilled
      student at identical 910 MFLOPs and 5.4 M parameters, with all gains p < 0.01.
    confidence: high
    relevance: low
  - claim_id: multi_stage_distillation_pipelines_do_not_provide_additive_quality_gains
    role: complicates
    claim: Multi-stage distillation pipelines do not provide additive quality gains over direct teacher-to-student
      distillation in neural codec compression.
    source: §IV.A, Table I
    evidence: Both indirect distillation schemes (NH→CH→CL and NH→NL→CL) underperform direct distillation across
      all objective metrics, with ViSQOL 4.294 and 4.305 vs. 4.313 for direct, suggesting that intermediate steps
      dilute rather than refine the knowledge transferred.
    confidence: high
    relevance: high
  - claim_id: fully_causal_neural_codec_architectures_incur_a_meaningful_quality_penalty
    role: complicates
    claim: Fully causal neural codec architectures incur a meaningful quality penalty relative to non-causal counterparts
      at the same bitrate, which knowledge distillation only partially closes.
    source: §IV.A, Table I
    evidence: Even with the best distillation strategy, StreamCodec2 (NH→CL) scores PESQ 2.744 vs. the teacher's
      3.132 and ViSQOL 4.313 vs. 4.463, leaving a gap that reflects the fundamental constraint of causal-only processing.
    confidence: high
    relevance: high
  - claim_id: distillation_loss_weighting_in_codec_training_requires_careful_calibration_excessively
    role: refines
    claim: Distillation loss weighting in codec training requires careful calibration; excessively large weights
      degrade reconstruction quality by shifting the learning objective toward teacher imitation.
    source: §IV.B, Figure 3
    evidence: ViSQOL for StreamCodec2 (NH→CL) peaks at lambda_KD = 0.01 and declines for weights above this value
      (tested at 0.002, 0.005, 0.01, 0.02, 0.05), consistent with the reconstruction objective being displaced by
      over-fitting the teacher's intermediate representations.
    confidence: high
    relevance: high
  limitations:
  - The evaluation uses only objective metrics (LSD, STOI, PESQ, ViSQOL); it is unclear whether the statistically
    significant gains over the undistilled student are perceptually meaningful. Comparisons are limited to the authors'
    own student and teacher variants with no benchmarking against published competing streamable codecs (SoundStream,
    EnCodec, or others) under matched latency and bitrate conditions, making it difficult to assess where StreamCodec2
    stands in the broader landscape. Experiments use a single dataset (LibriTTS at 16 kHz), leaving generalisation
    to other languages, domains, or sampling rates unconfirmed. Future directions noted by the authors include improving
    reconstruction quality further and evaluating on additional audio datasets.
  caveats: []
- id: '2509.14882'
  published_date: "2025-09-18"
  entry_date: '2026-07-28'
  year: 2025
  venue: arXiv
  task:
  - SCA
  architecture:
  - autoregressive-LM
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: active_evidence
  method_family:
  - autoregressive_codec_language_models
  claims:
  - claim_id: a_single_transformer_flattened_architecture_can_match_or_surpass_hierarchical
    role: supports
    claim: A single-Transformer flattened architecture can match or surpass hierarchical speech LM designs on acoustic
      consistency tasks when both use identical data and comparable parameter budgets.
    source: §4.3, Table 1
    evidence: Llama-Mimi-1.3B outperforms CSM-1.3B on all SALMon acoustic consistency dimensions and speaker similarity
      (0.346 vs. 0.320) under controlled training conditions.
    confidence: high
    relevance: medium
  - claim_id: flattening_rvq_tokens_into_a_single_autoregressive_sequence_creates_an
    role: complicates
    claim: 'Flattening RVQ tokens into a single autoregressive sequence creates an inherent acoustic-linguistic
      trade-off: strong acoustic performance comes at the cost of weaker linguistic benchmarks relative to SSL-based
      phonetic-token approaches.'
    source: §4.3, Table 1
    evidence: Llama-Mimi-1.3B achieves best acoustic consistency but underperforms TWIST-1.3B on sWUGGY (68.7 vs.
      71.7) and T-Story Cloze (64.0 vs. 69.9), attributed to the Q-fold sequence length increase from RVQ token
      flattening.
    confidence: high
    relevance: high
  - claim_id: increasing_the_number_of_rvq_quantizers_in_a_flattened_speech
    role: complicates
    claim: Increasing the number of RVQ quantizers in a flattened speech LM improves audio quality but degrades
      spoken content coherence, because longer token sequences shift modeling capacity toward acoustic reconstruction.
    source: §4.4, Table 5
    evidence: Ablation with Q∈{2,4,8} shows Q=8 achieves best Audiobox-Aesthetics scores and speaker similarity
      (0.474) but worst content quality (2.54), while Q=2 yields content quality (3.53) comparable to TWIST-1.3B.
    confidence: high
    relevance: high
  - claim_id: applying_a_higher_loss_weight_to_semantic_tokens_in_speech
    role: refines
    claim: Applying a higher loss weight to semantic tokens in speech LM training shifts the acoustic-linguistic
      balance toward linguistic accuracy, but causes measurable degradation in acoustic consistency and speaker
      similarity.
    source: §4.4, Table 3
    evidence: Llama-Mimi-1.3B with semantic weight λ=100 gains on sBLIMP (55.4 vs. 54.3) and T-Story Cloze (68.4
      vs. 64.0), but loses substantially on room consistency (74 vs. 92) and speaker similarity (0.196 vs. 0.346).
    confidence: high
    relevance: low
  - claim_id: scaling_model_size_in_flattened_speech_lms_consistently_improves_performance
    role: supports
    claim: Scaling model size in flattened speech LMs consistently improves performance across both acoustic and
      linguistic tasks, with the largest gains in spoken content quality.
    source: §4.4, Table 4
    evidence: Llama-Mimi-8B improves content quality (4.03 vs. 3.01) and T-Story Cloze (67.6 vs. 64.0) over the
      1.3B model, with qualitative analysis showing more semantically coherent long-form continuations.
    confidence: high
    relevance: medium
  limitations:
  - Evaluations are conducted exclusively on English speech using LibriSpeech prompts; generalisation to other languages,
    speakers, or acoustic conditions is untested. No human listening tests are reported; all acoustic and linguistic
    evaluations rely on automated metrics.
  - The paper evaluates only speech continuation, not text-conditioned generation or dialogue. Whether the acoustic-linguistic
    trade-off in flattened designs persists when text tokens are added to the sequence (as in Moshi's inner monologue
    approach) is left as an open question. The 8B model is evaluated only on a subset of tasks, and the cost of
    the Q-fold sequence length increase at large scale is not fully characterised. Training Mimi weights end-to-end
    jointly with the LM is not explored; the frozen codec is an architectural constraint, not a systematic choice.
  caveats: []
- id: '2509.15462'
  published_date: "2025-09-18"
  entry_date: '2026-07-28'
  year: 2025
  venue: arXiv
  task:
  - codec
  - VC
  architecture:
  - autoregressive-LM
  - flow-matching
  relevance: high
  evidence_role:
  - core_evidence
  - acceleration_evidence
  - control_evidence
  current_role: active_evidence
  method_family:
  - autoregressive_codec_language_models
  - flow_matching_codec_decoders
  claims:
  - claim_id: factorising_speech_into_semantic_components_content_style_timbre_enables_task
    role: supports
    claim: Factorising speech into semantic components (content, style, timbre) enables task-relevant transmission
      at lower bitrates than general-purpose codecs that encode all features uniformly.
    source: §3.3, Table 1
    evidence: Content-style tokens at 650 bps achieve WER 0.15 and sentiment accuracy 59%, within 1% of EnCodec
      at 1.5 kbps (WER 0.20, accuracy 60%), while using approximately half the bitrate.
    confidence: high
    relevance: medium
  - claim_id: reusing_a_compressed_speaker_reference_transmitted_once_per_speaker_can
    role: supports
    claim: Reusing a compressed speaker reference transmitted once per speaker can maintain speaker similarity at
      lower average bitrate than encoding full audio continuously.
    source: §3.4, Table 1
    evidence: Vevo with Zonos speaker embedding (1 sec timbre) achieves SpkrSim 0.54, matching EnCodec at 1.5 kbps
      (0.54), while long-run average bitrate approaches 650 bps as registered speakers accumulate.
    confidence: high
    relevance: high
  - claim_id: generative_reconstruction_from_semantic_tokens_improves_perceptual_quality_scores_but
    role: complicates
    claim: Generative reconstruction from semantic tokens improves perceptual quality scores but degrades low-level
      signal fidelity metrics relative to waveform-level codecs.
    source: §3.4, Table 1
    evidence: Vevo configurations outperform EnCodec on UTMOS and NISQA across all bitrates, but score lower on
      PESQ and STOI because the flow-matching decoder was not trained to preserve signal-level characteristics.
    confidence: high
    relevance: medium
  - claim_id: semantic_codec_approaches_face_a_latency_and_error_propagation_trade
    role: complicates
    claim: Semantic codec approaches face a latency and error-propagation trade-off that general-purpose codecs
      avoid.
    source: §2.2, §4
    evidence: Timbre transmission introduces per-speaker latency L = d_sample + d_transmit; errors in that one-time
      transmission cause permanent voice reconstruction inaccuracies until a correction is transmitted.
    confidence: high
    relevance: high
  - claim_id: neural_semantic_token_representations_are_moderately_robust_to_channel_bit
    role: supports
    claim: Neural semantic token representations are moderately robust to channel bit errors, maintaining acceptable
      downstream task performance at realistic noise levels.
    source: §3.5, Table 2
    evidence: At 0.1% bit-flip rate the system shows no measurable degradation; at 1% BER sentiment classification
      (0.63) and speaker verification (0.77) still exceed Opus at 5 kbps. Performance collapses only above 10% BER.
    confidence: high
    relevance: medium
  limitations:
  - 'Timbre transmission errors are permanent: a corrupted one-time speaker embedding results in incorrect voice
    reconstruction for all subsequent utterances from that speaker until a retransmission is triggered. The paper
    does not propose an error-recovery mechanism.'
  - The system does not handle overlapping speakers; the paper acknowledges this would require sender-side speaker
    separation at additional bitrate cost. Evaluation is restricted to English and does not address non-English
    languages, which are explicitly noted as out-of-distribution for several downstream models. The test set of
    1,000 clips from VoxCeleb1 is relatively small for claims about general-purpose voice communication performance.
    The real-time latency introduced by speaker embedding generation and transmission is uncharacterised across
    different network conditions.
  caveats: []
- id: '2509.15969'
  published_date: "2025-09-19"
  entry_date: '2026-07-28'
  year: 2025
  venue: arXiv
  task:
  - TTS
  architecture:
  - autoregressive-LM
  relevance: medium
  evidence_role:
  - architecture_variant
  - acceleration_evidence
  - control_evidence
  current_role: active_evidence
  method_family:
  - autoregressive_codec_language_models
  claims:
  - claim_id: fully_autoregressive_streaming_tts_can_achieve_first_packet_latencies_under
    role: supports
    claim: Fully autoregressive streaming TTS can achieve first-packet latencies under 150 ms without sacrificing
      intelligibility relative to non-streaming operation.
    source: §4, Table 1, Table 3
    evidence: Fully autoregressive streaming TTS can achieve first-packet latencies under 150 ms without sacrificing
      intelligibility relative to non-streaming operation.
    confidence: high
    relevance: low
  - claim_id: training_data_scale_is_the_primary_driver_of_speaker_similarity
    role: supports
    claim: Training data scale is the primary driver of speaker similarity in zero-shot TTS, and systems trained
      on an order-of-magnitude less data show measurable SPK-SIM gaps even when naturalness scores are competitive.
    source: §4, Table 1, Table 2
    evidence: Training data scale is the primary driver of speaker similarity in zero-shot TTS, and systems trained
      on an order-of-magnitude less data show measurable SPK-SIM gaps even when naturalness scores are competitive.
    confidence: high
    relevance: low
  - claim_id: borrowing_frozen_depth_transformer_weights_from_a_large_scale_pretrained
    role: supports
    claim: Borrowing frozen depth transformer weights from a large-scale pretrained model provides substantial quality
      improvement for a mid-scale system without requiring additional large-scale training.
    source: §3, Table 4
    evidence: Borrowing frozen depth transformer weights from a large-scale pretrained model provides substantial
      quality improvement for a mid-scale system without requiring additional large-scale training.
    confidence: high
    relevance: medium
  - claim_id: full_stream_input_processing_introduces_only_marginal_quality_degradation_relative
    role: supports
    claim: Full-stream input processing introduces only marginal quality degradation relative to output-streaming
      when a bounded phoneme look-ahead is used, suggesting that input latency and output quality are largely decoupled
      in autoregressive codec TTS.
    source: §4, Table 1
    evidence: Full-stream input processing introduces only marginal quality degradation relative to output-streaming
      when a bounded phoneme look-ahead is used, suggesting that input latency and output quality are largely decoupled
      in autoregressive codec TTS.
    confidence: high
    relevance: high
  - claim_id: non_autoregressive_flow_matching_decoders_used_in_chunk_based_streaming
    role: supports
    claim: Non-autoregressive flow-matching decoders used in chunk-based streaming systems incur first-packet latencies
      exceeding 1.5 seconds on standard hardware, which is prohibitive for real-time spoken conversational agents.
    source: §4, Table 3
    evidence: Non-autoregressive flow-matching decoders used in chunk-based streaming systems incur first-packet
      latencies exceeding 1.5 seconds on standard hardware, which is prohibitive for real-time spoken conversational
      agents.
    confidence: high
    relevance: low
  limitations:
  - Speaker similarity remains lower than large-scale systems trained on hundreds of thousands of hours, particularly
    in full-stream mode (SPK-SIM 0.458 on LibriSpeech test-clean vs. 0.587 for CosyVoice2 trained on 167k hours).
    The system is English-only; multilingual extension is not addressed. Prosody and speaking rate are not explicitly
    controllable at inference. The use of a frozen CSM depth transformer introduces a dependency on an external
    large-scale model. Long-form streaming beyond 10-15 second utterances is identified as future work. Performance
    in adverse or spontaneous speech conditions (outside the training domain) is untested.
  caveats: []
- id: '2509.16195'
  published_date: "2025-09-19"
  entry_date: '2026-07-28'
  year: 2025
  venue: arXiv
  task:
  - codec
  - VC
  architecture:
  - hybrid
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  - evaluation_caution
  current_role: active_evidence
  method_family:
  - hybrid_semantic_acoustic_tokenizers
  claims:
  - claim_id: single_codebook_binary_quantization_at_sub_1_kbps_bitrates_can
    role: supports
    claim: Single-codebook binary quantization at sub-1 kbps bitrates can match or exceed multi-codebook streaming
      codecs on naturalness and intelligibility in speech resynthesis.
    source: §4.1, Table 2
    evidence: FocalCodec-S@50-65k achieves UTMOS 3.85 and dWER 3.68% at 0.80 kbps with a single codebook of 65,536
      entries, outperforming Mimi6 (0.83 kbps, 6 codebooks, UTMOS 3.44, dWER 4.77%) on both metrics.
    confidence: high
    relevance: high
  - claim_id: multi_stage_causal_distillation_of_self_supervised_speech_encoders_preserves
    role: supports
    claim: Multi-stage causal distillation of self-supervised speech encoders preserves hybrid acoustic-semantic
      representations for downstream tasks under streaming constraints.
    source: §4.2, Table 3
    evidence: FocalCodec-Stream variants trained via four-stage WavLM distillation outperform acoustic streaming
      codecs (EnCodec, AudioDec, HILCodec) on ASR, keyword spotting, and intent classification despite operating
      at lower bitrates; the 65k variant matches or surpasses PAST on all discriminative and generative tasks except
      ASR.
    confidence: high
    relevance: low
  - claim_id: supervised_domain_specific_fine_tuning_of_hybrid_codecs_achieves_strong
    role: complicates
    claim: Supervised domain-specific fine-tuning of hybrid codecs achieves strong in-domain intelligibility at
      the cost of multilingual generalization.
    source: §4.1
    evidence: PAST, fine-tuned on English data, achieves the lowest English dWER among streaming codecs (4.04%)
      but degrades severely on multilingual MLS (49.35% dWER), whereas FocalCodec-Stream (trained on English-only
      Libri-Light but without supervised task fine-tuning) retains competitive multilingual performance (19.88%
      dWER).
    confidence: high
    relevance: medium
  - claim_id: voice_conversion_quality_in_streaming_codecs_requires_joint_optimization_of
    role: supports
    claim: Voice conversion quality in streaming codecs requires joint optimization of intelligibility and speaker
      fidelity; gains on one metric alone are insufficient for practical use.
    source: §4.1, Table 2
    evidence: Mimi6 achieves competitive speaker similarity (91.3%) in one-shot VC on VCTK but at 110% dWER, while
      PAST achieves lower dWER (18.28%) at only 68.5% speaker similarity; FocalCodec-S@50-65k is the only streaming
      codec to simultaneously achieve high values on both (dWER 22.71%, Sim 92.5%).
    confidence: high
    relevance: low
  - claim_id: a_lightweight_refiner_module_bridging_causal_and_full_context_feature
    role: supports
    claim: A lightweight refiner module bridging causal and full-context feature distributions substantially improves
      perceptual quality in distilled streaming codecs.
    source: §4.3, Table 4
    evidence: Ablation on FocalCodec-S@50-4k shows that removing the refiner degrades UTMOS from 3.87 to 3.84 and
      dWER from 4.39% to 4.65%; omitting Stage 4 fine-tuning (which jointly trains the refiner) has a larger effect,
      raising dWER to 5.05% and reducing speaker similarity from 96.3% to 95.8%.
    confidence: high
    relevance: low
  limitations:
  - Training data is limited to English (LibriTTS, Libri-Light), so the multilingual robustness observed on MLS
    reflects generalization rather than explicit multilingual training. A performance gap with the non-streaming
    FocalCodec@50 remains across most metrics, particularly in ASR WER (17% vs. 15.33%) and SI error rate (2.18%
    vs. 0.35%), reflecting the inherent cost of the 80 ms latency budget. Downstream generative tasks (TTS, speech
    language modeling) are left for future work, so performance of the codec's discrete representations in autoregressive
    modeling pipelines is not yet demonstrated.
  - The paper does not provide listening tests or crowd-sourced MOS, relying instead on the automatic UTMOS predictor
    for naturalness assessment.
  caveats: []
- id: '2509.17006'
  published_date: "2025-09-21"
  entry_date: '2026-07-28'
  year: 2025
  venue: arXiv
  task:
  - codec
  architecture:
  - GAN
  relevance: medium
  evidence_role:
  - architecture_variant
  current_role: active_evidence
  method_family:
  - gan_neural_codecs_and_decoders
  claims:
  - claim_id: explicit_semantic_acoustic_disentanglement_in_neural_codecs_improves_reconstruction_quality
    role: supports
    claim: Explicit semantic-acoustic disentanglement in neural codecs improves reconstruction quality at comparable
      bitrates relative to conventional residual quantization approaches.
    source: §3.2, Table 1
    evidence: MBCodec (16 codebooks, 50Hz, 8.8kbps) achieves PESQ 3.83 and MUSHRA 85.9, versus EnCodec at the same
      bitrate scoring PESQ 2.78 and MUSHRA 85.3; at 4.4kbps, MBCodec (16 codebooks, 25Hz) scores PESQ 3.64 versus
      SpeechTokenizer at PESQ 1.26 and MUSHRA 79.0.
    confidence: high
    relevance: medium
  - claim_id: pqmf_based_frequency_subband_supervision_substantially_improves_spectral_reconstruction_fidelity
    role: supports
    claim: PQMF-based frequency subband supervision substantially improves spectral reconstruction fidelity in multi-codebook
      codec training.
    source: §3.4, Table 3
    evidence: Ablation shows removing PQMF supervision degrades PESQ from 3.83 to 2.34 (a 39% drop) and increases
      Mel-spectrogram distance from 2.34 to 3.45 in the best MBCodec configuration.
    confidence: high
    relevance: high
  - claim_id: non_uniform_codebook_dropout_distributions_that_concentrate_sampling_on_early
    role: supports
    claim: Non-uniform codebook dropout distributions that concentrate sampling on early RVQ layers outperform uniform
      dropout, better matching the hierarchical information density of residual quantization.
    source: §2.2, §3.4, Table 2, Table 3
    evidence: Half-Gaussian adaptive dropout achieves PESQ 3.28 versus exponential decay at 3.02 and chi-squared
      at 2.75; removing adaptive dropout entirely drops SI-SDR from 7.94 to 7.32 and increases STFT distance from
      0.08 to 0.22.
    confidence: high
    relevance: high
  - claim_id: ultra_low_bitrate_neural_audio_compression_imposes_a_persistent_quality
    role: complicates
    claim: Ultra-low bitrate neural audio compression imposes a persistent quality gap relative to ground truth,
      even with strong disentanglement strategies.
    source: §3.2, Table 1
    evidence: MBCodec at 2.2kbps achieves MUSHRA 82.8 versus ground truth MUSHRA 90.7 (a gap of 7.9 points) and
      PESQ 2.98 versus 4.64, despite the 170x compression ratio and explicit semantic-acoustic disentanglement.
    confidence: high
    relevance: high
  limitations:
  - The paper does not evaluate MBCodec as a token representation for downstream TTS or speech LM inference, which
    is presented as a primary motivation. The codec is trained and evaluated on speech-dominant data; generalisation
    to music or general audio is not assessed. No parameter count is reported, making direct efficiency comparison
    with other codecs difficult. The comparison set is limited to DAC, EnCodec, and SpeechTokenizer; more recent
    codecs such as Mimi or BiCodec are not included. As a preprint, no independent reproduction or external listening
    evaluation is available.
  caveats: []
- id: '2509.17143'
  published_date: "2025-09-21"
  entry_date: '2026-07-28'
  year: 2025
  venue: arXiv
  task:
  - VC
  architecture:
  - autoregressive-LM
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: active_evidence
  method_family:
  - autoregressive_codec_language_models
  claims:
  - claim_id: in_zero_shot_voice_conversion_temporally_coarser_syllabic_representations_reduce
    role: supports
    claim: In zero-shot voice conversion, temporally coarser syllabic representations reduce pitch leakage from
      linguistic features more effectively than standard frame-aligned SSL features, enabling cleaner prosody control
      at the cost of intelligibility.
    source: §2.1, §4, Table 2
    evidence: MaskVCT-Spk using SylBoost syllabic tokens achieves the lowest FPC (0.167) among all tested systems,
      indicating near-complete pitch independence from the source, while MaskVCT-All with continuous features retains
      more pitch correlation (FPC 0.417).
    confidence: high
    relevance: low
  - claim_id: multiple_classifier_free_guidance_weights_applied_to_distinct_conditioning_factors
    role: supports
    claim: Multiple classifier-free guidance weights applied to distinct conditioning factors in a single masked
      generative model enable user-configurable inference-time trade-offs between intelligibility, pitch fidelity,
      and speaker similarity.
    source: §2.5, §3.4, Table 2
    evidence: MaskVCT defines triple CFG weights (w_all, w_spk, w_ling) over speaker, pitch, and linguistic conditions
      within one trained model; sweeping these weights continuously interpolates between MaskVCT-All (WER 4.68%,
      S-SIM 0.865) and MaskVCT-Spk (WER 6.47%, S-SIM 0.895).
    confidence: high
    relevance: low
  - claim_id: syllabic_speech_representations_that_suppress_pitch_leakage_in_voice_conversion
    role: complicates
    claim: Syllabic speech representations that suppress pitch leakage in voice conversion also degrade content
      intelligibility through syllable misreadings caused by coarse temporal quantisation.
    source: §4, §5, Table 2
    evidence: MaskVCT-Spk achieves the highest speaker similarity (S-SIM 0.895) but the highest WER (6.47%) among
      systems tested, substantially above FACodec (3.55%) and FreeVC (3.96%); the conclusion section attributes
      misreadings to K-means syllable mapping errors in SylBoost.
    confidence: high
    relevance: medium
  - claim_id: masked_non_autoregressive_codec_models_can_match_or_exceed_autoregressive
    role: supports
    claim: Masked non-autoregressive codec models can match or exceed autoregressive and diffusion-based baselines
      on speaker similarity and quality in zero-shot VC while operating with fewer discrete tokens per utterance.
    source: §3.4, §4, Table 2
    evidence: MaskVCT-Spk (2048 tokens) achieves higher S-SIM (0.895) and SS-MOS (3.69) than MaskGCT-S2A (8192 tokens,
      S-SIM 0.863, SS-MOS 3.02) and competitive UTMOS (3.17 vs. 3.24).
    confidence: high
    relevance: high
  limitations:
  - Syllabic tokens introduce misreadings that WER alone cannot fully diagnose; the authors acknowledge the K-means
    quantiser cannot recover from incorrect syllable boundary assignments, and propose future work to address this
    with a trainable VQ module.
  - The model is English-only. Accent conversion experiments are limited to L2-ARCTIC and test only two conversion
    directions; how well the syllabic pitch-stripping generalises to tonal languages (where pitch is phonemic) is
    unexplored. The 511-pair test set is relatively small for statistical confidence, especially given the reported
    confidence intervals overlap for several key comparisons. No code or trained checkpoint is publicly released,
    limiting reproducibility.
  caveats: []
- id: '2509.17765'
  published_date: "2025-09-22"
  entry_date: '2026-07-28'
  year: 2025
  venue: arXiv
  task:
  - SCA
  - TTS
  architecture:
  - autoregressive-LM
  - hybrid
  relevance: medium
  evidence_role:
  - architecture_variant
  - acceleration_evidence
  - control_evidence
  current_role: influential
  method_family:
  - autoregressive_codec_language_models
  - hybrid_semantic_acoustic_tokenizers
  claims:
  - claim_id: replacing_a_block_wise_diffusion_vocoder_with_a_lightweight_causal
    role: supports
    claim: Replacing a block-wise diffusion vocoder with a lightweight causal convolutional decoder, driven by a
      multi-codebook autoregressive token predictor, can substantially reduce first-packet latency in streaming
      speech generation without sacrificing competitiveness on content-consistency metrics.
    source: §2.4, §2.5, Table 1, Table 13
    evidence: The Talker's multi-codebook AR scheme plus a 200M-parameter causal ConvNet Code2Wav stage achieves
      a 234ms end-to-end first-packet latency at 1x concurrency and the lowest reported content-consistency error
      on SEED test-en (1.39) among all compared zero-shot TTS systems including flow-matching and diffusion-based
      baselines.
    confidence: high
    relevance: high
  - claim_id: mixing_unimodal_and_cross_modal_training_data_from_the_earliest
    role: supports
    claim: Mixing unimodal and cross-modal training data from the earliest stage of pretraining allows a language
      model to add new input/output modalities without degrading its original text, vision, or audio-specific capabilities
      relative to matched unimodal baselines.
    source: §6, Table 16
    evidence: A controlled comparison of parameter-matched text-only, vision-only, and Omni models trained on identical
      corpora, schedules, and compute shows the Omni model matches or exceeds the unimodal baselines on general,
      math/STEM, coding, and multilingual text benchmarks as well as vision and video benchmarks.
    confidence: high
    relevance: medium
  - claim_id: strong_zero_shot_voice_cloning_performance_in_one_or_two
    role: complicates
    claim: Strong zero-shot voice-cloning performance in one or two conditioning languages does not guarantee comparable
      speaker-similarity performance uniformly across all supported languages.
    source: §5.2.2, Table 14
    evidence: Against MiniMax-Speech and ElevenLabs Multilingual v2 on a 10-language multilingual test set, the
      system leads by a substantial margin on Chinese, English, and French but reports only "competitive," non-leading,
      speaker-similarity or content-consistency scores on several other languages such as Portuguese and Russian.
    confidence: high
    relevance: medium
  - claim_id: a_large_scale_purpose_built_supervised_audio_encoder_trained_from
    role: supports
    claim: A large-scale, purpose-built supervised audio encoder trained from scratch for a multimodal LLM's audio
      pathway can outperform reusing a general pretrained ASR encoder (e.g., Whisper) as the perceptual front-end
      for both speech understanding and downstream speech generation.
    source: §1, §2.2, Table 6, Table 7
    evidence: Replacing the Whisper-based audio encoder from the predecessor system with AuT, trained from scratch
      on 20 million hours of supervised audio at a 12.5 Hz token rate, is cited as a key driver of gains across
      ASR, lyric-ASR, and voice-interaction benchmarks relative to Qwen2.5-Omni.
    confidence: high
    relevance: medium
  limitations:
  - Speech generation quality is evaluated exclusively with automatic metrics (WER/CER-style content consistency,
    embedding-based speaker similarity, BLEU for translation); no human MOS or listening-test results are reported
    anywhere in the paper for the Talker's synthesized speech, so claims of "stable, naturalistic speech synthesis"
    in the conclusion are not directly supported by subjective evidence in this report.
  - The paper acknowledges suboptimal performance on long-video benchmarks, attributed to limited positional extrapolation
    and restricted context length, as an explicit architectural limitation left for future work. The reported 234ms
    first-packet latency is described as "theoretical," measured under a specific vLLM/torch.compile/CUDA-Graph
    deployment configuration rather than as an end-user-measured figure across arbitrary hardware or network conditions,
    and latency degrades substantially under higher concurrency (up to 1172ms at 6-way concurrency in the audio
    case). Several baselines used for comparison (ElevenLabs, MiniMax-Speech, Gemini-2.5-Pro, GPT-4o variants) are
    closed proprietary systems, so exact reproduction of the comparative numbers is not possible outside the authors'
    own evaluation pipeline. The non-degradation ablation study, while methodologically rigorous, was run at limited
    model scales due to computational cost, and the authors explicitly caution that they could not sweep across
    all model sizes.
  caveats: []
- id: '2501.04561'
  published_date: "2025-09-23"
  entry_date: '2026-07-28'
  year: 2025
  venue: arXiv
  task:
  - SCA
  - TTS
  architecture:
  - autoregressive-LM
  - hybrid
  relevance: medium
  evidence_role:
  - architecture_variant
  - acceleration_evidence
  - control_evidence
  current_role: active_evidence
  method_family:
  - autoregressive_codec_language_models
  - hybrid_semantic_acoustic_tokenizers
  claims:
  - claim_id: progressive_text_pivoted_alignment_across_modality_pairs_can_substitute_for
    role: supports
    claim: Progressive, text-pivoted alignment across modality pairs can substitute for paired tri-modal training
      data without sacrificing downstream omnimodal task performance.
    source: §4.2, Table 1
    evidence: OpenOmni trains only on speech-text and image-text pairs (no image-speech-text triples) yet outperforms
      VITA, which is trained on 5M tri-modal samples, by 4 points on OmniBench while using a 7B rather than 7×8B
      language model and roughly 5x fewer training samples.
    confidence: high
    relevance: medium
  - claim_id: non_autoregressive_discrete_unit_speech_decoding_trades_generation_quality_for
    role: complicates
    claim: Non-autoregressive discrete-unit speech decoding trades generation quality for latency relative to autoregressive
      decoding.
    source: §4.2, §D "AR mode"
    evidence: The paper reports that its AR mode (NTP loss, 16K-unit vocabulary) yields higher speech generation
      quality but slower streaming, while the NAR mode (CTC loss, 6K-unit vocabulary) achieves under-1-second latency
      for up to 30 seconds of speech (5x faster) at the cost of "slightly worse" generation quality.
    confidence: high
    relevance: low
  - claim_id: direct_preference_optimization_can_be_adapted_to_discrete_unit_ctc
    role: supports
    claim: Direct preference optimization can be adapted to discrete-unit, CTC-trained speech generators to improve
      emotional coherence without an auxiliary emotion-control module.
    source: §4.2, Table 4
    evidence: CTC-DPO training on the 9K-pair EO2S-9K preference dataset (Plutchik-based emotion categories, CosyVoice-synthesized
      positive/negative pairs) raises Emotion2Vec-classified accuracy from 57.9% to 70.4% on Chinese and 62.6% to
      65.4% on English test speech.
    confidence: high
    relevance: low
  - claim_id: mixture_of_experts_capacity_is_necessary_not_merely_beneficial_for
    role: complicates
    claim: Mixture-of-experts capacity is necessary, not merely beneficial, for stabilizing CTC-loss training of
      multilingual non-autoregressive speech decoders.
    source: §3.4, §C, Table 6
    evidence: Ablations show a single feed-forward decoder layer (1 expert) fails to converge on bilingual WeNetSpeech/LibriSpeech
      data (CER/WER of 113.6/129.7/87.8/96.5), while increasing to a 4-expert MoE layer brings these down to single
      digits (8.5/8.4/4.2/4.7); the paper states that "without this layer, the speech decoder fails to train effectively."
    confidence: high
    relevance: medium
  limitations:
  - The system is trained and validated only on Chinese and English; the authors explicitly note that multilingual
    speech data beyond these two languages was not used due to resource constraints, leaving the generalization
    of the alignment and speech-generation strategy to other languages untested (§Limitation). Separately, the paper
    acknowledges that because the speech decoder conditions on the LLM's internal hidden states rather than its
    final decoded text, occasional mismatches between the LLM's actual text answer and the conditioning features
    used for speech generation can still occur despite the text-guided fusion module designed to mitigate this (§Appendix
    D). Evaluation of emotional and omnimodal quality relies on automated classifiers (Emotion2Vec, Whisper-based
    WER) and the authors' own benchmarks rather than independent human listening tests, so subjective perceptual
    quality of the generated speech is not directly reported.
  caveats: []
- id: '2509.18823'
  published_date: "2025-09-23"
  entry_date: '2026-07-28'
  year: 2025
  venue: arXiv
  task:
  - evaluation
  - codec
  architecture:
  - GAN
  relevance: medium
  evidence_role:
  - architecture_variant
  - evaluation_caution
  current_role: active_evidence
  method_family:
  - gan_neural_codecs_and_decoders
  claims:
  - claim_id: fr_chet_distance_metrics_computed_on_audio_embeddings_correlate_more
    role: supports
    claim: Fréchet-distance metrics computed on audio embeddings correlate more reliably with human perceptual quality
      judgments than kernel-based maximum mean discrepancy metrics.
    source: §4.1, Table 1
    evidence: Across every embedding space tested (EnCodec, DAC, DACe, CLAP, CLAP LAION, OpenL3), FAD achieved equal
      or higher Pearson and Spearman correlation with MUSHRA scores than MMD on the same embeddings, e.g. DACe FAD
      Rp = 0.70 vs. DACe MMD Rp = 0.65.
    confidence: high
    relevance: medium
  - claim_id: higher_fidelity_neural_audio_codec_embeddings_track_human_perceived_audio
    role: supports
    claim: Higher-fidelity neural audio codec embeddings track human-perceived audio quality more accurately than
      lower-fidelity codec embeddings.
    source: §4.1, Table 1; §3, Figs. 1-2
    evidence: 'Correlation with MUSHRA scores rose monotonically with codec quality: EnCodec (Rp = 0.38) < DAC 16
      kb/s (Rp = 0.68) < DACe (Rp = 0.70) under FAD, mirrored by a separate subjective MUSHRA test confirming DACe
      > DAC > EnCodec.'
    confidence: high
    relevance: high
  - claim_id: codec_derived_embeddings_even_from_high_fidelity_codecs_are_not
    role: complicates
    claim: Codec-derived embeddings, even from high-fidelity codecs, are not automatically competitive with embeddings
      trained for perceptual or semantic similarity when used for generative audio quality prediction.
    source: §4.2
    evidence: General-purpose CLAP LAION Music and OpenL3-128M embeddings outperformed the best neural codec embedding
      (DACe) in both Pearson and Spearman correlation with MUSHRA scores, attributed to codecs' reconstruction-based
      training objective and roughly 10x less training data than the contrastively/self-supervised-trained alternatives.
    confidence: high
    relevance: high
  - claim_id: kernel_based_statistical_distance_metrics_for_audio_evaluation_are_sensitive
    role: complicates
    claim: Kernel-based statistical distance metrics for audio evaluation are sensitive to bandwidth hyperparameter
      choice in a way that moment-based distances are not.
    source: §4.1
    evidence: Fixed MMD kernel bandwidths (σ = 1, 10, 100, 1000, 10000) tested on DACe embeddings peaked in correlation
      near σ = 100 but never reached the correlation achieved with the median-distance heuristic bandwidth, with
      performance degrading more sharply for undersized than oversized bandwidths.
    confidence: high
    relevance: medium
  limitations:
  - Validation is performed in a signal-to-signal, reference-based setting (coded audio vs. its own uncoded reference
    under classical lossy coding artifacts), not on true generative-model outputs. The paper explicitly motivates
    this as a proxy because full-bandwidth subjective datasets for generative audio are scarce, but this means the
    reported correlations characterize embedding sensitivity to coding artifacts rather than to the artifact types
    produced by generative models, and the paper does not close that gap.
  - The mono-only USAC test excludes stereo perceptual effects, which the authors note most public codecs and embedding
    models do not model anyway. DACe's training data and exact architecture changes beyond codebook count and batch
    size are only partially disclosed (internal 720-hour dataset, internal MUSHRA test sets), limiting external
    reproducibility, and no code or demo release is mentioned for DACe itself.
  caveats: []
- id: '2509.19025'
  published_date: "2025-09-23"
  entry_date: '2026-07-28'
  year: 2025
  venue: arXiv
  task:
  - codec
  architecture:
  - GAN
  relevance: medium
  evidence_role:
  - architecture_variant
  current_role: active_evidence
  method_family:
  - gan_neural_codecs_and_decoders
  claims:
  - claim_id: deterministic_nearest_neighbor_codeword_selection_in_residual_vector_quantization_is
    role: supports
    claim: Deterministic nearest-neighbor codeword selection in residual vector quantization is fragile to small
      input perturbations, producing codeword reassignments that compound across RVQ stages into audible reconstruction
      artefacts.
    source: §2.1, Figure 1
    evidence: Codeword-shift analysis on Encodec's first RVQ stage (24 kHz, 6 kbps) shows that although the noisy
      codeword usually matches the clean top-1 choice, a pronounced long tail of shifts (k>1) occurs when 15 dB
      SNR DEMAND noise is added to 120 clean VCTK utterances.
    confidence: high
    relevance: high
  - claim_id: neural_speech_codec_noise_robustness_can_be_improved_without_any
    role: supports
    claim: Neural speech codec noise robustness can be improved without any paired noisy-clean training data by
      simulating perturbation-induced instability directly at the quantization step.
    source: §3.4, Table 1
    evidence: Fine-tuning Encodec and WavTokenizer with distance-weighted probabilistic top-K sampling, trained
      exclusively on clean speech, improves SI-SDR, PESQ, STOI, and UTMOS at both 15 dB and 10 dB SNR with statistically
      significant paired t-tests (e.g., Encodec UTMOS 3.475 to 3.586 at 15 dB).
    confidence: high
    relevance: high
  - claim_id: training_on_paired_noisy_clean_data_achieves_stronger_robustness_than
    role: complicates
    claim: Training on paired noisy-clean data achieves stronger robustness than perturbation-only training under
      the specific noise conditions it was exposed to, but this advantage does not transfer to unseen noise types
      and comes at the cost of clean-speech quality.
    source: §3.5.1, Tables 2-3
    evidence: A noise-exposed fine-tuning baseline (Closest*) outperforms the proposed clean-data-only method on
      some noisy-speech metrics (e.g., PESQ) under matched training/test noise, yet degrades clean-speech scores
      and is outperformed by the proposed method on all four metrics across three held-out noise types not present
      in either method's training.
    confidence: high
    relevance: medium
  - claim_id: introducing_training_time_quantization_perturbations_in_a_curriculum_like_schedule
    role: refines
    claim: Introducing training-time quantization perturbations in a curriculum-like schedule, from the finest residual
      quantizer stage toward the coarsest, stabilizes robustness training more effectively than perturbing all quantizer
      stages simultaneously.
    source: §3.5.3, Figure 2
    evidence: PESQ and UTMOS improve steadily as probabilistic top-K sampling is progressively rolled out from the
      6th to the 1st VQ of Encodec's RVQ, whereas applying the same sampling to all six VQs at once ("Direct Top-K")
      yields inferior results due to premature perturbation of core structural features.
    confidence: high
    relevance: medium
  limitations:
  - The evaluation is confined to a VCTK subset (a single-domain, studio-quality English read-speech corpus) and
    to relatively mild noise conditions (10-15 dB SNR from DEMAND). Whether the robustness gains extend to more
    severe noise, non-additive distortions (e.g., reverberation, codec cascading, packet loss), or multilingual/spontaneous
    speech is untested. The method is validated on only two codecs (Encodec, WavTokenizer); both use RVQ-style quantization,
    so its applicability to non-RVQ or single-codebook designs with different codebook geometries is unconfirmed.
    The paper itself notes extending the framework to more streamable architectures and integrating it with large
    speech-language models as future directions, but does not attempt either.
  caveats: []
- id: '2509.19186'
  published_date: "2025-09-23"
  entry_date: '2026-07-28'
  year: 2025
  venue: arXiv
  task:
  - codec
  architecture:
  - GAN
  relevance: high
  evidence_role:
  - core_evidence
  - acceleration_evidence
  current_role: active_evidence
  method_family:
  - gan_neural_codecs_and_decoders
  claims:
  - claim_id: greedy_per_layer_code_selection_in_residual_vector_quantization_leaves
    role: supports
    claim: Greedy per-layer code selection in residual vector quantization leaves quantization error on the table
      that can be recovered purely at test time, without retraining the codec.
    source: §5.2, Table 3
    evidence: Beam-search encoding (B=16) lowers average L2 quantization error on LibriTTS from 5.096 to 4.625 for
      EnCodec and from 22.29 to 21.81 for HiFi-Codec relative to greedy (B=1) encoding of the same pre-trained checkpoints.
    confidence: high
    relevance: high
  - claim_id: reductions_in_rvq_quantization_error_obtained_by_a_better_search
    role: supports
    claim: Reductions in RVQ quantization error obtained by a better search strategy translate into measurable gains
      on standard reconstruction-quality metrics, not just the raw error term.
    source: §5.3, Table 1
    evidence: Increasing beam size from 1 to 16 improves PESQ, STOI, NISQA, and SI-SNR simultaneously for both EnCodec
      and HiFi-Codec on LibriTTS test-clean.
    confidence: high
    relevance: high
  - claim_id: adding_more_rvq_codebooks_higher_bit_rate_does_not_by
    role: refines
    claim: Adding more RVQ codebooks (higher bit-rate) does not by itself eliminate the suboptimality of greedy
      encoding; the search-strategy gap persists even at high codebook counts.
    source: §5.5, Table 4
    evidence: At 24 kbps with 32 codebooks, EnCodec's greedy encoding still trails beam-search encoding (PESQ 3.670
      vs. 3.691, NISQA 3.992 vs. 4.001), a gap of comparable relative size to the one seen at 6 kbps with 8 codebooks.
    confidence: high
    relevance: high
  - claim_id: a_naive_implementation_of_wider_search_codec_encoding_introduces_a
    role: complicates
    claim: A naive implementation of wider-search codec encoding introduces a latency cost that scales with search
      width, and only a parallel hardware-aware implementation avoids this trade-off.
    source: §5.6, Table 5
    evidence: Sequential CPU beam-search encoding increases inference time by 285% from B=1 to B=16 (43.53 ms to
      167.6 ms per 5-second clip), while a GPU-parallelized implementation of the same algorithm increases by only
      about 9% (6.780 ms to 7.366 ms) over the same range.
    confidence: high
    relevance: high
  limitations:
  - The evaluation is confined to two pre-trained codec checkpoints (EnCodec, HiFi-Codec); the paper does not test
    the algorithm on more recent RVQ-GAN codec designs (e.g. DAC-style codecs) or on codecs with substantially larger
    codebook sizes, where the exponential O(S^L) exhaustive-search cost the method is designed to avoid is even
    more pronounced. Beam size B and the per-step candidate width k are always set equal in the experiments, so
    their individual contributions to the error reduction are not disentangled. All reported gains are measured
    on intrinsic reconstruction-quality metrics (PESQ, STOI, NISQA, SI-SNR, mel distance); the paper does not evaluate
    whether the improved codec reconstructions change downstream outcomes for systems that consume the codes, such
    as word error rate in a codec-token TTS or speech language model pipeline. The non-verbal vocalization test
    set is an internal, non-public dataset, limiting reproducibility of that portion of the evaluation.
  caveats: []
- id: '2509.19592'
  published_date: "2025-09-23"
  entry_date: '2026-07-28'
  year: 2025
  venue: arXiv
  task:
  - TTS
  architecture:
  - autoregressive-LM
  relevance: medium
  evidence_role:
  - architecture_variant
  current_role: active_evidence
  method_family:
  - autoregressive_codec_language_models
  claims:
  - claim_id: decoding_codebooks_of_a_multi_codebook_acoustic_frame_with_explicit
    role: supports
    claim: Decoding codebooks of a multi-codebook acoustic frame with explicit intra-frame dependencies (iteratively)
      yields a generated token distribution closer to the ground truth than decoding all codebooks in parallel under
      an independence assumption.
    source: §3.3.1, Fig. 2e
    evidence: Across all tested frame-stacking factors (1, 2, 4), every autoregressive- or MaskGIT-local-transformer
      configuration achieves lower Fréchet Distance than every parallel-sampled configuration, including the unstacked
      parallel baseline.
    confidence: high
    relevance: high
  - claim_id: offloading_intra_frame_codebook_decoding_to_a_small_auxiliary_transformer
    role: supports
    claim: Offloading intra-frame codebook decoding to a small auxiliary transformer lets a primary acoustic decoder
      predict multiple codec frames per generation step, substantially increasing throughput without retraining
      the underlying codec at a lower frame rate.
    source: §2.4, §3.3.2, Table 1, Fig. 2f
    evidence: At a frame-stacking factor of 2, the autoregressive local-transformer model reaches 2.1x throughput
      over the unstacked parallel baseline while improving Fréchet Distance and keeping WER, speaker similarity,
      and MOS within or better than baseline; the MaskGIT variant reaches 3.1x throughput at comparable quality.
    confidence: high
    relevance: high
  - claim_id: parallel_independent_codebook_prediction_degrades_disproportionately_not_just_proportionally_as
    role: complicates
    claim: Parallel independent codebook prediction degrades disproportionately, not just proportionally, as more
      codebook information is packed into a single decoding step.
    source: §3.3.2
    evidence: Applying parallel sampling to a 2x frame-stacked model (instead of routing through the local transformer)
      increases unseen-speaker Fréchet Distance by 67% relative to the unstacked parallel baseline and lowers MOS.
    confidence: high
    relevance: high
  - claim_id: the_throughput_gains_of_iterative_masked_prediction_decoding_for_acoustic
    role: complicates
    claim: The throughput gains of iterative masked-prediction decoding for acoustic codebooks come at a quality
      cost that grows sharply once the number of sampling steps is small relative to the number of tokens being
      resolved per step.
    source: §3.3.2, Fig. 2a
    evidence: At a stacking factor of 4, the MaskGIT local transformer with 3 sampling steps decoding 32 tokens
      per step (8 codebooks × 4 stacked frames) shows a significant MOS drop relative to baseline, while the autoregressive
      local transformer at the same stacking factor does not exhibit this drop.
    confidence: high
    relevance: medium
  limitations:
  - 'Robustness to unseen speakers degrades substantially at higher frame-stacking factors: unseen-speaker speaker
    similarity falls from 0.765 at stacking factor 1 to 0.642 (AR LT) and 0.624 (MaskGIT LT) at stacking factor
    4, which the authors themselves flag by recommending high stacking only "when not needing zero-shot functionality."'
  - All experiments build on a single base system (Koel-TTS) and a single codec (NanoCodec, FSQ-based, 8 codebooks
    at 21.5 fps); it is untested whether the same tradeoffs hold for RVQ-based codecs, different codebook counts,
    or other primary-decoder architectures. The MaskGIT local transformer's degradation at high stacking is attributed
    to using only 3 sampling steps, but the paper does not run the ablation that would confirm more steps recover
    quality, leaving the speed-quality Pareto frontier for MaskGIT only partially characterized. Training data is
    described only as "the same 18k hours of data as in the Koel-TTS paper," with no further specification of language,
    speaker count, or domain in this paper itself.
  caveats: []
- id: '2509.19883'
  published_date: "2025-09-24"
  entry_date: '2026-07-28'
  year: 2025
  venue: arXiv
  task:
  - singing
  architecture:
  - hybrid
  relevance: high
  evidence_role:
  - core_evidence
  - control_evidence
  current_role: active_evidence
  method_family:
  - hybrid_semantic_acoustic_tokenizers
  claims:
  - claim_id: prompt_based_conditioning_in_masked_generative_or_codec_language_speech
    role: supports
    claim: Prompt-based conditioning in masked generative or codec-language speech synthesis models causes measurable
      leakage of prosodic attributes from the acoustic prompt into the synthesized output, independent of the target
      language.
    source: §V.A, Table I
    evidence: Paired-prompt outputs show consistently lower pitch/energy/jitter differences than unpaired-prompt
      outputs from the same speaker on both LibriTTS (English) and AISHELL-3 (Mandarin) when synthesizing with MaskGCT.
    confidence: high
    relevance: high
  - claim_id: explicit_contrastive_regularization_between_the_acoustic_prompt_and_an_external
    role: supports
    claim: Explicit contrastive regularization between the acoustic prompt and an external control signal (e.g.,
      melody/pitch) reduces attribute leakage and improves controllability in prompt-based zero-shot synthesis.
    source: §V.D, Table V
    evidence: Removing the coarse-to-fine (sequence + frame level) contrastive loss increases F0-RMSE from 0.042
      to 0.08 and lowers SingMOS from 4.32 to 4.12 on the seen-singer test set, with sequence-level and frame-level
      components independently ablated to show complementary effects on speaker-identity and pitch-detail metrics
      respectively.
    confidence: high
    relevance: medium
  - claim_id: integrating_auxiliary_transcription_derived_frame_level_supervision_directly_into_a
    role: supports
    claim: Integrating auxiliary transcription-derived frame-level supervision directly into a synthesis model's
      training loop, rather than using it only as an offline data-cleaning step, improves fine-grained attribute
      alignment.
    source: §V.D, Table V
    evidence: Removing the in-loop SVT auxiliary loss produces the largest single-component degradation in the ablation,
      raising F0-RMSE from 0.042 to 0.194 and lowering SingMOS from 4.32 to 3.95.
    confidence: high
    relevance: medium
  - claim_id: parameter_efficient_fine_tuning_can_match_or_exceed_full_fine
    role: refines
    claim: Parameter-efficient fine-tuning can match or exceed full fine-tuning when adapting a large pretrained
      codec-based speech model to a lower-resource downstream domain, provided the low-rank capacity is placed appropriately.
    source: §V.D, Table VII
    evidence: LoRA fine-tuning of the S2A diffusion estimator (6.51% trainable parameters) achieves lower F0-RMSE
      (0.053) and higher SECS (0.92) than fully fine-tuning the same backbone (100% trainable, F0-RMSE 0.099, SECS
      0.859) when adapting MaskGCT to singing voice synthesis.
    confidence: high
    relevance: high
  limitations:
  - The zero-shot evaluation set (OpenSinger) has no native music-score annotations, so the authors pair OpenSinger
    audio with M4Singer's pitch/duration sequences to construct test inputs. This means the "unseen singer" evaluation
    is not evaluating on musically native score-audio pairs, which could understate or overstate melody-control
    difficulty relative to a genuinely paired unseen-singer benchmark.
  - Beyond that, evaluation is confined to Mandarin singing corpora (M4Singer, Opencpop, OpenSinger); no cross-lingual
    or multilingual SVS results are reported, despite the underlying MaskGCT backbone being multilingual-capable.
    Overall S2A/T2S model size is not reported (only the small SVT module's dimensions are given), limiting reproducibility
    of compute requirements. The subjective evaluation panel is modest (20 musically trained participants), typical
    for SVS papers but smaller than typical large-scale TTS MOS studies. Finally, the sequence-level contrastive
    objective depends on having multiple same-singer utterances available per training batch, which the curated
    corpora used here provide but which may not hold for less structured or lower-resource singing data.
  caveats: []
- id: '2509.20485'
  published_date: "2025-09-24"
  entry_date: '2026-07-28'
  year: 2025
  venue: arXiv
  task:
  - evaluation
  - TTS
  architecture: []
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  - evaluation_caution
  current_role: active_evidence
  method_family: []
  claims:
  - claim_id: conditioning_a_discrete_token_likelihood_model_on_the_input_text
    role: supports
    claim: Conditioning a discrete-token likelihood model on the input text produces a substantially more discriminative
      and human-aligned intelligibility signal than unconditional token language modeling.
    source: §V.A, Table 1
    evidence: TTScore-int reaches utterance-level correlation with WER up to 0.78 on VoiceMOS22 (system-level up
      to 0.96), compared to at most 0.44/0.55 for an identically-trained unconditional token-LM (uLM) baseline and
      0.32/0.42 for SpeechLMScore.
    confidence: high
    relevance: medium
  - claim_id: mos_predictors_trained_directly_on_mos_labeled_data_generalize_poorly
    role: supports
    claim: MOS predictors trained directly on MOS-labeled data generalize poorly across benchmarks relative to reference-free
      metrics with a narrower, aspect-specific training objective.
    source: §V.B, Table 2
    evidence: UTMOS, trained on VoiceMOS22 data, achieves high correlation with MOS on VoiceMOS22 but notably lower
      correlation on the out-of-distribution SOMOS dataset, while TTScore-int retains stronger correlation with
      MOS on SOMOS.
    confidence: high
    relevance: medium
  - claim_id: reference_free_text_conditioned_prosody_scoring_can_track_prosodic_appropriateness
    role: supports
    claim: Reference-free, text-conditioned prosody scoring can track prosodic appropriateness where reference-dependent
      F0-based metrics fail, including on benchmarks lacking aligned reference audio.
    source: §VI.C, Tables 3-5
    evidence: TTScore-pro shows positive correlation with MOS on both SOMOS and VoiceMOS22 and correct-sign correlation
      with TTS Arena ELO scores, while F0-RMSE shows near-zero or negative correlation with MOS and an unexpected
      positive (wrong-sign) correlation with ELO.
    confidence: high
    relevance: low
  - claim_id: objective_prosody_metrics_that_isolate_a_single_acoustic_dimension_pitch
    role: complicates
    claim: Objective prosody metrics that isolate a single acoustic dimension (pitch) correlate only moderately
      with overall perceived naturalness, since naturalness depends on additional prosodic factors beyond F0.
    source: §VI.C, Table 3
    evidence: TTScore-pro's correlations with MOS remain in the 0.2-0.5 range even where it clearly outperforms
      F0-RMSE and F0-correlation baselines, and the paper attributes the residual gap to prosody encompassing rhythm,
      energy, and phrasing beyond the F0-based prosody tokens used here.
    confidence: high
    relevance: low
  limitations:
  - The method targets only F0-related prosody; rhythm, energy, and phrasing, which the authors acknowledge are
    also core prosodic dimensions, are not modeled or evaluated.
  - The text-to-token generators are trained exclusively on a large-scale, clean, English-only corpus (LibriSpeech-960),
    so the metric's reliability in noisy, spontaneous, or non-English settings is untested. Because the token distributions
    are learned from data, the authors note the approach may inherit whatever biases exist in that training corpus.
    The paper also does not report whether code will be released, and the demo/code availability fields could not
    be confirmed from the paper text.
  caveats: []
- id: '2509.20802'
  published_date: "2025-09-25"
  entry_date: '2026-07-28'
  year: 2025
  venue: arXiv
  task:
  - TTS
  architecture:
  - autoregressive-LM
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: active_evidence
  method_family:
  - autoregressive_codec_language_models
  claims:
  - claim_id: many_transformer_layers_in_autoregressive_llm_tts_backbones_contribute_little
    role: supports
    claim: Many transformer layers in autoregressive LLM-TTS backbones contribute little to synthesis quality and
      can be removed with minimal loss in naturalness and speaker similarity.
    source: §4.1, Table 1
    evidence: Halving CosyVoice 2's depth to 12 layers (39.7% fewer parameters) increases Seed-TTS WER by only 0.68
      and decreases NMOS by 0.13, while speaker similarity and UTMOS remain essentially unchanged.
    confidence: high
    relevance: low
  - claim_id: layer_importance_criteria_developed_for_pruning_text_only_llms_do
    role: refines
    claim: Layer-importance criteria developed for pruning text-only LLMs do not transfer directly to speech generation
      backbones; intelligibility-grounded criteria are needed to identify prunable layers correctly.
    source: §2.1, §4.2, Table 2
    evidence: Cosine-based layer importance (input/output latent similarity) diverges from WER-based importance
      in TTS backbones, and substituting it for the proposed WER-based criterion increases WER from 1.59 to 1.74
      and CER from 0.54 to 0.61 on LibriTTS test-clean.
    confidence: high
    relevance: medium
  - claim_id: knowledge_distillation_can_recover_most_of_the_performance_lost_from
    role: supports
    claim: Knowledge distillation can recover most of the performance lost from aggressive layer pruning in speech-generation
      LLMs using only a small fraction of the original pretraining data.
    source: §3, §4.1, Table 1b
    evidence: Fine-tuning pruned variants required under 5% of the original pretraining data (25% of LibriTTS for
      CosyVoice 2, an upper-bounded 12.5% of LibriHeavy for LLaSA) yet speaker similarity and UTMOS changed by at
      most 0.045 and 0.04 respectively relative to the uncompressed backbones.
    confidence: high
    relevance: medium
  - claim_id: robustness_to_layer_pruning_varies_substantially_across_llm_tts_backbones
    role: complicates
    claim: Robustness to layer pruning varies substantially across LLM-TTS backbones depending on how redundant
      their transformer layers are, so a single pruning ratio does not generalize uniformly.
    source: §4.1
    evidence: LLaSA showed a larger relative quality drop after 50% layer pruning (speaker similarity −0.045, UTMOS
      −0.04) than CosyVoice 2 at the same pruning ratio, attributed to LLaSA's WLI values being more uniformly high
      across layers, indicating less exploitable redundancy.
    confidence: high
    relevance: medium
  limitations:
  - The evaluation covers only two backbones (CosyVoice 2 and LLaSA-1B) and English-only test sets (LibriTTS test-clean,
    Seed-TTS test-en), so it is untested whether the WER-based pruning criterion and dynamic distillation scheme
    generalize to other LLM-TTS architectures, multilingual settings, or streaming inference. The more aggressive
    pruning configurations (9-layer CosyVoice 2) trade a larger, unquantified increase in WER for additional speed
    and memory gains, and the paper does not characterize where this trade-off becomes unacceptable for deployment.
    Computing WLI itself requires running WER evaluation over a data subset for each candidate layer removal, adding
    an upfront cost to the pruning procedure that is not reported in wall-clock terms.
  caveats: []
- id: '2509.21968'
  published_date: "2025-09-26"
  entry_date: '2026-07-28'
  year: 2025
  venue: arXiv
  task:
  - codec
  architecture:
  - GAN
  - VAE
  relevance: medium
  evidence_role:
  - architecture_variant
  - evaluation_caution
  current_role: active_evidence
  method_family:
  - gan_neural_codecs_and_decoders
  - vae_vector_quantized_codecs
  claims:
  - claim_id: a_shared_single_codebook_can_be_structured_with_overlapping_nested
    role: supports
    claim: A shared single codebook can be structured with overlapping, nested domain partitions rather than rigid
      disjoint splits, improving both reconstruction and downstream generation quality relative to a rigid-split
      design.
    source: §4.4, Tables 3–4
    evidence: Ablating codebook design at fixed codebook size and data scale, the nested codebook reduces reconstruction
      WER from 4.21 (rigid-split) to 3.99 and downstream TTS generation WER from 6.26 to 4.99 on LibriSpeech-PC
      test-clean.
    confidence: high
    relevance: high
  - claim_id: distilling_frame_level_representations_from_multiple_domain_specific_self_supervised
    role: supports
    claim: Distilling frame-level representations from multiple domain-specific self-supervised teacher models into
      one acoustic codec can improve reconstruction and generation quality across all the covered domains simultaneously,
      not just the domain of a single teacher.
    source: §3.3, §4.4, Tables 3–4
    evidence: Adding multi-domain distillation (WavLM for speech, MuQ for vocal/music, BEATs for sound) on top of
      the nested codebook further raises the speech-partition code-usage ratio from 37.1% to 38.2% and improves
      downstream TTS generation WER from 4.99 to 4.51.
    confidence: high
    relevance: high
  - claim_id: unifying_multiple_audio_domains_into_a_single_shared_quantization_codebook
    role: complicates
    claim: Unifying multiple audio domains into a single, shared quantization codebook does not close the gap with
      domain-specific single-layer codecs on every reconstruction metric, even when the unified model uses a larger
      codebook and lower token rate.
    source: §4.2, Table 1
    evidence: On LibriSpeech test-clean, AUV's PESQ-WB (2.40) and SPK-SIM (0.81) trail dedicated speech codecs such
      as DAC (4.01, 0.95) and are roughly on par with, not clearly better than, BigCodec and UniCodec.
    confidence: high
    relevance: high
  - claim_id: larger_unified_codebooks_intended_to_accommodate_more_audio_domains_can
    role: complicates
    claim: Larger unified codebooks intended to accommodate more audio domains can hurt downstream autoregressive
      generation quality by increasing the modeling burden on the generative model consuming the tokens, even when
      reconstruction quality is unaffected.
    source: §4.3, Table 4
    evidence: Scaling the codebook from 16,384 to 20,480 entries (C0 vs. C2) yields slightly worse downstream generation
      WER (4.51 vs. 4.89) despite improved reconstruction metrics; training an autoregressive model on codes from
      a still-larger 131,072-entry codebook (MagiCodec) reportedly failed outright.
    confidence: high
    relevance: medium
  limitations:
  - The paper reports no total parameter count for the AUV encoder-decoder, limiting direct efficiency comparison
    with baselines whose sizes are known. Downstream generative evaluation is restricted to a single autoregressive
    TTS backbone (EmoVoice) trained on a comparatively small 1K-hour subset of LibriSpeech, so it is unclear whether
    the reconstruction and generation gains hold at larger downstream training scales or with non-autoregressive
    generators. Speaker similarity in the downstream TTS setting remains low in absolute terms (SPK-SIM ≈ 0.43–0.44)
    across all codecs tested, including AUV, suggesting the codec-level improvements shown here do not yet translate
    into strong speaker fidelity for generated speech. The domain-label input used during training but withheld
    at inference creates a train/inference mismatch whose effect on codebook index selection is only indirectly
    probed via the reported index-distribution statistics, not directly ablated.
  caveats: []
- id: '2509.22062'
  published_date: "2025-09-26"
  entry_date: '2026-07-28'
  year: 2025
  venue: arXiv
  task:
  - TTS
  architecture:
  - autoregressive-LM
  relevance: high
  evidence_role:
  - core_evidence
  - acceleration_evidence
  - control_evidence
  current_role: active_evidence
  method_family:
  - autoregressive_codec_language_models
  claims:
  - claim_id: injecting_explicit_linguistic_structure_into_the_primary_codebook_of_a
    role: supports
    claim: Injecting explicit linguistic structure into the primary codebook of a neural speech codec reduces the
      downstream language model's learning burden and improves synthesis intelligibility.
    source: §4.3, Table 4
    evidence: Removing the semantic distillation loss during codec-conditioned TTS training raises WER from 3.31%
      to 3.97% on SeedTTS-test, 9.74% to 11.83% on PGC-Hard, and 16.57% to 18.34% on PGC-Poly, with SIM and UTMOS
      also degrading.
    confidence: high
    relevance: high
  - claim_id: an_automatic_speech_recognition_model_can_serve_as_an_effective
    role: supports
    claim: An automatic speech recognition model can serve as an effective semantic teacher for codec distillation,
      as an alternative to self-supervised speech representation models.
    source: §E.2, Table 5
    evidence: S3Codec distills Whisper encoder embeddings (rather than HuBERT/SSL features) into the first RVQ level;
      a small model trained with S3Codec reaches 3.30% WER on SeedTTS-test vs. 4.21% for an otherwise identical
      model using undistilled DAC tokens.
    confidence: high
    relevance: high
  - claim_id: fully_autoregressive_tts_systems_that_omit_an_explicit_continuous_acoustic
    role: complicates
    claim: Fully autoregressive TTS systems that omit an explicit continuous acoustic-feature conditioning stage
      (e.g., mel-spectrogram or speaker-similarity-vector guidance) tend to underperform hybrid AR+NAR or flow-matching
      systems on speaker similarity even when intelligibility is competitive.
    source: §4.2, Table 2-3
    evidence: CaT-TTS reports SIM of 0.668-0.678 across test sets versus 0.71-0.80 for Seed-TTS, CosyVoice 2/3,
      and F5-TTS, despite comparable or better WER among AR-only baselines.
    confidence: high
    relevance: low
  - claim_id: test_time_parallel_decoding_with_learned_input_dependent_aggregation_weights
    role: refines
    claim: Test-time parallel decoding with learned, input-dependent aggregation weights can reduce autoregressive
      error accumulation at near-zero added latency, but the achievable robustness gain is bounded by how many parallel
      streams are used, trading GPU utilization against benefit.
    source: §4.3, §3.3, Figure 3-4
    evidence: MAPI ablation across increasing parallel-stream counts shows WER improving and becoming more stable
      across 10 repeated inferences per sample, while the authors note GPU resource utilization rises correspondingly
      and stream count must be tuned per deployment scenario.
    confidence: high
    relevance: low
  limitations:
  - Training relies on an unreleased proprietary corpus (~200k hours, ~85% Chinese / ~15% English), and neither
    code nor a demo is available, which limits independent verification of the reported results.
  - The semantic-distillation ablation (removing the loss) and the MAPI ablation are both run on smaller sub-datasets
    and reduced-size "CaT-TTS-small" models rather than the full 0.4B system, so it is not established that the
    same magnitude of gains transfers to the full-scale model. Speaker similarity remains a clear weak point relative
    to hybrid and NAR baselines, which the authors attribute to the deliberate absence of continuous acoustic conditioning
    rather than treat as a target for improvement. The evaluation is also dominated by Chinese-language and Chinese-out-of-domain
    test sets (PGC-Hard, PGC-Poly, Seed-TTS test-zh/test-hard), with comparatively less English-language evidence.
  caveats: []
- id: '2509.22167'
  published_date: "2025-09-26"
  entry_date: '2026-07-28'
  year: 2025
  venue: arXiv
  task:
  - TTS
  - codec
  architecture:
  - VAE
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: active_evidence
  method_family:
  - vae_vector_quantized_codecs
  claims:
  - claim_id: regularizing_a_continuous_vae_latent_space_toward_semantic_structure_derived
    role: supports
    claim: Regularizing a continuous VAE latent space toward semantic structure derived from self-supervised speech
      features mitigates the reconstruction-generation trade-off in latent-based non-autoregressive TTS.
    source: §4.1, Table 1
    evidence: At a fixed 64-dimensional latent size, adding cosine-similarity alignment to WavLM features reduces
      WER from 2.65% (vanilla VAE) to 2.10% while raising speaker similarity from 0.59 to 0.64 on LibriSpeech-PC
      test-clean.
    confidence: high
    relevance: medium
  - claim_id: semantic_alignment_regularization_of_the_generation_target_accelerates_training_convergence
    role: supports
    claim: Semantic alignment regularization of the generation target accelerates training convergence of latent
      diffusion/flow-matching TTS models.
    source: §4.1, Figure 3
    evidence: Training-step comparisons show Semantic-VAE-based F5-TTS reaching lower WER and higher SIM than both
      the vanilla-VAE and mel-spectrogram baselines at the same number of training steps.
    confidence: high
    relevance: medium
  - claim_id: the_intelligibility_versus_speaker_similarity_trade_off_in_vae_latent
    role: refines
    claim: The intelligibility-versus-speaker-similarity trade-off in VAE latent representations is governed by
      which layer of a self-supervised model is used for semantic supervision, not eliminated outright by adding
      semantic alignment.
    source: §4.3, Table 3
    evidence: Ablating SSL layer choice shows the final layer of WavLM/HuBERT substantially degrades SIM (0.50-0.58)
      despite comparable WER, while an intermediate layer (WavLM layer 23) gives the best joint WER/SIM balance.
    confidence: high
    relevance: medium
  - claim_id: cosine_similarity_alignment_losses_to_self_supervised_speech_features_preserve
    role: supports
    claim: Cosine-similarity alignment losses to self-supervised speech features preserve useful semantic structure
      in a latent space more effectively than L1 or L2 distance losses to the same features.
    source: §4.3, Table 3
    evidence: Negative cosine alignment achieves 2.10% WER / 0.64 SIM, compared to 3.12% WER / 0.47 SIM for L1 alignment
      and 4.37% WER / 0.48 SIM for L2 alignment under otherwise identical settings.
    confidence: high
    relevance: medium
  - claim_id: adding_a_semantic_regularization_objective_to_a_vae_s_training
    role: complicates
    claim: Adding a semantic regularization objective to a VAE's training loss does not necessarily degrade the
      representation's reconstruction fidelity, contrary to the general expectation that additional regularization
      terms cost reconstruction quality.
    source: §4.2, Table 2
    evidence: Semantic-VAE reconstruction metrics on LibriTTS test-other (PESQ 3.74, STOI 0.96, UTMOS 3.56) are
      nearly identical to an unregularized vanilla VAE of the same architecture (PESQ 3.75, STOI 0.97, UTMOS 3.57).
    confidence: high
    relevance: high
  limitations:
  - The high-resource baseline numbers in Table 1 (CosyVoice, FireRedTTS, E2 TTS, F5-TTS at 100k-580k training hours)
    are quoted directly from their original papers rather than re-run under matched data and compute, so any comparison
    against Semantic-VAE's low-resource (0.6k-hour) results is illustrative context only, not a controlled comparison.
    The method is validated on English speech only, at a single latent configuration (64 dimensions, 40Hz), and
    on two NAR backbones (F5-TTS, E2 TTS); generalization to AR codec-token TTS systems, other languages, or other
    latent dimensionalities/frame rates is untested. The paper also does not report inference latency or discuss
    how the added SSL model changes deployment cost of the VAE training pipeline (inference itself is unaffected,
    since the SSL branch is training-only).
  caveats: []
- id: '2509.24457'
  published_date: "2025-09-29"
  entry_date: '2026-07-28'
  year: 2025
  venue: arXiv
  task:
  - codec
  - evaluation
  architecture: []
  relevance: high
  evidence_role:
  - core_evidence
  - evaluation_caution
  current_role: active_evidence
  method_family: []
  claims:
  - claim_id: neural_network_based_objective_quality_metrics_correlate_more_strongly_with
    role: supports
    claim: Neural-network-based objective quality metrics correlate more strongly with human subjective judgments
      than classical DSP-based metrics when evaluating neural audio codec output.
    source: §4
    evidence: Across 1,700 data points from 17 codec conditions, scoreq_ref reaches a Pearson correlation of 0.87
      with MUSHRA-1S scores and utmos, nomad, scoreq_nr, sheet_ssqa, and audiobox ce all exceed 0.80, while the
      strongest classical baselines (warpq, pesq) reach only 0.73; all metrics with correlation above 0.8 are neural-network-based.
    confidence: high
    relevance: high
  - claim_id: non_intrusive_objective_metrics_trained_on_mos_based_subjective_ratings
    role: complicates
    claim: Non-intrusive objective metrics trained on MOS-based subjective ratings saturate near the top of the
      quality scale, weakening their ability to discriminate between near-transparent codec conditions.
    source: §4
    evidence: Spearman and Kendall correlations for utmos and scoreq_nr drop sharply relative to their Pearson correlation,
      and condition-wise plots show scoreq_nr, ce, sheet_ssqa, and utmos producing near-constant scores in the high-MUSHRA
      range (red-ellipse regions of Fig. 3), attributed to the coarse 5-point ACR scale underlying the MOS targets
      these metrics are trained on.
    confidence: high
    relevance: high
  - claim_id: reference_based_intrusive_objective_metrics_remain_more_statistically_stable_and
    role: refines
    claim: Reference-based (intrusive) objective metrics remain more statistically stable and discriminative than
      reference-free (non-intrusive) metrics for high-quality neural codec conditions, even though non-intrusive
      metrics are the only option when no reference signal is available.
    source: §4
    evidence: Confidence-interval analysis across a high-, medium-, and low-quality codec condition (DAC-8kbps,
      SemantiCodec-1.4kbps, EnCodec-1.5kbps) shows intrusive scoreq_ref producing consistently smaller confidence
      intervals than non-intrusive scoreq_nr across sample sizes, an effect most pronounced for the high-quality
      condition.
    confidence: high
    relevance: high
  - claim_id: reference_based_classical_metrics_can_also_fail_to_discriminate_between
    role: complicates
    claim: Reference-based classical metrics can also fail to discriminate between codec conditions that differ
      substantially in perceived subjective quality, despite having access to the clean reference signal.
    source: §4
    evidence: pesq shows the widest observed blind spot among all evaluated metrics, mapping codecs spanning roughly
      35 MUSHRA points to nearly identical objective scores in the low-to-medium quality range (green-ellipse regions
      of Fig. 3), a limitation the authors note is a previously known phenomenon for pesq.
    confidence: high
    relevance: high
  limitations:
  - All results are obtained under clean speech conditions only; the authors explicitly leave analysis of noise
    and reverberation for future work, so the metric rankings and saturation findings here should not be assumed
    to transfer to noisy or reverberant deployment scenarios.
  - The study is restricted to 12 codecs and 17 conditions selected by the authors, and to a single test set of
    100 files drawn from the LRAC challenge Track 1 blind set; generalization to other speech domains, languages,
    or codec families is untested. Sample-size guidance is derived from only three representative codec conditions
    and two metrics (scoreq_ref, scoreq_nr), leaving open how confidence-interval behavior generalizes across the
    full metric suite. The authors also note unexplained run-to-run variability for utmosv2, possibly from internal
    random frame selection not currently exposed by the toolkit used, which they flag rather than resolve.
  caveats: []
- id: '2509.24570'
  published_date: "2025-09-29"
  entry_date: '2026-07-28'
  year: 2025
  venue: arXiv
  task:
  - TTS
  - VC
  - evaluation
  architecture:
  - autoregressive-LM
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  - infrastructure
  current_role: active_evidence
  method_family:
  - autoregressive_codec_language_models
  claims:
  - claim_id: automated_pipelines_combining_expressive_tts_voice_conversion_and_llm_based
    role: supports
    claim: Automated pipelines combining expressive TTS, voice conversion, and LLM-based instruction generation
      can produce large-scale paired speech style editing data without manual recording or annotation, provided
      a multi-criterion filtering step is applied.
    source: §2.1, §2.2, Fig. 3
    evidence: The three-stage pipeline (EmoCapTTS synthesis + Chatterbox voice conversion + Qwen3-8B instruction
      generation) yields 382 hours and ~100,000 pairs from EARS and Expresso source material, filtered to WER <
      10, style similarity > 0.5, and speaker similarity > 0.5 *(§2.1, §2.2, Fig. 3)*.
    confidence: high
    relevance: medium
  - claim_id: fine_grained_diverse_natural_language_instructions_improve_both_in_domain
    role: supports
    claim: Fine-grained, diverse natural-language instructions improve both in-domain accuracy and cross-domain
      generalization of instruction-guided speech style editing models relative to coarse, templated instruction
      sets.
    source: §4.2, Table 2
    evidence: LlasaEdit trained on ISSE outperforms the same architecture trained on ESD across WER, style similarity,
      speaker similarity, and UTMOS in-domain (8.06 vs. 10.07 WER; 0.68 vs. 0.64 style-sim), and the ISSE-trained
      model's cross-domain performance on ESD exceeds the ESD-trained model's in-domain performance on several metrics
      *(§4.2, Table 2)*.
    confidence: high
    relevance: medium
  - claim_id: instruction_guided_style_editing_models_trained_on_narrow_templated_instruction
    role: complicates
    claim: Instruction-guided style editing models trained on narrow, templated-instruction datasets fail catastrophically
      when evaluated on more diverse, fine-grained instruction distributions.
    source: §4.2, Table 2
    evidence: The ESD-trained LlasaEdit model, when evaluated on the ISSE test set, produces a WER of 68.17, compared
      to 10.07 on its own in-domain ESD test set, indicating the model does not generalize beyond the coarse single-attribute
      instructions it was trained on *(§4.2, Table 2)*.
    confidence: high
    relevance: medium
  - claim_id: isolating_style_variation_from_speaker_identity_in_synthetically_generated_paired
    role: complicates
    claim: Isolating style variation from speaker identity in synthetically generated paired training data requires
      an explicit voice-conversion correction step, since expressive TTS models conditioned on style descriptions
      alone conflate style and timbre changes.
    source: §2.1
    evidence: EmoCapTTS-generated stylized speech differs from the anchor speech in timbre because the model lacks
      explicit speaker control; a separate voice conversion stage (Chatterbox) is needed to re-align target timbre
      to the anchor speaker before the pair can be used to define a style-only edit *(§2.1)*.
    confidence: high
    relevance: medium
  limitations:
  - The dataset and benchmark are limited to English, which the authors explicitly flag as constraining applicability
    to multilingual editing scenarios. The generated portion of ISSE (292 of 382 hours) is itself the product of
    a TTS+VC synthesis pipeline rather than real recordings, so any systematic biases or artifacts introduced by
    EmoCapTTS or Chatterbox could propagate into models trained on it; the quality-filtering thresholds (WER < 10,
    similarity > 0.5) are relatively loose and their effect on downstream editing fidelity is not separately ablated.
    The benchmark comparison is against a single alternative dataset (ESD) and a single model architecture (LlasaEdit);
    no comparison is made against other instruction-guided editing systems such as InstructSpeech, and no ablation
    isolates the individual contribution of instruction diversity versus raw data scale.
  caveats: []
- id: '2509.24650'
  published_date: "2025-09-29"
  entry_date: '2026-07-28'
  year: 2025
  venue: arXiv
  task:
  - TTS
  - VC
  architecture:
  - autoregressive-LM
  - diffusion
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: influential
  method_family:
  - autoregressive_codec_language_models
  - diffusion_latent_codec_generators
  claims:
  - claim_id: a_differentiable_scalar_quantization_bottleneck_applied_to_hidden_states_rather
    role: supports
    claim: A differentiable scalar-quantization bottleneck applied to hidden states, rather than used as a discrete
      prediction target, can induce semantic/acoustic task separation inside a continuous autoregressive TTS model
      without an external speech tokenizer.
    source: §4.3, Table 6
    evidence: Removing the FSQ bottleneck from an otherwise identical hierarchical architecture increases ZH-hard-case
      CER from 18.19% to 24.92%, while FSQ dimensionality shows a non-monotonic optimum around 128-256 dimensions
      rather than monotonic improvement with capacity.
    confidence: high
    relevance: medium
  - claim_id: explicitly_separating_acoustic_detail_recovery_into_a_dedicated_residual_module
    role: supports
    claim: Explicitly separating acoustic detail recovery into a dedicated residual module improves robustness on
      complex inputs beyond what a single semantic language model plus a diffusion decoder achieves.
    source: §4.4, Table 7
    evidence: Removing the RALM (TSLM output feeding the diffusion decoder directly, architecturally close to DiTAR)
      degrades EN-WER from 2.98% to 4.34% and ZH-hard-case CER from 18.19% to 25.0%; removing the historical acoustic
      embedding from the RALM input degrades results further.
    confidence: high
    relevance: medium
  - claim_id: learning_rate_schedule_design_not_just_architecture_materially_affects_zero
    role: complicates
    claim: Learning-rate schedule design, not just architecture, materially affects zero-shot speaker similarity
      in large-scale continuous TTS training.
    source: §4.5, Table 8
    evidence: A two-phase Warmup-Stable-Decay schedule's decay phase alone improves ZH-hard-case CER from 13.22%
      to 8.87% and SIM by 4.4 points over the stable-phase-only checkpoint on an otherwise identical model.
    confidence: high
    relevance: low
  - claim_id: classifier_free_guidance_strength_in_diffusion_based_tts_decoders_trades
    role: complicates
    claim: Classifier-free guidance strength in diffusion-based TTS decoders trades off intelligibility against
      speaker similarity non-monotonically, with both very low and very high guidance scales degrading both metrics
      simultaneously.
    source: §4.6, Table 9
    evidence: CFG scale 1.0 (no guidance) yields EN-WER 16.32% and SIM 55.1%, while scale 5.0 yields EN-WER 12.78%
      and SIM 60.7%; the optimum at scale 2.0 achieves EN-WER 1.85% and SIM 72.9%, with degradation on both sides
      of the optimum.
    confidence: high
    relevance: low
  - claim_id: removing_dependency_on_a_pre_trained_discrete_speech_tokenizer_does
    role: refines
    claim: Removing dependency on a pre-trained discrete speech tokenizer does not require sacrificing zero-shot
      voice cloning quality relative to discrete-token-based open-source TTS systems.
    source: §4.2, Table 3
    evidence: On SEED-TTS-EVAL, the fully continuous VoxCPM reports SIM of 72.9% (EN) and 77.2% (ZH), exceeding
      the discrete-token-based IndexTTS2 and CosyVoice2 baselines on the same benchmark.
    confidence: high
    relevance: high
  limitations:
  - Multilingual capability is limited to Chinese and English by construction; the paper explicitly reports uncertain
    generalization to other languages, and prosody/emotion control lacks any intuitive or precise user-facing conditioning
    mechanism. The causal audio VAE operates at 16kHz, which the authors acknowledge falls short of the 24kHz or
    44.1kHz sampling rates typically expected for high-fidelity applications. Baseline comparisons draw on official
    implementations or numbers reported in prior papers rather than a uniformly controlled re-evaluation, so cross-system
    rankings on tables that mix reproduced and self-reported numbers should be read cautiously. The training corpus
    (1.8M hours) is internal and not released, which limits independent reproduction of the full-scale result even
    though code and weights for the trained model are public.
  caveats: []
- id: '2509.25131'
  published_date: "2025-09-29"
  entry_date: '2026-07-28'
  year: 2025
  venue: arXiv
  task:
  - SCA
  - TTS
  architecture:
  - autoregressive-LM
  - flow-matching
  - hybrid
  relevance: medium
  evidence_role:
  - architecture_variant
  - acceleration_evidence
  current_role: active_evidence
  method_family:
  - autoregressive_codec_language_models
  - flow_matching_codec_decoders
  - hybrid_semantic_acoustic_tokenizers
  claims:
  - claim_id: chunking_text_into_aligned_segments_with_a_short_token_delay
    role: supports
    claim: Chunking text into aligned segments with a short token-delay before speech decoding reduces error accumulation
      in long-form autoregressive speech generation.
    source: §4.2, Table 6
    evidence: Removing chunk-based decoding raises Long-TTS-Eval error rates above those of concurrent long-form
      TTS baselines, and with it enabled MGM-Omni-TTS-2B achieves EN-hard WER 26.26 versus 42.48-98.61 for CosyVoice2,
      MOSS-TTSD-v0.5, and Higgs-Audio-v2.
    confidence: high
    relevance: medium
  - claim_id: multi_token_parallel_decoding_is_not_restricted_to_rvq_speech
    role: supports
    claim: Multi-token parallel decoding is not restricted to RVQ speech tokenizers and can be applied effectively
      to finite scalar quantization (FSQ) tokenizers.
    source: §4.2, Table 6
    evidence: Increasing parallel decoding size on the CosyVoice2 FSQ tokenizer maintains TTS quality on Seed-TTS-Eval
      while cutting inference RTF by roughly 3x at parallel size 4.
    confidence: high
    relevance: high
  - claim_id: increasing_the_parallel_decoding_size_trades_off_synthesis_error_rate
    role: complicates
    claim: Increasing the parallel decoding size trades off synthesis error rate against inference speed rather
      than improving both simultaneously.
    source: §4.2
    evidence: Larger parallel sizes in the ablation slightly raise audio error rate even as they substantially accelerate
      inference, leading the authors to select a parallel size of 4 as a balance point.
    confidence: high
    relevance: medium
  - claim_id: separating_multimodal_reasoning_from_speech_synthesis_into_distinct_model_components
    role: refines
    claim: Separating multimodal reasoning from speech synthesis into distinct model components can improve long-form
      audio understanding without sacrificing speech generation efficiency.
    source: §4.1.1, §4.1.3, Figure 5, Table 5b
    evidence: The dual-track brain-mouth design lets the MLLM handle needle-in-the-haystack audio inputs up to 4,500
      seconds while the SpeechLM independently achieves the lowest RTF among compared long-form TTS systems.
    confidence: high
    relevance: medium
  limitations:
  - 'The long-form evaluation itself is partly self-authored: Long-TTS-Eval is introduced by this paper, and while
    its construction and normalized-text scoring procedure are documented, results on it cannot yet be cross-checked
    against independent replications. The comparison in Table 5b is limited to three baseline systems, and the qualitative
    long-speech examples in the appendix (a classical Chinese poem and a code-switched English-Chinese poem) are
    illustrative rather than a systematic error analysis. The paper does not report results on emotion or prosody
    control, nor does it evaluate robustness to reference audio recorded in noisy or far-field conditions. The 32B
    MLLM variant''s long-form and vision-speech results are mixed relative to the 7B variant (e.g., lower TextVQA-Speech
    and EN-hard performance context is not directly reported for 32B TTS), suggesting scaling benefits are not uniform
    across all sub-tasks.'
  caveats: []
- id: '2509.26276'
  published_date: "2025-09-30"
  entry_date: '2026-07-28'
  year: 2025
  venue: arXiv
  task:
  - TTS
  - SCA
  architecture:
  - autoregressive-LM
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: active_evidence
  method_family:
  - autoregressive_codec_language_models
  claims:
  - claim_id: acoustic_consistency_in_autoregressive_speech_language_models_can_be_improved
    role: supports
    claim: Acoustic consistency in autoregressive speech language models can be improved through LM-side embedding
      initialization and auxiliary training objectives, independent of parameter count.
    source: §4, Table 1
    evidence: The 0.7B speech-only CAST model scores 90.8 speaker and 90.0 gender consistency on SALMON, exceeding
      the 7B SpiritLM (81.0/85.0) and 7B Twist (71.0/70.0), despite using an order of magnitude fewer parameters
      and the same evaluation protocol.
    confidence: high
    relevance: medium
  - claim_id: interleaving_text_and_speech_tokens_in_a_shared_vocabulary_language
    role: complicates
    claim: Interleaving text and speech tokens in a shared-vocabulary language model improves semantic and lexical
      grounding at the cost of acoustic consistency.
    source: §4, Tables 1-2
    evidence: The 1.0B interleaved model drops SALMON consistency by 5-9 points across every measured factor relative
      to the 1.0B speech-only model (e.g., speaker 83.5 vs. 90.0) while sWUGGY rises from 67.0 to 73.7 and semantic-acoustic
      alignment scores improve by 6-7.5 points.
    confidence: high
    relevance: medium
  - claim_id: auxiliary_objectives_that_require_the_model_to_plan_coarse_content
    role: supports
    claim: Auxiliary objectives that require the model to plan coarse content structure before predicting fine acoustic
      detail materially contribute to speaker- and gender-identity stability beyond what embedding initialization
      alone provides.
    source: §4, Table 4
    evidence: Removing the delayed coarse-label and next-code auxiliary losses drops speaker consistency from 90.8
      to 83.5 in the 0.7B model and from 90.0 to 83.0 in the 1.0B model, with background and room consistency declining
      more moderately.
    confidence: high
    relevance: medium
  - claim_id: initializing_speech_token_embeddings_from_self_supervised_acoustic_features_biases
    role: refines
    claim: Initializing speech-token embeddings from self-supervised acoustic features biases the resulting representation
      toward content structure at the expense of fine-grained prosodic detail.
    source: §4, Table 3
    evidence: Linear probes on frozen LM features show the SSL-initialized variant gains on average +16% relative
      accuracy on content-leaning tasks (ESC-50, US8K, RAVDESS) while losing about 7% relative accuracy on prosody-leaning
      tasks (VIVAE, EMOVO), compared to a randomly initialized baseline.
    confidence: high
    relevance: high
  limitations:
  - The paper evaluates consistency exclusively through likelihood-based pairwise preference scoring (SALMON) rather
    than through generated-sample quality or human listening tests; no MOS or subjective naturalness evaluation
    is reported, so it is untested whether the consistency gains translate into perceptibly better generations.
    The comparison against Twist, SpiritLM, LAST, and Flow-SLM baselines is confounded by differences in tokenizer,
    training data, and architecture beyond scale, which the paper acknowledges only implicitly by noting these systems
    "differ in tokenization, architecture, and training objectives." The 1.0B ablation shows a single anomalous
    result for room consistency that the authors attribute to run instability rather than investigate further. All
    experiments use English-only, largely read-speech and audiobook data (LibriLight, People's Speech), leaving
    open whether the training recipe generalizes to conversational, multilingual, or noisier real-world speech.
  caveats: []
- id: '2510.00264'
  published_date: "2025-09-30"
  entry_date: '2026-07-28'
  year: 2025
  venue: arXiv
  task:
  - codec
  - evaluation
  architecture:
  - GAN
  relevance: high
  evidence_role:
  - core_evidence
  - acceleration_evidence
  - evaluation_caution
  current_role: minor
  method_family:
  - gan_neural_codecs_and_decoders
  claims:
  - claim_id: residual_vector_quantized_convolutional_codecs_trained_end_to_end_with
    role: supports
    claim: Residual-vector-quantized convolutional codecs trained end-to-end with adversarial and multi-scale mel-spectrogram
      losses can be configured to satisfy joint sub-1 kbps-to-6 kbps bitrate, sub-50 ms latency, and sub-3000 MFLOPS
      compute budgets simultaneously.
    source: §3.1, §3.2, Table 2, Table 3
    evidence: The Track 1 baseline totals 691.35 MFLOPS and 20 ms latency (within a 700 MFLOPS / 30 ms budget) and
      the Track 2 baseline totals 2546.2 MFLOPS and 40 ms latency (within a 2600 MFLOPS / 50 ms budget), both operating
      across the full 1-6 kbps range via quantizer dropout.
    confidence: high
    relevance: high
  - claim_id: joint_denoising_and_dereverberation_enhancement_codecs_require_substantially_more_compute
    role: complicates
    claim: Joint denoising-and-dereverberation ("enhancement") codecs require substantially more compute than transparency-only
      codecs even when the coding architecture and bitrate range are otherwise matched.
    source: §3.1, §3.2, Table 2, Table 3
    evidence: The Track 2 (enhancement) baseline requires 2546.2 MFLOPS overall versus 691.35 MFLOPS for the Track
      1 (transparency-only) baseline, roughly 3.7x more compute, despite sharing the same 6-layer, 1,024-codeword
      RVQ design and 1-6 kbps target range.
    confidence: high
    relevance: high
  - claim_id: automatic_objective_codec_quality_metrics_can_disagree_substantially_with_one
    role: complicates
    claim: Automatic objective codec-quality metrics can disagree substantially with one another at low bitrates,
      complicating single-metric benchmarking of low-resource codecs.
    source: §4, Table 4, Table 5
    evidence: At 1 kbps under clean conditions, SCOREQ_Ref reports 1.15/1.01 (Track 1/2, near floor) while Audiobox_AE_CE
      reports 3.9/3.96 for the same systems and conditions in the same evaluation table.
    confidence: high
    relevance: high
  - claim_id: filtering_public_speech_corpora_by_estimated_signal_quality_snr_reverberation
    role: supports
    claim: Filtering public speech corpora by estimated signal quality (SNR, reverberation, bandwidth) before training
      a low-resource codec substantially reduces the usable data volume relative to the source corpora.
    source: §2, Table 1
    evidence: Quality- and diversity-based curation across LibriTTS, VCTK, EARS, Librivox (DNS5), MLS, and GLOBE
      retains 702.7 of 1,641.7 total hours (42.8% overall retention), ranging from 24.2% (LibriTTS) to 100% (EARS)
      by dataset.
    confidence: high
    relevance: high
  limitations:
  - All reported results (Table 4, Table 5) are automatic objective metrics on the challenge's open development
    test set; the paper reports no human listening-test (MOS/subjective) results, and the blind test set used for
    official challenge scoring had not yet been released at the time of writing, so these baseline numbers may not
    reflect final challenge-condition performance.
  - Checkpoint selection relies solely on validation multi-scale mel-spectrogram loss; the authors themselves note
    this prioritizes implementation simplicity over correlation with subjective quality, and suggest combining objective
    metrics that correlate better with listening tests as a likely improvement. The Track 2 validation set uses
    on-the-fly (rather than fixed, offline) noise and reverberation augmentation, which the authors acknowledge
    may increase variance in validation loss estimates. No ablations isolate the individual contributions of the
    loss terms, the data curation pipeline, or the quantizer-dropout bitrate-scalability scheme.
  caveats: []
claim_clusters:
- id: learned_codec_low_bitrate_quality
  claim: Learned neural codecs can preserve intelligible and perceptually useful speech at bitrates far below conventional
    waveform representations.
  status: strongly_supported
  confidence: high
  supporting_papers:
  - '2104.00355'
  - '2506.23325'
  - '2508.02849'
  - interspeech-2025-0355
  - interspeech-2025-1440
  - '2509.15462'
  contradicting_papers: []
  refining_papers:
  - '2210.13438'
  - '2508.20660'
  - '2509.17006'
  caveats:
  - Reported bitrate and quality figures are rarely measured on a shared corpus with matched listeners and decoder
    capacity.
  last_reviewed: '2026-07-28'
- id: residual_quantization_supports_scalable_bitrate
  claim: Residual vector quantization provides a hierarchical and scalable representation whose active codebooks
    can trade bitrate against reconstruction detail.
  status: strongly_supported
  confidence: high
  supporting_papers:
  - '2210.13438'
  - '2301.02111'
  - '2305.02765'
  - '2305.09636'
  - '2308.16692'
  - '2310.00704'
  - '2402.13236'
  - '2406.07855'
  - '2411.01156'
  - '2411.19842'
  - '2502.04128'
  - '2502.17239'
  - iclr-2025-868masI331
  - 2025.findings-naacl.184
  - 2025.naacl-srw.6
  - '2505.13000'
  - '2506.10274'
  - '2507.12197'
  - 2025.acl-long.1498
  - 2025.acl-long.654
  - 2025.acl-long.937
  - 2025.findings-acl.1051
  - interspeech-2025-0115
  - interspeech-2025-0319
  - interspeech-2025-0468
  - interspeech-2025-0669
  - interspeech-2025-1084
  - interspeech-2025-1639
  - interspeech-2025-2075
  - '2508.20660'
  - '2509.02020'
  - '2509.02244'
  - '2509.09550'
  - '2509.13068'
  - '2509.16195'
  - '2509.17006'
  - '2509.19025'
  - '2509.19186'
  - '2509.25131'
  contradicting_papers: []
  refining_papers:
  - '2210.13438'
  - 2025.naacl-srw.6
  - '2508.20660'
  - '2509.09550'
  - '2509.14882'
  - '2509.19186'
  caveats:
  - RVQ layer semantics differ across objectives and cannot be assumed to order cleanly from linguistic to acoustic
    information.
  last_reviewed: '2026-07-28'
- id: semantic_acoustic_hierarchy_balances_coherence_fidelity
  claim: Separating semantic and acoustic token layers helps balance linguistic coherence against fine-grained waveform
    fidelity.
  status: strongly_supported
  confidence: high
  supporting_papers:
  - '2209.03143'
  - '2301.11325'
  - '2410.03751'
  - '2411.01156'
  - '2502.06490'
  - '2502.17239'
  - '2505.13000'
  - '2506.23325'
  - 2025.acl-long.1043
  - 2025.acl-long.1498
  - 2025.acl-long.682
  - '2508.02849'
  - '2508.04141'
  - '2508.19205'
  - '2508.20660'
  - '2509.09174'
  - '2509.17006'
  - '2509.24650'
  contradicting_papers: []
  refining_papers:
  - '2410.03751'
  - '2502.06490'
  - interspeech-2025-1776
  - '2509.26276'
  caveats:
  - The boundary between semantic and acoustic information is porous, especially for prosody and speaker identity.
  last_reviewed: '2026-07-28'
- id: ssl_guidance_improves_codec_semantics
  claim: Self-supervised speech representations can guide codec quantization toward tokens that carry stronger linguistic
    content for downstream generation.
  status: strongly_supported
  confidence: high
  supporting_papers:
  - '2104.00355'
  - '2209.03143'
  - '2402.08093'
  - '2402.13236'
  - '2505.13000'
  - '2506.10274'
  - 2025.acl-long.937
  - '2508.02849'
  - '2508.04141'
  - '2508.11224'
  - interspeech-2025-0246
  - interspeech-2025-0468
  - '2509.00503'
  - '2509.09201'
  - '2509.11425'
  - '2509.21968'
  - '2509.22167'
  contradicting_papers: []
  refining_papers:
  - '2506.23325'
  - 2025.acl-long.937
  - interspeech-2025-1106
  - '2509.14882'
  - '2509.22167'
  caveats:
  - Stronger linguistic alignment may remove acoustic detail needed for expressive or speaker-faithful synthesis.
  last_reviewed: '2026-07-28'
- id: codec_tokens_enable_language_model_generation
  claim: Discrete codec tokens make speech generation compatible with language-model training and in-context conditioning.
  status: strongly_supported
  confidence: high
  supporting_papers:
  - '2301.02111'
  - '2402.08093'
  - '2403.03100'
  - '2404.03204'
  - '2408.16532'
  - '2502.17239'
  - iclr-2025-tQ1PmLfPBL
  - 2025.findings-naacl.184
  - '2507.16632'
  - 2025.acl-long.1498
  - 2025.acl-long.654
  - interspeech-2025-0246
  - '2509.11425'
  contradicting_papers: []
  refining_papers:
  - '2402.13236'
  - '2504.08528'
  - iclr-2025-868masI331
  - '2509.26276'
  caveats:
  - Language-model compatibility does not by itself establish efficient decoding or faithful acoustic reconstruction.
  last_reviewed: '2026-07-28'
- id: quantizer_structure_affects_downstream_modelability
  claim: Quantizer and codebook structure materially affects both reconstruction quality and the difficulty of downstream
    token prediction.
  status: strongly_supported
  confidence: high
  supporting_papers:
  - '2210.13438'
  - '2301.02111'
  - '2305.02765'
  - '2308.16692'
  - '2406.04904'
  - '2408.16532'
  - '2411.19842'
  - '2502.06490'
  - '2502.11946'
  - '2504.10344'
  - iclr-2025-cuFzE8Jlvb
  - '2505.13000'
  - '2507.18897'
  - 2025.acl-long.654
  - 2025.acl-long.937
  - '2508.05207'
  - interspeech-2025-0246
  - interspeech-2025-1020
  - interspeech-2025-1106
  - interspeech-2025-2726
  - '2508.19205'
  - '2509.02244'
  - '2509.17765'
  - '2509.19592'
  - '2509.22062'
  contradicting_papers: []
  refining_papers:
  - '2507.12197'
  caveats:
  - Many comparisons change codebook size, token rate, decoder, and training loss simultaneously.
  last_reviewed: '2026-07-28'
- id: codebook_utilization_requires_explicit_control
  claim: Codebook collapse and uneven token utilization are recurring bottlenecks that require explicit training
    or pruning strategies.
  status: strongly_supported
  confidence: high
  supporting_papers:
  - '2308.16692'
  - '2406.04904'
  - '2408.16532'
  - '2409.05377'
  - '2411.01156'
  - '2411.19842'
  - '2412.10117'
  - '2502.05512'
  - '2507.18897'
  - '2508.02849'
  contradicting_papers: []
  refining_papers: []
  caveats:
  - Improved utilization does not necessarily translate into perceptual or downstream gains.
  last_reviewed: '2026-07-28'
- id: adversarial_objectives_improve_perceptual_reconstruction
  claim: Adversarial and multi-scale spectral objectives improve perceptual codec reconstruction under aggressive
    compression.
  status: strongly_supported
  confidence: high
  supporting_papers:
  - '2210.13438'
  - '2305.07243'
  - '2306.00814'
  - '2410.00037'
  - '2502.05512'
  - iclr-2025-uxDFlPGRLX
  - '2509.02244'
  contradicting_papers: []
  refining_papers:
  - '2507.07799'
  - '2507.18897'
  caveats:
  - Adversarial gains are sensitive to discriminator design and may be poorly reflected by standard objective metrics.
  last_reviewed: '2026-07-28'
- id: decoder_capacity_compensates_for_compression
  claim: Stronger codec decoders and longer receptive context can recover acoustic detail that compact token streams
    do not represent explicitly.
  status: strongly_supported
  confidence: high
  supporting_papers:
  - '2306.00814'
  - '2406.07855'
  - '2408.16532'
  - iclr-2025-uxDFlPGRLX
  - 2025.acl-long.654
  - interspeech-2025-2726
  - '2508.16790'
  contradicting_papers: []
  refining_papers: []
  caveats:
  - Decoder compensation can hallucinate plausible detail rather than faithfully preserve the source signal.
  last_reviewed: '2026-07-28'
- id: speaker_prosody_entanglement_is_persistent
  claim: Codec tokens frequently entangle linguistic content with speaker, pitch, and prosodic information, complicating
    controllable generation.
  status: strongly_supported
  confidence: high
  supporting_papers:
  - '2104.00355'
  - '2301.02111'
  - '2402.08093'
  - '2407.08551'
  - '2502.04128'
  - '2503.01710'
  - 2025.findings-naacl.184
  - '2507.09070'
  - '2505.15670'
  - 2025.acl-long.346
  - 2025.acl-long.654
  - '2508.04141'
  - interspeech-2025-0319
  - interspeech-2025-0347
  - interspeech-2025-0575
  - interspeech-2025-0874
  - interspeech-2025-1106
  - interspeech-2025-1440
  - '2412.16846'
  contradicting_papers: []
  refining_papers:
  - '2402.13236'
  - '2507.21138'
  - '2508.08399'
  - interspeech-2025-0115
  - interspeech-2025-1538
  caveats:
  - The desired degree of invariance depends on whether the downstream task is TTS, voice conversion, dialogue,
    or reconstruction.
  last_reviewed: '2026-07-28'
- id: continuous_discrete_representation_tradeoff
  claim: Continuous and discrete speech representations trade modeling convenience and compression against retained
    acoustic detail and generation stability.
  status: strongly_supported
  confidence: high
  supporting_papers:
  - '2305.07243'
  - '2410.11190'
  - '2502.18924'
  - iclr-2025-cuFzE8Jlvb
  - 2025.findings-naacl.184
  - '2506.10274'
  - 2025.acl-long.65
  - '2507.22746'
  - '2509.24650'
  contradicting_papers: []
  refining_papers:
  - '2508.08399'
  caveats:
  - Comparisons are often confounded by different generators, sampling procedures, and representation rates.
  last_reviewed: '2026-07-28'
- id: token_rate_drives_sequence_cost
  claim: Codec frame rate and token hierarchy directly govern language-model sequence length, inference cost, and
    streaming latency.
  status: strongly_supported
  confidence: high
  supporting_papers:
  - '2402.08093'
  - '2502.04128'
  - '2504.10344'
  - iclr-2025-868masI331
  - iclr-2025-dGSOn7sdWg
  - '2505.13000'
  - interspeech-2025-0115
  - interspeech-2025-0468
  - interspeech-2025-1289
  - '2509.19592'
  contradicting_papers: []
  refining_papers:
  - iclr-2025-868masI331
  - interspeech-2025-1106
  - '2509.21968'
  caveats:
  - Lower token rates can hide losses in transient detail, timing precision, or speaker similarity.
  last_reviewed: '2026-07-28'
- id: parallel_latent_generation_improves_robustness
  claim: Diffusion, flow-matching, and masked non-autoregressive generation over codec latents can reduce sequential
    decoding cost and autoregressive failure modes.
  status: strongly_supported
  confidence: high
  supporting_papers:
  - '2304.09116'
  - '2305.07243'
  - '2403.03100'
  - iclr-2025-uxDFlPGRLX
  - 2025.acl-long.65
  - interspeech-2025-1776
  - '2509.17143'
  - '2509.17765'
  contradicting_papers: []
  refining_papers: []
  caveats:
  - Speed and robustness claims depend strongly on step count, decoder implementation, and matched compute.
  last_reviewed: '2026-07-28'
- id: codec_bottlenecks_shape_intelligibility
  claim: Codec information loss and token prediction errors propagate into intelligibility, repetition, omission,
    and long-form robustness failures.
  status: strongly_supported
  confidence: high
  supporting_papers:
  - '2304.09116'
  - '2406.05370'
  - '2505.02625'
  - 2025.acl-long.654
  - 2025.findings-acl.115
  - interspeech-2025-0246
  - interspeech-2025-1639
  - interspeech-2025-1641
  - '2509.15462'
  - '2509.25131'
  contradicting_papers: []
  refining_papers:
  - '2509.15462'
  caveats:
  - Observed errors may arise from the generator or alignment model rather than the codec alone.
  last_reviewed: '2026-07-28'
- id: codec_representations_enable_multilingual_transfer
  claim: Shared codec and semantic representations support multilingual and cross-lingual speech generation, but
    language coverage can trade against speaker fidelity.
  status: strongly_supported
  confidence: high
  supporting_papers:
  - '2301.02111'
  - '2303.03926'
  - '2306.12925'
  - '2310.00704'
  - '2402.01912'
  - '2404.03204'
  - '2406.04904'
  - '2411.01156'
  - '2411.17607'
  - '2504.10344'
  - 2025.naacl-srw.6
  - '2505.02625'
  - 2025.findings-acl.1051
  - 2025.ccl-1.80
  - interspeech-2025-1641
  - '2509.05863'
  - '2509.11425'
  contradicting_papers: []
  refining_papers:
  - '2402.13236'
  caveats:
  - Evidence is uneven across languages and often dominated by high-resource evaluation conditions.
  last_reviewed: '2026-07-28'
- id: codec_tokens_support_unified_speech_systems
  claim: Codec token interfaces enable unified models to combine speech understanding, generation, translation,
    and dialogue within one sequence framework.
  status: strongly_supported
  confidence: high
  supporting_papers:
  - 2025.findings-acl.115
  - '2508.19205'
  - '2509.11425'
  contradicting_papers: []
  refining_papers: []
  caveats:
  - Unified task coverage often comes with weaker task-specific quality or relies on heterogeneous proprietary data.
  last_reviewed: '2026-07-28'
- id: codec_evaluation_lacks_metric_consensus
  claim: Neural codec evaluation lacks a single reliable proxy for listener-perceived quality and downstream generation
    utility.
  status: strongly_supported
  confidence: high
  supporting_papers:
  - '2305.02765'
  - '2308.16692'
  - '2410.00037'
  - iclr-2025-uxDFlPGRLX
  - '2508.02849'
  - '2508.08961'
  - interspeech-2025-0984
  - '2508.16790'
  - '2509.02244'
  - '2509.24457'
  contradicting_papers: []
  refining_papers:
  - interspeech-2025-1538
  - '2509.05863'
  - '2509.22167'
  - '2510.00264'
  caveats:
  - Objective, subjective, and downstream-task evaluations measure different failure modes and should not be collapsed
    into one ranking.
  last_reviewed: '2026-07-28'
method_families:
- id: autoregressive_codec_language_models
  name: Autoregressive codec language models
  summary: Decoder-only language models generate discrete acoustic or semantic codec streams, often using hierarchical
    AR+NAR prediction to control sequence length and reconstruction detail.
  papers:
  - '1609.03499'
  - '2209.03143'
  - '2301.02111'
  - '2301.11325'
  - '2303.03926'
  - '2305.07243'
  - '2305.09636'
  - '2305.11000'
  - '2306.12925'
  - '2310.00704'
  - '2401.07333'
  - '2402.01912'
  - '2402.08093'
  - '2403.16973'
  - '2404.03204'
  - '2406.00654'
  - '2406.04904'
  - '2406.05370'
  - '2406.07855'
  - '2407.05407'
  - '2407.08551'
  - '2408.16725'
  - '2409.00750'
  - '2409.03283'
  - '2410.00037'
  - '2410.11190'
  - '2410.17799'
  - '2411.00774'
  - '2411.01156'
  - '2411.13577'
  - '2411.17607'
  - '2412.02612'
  - '2412.10117'
  - '2412.15649'
  - '2502.04128'
  - '2502.05512'
  - '2502.11946'
  - '2502.17239'
  - '2503.01710'
  - '2503.14345'
  - '2504.10344'
  - iclr-2025-868masI331
  - iclr-2025-cuFzE8Jlvb
  - iclr-2025-dGSOn7sdWg
  - 2025.findings-naacl.184
  - 2025.naacl-demo.12
  - 2025.naacl-srw.6
  - '2505.02625'
  - '2505.13000'
  - '2507.02380'
  - '2507.01348'
  - '2412.18603'
  - '2507.07799'
  - '2507.09070'
  - '2507.12197'
  - '2507.16632'
  - '2507.21138'
  - '2505.15670'
  - 2025.acl-long.1498
  - 2025.acl-long.65
  - 2025.acl-long.682
  - 2025.acl-long.817
  - 2025.acl-long.912
  - 2025.acl-short.81
  - 2025.findings-acl.1051
  - 2025.findings-acl.115
  - 2025.findings-ijcnlp.49
  - 2025.ccl-1.80
  - '2504.10352'
  - '2508.04141'
  - '2508.04585'
  - '2508.06262'
  - '2508.07302'
  - '2508.08961'
  - '2504.12867'
  - '2508.11326'
  - interspeech-2025-0253
  - interspeech-2025-0310
  - interspeech-2025-0319
  - interspeech-2025-0464
  - interspeech-2025-0874
  - interspeech-2025-0989
  - interspeech-2025-1084
  - interspeech-2025-1538
  - interspeech-2025-1641
  - interspeech-2025-1993
  - interspeech-2025-2447
  - interspeech-2025-2564
  - '2508.15827'
  - '2508.16790'
  - '2509.00503'
  - '2509.02020'
  - '2509.05863'
  - '2509.09174'
  - '2509.11425'
  - '2509.13068'
  - '2412.16846'
  - '2509.14882'
  - '2509.15462'
  - '2509.15969'
  - '2509.17143'
  - '2509.17765'
  - '2501.04561'
  - '2509.19592'
  - '2509.20802'
  - '2509.22062'
  - '2509.24570'
  - '2509.24650'
  - '2509.25131'
  - '2509.26276'
  open_questions:
  - Which token hierarchy best preserves intelligibility and speaker identity without making autoregressive decoding
    prohibitively long?
- id: gan_neural_codecs_and_decoders
  name: GAN neural codecs and adversarial decoders
  summary: Convolutional codec encoders and decoders use adversarial and spectral objectives to preserve perceptual
    detail under aggressive quantization.
  papers:
  - '2104.00355'
  - '2210.13438'
  - '2305.02765'
  - '2306.00814'
  - '2308.16692'
  - '2406.04904'
  - '2408.16532'
  - '2409.05377'
  - '2411.01156'
  - '2411.18803'
  - '2502.05512'
  - iclr-2025-uxDFlPGRLX
  - '2507.01348'
  - '2506.23325'
  - '2507.12197'
  - '2507.18897'
  - 2025.acl-long.1498
  - 2025.acl-long.654
  - 2025.acl-long.682
  - '2508.05207'
  - interspeech-2025-0196
  - interspeech-2025-0347
  - interspeech-2025-0468
  - interspeech-2025-1106
  - interspeech-2025-1639
  - interspeech-2025-2726
  - '2509.02244'
  - '2509.04685'
  - '2509.09201'
  - '2509.09550'
  - '2509.11425'
  - '2509.13670'
  - '2509.17006'
  - '2509.18823'
  - '2509.19025'
  - '2509.19186'
  - '2509.21968'
  - '2510.00264'
  open_questions:
  - How can adversarially trained codecs be compared reliably when objective quality metrics diverge from listener
    judgments?
- id: vae_vector_quantized_codecs
  name: VAE and vector-quantized codecs
  summary: VAE-derived systems compress speech through learned latent bottlenecks, commonly combining vector quantization
    with reconstruction and regularization objectives.
  papers:
  - '2104.00355'
  - '2304.09116'
  - '2305.02765'
  - '2305.07243'
  - '2408.16532'
  - '2409.05377'
  - '2411.19842'
  - '2502.18924'
  - '2504.10344'
  - iclr-2025-cuFzE8Jlvb
  - 2025.findings-naacl.184
  - '2507.01348'
  - 2025.acl-long.65
  - 2025.acl-long.654
  - 2025.acl-long.937
  - '2508.02849'
  - '2508.08399'
  - interspeech-2025-0196
  - interspeech-2025-0347
  - interspeech-2025-0468
  - interspeech-2025-0575
  - interspeech-2025-1020
  - interspeech-2025-1106
  - interspeech-2025-1440
  - '2509.02244'
  - '2509.13068'
  - '2412.16846'
  - '2509.21968'
  - '2509.22167'
  open_questions:
  - What allocation of latent capacity best balances reconstruction fidelity, codebook utilization, and downstream
    modelability?
- id: hybrid_semantic_acoustic_tokenizers
  name: Hybrid semantic–acoustic tokenizers
  summary: Hybrid systems combine self-supervised semantic representations with acoustic codec detail to separate
    linguistic coherence from waveform fidelity.
  papers:
  - '2310.00704'
  - '2403.03100'
  - '2407.05407'
  - '2409.03283'
  - '2409.06666'
  - '2410.00037'
  - '2411.00774'
  - '2411.13577'
  - '2412.10117'
  - '2502.11946'
  - '2502.17239'
  - 2025.naacl-srw.6
  - '2505.02625'
  - '2506.23325'
  - '2507.16632'
  - 2025.acl-long.1043
  - 2025.acl-long.346
  - 2025.acl-long.65
  - 2025.acl-long.682
  - 2025.acl-long.912
  - 2025.ccl-1.80
  - '2507.22746'
  - '2504.10352'
  - '2508.04141'
  - interspeech-2025-0253
  - interspeech-2025-0669
  - interspeech-2025-0874
  - interspeech-2025-1289
  - interspeech-2025-1776
  - interspeech-2025-2726
  - '2508.19205'
  - '2509.09550'
  - '2509.16195'
  - '2509.17765'
  - '2501.04561'
  - '2509.19883'
  - '2509.25131'
  open_questions:
  - Can semantic and acoustic information be separated cleanly without discarding prosody, speaker identity, or
    non-verbal events?
- id: flow_matching_codec_decoders
  name: Flow-matching codec decoders
  summary: Flow-matching models reconstruct waveforms or continuous acoustic latents from compact semantic or codec
    conditioning with parallel generation.
  papers:
  - '2312.15821'
  - '2407.05407'
  - '2409.03283'
  - '2411.17607'
  - '2412.02612'
  - '2412.10117'
  - '2412.15649'
  - '2502.11946'
  - '2502.17239'
  - '2503.14345'
  - iclr-2025-tQ1PmLfPBL
  - iclr-2025-uxDFlPGRLX
  - 2025.findings-naacl.184
  - '2505.02625'
  - '2507.02380'
  - '2507.09070'
  - '2507.16632'
  - 2025.acl-long.1043
  - 2025.acl-long.87
  - 2025.acl-long.912
  - '2508.04585'
  - '2508.07302'
  - '2504.12867'
  - '2509.09631'
  - '2509.15462'
  - '2509.25131'
  open_questions:
  - When does continuous flow decoding outperform direct discrete-code prediction under matched latency and model
    capacity?
- id: diffusion_latent_codec_generators
  name: Diffusion latent-codec generators
  summary: Diffusion systems model continuous codec latents or masked discrete representations to trade iterative
    decoding for stable non-autoregressive generation.
  papers:
  - '2304.09116'
  - '2305.07243'
  - '2403.03100'
  - '2502.18924'
  - iclr-2025-hQvX9MBowC
  - '2508.11326'
  - '2508.16790'
  - '2509.24650'
  open_questions:
  - How low can diffusion step counts fall before codec reconstruction and prosodic quality degrade materially?
- id: transformer_encoder_decoder_tokenizers
  name: Transformer encoder–decoder tokenizers
  summary: Transformer encoder–decoder systems learn contextual speech tokens or reconstruct acoustic representations
    with explicit bidirectional encoding and conditioned decoding.
  papers:
  - '2411.13577'
  - '2411.19842'
  - iclr-2025-hQvX9MBowC
  - 2025.naacl-long.591
  - 2025.acl-long.346
  - 2025.acl-short.81
  - interspeech-2025-0115
  - interspeech-2025-0246
  - interspeech-2025-0506
  - '2508.15931'
  - '2508.16790'
  - '2509.00503'
  - '2509.09631'
  open_questions:
  - Do contextual tokenizers retain their downstream advantage when codec bitrate, data scale, and decoder capacity
    are controlled?
reassessment_queue:
- id: continuous_discrete_representation_tradeoff
  type: claim_status
  reason: The comparison spans different generators, rates, and sampling procedures.
  trigger: A matched study compares continuous and discrete representations with the same generator, data, rate,
    and evaluation protocol.
  due: 2026-10
  current_assessment: strongly_supported
  watch_for:
  - Matched representation ablations
  - Long-form robustness comparisons
- id: codec_evaluation_lacks_metric_consensus
  type: benchmark_validity
  reason: Metric reliability varies across reconstruction, TTS, voice conversion, and dialogue use cases.
  trigger: A shared multi-task codec benchmark validates objective metrics against controlled listener studies.
  due: 2026-10
  current_assessment: strongly_supported
  watch_for:
  - Cross-dataset metric correlations
  - Listener-calibrated codec benchmarks
- id: hybrid_semantic_acoustic_tokenizers
  type: method_family
  reason: This broad family combines several incompatible placements of semantic supervision.
  trigger: Enough matched systems exist to split pre-quantization guidance, first-layer distillation, and parallel
    semantic streams.
  due: 2026-10
  current_assessment: active_evidence
  watch_for:
  - Matched semantic-integration ablations
  - Layer-wise information probes
- id: learned_codec_low_bitrate_quality
  type: claim_status
  reason: Low-bitrate quality claims are strong but rarely use matched datasets and listening protocols.
  trigger: Independent reproduction compares neural and conventional codecs at matched rates on shared speech data.
  due: 2026-10
  current_assessment: strongly_supported
  watch_for:
  - Independent codec reproductions
  - Rate-controlled listening tests
- id: speaker_prosody_entanglement_is_persistent
  type: claim_status
  reason: Invariance desirable for semantic modeling may be harmful for expressive reconstruction.
  trigger: Task-conditioned evaluations quantify the optimal information split for TTS, VC, and dialogue separately.
  due: 2026-10
  current_assessment: strongly_supported
  watch_for:
  - Task-specific disentanglement studies
  - Prosody and identity leakage probes
open_questions:
- What codec representation rate minimizes sequence cost without losing transient, prosodic, and speaker information
  needed for high-fidelity generation?
- How should semantic and acoustic information be distributed across codebooks for TTS, voice conversion, and spoken
  dialogue respectively?
- Which objective and subjective measures jointly predict both reconstruction quality and downstream generation
  performance?
- Can a single universal codec remain competitive across speech, music, environmental audio, multilingual speech,
  and non-verbal vocalization?
- How much of codec-language-model robustness is determined by the tokenizer versus the generator and alignment
  mechanism?
- Can continuous or hybrid representations retain their quality advantage while matching the storage, streaming,
  and language-model compatibility of discrete tokens?
trend_notes:
- The field moved from waveform-level autoregression toward discrete codec language modeling after 2022.
- Residual quantization remains common, but 2024–2025 work increasingly tests single-codebook and semantically guided
  alternatives.
- Semantic–acoustic hybrid tokenizers became a major design axis as unified speech models expanded beyond TTS.
- Flow-matching and diffusion decoders increasingly reconstruct audio from compact semantic or codec representations
  without fully autoregressive acoustic decoding.
- Codec evaluation is shifting from reconstruction-only scores toward downstream intelligibility, speaker similarity,
  dialogue capability, and listener studies.
