concept: self-supervised-speech
last_updated: '2026-07-29'
paper_count: 146
papers:
- id: '2104.00355'
  published_date: "2021-04-01"
  entry_date: '2026-07-29'
  year: 2021
  venue: arXiv
  task:
  - TTS
  - VC
  architecture:
  - GAN
  - VAE
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  - evaluation_caution
  current_role: influential
  method_family:
  - gan_decoders_for_ssl_units
  - vae_and_quantized_ssl_representations
  claims:
  - claim_id: ssl_content_representations_that_are_well_disentangled_from_speaker_identity
    role: supports
    claim: SSL content representations that are well-disentangled from speaker identity also exhibit stronger voice
      conversion performance, while representations that entangle speaker information perform worse at conversion
      but better at pitch reconstruction.
    source: §4, Table 1, Table 2
    evidence: SSL content representations that are well-disentangled from speaker identity also exhibit stronger
      voice conversion performance, while representations that entangle speaker information perform worse at conversion
      but better at pitch reconstruction.
    confidence: high
    relevance: high
  - claim_id: discrete_speech_units_learned_by_ssl_models_can_form_the
    role: supports
    claim: Discrete speech units learned by SSL models can form the basis of an ultra-low-bitrate speech codec that
      outperforms classical parametric codecs in subjective quality.
    source: §4, Figure 2
    evidence: Discrete speech units learned by SSL models can form the basis of an ultra-low-bitrate speech codec
      that outperforms classical parametric codecs in subjective quality.
    confidence: high
    relevance: high
  - claim_id: among_self_supervised_content_encoders_hubert_units_carry_less_speaker
    role: supports
    claim: Among self-supervised content encoders, HuBERT units carry less speaker and pitch information than VQ-VAE
      units, making them better suited for downstream controllable synthesis.
    source: §4, Table 2
    evidence: Among self-supervised content encoders, HuBERT units carry less speaker and pitch information than
      VQ-VAE units, making them better suited for downstream controllable synthesis.
    confidence: high
    relevance: high
  - claim_id: pitch_and_speaker_identity_can_be_independently_conditioned_in_a
    role: supports
    claim: Pitch and speaker identity can be independently conditioned in a GAN vocoder through separate discrete
      token streams, enabling controllable F0 manipulation without retraining.
    source: §3, §4
    evidence: Pitch and speaker identity can be independently conditioned in a GAN vocoder through separate discrete
      token streams, enabling controllable F0 manipulation without retraining.
    confidence: high
    relevance: medium
  limitations:
  - The codec evaluation uses only 20 utterances from 5 VCTK speakers, all unseen during training but from the same
    corpus. Generalization to out-of-domain speech (conversational, noisy, or non-English) is untested.
  - The resynthesis MOS scores remain well below ground truth on both LJSpeech (3.66 vs. 4.33) and VCTK (3.41 vs.
    4.08), indicating a quality gap the system does not close. Disentanglement is evaluated indirectly through proxy
    metrics (EER, VDE, FFE) rather than a direct information-theoretic measure. The speaker encoder requires speaker
    embeddings from training-set speakers for the lookup-table variant; the d-vector approach generalizes but relies
    on a separately trained verification model. No ablation isolates the contribution of the F0 conditioning stream
    to final MOS. The MUSHRA scores in Figure 2 are visual only, making exact numerical comparison to baselines
    difficult to reproduce from the paper text alone.
  caveats: []
- id: '2204.02152'
  published_date: "2022-04-05"
  entry_date: '2026-07-29'
  year: 2022
  venue: arXiv
  task:
  - evaluation
  architecture: []
  relevance: medium
  evidence_role:
  - architecture_variant
  - evaluation_caution
  current_role: influential
  method_family: []
  claims:
  - claim_id: listener_dependent_modeling_substantially_improves_mos_prediction_accuracy_particularly_for
    role: complicates
    claim: Listener-dependent modeling substantially improves MOS prediction accuracy, particularly for out-of-domain
      data with limited labeled samples.
    source: §3.1.3, Table 2b, Table 3b
    evidence: Listener-dependent modeling substantially improves MOS prediction accuracy, particularly for out-of-domain
      data with limited labeled samples.
    confidence: high
    relevance: medium
  - claim_id: contrastive_pairwise_ranking_losses_improve_spearman_correlation_in_mos_prediction
    role: supports
    claim: Contrastive pairwise ranking losses improve Spearman correlation in MOS prediction more than standard
      regression losses alone.
    source: §3.1.2, §4.3, Table 2
    evidence: Contrastive pairwise ranking losses improve Spearman correlation in MOS prediction more than standard
      regression losses alone.
    confidence: high
    relevance: medium
  - claim_id: ssl_based_features_from_multiple_heterogeneous_pretrained_models_wav2vec_2
    role: supports
    claim: SSL-based features from multiple heterogeneous pretrained models (wav2vec 2.0, HuBERT, WavLM) combined
      via ensemble stacking yield more robust MOS predictions than any single model.
    source: §3.3, §4.4, Tables 4–5
    evidence: SSL-based features from multiple heterogeneous pretrained models (wav2vec 2.0, HuBERT, WavLM) combined
      via ensemble stacking yield more robust MOS predictions than any single model.
    confidence: high
    relevance: high
  - claim_id: phoneme_level_linguistic_inputs_improve_mos_prediction_robustly_in_low
    role: supports
    claim: Phoneme-level linguistic inputs improve MOS prediction robustly in low-resource conditions but may not
      benefit in-domain prediction when sufficient labeled data is available.
    source: §3.1.4, §4.3, Table 2
    evidence: Phoneme-level linguistic inputs improve MOS prediction robustly in low-resource conditions but may
      not benefit in-domain prediction when sufficient labeled data is available.
    confidence: high
    relevance: medium
  limitations:
  - The UTMOS model is trained and evaluated exclusively on synthetic speech data from Blizzard and Voice Conversion
    Challenges; its generalization to modern neural TTS systems (LLM-based, flow-matching) with different quality
    failure modes is not characterized.
  - The OOD evaluation uses Chinese synthetic speech while the strong learner SSL backbone (wav2vec 2.0) was pretrained
    primarily on English. Generalization to other languages depends on the SSL model's cross-lingual transfer capability,
    which is not analyzed here. The external data collection for OOD (32 Chinese listeners, ~2 ratings per utterance)
    is low-coverage and introduces annotation noise. The paper proposes larger-scale general-purpose data collection
    as future work but does not address it here. Additionally, data augmentation (pitch-shifting, time-stretching)
    is assumed not to change perceptual quality, but this assumption is not verified through listening tests.
  caveats: []
- id: '2209.03143'
  published_date: "2022-09-07"
  entry_date: '2026-07-29'
  year: 2022
  venue: arXiv
  task:
  - TTS
  - SCA
  architecture:
  - autoregressive-LM
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: foundational
  method_family:
  - autoregressive_ssl_conditioned_models
  claims:
  - claim_id: combining_self_supervised_semantic_tokens_with_codec_acoustic_tokens_in
    role: supports
    claim: Combining self-supervised semantic tokens with codec acoustic tokens in a hierarchical language model
      resolves the quality-versus-coherence tension that affects single-tokenizer audio language models.
    source: §III-B, §III-C, Table I
    evidence: Combining self-supervised semantic tokens with codec acoustic tokens in a hierarchical language model
      resolves the quality-versus-coherence tension that affects single-tokenizer audio language models.
    confidence: high
    relevance: high
  - claim_id: semantic_and_acoustic_tokens_in_speech_carry_complementary_information_semantic
    role: supports
    claim: 'Semantic and acoustic tokens in speech carry complementary information: semantic tokens primarily encode
      linguistic content and prosody, while acoustic tokens primarily encode speaker identity and recording conditions.'
    source: §IV-C, §IV-D, Tables II–III
    evidence: 'Semantic and acoustic tokens in speech carry complementary information: semantic tokens primarily
      encode linguistic content and prosody, while acoustic tokens primarily encode speaker identity and recording
      conditions.'
    confidence: high
    relevance: medium
  - claim_id: autoregressive_language_modeling_over_discrete_audio_tokens_can_produce_speech
    role: supports
    claim: Autoregressive language modeling over discrete audio tokens can produce speech continuations indistinguishable
      from real speech to human listeners in an unpaired forced-choice test.
    source: §IV-G
    evidence: Autoregressive language modeling over discrete audio tokens can produce speech continuations indistinguishable
      from real speech to human listeners in an unpaired forced-choice test.
    confidence: high
    relevance: medium
  - claim_id: the_semantic_to_acoustic_hierarchical_generation_pattern_transfers_across_audio
    role: supports
    claim: 'The semantic-to-acoustic hierarchical generation pattern transfers across audio domains: a model trained
      on piano music without symbolic notation also benefits from the two-tier tokenization.'
    source: §IV-I
    evidence: 'The semantic-to-acoustic hierarchical generation pattern transfers across audio domains: a model
      trained on piano music without symbolic notation also benefits from the two-tier tokenization.'
    confidence: high
    relevance: medium
  - claim_id: self_supervised_speech_representations_trained_with_masked_language_modeling_objectives
    role: supports
    claim: Self-supervised speech representations trained with masked language modeling objectives encode sufficient
      lexical and syntactic information to outperform earlier causal spoken language models on zero-resource linguistic
      benchmarks.
    source: §IV-E, Table IV
    evidence: Self-supervised speech representations trained with masked language modeling objectives encode sufficient
      lexical and syntactic information to outperform earlier causal spoken language models on zero-resource linguistic
      benchmarks.
    confidence: high
    relevance: high
  limitations:
  - 'AudioLM is a continuation model only: it generates continuations of an audio prompt but cannot synthesise speech
    from a specified transcript. The paper explicitly frames TTS integration (encoder-decoder with text conditioning)
    as future work, which means the WER/CER results reflect acoustic fidelity to a given semantic token sequence,
    not instruction-following capability.'
  - The system requires three separately trained transformer models totaling approximately 0.9B parameters, plus
    frozen SoundStream and w2v-BERT models, making inference substantially more expensive than single-model TTS
    systems. Inference latency is not reported.
  - All speech experiments use English only; the model is trained on Libri-Light (audiobook speech), which is a
    clean, single-language, read-speech corpus. Generalisability to spontaneous speech, noise, or other languages
    is untested.
  - The paper does not evaluate unconditional generation quality against baselines with matched compute or training
    data, so the contribution of scale versus architecture is not fully disentangled.
  - The anti-spoofing classifier achieving 98.6% accuracy is trained on the same generative model it detects; its
    performance against other generators, or against an adversarially optimised version of AudioLM, is not assessed.
  caveats: []
- id: '2212.04356'
  published_date: "2022-12-06"
  entry_date: '2026-07-29'
  year: 2022
  venue: arXiv
  task:
  - evaluation
  architecture:
  - transformer-enc-dec
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  - evaluation_caution
  current_role: influential
  method_family:
  - transformer_ssl_encoders_and_adapters
  claims:
  - claim_id: weakly_supervised_asr_training_at_sufficient_scale_produces_models_that
    role: supports
    claim: Weakly supervised ASR training at sufficient scale produces models that generalise robustly across recording
      conditions, speakers, and languages without dataset-specific fine-tuning.
    source: §3.3, Table 2
    evidence: Weakly supervised ASR training at sufficient scale produces models that generalise robustly across
      recording conditions, speakers, and languages without dataset-specific fine-tuning.
    confidence: high
    relevance: medium
  - claim_id: models_evaluated_solely_on_in_distribution_benchmarks_can_appear_superhuman
    role: supports
    claim: Models evaluated solely on in-distribution benchmarks can appear superhuman while remaining significantly
      below human performance out-of-distribution, indicating that benchmark-focused evaluation overstates robustness.
    source: §3.3, Figure 2
    evidence: Models evaluated solely on in-distribution benchmarks can appear superhuman while remaining significantly
      below human performance out-of-distribution, indicating that benchmark-focused evaluation overstates robustness.
    confidence: high
    relevance: low
  - claim_id: multitask_and_multilingual_joint_training_introduces_negative_transfer_in_small
    role: supports
    claim: Multitask and multilingual joint training introduces negative transfer in small models but confers benefits
      at large scale.
    source: §4.3, Figure 9
    evidence: Multitask and multilingual joint training introduces negative transfer in small models but confers
      benefits at large scale.
    confidence: high
    relevance: medium
  - claim_id: per_language_zero_shot_asr_performance_scales_predictably_with_per
    role: supports
    claim: Per-language zero-shot ASR performance scales predictably with per-language training data, with WER roughly
      halving for every 16-fold increase in supervision hours.
    source: §3.4, Figure 3
    evidence: Per-language zero-shot ASR performance scales predictably with per-language training data, with WER
      roughly halving for every 16-fold increase in supervision hours.
    confidence: high
    relevance: medium
  limitations:
  - Long-form transcription remains unreliable without explicit decoding heuristics (beam search, temperature fallback,
    voice activity detection, initial timestamp constraints) to prevent repetition loops, truncation of edge segments,
    and hallucination. These are workarounds rather than principled solutions (§4.5). Performance on languages with
    non-Indo-European scripts (Hebrew, Telugu, Chinese, Korean) is worse than the training-data trend predicts,
    likely reflecting tokenisation mismatch and linguistic distance. Language identification accuracy on Fleurs
    is penalised by having no training data for 20 of 102 Fleurs languages, upper-bounding accuracy at 80.4% regardless
    of model capacity (§3.6).
  - The text normalisation procedure carries a risk of being overfitted to Whisper's transcription style, which
    could bias WER comparisons in Whisper's favour (§4.4). Fine-tuning properties are not studied in this paper,
    leaving open how the model's robustness advantage transfers to supervised fine-tuning settings where prior work
    performs.
  caveats: []
- id: '2301.11325'
  published_date: "2023-01-26"
  entry_date: '2026-07-29'
  year: 2023
  venue: arXiv
  task: []
  architecture:
  - autoregressive-LM
  relevance: medium
  evidence_role:
  - architecture_variant
  - evaluation_caution
  current_role: influential
  method_family:
  - autoregressive_ssl_conditioned_models
  claims:
  - claim_id: hierarchical_autoregressive_modeling_over_semantic_and_acoustic_tokens_enables_long
    role: supports
    claim: Hierarchical autoregressive modeling over semantic and acoustic tokens enables long-form music generation
      (several minutes) with temporal coherence at 24 kHz.
    source: §3.2, §6
    evidence: Hierarchical autoregressive modeling over semantic and acoustic tokens enables long-form music generation
      (several minutes) with temporal coherence at 24 kHz.
    confidence: high
    relevance: medium
  - claim_id: a_joint_audio_text_embedding_space_can_substitute_for_paired
    role: supports
    claim: A joint audio-text embedding space can substitute for paired text-audio supervision at training time,
      allowing generative models to be trained on audio-only corpora and conditioned on text at inference.
    source: §3.1, §4.2
    evidence: A joint audio-text embedding space can substitute for paired text-audio supervision at training time,
      allowing generative models to be trained on audio-only corpora and conditioned on text at inference.
    confidence: high
    relevance: medium
  - claim_id: semantic_token_intermediaries_improve_adherence_to_text_descriptions_in_hierarchical
    role: supports
    claim: Semantic token intermediaries improve adherence to text descriptions in hierarchical audio generation
      beyond what direct acoustic token prediction achieves.
    source: §5, Table 1
    evidence: Semantic token intermediaries improve adherence to text descriptions in hierarchical audio generation
      beyond what direct acoustic token prediction achieves.
    confidence: high
    relevance: high
  - claim_id: for_text_conditioned_music_generation_perceptual_audio_quality_fad_and
    role: supports
    claim: For text-conditioned music generation, perceptual audio quality (FAD) and semantic text alignment (MCC,
      KLD) are complementary evaluation axes that do not always correlate with each other.
    source: §4.4, Table 1
    evidence: For text-conditioned music generation, perceptual audio quality (FAD) and semantic text alignment
      (MCC, KLD) are complementary evaluation axes that do not always correlate with each other.
    confidence: high
    relevance: low
  - claim_id: large_autoregressive_audio_lms_trained_on_extensive_unlabeled_corpora_memorize
    role: supports
    claim: Large autoregressive audio LMs trained on extensive unlabeled corpora memorize only a small fraction
      of training sequences exactly, but approximate semantic matches affect a higher proportion of generated outputs
      under targeted prompting.
    source: §5, Figure 3
    evidence: Large autoregressive audio LMs trained on extensive unlabeled corpora memorize only a small fraction
      of training sequences exactly, but approximate semantic matches affect a higher proportion of generated outputs
      under targeted prompting.
    confidence: high
    relevance: medium
  limitations:
  - MCC, one of the two primary text-adherence metrics, is computed using MuLan itself, the same model used for
    conditioning. This circularity biases the metric in MusicLM's favour relative to baselines that do not use MuLan
    representations.
  - 'The system inherits MuLan''s known weaknesses: negations in text prompts are not well handled, and temporal
    ordering of described events is not captured. Vocal quality and lyrics generation are not supported and are
    noted as future work. The 280,000-hour training corpus is proprietary and not released, limiting reproducibility.
    The model weights are not released. Evaluation is conducted against two relatively weak baselines (Mubert, a
    rule-based tag-to-audio API; Riffusion, a fine-tuned image diffusion model on spectrograms); no contemporary
    neural music generation system of comparable scale is included in comparisons.'
  caveats: []
- id: '2305.02765'
  published_date: "2023-05-04"
  entry_date: '2026-07-29'
  year: 2023
  venue: arXiv
  task:
  - codec
  architecture:
  - GAN
  - VAE
  relevance: medium
  evidence_role:
  - architecture_variant
  - evaluation_caution
  current_role: influential
  method_family:
  - gan_decoders_for_ssl_units
  - vae_and_quantized_ssl_representations
  claims:
  - claim_id: grouping_residual_vector_quantization_into_parallel_chains_rather_than_a
    role: supports
    claim: Grouping residual vector quantization into parallel chains rather than a single sequential chain improves
      reconstruction quality per codebook, enabling competitive fidelity with fewer total quantizers.
    source: §3.3, Table 1
    evidence: Grouping residual vector quantization into parallel chains rather than a single sequential chain improves
      reconstruction quality per codebook, enabling competitive fidelity with fewer total quantizers.
    confidence: high
    relevance: medium
  - claim_id: the_burden_that_codec_codebook_count_imposes_on_downstream_generation
    role: supports
    claim: The burden that codec codebook count imposes on downstream generation models is a practical constraint
      that drives codec architecture choices independently of raw reconstruction quality.
    source: §1 Introduction, §3.3
    evidence: The burden that codec codebook count imposes on downstream generation models is a practical constraint
      that drives codec architecture choices independently of raw reconstruction quality.
    confidence: high
    relevance: low
  - claim_id: objective_speech_quality_metrics_such_as_pesq_and_stoi_are
    role: supports
    claim: Objective speech quality metrics such as PESQ and STOI are insufficient alone to characterise codec reconstruction
      quality, and subjective evaluation is necessary but often omitted in codec research.
    source: §6 Limitations
    evidence: Objective speech quality metrics such as PESQ and STOI are insufficient alone to characterise codec
      reconstruction quality, and subjective evaluation is necessary but often omitted in codec research.
    confidence: high
    relevance: low
  - claim_id: publicly_available_training_code_and_pre_trained_baselines_for_neural
    role: supports
    claim: Publicly available training code and pre-trained baselines for neural audio codecs are necessary for
      reproducible research, as previously these were unavailable for EnCodec and SoundStream.
    source: §5 Conclusion, §6 Limitations
    evidence: Publicly available training code and pre-trained baselines for neural audio codecs are necessary for
      reproducible research, as previously these were unavailable for EnCodec and SoundStream.
    confidence: high
    relevance: medium
  limitations:
  - 'No subjective evaluation is included. The paper''s own §6 acknowledges this as a limitation: "Subjective evaluation
    is always the best choice, but this part is missed in this study." All quality comparisons rest solely on PESQ
    and STOI, which the authors themselves note may not accurately reflect perceptual quality.'
  - Beyond the missing subjective evaluation, the paper does not validate GRVQ on downstream generation tasks. The
    claim that 4 codebooks reduce burden on generation models is plausible and consistent with the motivation, but
    no TTS or audio LM experiments are included. The training data is limited to English and Chinese speech from
    public TTS corpora; generalisation to music, environmental sound, or noisy/spontaneous speech is not tested.
    Finally, the paper trains at 16kHz and 24kHz only; higher sample rates (44.1kHz, 48kHz) used in music and high-fidelity
    audio applications are not addressed.
  caveats: []
- id: '2305.09636'
  published_date: "2023-05-16"
  entry_date: '2026-07-29'
  year: 2023
  venue: arXiv
  task:
  - TTS
  - SCA
  architecture:
  - autoregressive-LM
  relevance: medium
  evidence_role:
  - architecture_variant
  - acceleration_evidence
  - control_evidence
  current_role: influential
  method_family:
  - autoregressive_ssl_conditioned_models
  claims:
  - claim_id: parallel_iterative_masked_decoding_adapted_to_rvq_structure_enables_acoustic
    role: supports
    claim: Parallel, iterative masked decoding adapted to RVQ structure enables acoustic token generation two orders
      of magnitude faster than autoregressive generation at matched perceptual quality.
    source: §4.3, Figure 3
    evidence: Parallel, iterative masked decoding adapted to RVQ structure enables acoustic token generation two
      orders of magnitude faster than autoregressive generation at matched perceptual quality.
    confidence: high
    relevance: medium
  - claim_id: non_autoregressive_rvq_level_by_level_decoding_maintains_better_voice
    role: supports
    claim: Non-autoregressive RVQ-level-by-level decoding maintains better voice and acoustic consistency over long
      sequences than autoregressive chunk-and-prompt approaches.
    source: §4.2, Table 1, Figure 2
    evidence: Non-autoregressive RVQ-level-by-level decoding maintains better voice and acoustic consistency over
      long sequences than autoregressive chunk-and-prompt approaches.
    confidence: high
    relevance: medium
  - claim_id: fine_level_rvq_tokens_are_conditionally_independent_given_coarser_tokens
    role: supports
    claim: Fine-level RVQ tokens are conditionally independent given coarser tokens and can be decoded greedily
      in a single pass without measurable quality loss.
    source: §3.3, §4.3
    evidence: Fine-level RVQ tokens are conditionally independent given coarser tokens and can be decoded greedily
      in a single pass without measurable quality loss.
    confidence: high
    relevance: medium
  - claim_id: confidence_based_iterative_decoding_provides_a_meaningful_quality_gain_over
    role: supports
    claim: Confidence-based iterative decoding provides a meaningful quality gain over greedy decoding at the coarsest
      RVQ level, but additional iterations at finer levels yield no significant improvement for speech.
    source: §4.3, Figure 4
    evidence: Confidence-based iterative decoding provides a meaningful quality gain over greedy decoding at the
      coarsest RVQ level, but additional iterations at finer levels yield no significant improvement for speech.
    confidence: high
    relevance: medium
  - claim_id: coupling_a_text_to_semantic_token_model_with_an_efficient
    role: supports
    claim: Coupling a text-to-semantic token model with an efficient acoustic generator enables real-time synthesis
      of controllable multi-speaker dialogue at 30-second horizons.
    source: §5
    evidence: Coupling a text-to-semantic token model with an efficient acoustic generator enables real-time synthesis
      of controllable multi-speaker dialogue at 30-second horizons.
    confidence: high
    relevance: high
  limitations:
  - Evaluation uses a DNSMOS-style estimator rather than human listening tests for audio quality comparisons, and
    the subjective baseline is carried over from earlier AudioLM papers rather than re-run. Direct perceptual comparisons
    between SoundStorm and AudioLM on the same conditions by human raters are not reported.
  - The conditioning mechanism requires time-aligned semantic tokens, restricting drop-in use to AudioLM-family
    pipelines with compatible semantic tokenisers. Extension to cross-attention conditioning or unconditional generation
    is left as future work. The 350M acoustic model trained on LibriLight is English-only; no multilingual evaluation
    is presented. The dialogue synthesis component depends on a proprietary 100k-hour dialogue corpus, limiting
    reproducibility of that subsystem. The ablation on decoding iterations is limited to speech; the authors hypothesise
    that multiple fine-level iterations may matter for non-speech audio, but this is not tested.
  caveats: []
- id: '2305.11000'
  published_date: "2023-05-18"
  entry_date: '2026-07-29'
  year: 2023
  venue: arXiv
  task:
  - SCA
  - TTS
  architecture:
  - autoregressive-LM
  relevance: medium
  evidence_role:
  - architecture_variant
  current_role: influential
  method_family:
  - autoregressive_ssl_conditioned_models
  claims:
  - claim_id: expanding_an_llm_s_token_vocabulary_with_discretised_speech_units
    role: supports
    claim: Expanding an LLM's token vocabulary with discretised speech units enables a single model to perform both
      speech comprehension and speech generation without a cascade pipeline.
    source: §4.1
    evidence: Expanding an LLM's token vocabulary with discretised speech units enables a single model to perform
      both speech comprehension and speech generation without a cascade pipeline.
    confidence: high
    relevance: medium
  - claim_id: a_multi_stage_training_curriculum_separating_modality_adaptation_cross_modal
    role: supports
    claim: A multi-stage training curriculum, separating modality adaptation, cross-modal instruction tuning, and
      chain-of-modality alignment, is necessary to acquire reliable cross-modal instruction-following from an LLM
      backbone.
    source: §4.2
    evidence: A multi-stage training curriculum, separating modality adaptation, cross-modal instruction tuning,
      and chain-of-modality alignment, is necessary to acquire reliable cross-modal instruction-following from an
      LLM backbone.
    confidence: high
    relevance: medium
  - claim_id: the_chain_of_modality_pattern_generating_a_text_intermediate_before
    role: supports
    claim: The chain-of-modality pattern, generating a text intermediate before the speech response, is a practical
      mechanism for transferring LLM reasoning capability to speech output.
    source: §3.2, §4.2
    evidence: The chain-of-modality pattern, generating a text intermediate before the speech response, is a practical
      mechanism for transferring LLM reasoning capability to speech output.
    confidence: high
    relevance: medium
  - claim_id: large_scale_instruction_dataset_construction_via_gpt_4_generated_task
    role: supports
    claim: Large-scale instruction dataset construction via GPT-4-generated task descriptions applied to existing
      ASR corpora is a scalable approach to bootstrapping cross-modal training data.
    source: §3.1
    evidence: Large-scale instruction dataset construction via GPT-4-generated task descriptions applied to existing
      ASR corpora is a scalable approach to bootstrapping cross-modal training data.
    confidence: high
    relevance: medium
  limitations:
  - 'The paper provides no quantitative evaluation: no MOS, WER, or speaker similarity scores are reported, and
    no comparison to cascade baselines is made. All results are case studies. Claims about spoken dialogue quality
    and instruction-following capability cannot be independently verified from the paper alone.'
  - 'Additional limitations acknowledged by the authors:'
  - '- The system does not model paralinguistic information; it cannot generate responses with different emotional
    tones or speaking styles. - The chain-of-modality design requires generating a full text response before producing
    speech, introducing latency incompatible with real-time interaction. - Context length limitations (2048 tokens)
    prevent true multi-turn dialogue; only single-turn exchanges are demonstrated. - HuBERT units discard fine-grained
    acoustic detail (prosody, speaker identity beyond the vocoder''s speaker embedding), limiting expressiveness
    relative to audio codec approaches. - The unit vocoder is trained separately and is not end-to-end jointly optimised
    with the LLM, leaving a quality gap between unit-based and codec-based approaches.'
  caveats: []
- id: '2306.12925'
  published_date: "2023-06-22"
  entry_date: '2026-07-29'
  year: 2023
  venue: arXiv
  task:
  - TTS
  - SCA
  architecture:
  - autoregressive-LM
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: influential
  method_family:
  - autoregressive_ssl_conditioned_models
  claims:
  - claim_id: initializing_a_speech_text_llm_from_a_pretrained_text_only
    role: supports
    claim: Initializing a speech-text LLM from a pretrained text-only checkpoint substantially outperforms training
      from scratch at equivalent model scale.
    source: §5.4.2, Table 6
    evidence: Initializing a speech-text LLM from a pretrained text-only checkpoint substantially outperforms training
      from scratch at equivalent model scale.
    confidence: high
    relevance: medium
  - claim_id: audio_tokenizer_quality_is_a_primary_bottleneck_in_llm_based
    role: supports
    claim: 'Audio tokenizer quality is a primary bottleneck in LLM-based speech generation and understanding: stronger
      semantic tokenizers yield large downstream gains independent of LM scale.'
    source: §5.4.3, Table 7
    evidence: 'Audio tokenizer quality is a primary bottleneck in LLM-based speech generation and understanding:
      stronger semantic tokenizers yield large downstream gains independent of LM scale.'
    confidence: high
    relevance: medium
  - claim_id: a_unified_multimodal_vocabulary_that_interleaves_text_and_audio_tokens
    role: supports
    claim: A unified multimodal vocabulary that interleaves text and audio tokens enables zero-shot speech translation
      to language pairs not seen during speech training, by inheriting translation capability from text pretraining.
    source: §5.2, Table 3
    evidence: A unified multimodal vocabulary that interleaves text and audio tokens enables zero-shot speech translation
      to language pairs not seen during speech training, by inheriting translation capability from text pretraining.
    confidence: high
    relevance: medium
  - claim_id: training_on_combined_tasks_that_decompose_complex_speech_operations_into
    role: supports
    claim: Training on combined tasks that decompose complex speech operations into intermediate text steps improves
      performance over direct end-to-end decoding.
    source: §5.4.4, Table 8
    evidence: Training on combined tasks that decompose complex speech operations into intermediate text steps improves
      performance over direct end-to-end decoding.
    confidence: high
    relevance: medium
  - claim_id: voice_identity_preservation_in_cross_lingual_speech_synthesis_can_exceed
    role: supports
    claim: Voice identity preservation in cross-lingual speech synthesis can exceed high-quality TTS-based references
      when an audio LM is conditioned on a short spoken prompt.
    source: §5.3, Table 4
    evidence: Voice identity preservation in cross-lingual speech synthesis can exceed high-quality TTS-based references
      when an audio LM is conditioned on a short spoken prompt.
    confidence: high
    relevance: medium
  limitations:
  - The entire system depends on the quality of the audio tokenizer, which is not released and requires access to
    Google-internal USM models. The best-performing configuration (AudioPaLM-2 with USM-v2 tokens) is not reproducible
    externally; the published ablations use the multilingual w2v-BERT tokenizer as the weakest condition, suggesting
    that reported performance at USM-v2 quality cannot be independently verified.
  - Adding S2ST tasks modestly degrades ASR and AST performance, suggesting that sharing model capacity across output
    modalities introduces trade-offs that are not fully resolved by the training mixture design (§5.4.5, Table 9).
    The paper evaluates primarily on speech translation tasks; generative speech quality at naturalness is only
    assessed in the S2ST with voice transfer setting, not for open-ended TTS. The subjective evaluations were conducted
    on an earlier version of AudioPaLM using AudioLM decoding rather than SoundStorm, meaning the best-performing
    decoder was not evaluated subjectively. Evaluation benchmarks for generative audio tasks more generally remain
    underdeveloped relative to text, a limitation the paper explicitly notes.
  caveats: []
- id: '2308.16692'
  published_date: "2023-08-31"
  entry_date: '2026-07-29'
  year: 2023
  venue: arXiv
  task:
  - TTS
  - codec
  architecture:
  - GAN
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: influential
  method_family:
  - gan_decoders_for_ssl_units
  claims:
  - claim_id: separating_semantic_content_from_paralinguistic_information_across_rvq_layers_within
    role: supports
    claim: Separating semantic content from paralinguistic information across RVQ layers within a single codec improves
      both reconstruction quality and speech language model coherence compared to undifferentiated acoustic tokenisation.
    source: §4.4, Tables 2 and 4
    evidence: Separating semantic content from paralinguistic information across RVQ layers within a single codec
      improves both reconstruction quality and speech language model coherence compared to undifferentiated acoustic
      tokenisation.
    confidence: high
    relevance: low
  - claim_id: acoustic_tokens_from_standard_neural_codecs_encode_content_and_speaker
    role: supports
    claim: Acoustic tokens from standard neural codecs encode content and speaker identity in an entangled form
      that causes systematic word errors in autoregressive language model generation.
    source: §2.3, Table 3
    evidence: Acoustic tokens from standard neural codecs encode content and speaker identity in an entangled form
      that causes systematic word errors in autoregressive language model generation.
    confidence: high
    relevance: medium
  - claim_id: a_distillation_objective_computed_per_feature_dimension_d_axis_produces
    role: supports
    claim: A distillation objective computed per feature dimension (D-axis) produces stronger semantic guidance
      to a codec's first quantizer than the conventional per-timestep (T-axis) formulation.
    source: Appendix C, Table 7
    evidence: A distillation objective computed per feature dimension (D-axis) produces stronger semantic guidance
      to a codec's first quantizer than the conventional per-timestep (T-axis) formulation.
    confidence: high
    relevance: low
  - claim_id: the_first_layer_tokens_of_a_hierarchically_disentangled_codec_can
    role: supports
    claim: The first-layer tokens of a hierarchically disentangled codec can serve as a zero-shot voice conversion
      mechanism by swapping higher-layer tokens from a reference speaker, without requiring a separate conversion
      model.
    source: §5.2, Table 5
    evidence: The first-layer tokens of a hierarchically disentangled codec can serve as a zero-shot voice conversion
      mechanism by swapping higher-layer tokens from a reference speaker, without requiring a separate conversion
      model.
    confidence: high
    relevance: low
  - claim_id: codec_tokens_trained_without_explicit_content_supervision_exhibit_poor_codebook
    role: supports
    claim: Codec tokens trained without explicit content supervision exhibit poor codebook utilisation and weak
      phoneme-code correspondence, increasing the modelling burden on downstream language models.
    source: Appendix F, Table 8
    evidence: Codec tokens trained without explicit content supervision exhibit poor codebook utilisation and weak
      phoneme-code correspondence, increasing the modelling burden on downstream language models.
    confidence: high
    relevance: low
  limitations:
  - SpeechTokenizer is trained solely on English LibriSpeech. While preliminary results in Appendix G suggest cross-lingual
    token transfer is plausible, the codec is not validated for multilingual speech language models and the text-alignment
    properties of RVQ-1 may not hold for typologically distant languages.
  - The SLMTokBench mutual information metric relies on a variational upper bound and a fixed BLSTM downstream model;
    its absolute values are not directly comparable across evaluation frameworks. The USLM is evaluated only on
    VCTK with a single 3-second prompt per speaker, which does not represent the range of zero-shot conditions used
    in contemporaneous benchmarks. Model size is not reported for SpeechTokenizer itself. The voice conversion application
    (§5.2) is characterised as a single-shot demonstration rather than a full evaluation against dedicated VC baselines.
  caveats: []
- id: '2402.05755'
  published_date: "2024-02-08"
  entry_date: '2026-07-29'
  year: 2024
  venue: arXiv
  task:
  - SCA
  - TTS
  architecture:
  - autoregressive-LM
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: influential
  method_family:
  - autoregressive_ssl_conditioned_models
  claims:
  - claim_id: continuous_pretraining_of_a_text_llm_on_interleaved_speech_and
    role: supports
    claim: Continuous pretraining of a text LLM on interleaved speech and text tokens transfers the text model's
      few-shot learning and semantic reasoning abilities to the speech modality.
    source: §4.2, §4.3, Table 4
    evidence: Continuous pretraining of a text LLM on interleaved speech and text tokens transfers the text model's
      few-shot learning and semantic reasoning abilities to the speech modality.
    confidence: high
    relevance: medium
  - claim_id: word_level_interleaving_of_speech_and_text_during_training_is
    role: supports
    claim: Word-level interleaving of speech and text during training is more effective than parallel ASR/TTS training
      or speech-only fine-tuning for cross-modal semantic understanding.
    source: §4.2, Table 6
    evidence: Word-level interleaving of speech and text during training is more effective than parallel ASR/TTS
      training or speech-only fine-tuning for cross-modal semantic understanding.
    confidence: high
    relevance: medium
  - claim_id: expressive_speech_properties_sentiment_pitch_style_can_be_modeled_in
    role: supports
    claim: Expressive speech properties (sentiment, pitch, style) can be modeled in a language model through discrete
      token streams that supplement phonetic tokens, enabling cross-modal sentiment preservation.
    source: §5, Table 3
    evidence: Expressive speech properties (sentiment, pitch, style) can be modeled in a language model through
      discrete token streams that supplement phonetic tokens, enabling cross-modal sentiment preservation.
    confidence: high
    relevance: medium
  - claim_id: adding_expressive_speech_tokens_to_a_speech_lm_improves_expressivity
    role: complicates
    claim: Adding expressive speech tokens to a speech LM improves expressivity at the cost of moderate degradation
      in lexical and grammatical speech understanding.
    source: §4.2, Table 4
    evidence: Adding expressive speech tokens to a speech LM improves expressivity at the cost of moderate degradation
      in lexical and grammatical speech understanding.
    confidence: high
    relevance: medium
  - claim_id: cascade_speech_pipelines_remain_substantially_stronger_than_end_to_end
    role: supports
    claim: Cascade speech pipelines remain substantially stronger than end-to-end unified models on task-specific
      metrics such as ASR WER and TTS intelligibility at equivalent model scale.
    source: §4.3, Table 5
    evidence: Cascade speech pipelines remain substantially stronger than end-to-end unified models on task-specific
      metrics such as ASR WER and TTS intelligibility at equivalent model scale.
    confidence: high
    relevance: medium
  limitations:
  - Spirit LM's vocoder is conditioned on only 4 speaker voices from the Expresso dataset, which severely constrains
    the diversity and quality of synthesized speech; the model cannot generalize to arbitrary target speakers at
    inference without retraining the vocoder.
  - The STSP benchmark is evaluated using fine-tuned automatic classifiers rather than human listeners, which may
    not capture perceptual sentiment fidelity accurately. The model was evaluated only in English, leaving multilingual
    capabilities untested. The 7B scale represents a reasonable starting point, but the authors note that scaling
    beyond 7B could substantially improve both semantic and expressive performance. Toxicity analysis shows the
    model can add harmful speech content at levels comparable to cascades, with higher MUTOX scores in S→S, warranting
    safety work before deployment.
  caveats: []
- id: '2402.08093'
  published_date: "2024-02-12"
  entry_date: '2026-07-29'
  year: 2024
  venue: arXiv
  task:
  - TTS
  architecture:
  - autoregressive-LM
  relevance: medium
  evidence_role:
  - architecture_variant
  - acceleration_evidence
  - control_evidence
  - evaluation_caution
  current_role: influential
  method_family:
  - autoregressive_ssl_conditioned_models
  claims:
  - claim_id: scaling_autoregressive_codec_tts_to_500m_parameters_and_10k_hours
    role: supports
    claim: Scaling autoregressive codec TTS to 500M+ parameters and 10K+ hours of training data produces qualitatively
      different prosody rendering on linguistically complex inputs compared to smaller models trained on less data.
    source: §4.3, Figure 4, Table 5
    evidence: Scaling autoregressive codec TTS to 500M+ parameters and 10K+ hours of training data produces qualitatively
      different prosody rendering on linguistically complex inputs compared to smaller models trained on less data.
    confidence: high
    relevance: low
  - claim_id: ssl_based_speech_representations_with_explicit_speaker_disentanglement_outperform_purely
    role: supports
    claim: SSL-based speech representations with explicit speaker disentanglement outperform purely acoustic codec
      representations for zero-shot TTS, particularly in lower-resource languages.
    source: §4.1, Table 3
    evidence: SSL-based speech representations with explicit speaker disentanglement outperform purely acoustic
      codec representations for zero-shot TTS, particularly in lower-resource languages.
    confidence: high
    relevance: high
  - claim_id: a_streamable_convolutional_decoder_can_match_or_exceed_a_diffusion
    role: supports
    claim: A streamable convolutional decoder can match or exceed a diffusion-based spectrogram decoder in subjective
      naturalness while reducing synthesis compute by approximately 3x and enabling low-latency streaming.
    source: §4.2, §4.5, Table 4
    evidence: A streamable convolutional decoder can match or exceed a diffusion-based spectrogram decoder in subjective
      naturalness while reducing synthesis compute by approximately 3x and enabling low-latency streaming.
    confidence: high
    relevance: low
  - claim_id: applying_bpe_to_discrete_speech_tokens_reduces_autoregressive_sequence_length
    role: supports
    claim: Applying BPE to discrete speech tokens reduces autoregressive sequence length by approximately 40% without
      degrading downstream synthesis quality, enabling longer-context training.
    source: §2.2.3
    evidence: Applying BPE to discrete speech tokens reduces autoregressive sequence length by approximately 40%
      without degrading downstream synthesis quality, enabling longer-context training.
    confidence: high
    relevance: medium
  - claim_id: autoregressive_tts_trained_at_scale_generalises_to_a_wide_range
    role: supports
    claim: Autoregressive TTS trained at scale generalises to a wide range of textual phenomena without any explicit
      prosody annotation or task-specific supervision.
    source: §4.3, §6
    evidence: Autoregressive TTS trained at scale generalises to a wide range of textual phenomena without any explicit
      prosody annotation or task-specific supervision.
    confidence: high
    relevance: medium
  limitations:
  - Model weights are not released, and evaluation uses proprietary test speakers. The MUSHRA baselines (YourTTS,
    Bark, TortoiseTTS) are not trained on comparable data or compute, making architecture-level conclusions difficult
    to separate from scale effects.
  - The speechcode decoder is tightly coupled to a specific frozen SpeechGPT checkpoint via hidden-state conditioning,
    preventing modular updates and complicating experimentation. The paper identifies this as a limitation requiring
    future work.
  - Hallucinations and cutoffs remain an inherent issue of the autoregressive formulation, worsened by misalignment
    between noisy web audio and ASR-generated transcripts. The authors avoid denoising during training to test robustness
    but acknowledge this makes the alignment problem harder.
  - Emotions and paralinguistics remain below ceiling even for BASE-large, suggesting that 100K hours and 980M parameters
    are not sufficient for reliable rendering of these categories. Formal scaling laws for TTS (analogous to Chinchilla
    for text LMs) are proposed as future work but not established here.
  - The "emergent abilities" phenomenon is characterised across only three data-scale points and assessed by a single
    expert linguist, leaving open whether it is a smooth or discontinuous function of scale and how sensitive it
    is to tokenization and architecture choices.
  caveats: []
- id: '2402.13236'
  published_date: "2024-02-20"
  entry_date: '2026-07-29'
  year: 2024
  venue: arXiv
  task:
  - TTS
  - SCA
  - codec
  architecture: []
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: active_evidence
  method_family: []
  claims:
  - claim_id: residual_vector_quantisation_is_the_dominant_quantisation_strategy_across_neural
    role: supports
    claim: Residual vector quantisation is the dominant quantisation strategy across neural audio codec models,
      with variation concentrated in discriminator design, bitrate, and semantic token integration rather than in
      the core compression mechanism.
    source: §II.A, Table II
    evidence: Residual vector quantisation is the dominant quantisation strategy across neural audio codec models,
      with variation concentrated in discriminator design, bitrate, and semantic token integration rather than in
      the core compression mechanism.
    confidence: high
    relevance: high
  - claim_id: codec_based_audio_language_models_increasingly_target_multi_task_coverage
    role: supports
    claim: Codec-based audio language models increasingly target multi-task coverage rather than single-task specialisation,
      with several systems spanning TTS, voice conversion, speech editing, speech enhancement, and translation in
      a single framework.
    source: §III.B, Table III
    evidence: Codec-based audio language models increasingly target multi-task coverage rather than single-task
      specialisation, with several systems spanning TTS, voice conversion, speech editing, speech enhancement, and
      translation in a single framework.
    confidence: high
    relevance: low
  - claim_id: integrating_semantic_tokens_from_self_supervised_speech_representations_into_the
    role: supports
    claim: Integrating semantic tokens from self-supervised speech representations into the codec quantisation process
      improves audio quality at low bitrates, with HuBERT-guided RVQ being the most common approach.
    source: §II.B
    evidence: Integrating semantic tokens from self-supervised speech representations into the codec quantisation
      process improves audio quality at low bitrates, with HuBERT-guided RVQ being the most common approach.
    confidence: high
    relevance: high
  - claim_id: discrete_units_derived_from_self_supervised_representations_enable_textless_speech
    role: complicates
    claim: Discrete units derived from self-supervised representations enable textless speech language modelling
      but sacrifice speaker and paralinguistic information relative to codec-based approaches.
    source: §III.A
    evidence: Discrete units derived from self-supervised representations enable textless speech language modelling
      but sacrifice speaker and paralinguistic information relative to codec-based approaches.
    confidence: high
    relevance: high
  limitations:
  - 'The survey covers only open-source codec models and does not include proprietary codecs used in industry systems.
    Evaluation methodology is not addressed: the paper does not compare codecs on shared benchmarks or report reproduction
    numbers, making it difficult to assess quality claims from the original papers in a unified way. The codec-based
    LM survey relies entirely on self-reported results from each system''s own paper, with no cross-system comparison
    on shared evaluation sets. Coverage is limited to models available as of early 2024; the field has moved substantially
    since, with systems such as SpeechTokenizer and more recent codec variants not included in the analysis.'
  - 'Open questions the survey surfaces: What is the quality impact of codec design choices (bitrate, discriminator
    type, semantic integration) on downstream speech generation systems? Can a universal codec serve all audio types
    equally well? How should speech LMs be prompted efficiently in zero-shot settings?'
  caveats: []
- id: '2408.02622'
  published_date: "2024-08-05"
  entry_date: '2026-07-29'
  year: 2024
  venue: arXiv
  task:
  - SCA
  - TTS
  architecture:
  - autoregressive-LM
  relevance: medium
  evidence_role:
  - architecture_variant
  - acceleration_evidence
  - control_evidence
  current_role: influential
  method_family:
  - autoregressive_ssl_conditioned_models
  claims:
  - claim_id: full_duplex_speech_generation_requires_that_the_listening_channel_information
    role: supports
    claim: Full-duplex speech generation requires that the listening channel information be injected at intermediate
      representation layers, not at the input or output level, to avoid degrading speech generation quality.
    source: §6.2, Table 2
    evidence: Full-duplex speech generation requires that the listening channel information be injected at intermediate
      representation layers, not at the input or output level, to avoid degrading speech generation quality.
    confidence: high
    relevance: medium
  - claim_id: a_single_layer_discrete_token_autoregressive_tts_backbone_can_integrate
    role: supports
    claim: A single-layer discrete token autoregressive TTS backbone can integrate real-time audio input streams
      with minimal WER degradation relative to a no-listening baseline in controlled conditions.
    source: §6.2, Table 2
    evidence: A single-layer discrete token autoregressive TTS backbone can integrate real-time audio input streams
      with minimal WER degradation relative to a no-listening baseline in controlled conditions.
    confidence: high
    relevance: medium
  - claim_id: robustness_to_noise_and_sensitivity_to_unseen_speaker_interruptions_are
    role: supports
    claim: Robustness to noise and sensitivity to unseen speaker interruptions are distinct challenges in full-duplex
      SLMs, and voice-based generalisation introduces significantly higher error rates than command-based triggering.
    source: §6.2, Table 3
    evidence: Robustness to noise and sensitivity to unseen speaker interruptions are distinct challenges in full-duplex
      SLMs, and voice-based generalisation introduces significantly higher error rates than command-based triggering.
    confidence: high
    relevance: medium
  - claim_id: joint_fine_tuning_of_both_the_speech_generation_backbone_and
    role: supports
    claim: Joint fine-tuning of both the speech generation backbone and the streaming SSL encoder is necessary for
      full-duplex models to reach peak interactive capability; freezing either component degrades turn-taking recall.
    source: §6.3, Table 4
    evidence: Joint fine-tuning of both the speech generation backbone and the streaming SSL encoder is necessary
      for full-duplex models to reach peak interactive capability; freezing either component degrades turn-taking
      recall.
    confidence: high
    relevance: high
  limitations:
  - 'The system produces speech tokens but not semantic speech responses: it stops speaking when interrupted but
    does not generate a contextually appropriate spoken reply. The paper evaluates TTS output and turn-taking accuracy,
    not full dialogue capability. The "full duplex" claim is therefore limited to interruption detection and cessation,
    not conversational back-and-forth.'
  - Voice-based FDM shows a meaningful WER increase (5.33% vs 4.28% baseline in clean conditions, 8.50% under noise),
    suggesting real costs from the dual-channel architecture in harder generalisation settings. The evaluation relies
    entirely on automatic metrics (WER, precision/recall for turn-taking); no subjective MOS or naturalness assessment
    is reported, making quality comparisons to cascade systems difficult. The 106M parameter model trained on LibriTTS
    is modest in scale, and it is unclear whether the fusion findings generalise to larger LLM backbones. Speaker-following
    (identifying which interrupting speaker to respond to) and speech-in/speech-out dialogue generation with full-duplex
    capability are explicitly left as future work.
  caveats: []
- id: '2409.00750'
  published_date: "2024-09-01"
  entry_date: '2026-07-29'
  year: 2024
  venue: arXiv
  task:
  - TTS
  - VC
  architecture:
  - autoregressive-LM
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: influential
  method_family:
  - autoregressive_ssl_conditioned_models
  claims:
  - claim_id: non_autoregressive_masked_generative_transformers_can_achieve_human_level_speaker
    role: supports
    claim: Non-autoregressive masked generative transformers can achieve human-level speaker similarity in zero-shot
      TTS without requiring explicit text-speech alignment or phone-level duration supervision.
    source: §4.2.1, Table 2
    evidence: Non-autoregressive masked generative transformers can achieve human-level speaker similarity in zero-shot
      TTS without requiring explicit text-speech alignment or phone-level duration supervision.
    confidence: high
    relevance: low
  - claim_id: replacing_k_means_quantisation_of_ssl_features_with_vq_vae
    role: supports
    claim: Replacing k-means quantisation of SSL features with VQ-VAE vector quantisation reduces information loss
      in tonal languages and improves downstream acoustic token prediction.
    source: §3.2.1
    evidence: Replacing k-means quantisation of SSL features with VQ-VAE vector quantisation reduces information
      loss in tonal languages and improves downstream acoustic token prediction.
    confidence: high
    relevance: high
  - claim_id: masked_generative_tts_substantially_outperforms_autoregressive_tts_on_hard_text
    role: supports
    claim: Masked generative TTS substantially outperforms autoregressive TTS on hard-text robustness (tongue twisters,
      repeating phrases) while maintaining competitive naturalness on standard benchmarks.
    source: §4.2.2, Appendix J, Table 13
    evidence: Masked generative TTS substantially outperforms autoregressive TTS on hard-text robustness (tongue
      twisters, repeating phrases) while maintaining competitive naturalness on standard benchmarks.
    confidence: high
    relevance: medium
  - claim_id: parallel_iterative_decoding_in_masked_generative_models_yields_constant_inference
    role: supports
    claim: Parallel iterative decoding in masked generative models yields constant inference cost regardless of
      output length, in contrast to autoregressive decoding whose cost scales linearly with utterance duration.
    source: §4.2.2
    evidence: Parallel iterative decoding in masked generative models yields constant inference cost regardless
      of output length, in contrast to autoregressive decoding whose cost scales linearly with utterance duration.
    confidence: high
    relevance: medium
  - claim_id: zero_shot_style_cloning_via_in_context_learning_extends_to
    role: supports
    claim: Zero-shot style cloning via in-context learning extends to accent and emotion transfer without task-specific
      architectural changes.
    source: §4.3, Tables 4–5
    evidence: Zero-shot style cloning via in-context learning extends to accent and emotion transfer without task-specific
      architectural changes.
    confidence: high
    relevance: medium
  limitations:
  - Speech content editing is acknowledged as "not very robust" by the authors, who attribute this to a training
    objective mismatch (mask-and-predict vs. fill-in-mask). The editing capability is demonstrated qualitatively
    only, with no quantitative evaluation reported.
  - 'Training uses 100K hours of English and Chinese speech from Emilia, with multilingual extension at far smaller
    data budgets (2,500–8,200 hours per language). Multilingual performance is uneven: French and German show higher
    WER in cross-lingual dubbing, and the authors note limitations from insufficient retraining of all components
    on expanded data.'
  - Duration control requires either a ground-truth length or the flow-matching duration predictor; errors in predicted
    duration propagate to WER. The gap between predicted-length and ground-truth-length WER is measurable (e.g.,
    2.634 vs. 2.012 on LibriSpeech test-clean).
  - Inference steps of 25-50 for T2S plus the S2A step schedule add latency compared to single-pass systems, though
    the paper does not report real-time factor or wall-clock comparisons.
  - Emotion control requires post-training fine-tuning on labelled data; it is not available zero-shot from the
    base model alone.
  caveats: []
- id: '2409.03283'
  published_date: "2024-09-05"
  entry_date: '2026-07-29'
  year: 2024
  venue: arXiv
  task:
  - TTS
  architecture:
  - autoregressive-LM
  - flow-matching
  - hybrid
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: active_evidence
  method_family:
  - autoregressive_ssl_conditioned_models
  - flow_matching_with_ssl_conditioning
  - hybrid_semantic_acoustic_ssl_systems
  claims:
  - claim_id: separating_the_waveform_generation_stage_into_a_low_sampling_rate
    role: supports
    claim: Separating the waveform generation stage into a low-sampling-rate Mel decoder and a super-resolution
      vocoder allows a system trained predominantly on low-sampling-rate data to produce high-fidelity 48kHz output.
    source: §3.3
    evidence: Separating the waveform generation stage into a low-sampling-rate Mel decoder and a super-resolution
      vocoder allows a system trained predominantly on low-sampling-rate data to produce high-fidelity 48kHz output.
    confidence: high
    relevance: medium
  - claim_id: few_shot_fine_tuning_of_a_large_foundation_tts_model
    role: supports
    claim: Few-shot fine-tuning of a large foundation TTS model substantially outperforms zero-shot in-context learning
      for highly expressive, distinctive target voices, even with only one hour of data.
    source: §5.2.1, Table 5
    evidence: Few-shot fine-tuning of a large foundation TTS model substantially outperforms zero-shot in-context
      learning for highly expressive, distinctive target voices, even with only one hour of data.
    confidence: high
    relevance: medium
  - claim_id: prompt_audio_enhancement_improves_voice_cloning_quality_for_noisy_prompts
    role: supports
    claim: Prompt audio enhancement improves voice cloning quality for noisy prompts but can slightly degrade performance
      when prompts are already clean.
    source: §5.2.2, Table 6
    evidence: Prompt audio enhancement improves voice cloning quality for noisy prompts but can slightly degrade
      performance when prompts are already clean.
    confidence: high
    relevance: medium
  - claim_id: instruction_tuning_with_a_small_domain_specific_dataset_dramatically_improves
    role: supports
    claim: Instruction tuning with a small domain-specific dataset dramatically improves emotion controllability
      in a pre-trained TTS language model, raising accuracy from near-chance to near-ceiling.
    source: §5.3, Table 7
    evidence: Instruction tuning with a small domain-specific dataset dramatically improves emotion controllability
      in a pre-trained TTS language model, raising accuracy from near-chance to near-ceiling.
    confidence: high
    relevance: medium
  - claim_id: autoregressive_tts_systems_trained_on_predominantly_one_language_show_markedly
    role: supports
    claim: Autoregressive TTS systems trained on predominantly one language show markedly higher pronunciation error
      rates on under-represented languages, even at large data scales.
    source: §5.1.2, Table 3
    evidence: Autoregressive TTS systems trained on predominantly one language show markedly higher pronunciation
      error rates on under-represented languages, even at large data scales.
    confidence: high
    relevance: medium
  limitations:
  - All evaluations are conducted on proprietary internal test sets with no publicly released benchmarks, data,
    or model weights. This makes direct comparison with other systems difficult to reproduce and limits the generalisability
    of the reported numbers.
  - The streamable decoder incurs a measurable quality penalty (0.07 CoMOS) and the paper notes that Mel codec quality
    is a bottleneck, which the authors flag for future work. The English and code-switch pronunciation error rates
    remain high (12% and 8.5%), driven by limited language diversity in training data. The paralinguistic behaviour
    framework currently supports 13 types targeting primarily Chinese conversational speech; coverage of other languages
    and more complex prosodic phenomena is not addressed.
  caveats: []
- id: '2409.06666'
  published_date: "2024-09-10"
  entry_date: '2026-07-29'
  year: 2024
  venue: arXiv
  task:
  - SCA
  architecture:
  - hybrid
  relevance: medium
  evidence_role:
  - architecture_variant
  - acceleration_evidence
  - control_evidence
  current_role: influential
  method_family:
  - hybrid_semantic_acoustic_ssl_systems
  claims:
  - claim_id: end_to_end_speech_llms_with_parallel_text_and_speech
    role: supports
    claim: End-to-end speech LLMs with parallel text and speech generation can achieve lower response latency than
      cascaded ASR-LLM-TTS pipelines without sacrificing prosody coherence under low-latency streaming conditions.
    source: §4.5
    evidence: End-to-end speech LLMs with parallel text and speech generation can achieve lower response latency
      than cascaded ASR-LLM-TTS pipelines without sacrificing prosody coherence under low-latency streaming conditions.
    confidence: high
    relevance: low
  - claim_id: aligning_llm_output_to_speech_interaction_conventions_through_targeted_instruction
    role: supports
    claim: Aligning LLM output to speech interaction conventions through targeted instruction data rewriting substantially
      improves response style suitability, independently of model architecture.
    source: §3, §4.4
    evidence: Aligning LLM output to speech interaction conventions through targeted instruction data rewriting
      substantially improves response style suitability, independently of model architecture.
    confidence: high
    relevance: medium
  - claim_id: non_autoregressive_ctc_decoding_from_llm_hidden_states_enables_streaming
    role: supports
    claim: Non-autoregressive CTC decoding from LLM hidden states enables streaming speech synthesis whose speech
      rate and naturalness are robust to chunk size variation, unlike word-level streaming TTS cascades.
    source: §4.5, Table 4
    evidence: Non-autoregressive CTC decoding from LLM hidden states enables streaming speech synthesis whose speech
      rate and naturalness are robust to chunk size variation, unlike word-level streaming TTS cascades.
    confidence: high
    relevance: low
  - claim_id: training_an_end_to_end_speech_interaction_model_on_a
    role: supports
    claim: Training an end-to-end speech interaction model on a small, carefully curated speech instruction dataset
      is sufficient to significantly close the gap with models trained on orders of magnitude more data, provided
      the LLM backbone is sufficiently capable.
    source: §4.4, §5
    evidence: Training an end-to-end speech interaction model on a small, carefully curated speech instruction dataset
      is sufficient to significantly close the gap with models trained on orders of magnitude more data, provided
      the LLM backbone is sufficiently capable.
    confidence: high
    relevance: medium
  limitations:
  - The ASR-WER of 10.82% is notably higher than cascaded baselines (3.78% for SALMONN+Orca), reflecting that the
    speech decoder is trained on only approximately 1K hours of response speech — far below industrial TTS scale.
    Intelligibility limitations restrict applicability in domains requiring precise spoken content.
  - The evaluation benchmark (InstructS2S-Eval) is derived from AlpacaEval with math and code questions removed,
    which skews toward conversational helpfulness and may not represent more demanding speech interaction tasks.
    The speech encoder relies on Whisper, which is optimised for ASR rather than general speech understanding, potentially
    limiting response to prosodic or para-linguistic cues in the user's speech. The current architecture does not
    support full-duplex interaction (interruption, turn-taking) — speech responses are generated after the full
    instruction is received. The training data is synthesised from text corpora, which may not capture the naturalness
    and variability of real spoken dialogue.
  caveats: []
- id: '2410.00037'
  published_date: "2024-09-17"
  entry_date: '2026-07-29'
  year: 2024
  venue: arXiv
  task:
  - SCA
  - TTS
  architecture:
  - autoregressive-LM
  - hybrid
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  - evaluation_caution
  current_role: foundational
  method_family:
  - autoregressive_ssl_conditioned_models
  - hybrid_semantic_acoustic_ssl_systems
  claims:
  - claim_id: eliminating_the_text_bottleneck_in_spoken_dialogue_requires_modeling_acoustic
    role: supports
    claim: Eliminating the text bottleneck in spoken dialogue requires modeling acoustic tokens jointly with semantic
      tokens in a single generative model, as purely semantic approaches cannot capture paralinguistic information
      or generate in arbitrary voices.
    source: §3.4, §5.4
    evidence: Eliminating the text bottleneck in spoken dialogue requires modeling acoustic tokens jointly with
      semantic tokens in a single generative model, as purely semantic approaches cannot capture paralinguistic
      information or generate in arbitrary voices.
    confidence: high
    relevance: low
  - claim_id: predicting_time_aligned_text_tokens_as_a_per_frame_prefix
    role: supports
    claim: Predicting time-aligned text tokens as a per-frame prefix to audio tokens substantially improves the
      linguistic quality and factual accuracy of speech generated by audio language models, with minimal inference
      overhead.
    source: §3.4.4, §5.3, Table 6
    evidence: Predicting time-aligned text tokens as a per-frame prefix to audio tokens substantially improves the
      linguistic quality and factual accuracy of speech generated by audio language models, with minimal inference
      overhead.
    confidence: high
    relevance: medium
  - claim_id: modeling_conversation_as_parallel_autoregressive_streams_for_each_speaker_without
    role: supports
    claim: Modeling conversation as parallel autoregressive streams for each speaker, without explicit turn boundaries,
      enables full-duplex spoken interaction and allows training on naturally overlapping speech.
    source: §3.4.3, §5.6, Table 9
    evidence: Modeling conversation as parallel autoregressive streams for each speaker, without explicit turn boundaries,
      enables full-duplex spoken interaction and allows training on naturally overlapping speech.
    confidence: high
    relevance: medium
  - claim_id: adversarial_only_training_of_neural_audio_codecs_substantially_improves_subjectively
    role: supports
    claim: Adversarial-only training of neural audio codecs substantially improves subjectively rated audio quality
      relative to mixed reconstruction-adversarial objectives, despite degrading objective metrics such as VisQOL.
    source: §3.3, §5.2, Table 4
    evidence: Adversarial-only training of neural audio codecs substantially improves subjectively rated audio quality
      relative to mixed reconstruction-adversarial objectives, despite degrading objective metrics such as VisQOL.
    confidence: high
    relevance: medium
  - claim_id: standard_objective_audio_quality_metrics_visqol_mosnet_are_unreliable_proxies
    role: supports
    claim: Standard objective audio quality metrics (VisQOL, MOSNet) are unreliable proxies for perceived quality
      when the training objective changes, making human evaluation indispensable for codec comparison.
    source: §5.2, §5.8
    evidence: Standard objective audio quality metrics (VisQOL, MOSNet) are unreliable proxies for perceived quality
      when the training objective changes, making human evaluation indispensable for codec comparison.
    confidence: high
    relevance: low
  limitations:
  - 'Moshi''s spoken factual question answering performance lags substantially behind its Helium text baseline,
    particularly on multi-sentence or syntactically complex questions (TriviaQA: 22.8 vs. 56.4 for text-only Helium).
    This indicates that audio training causes significant forgetting of factual knowledge, and the instruct fine-tuning
    data does not cover the syntactic diversity needed to recover it.'
  - 'Signal-based watermarking (Audioseal) is ineffective against codec compression: Mimi''s own lossy coding removes
    the watermark to below detection threshold. The generative watermarking alternatives explored in §6.4 are blocked
    by the non-idempotence of audio codecs, leaving no robust content attribution mechanism available at release.'
  - The instruction fine-tuning pipeline relies heavily on synthetic TTS-generated speech for both conversation
    transcripts and user voice diversity. This introduces a distribution mismatch with real conversational speech
    that likely limits robustness to unusual acoustic conditions and speaking styles. The paper notes this but leaves
    more realistic instruct data collection as future work.
  - Quantization below 4-bit precision causes noticeable audio artifacts (repetitive generation, noisy voice) that
    current automatic metrics fail to detect, requiring entropy-spectrum analysis as a surrogate. This underscores
    a general gap in speech quality evaluation tooling for generative dialogue models.
  caveats: []
- id: '2410.03751'
  published_date: "2024-10-01"
  entry_date: '2026-07-29'
  year: 2024
  venue: arXiv
  task:
  - SCA
  - TTS
  architecture: []
  relevance: medium
  evidence_role:
  - architecture_variant
  - acceleration_evidence
  - infrastructure
  current_role: influential
  method_family: []
  claims:
  - claim_id: end_to_end_speech_generation_models_avoid_the_information_loss
    role: supports
    claim: End-to-end speech generation models avoid the information loss, latency, and cumulative error introduced
      by cascaded ASR-LLM-TTS pipelines, but require integrating speech tokenisation and synthesis into a unified
      training regime.
    source: §I, §II
    evidence: End-to-end speech generation models avoid the information loss, latency, and cumulative error introduced
      by cascaded ASR-LLM-TTS pipelines, but require integrating speech tokenisation and synthesis into a unified
      training regime.
    confidence: high
    relevance: low
  - claim_id: semantic_tokenizers_and_acoustic_tokenizers_impose_an_inherent_trade_off
    role: complicates
    claim: 'Semantic tokenizers and acoustic tokenizers impose an inherent trade-off: semantic tokens produce coherent
      content but poor acoustic quality, while acoustic tokens enable high-fidelity reconstruction but risk content
      inaccuracies.'
    source: §III-A, §IV-A1
    evidence: 'Semantic tokenizers and acoustic tokenizers impose an inherent trade-off: semantic tokens produce
      coherent content but poor acoustic quality, while acoustic tokens enable high-fidelity reconstruction but
      risk content inaccuracies.'
    confidence: high
    relevance: medium
  - claim_id: initialising_a_speech_lm_from_a_text_pretrained_checkpoint_accelerates
    role: supports
    claim: Initialising a speech LM from a text-pretrained checkpoint accelerates convergence and improves speech
      understanding, whereas initialisation from image-pretrained checkpoints yields worse results than random initialisation.
    source: §IV-B1
    evidence: Initialising a speech LM from a text-pretrained checkpoint accelerates convergence and improves speech
      understanding, whereas initialisation from image-pretrained checkpoints yields worse results than random initialisation.
    confidence: high
    relevance: medium
  - claim_id: interleaving_speech_and_text_tokens_during_pre_training_measurably_improves
    role: supports
    claim: Interleaving speech and text tokens during pre-training measurably improves cross-modal representation
      alignment compared to training on speech tokens alone.
    source: §IV-B1
    evidence: Interleaving speech and text tokens during pre-training measurably improves cross-modal representation
      alignment compared to training on speech tokens alone.
    confidence: high
    relevance: medium
  - claim_id: post_alignment_techniques_rlhf_dpo_for_speech_lms_remain_substantially
    role: supports
    claim: Post-alignment techniques (RLHF, DPO) for speech LMs remain substantially underexplored relative to their
      established role in text LM development, leaving semantic consistency and acoustic quality gaps in deployed
      systems.
    source: §IV-B3, §VII
    evidence: Post-alignment techniques (RLHF, DPO) for speech LMs remain substantially underexplored relative to
      their established role in text LM development, leaving semantic consistency and acoustic quality gaps in deployed
      systems.
    confidence: high
    relevance: medium
  limitations:
  - The survey's arXiv version was submitted in October 2024 and the rapidly evolving SpeechLM landscape means several
    systems surveyed (notably Moshi, Mini-Omni, Llama-Omni) were still very recent preprints without peer-reviewed
    evaluation. The coverage of full-duplex systems and post-alignment techniques is acknowledged by the authors
    as incomplete.
  - 'The survey identifies several open questions: whether end-to-end joint training of all three components (tokenizer,
    LM, vocoder) outperforms separately trained pipelines; how to enable real-time speech generation with acceptable
    latency; the unique safety risks of SpeechLMs (toxicity, acoustic inappropriate content, speaker identity leakage)
    that differ from text LM safety challenges; and the potential of SpeechLMs for low-resource spoken languages
    where audio data is more available than text. The question of whether text alignment universally helps or degrades
    paralinguistic modelling (by anchoring the model too closely to textual semantics) is flagged but unresolved.'
  caveats: []
- id: '2411.13577'
  published_date: "2024-11-15"
  entry_date: '2026-07-29'
  year: 2024
  venue: arXiv
  task:
  - SCA
  architecture:
  - autoregressive-LM
  - transformer-enc-dec
  - hybrid
  relevance: medium
  evidence_role:
  - architecture_variant
  - acceleration_evidence
  - evaluation_caution
  - infrastructure
  current_role: influential
  method_family:
  - autoregressive_ssl_conditioned_models
  - transformer_ssl_encoders_and_adapters
  - hybrid_semantic_acoustic_ssl_systems
  claims:
  - claim_id: the_cascaded_and_end_to_end_paradigms_for_spoken_dialogue
    role: supports
    claim: The cascaded and end-to-end paradigms for spoken dialogue impose fundamentally different trade-offs between
      latency, paralinguistic fidelity, and intelligibility, and the appropriate choice depends on the target interaction
      scenario.
    source: §2.2, §2.3
    evidence: The cascaded and end-to-end paradigms for spoken dialogue impose fundamentally different trade-offs
      between latency, paralinguistic fidelity, and intelligibility, and the appropriate choice depends on the target
      interaction scenario.
    confidence: high
    relevance: low
  - claim_id: semantic_speech_representations_offer_higher_compression_rates_and_better_llm
    role: complicates
    claim: Semantic speech representations offer higher compression rates and better LLM compatibility than acoustic
      representations, but sacrifice expressiveness, timbre, and style fidelity, necessitating additional vocoders
      in pipeline-based generation.
    source: §3.3.1, Table 1
    evidence: Semantic speech representations offer higher compression rates and better LLM compatibility than acoustic
      representations, but sacrifice expressiveness, timbre, and style fidelity, necessitating additional vocoders
      in pipeline-based generation.
    confidence: high
    relevance: medium
  - claim_id: achieving_genuine_full_duplex_spoken_dialogue_simultaneous_listening_and_speaking
    role: supports
    claim: Achieving genuine full-duplex spoken dialogue (simultaneous listening and speaking with interrupt handling)
      requires architecturally causal models throughout the full pipeline, a constraint that current systems largely
      satisfy only on the output side.
    source: §5.1, §5.2
    evidence: Achieving genuine full-duplex spoken dialogue (simultaneous listening and speaking with interrupt
      handling) requires architecturally causal models throughout the full pipeline, a constraint that current systems
      largely satisfy only on the output side.
    confidence: high
    relevance: low
  - claim_id: speech_text_modality_alignment_in_current_spoken_dialogue_systems_relies
    role: supports
    claim: Speech-text modality alignment in current spoken dialogue systems relies heavily on paired data, introducing
      catastrophic forgetting risk and creating a structural dependency on the availability of labelled speech corpora.
    source: §4.4.1
    evidence: Speech-text modality alignment in current spoken dialogue systems relies heavily on paired data, introducing
      catastrophic forgetting risk and creating a structural dependency on the availability of labelled speech corpora.
    confidence: high
    relevance: low
  - claim_id: evaluation_infrastructure_for_spoken_dialogue_lags_substantially_behind_system_capabilities
    role: supports
    claim: 'Evaluation infrastructure for spoken dialogue lags substantially behind system capabilities: interaction,
      streaming latency, and audio generation are either absent from or severely underrepresented in existing benchmarks.'
    source: §6.2, §6.3, Table 3
    evidence: 'Evaluation infrastructure for spoken dialogue lags substantially behind system capabilities: interaction,
      streaming latency, and audio generation are either absent from or severely underrepresented in existing benchmarks.'
    confidence: high
    relevance: low
  limitations:
  - The survey was produced concurrently with the primary wave of open-source spoken dialogue models it covers (late
    2024), meaning that some systems are described in early form and the field will have evolved by the time readers
    encounter the paper. The coverage of music and sound understanding and generation within dialogue systems is
    acknowledged as thin (the authors defer to an appendix), and security evaluation for spoken dialogue receives
    less treatment than its importance warrants. A number of prominent systems (Westlake-Omni, Hertz-dev, SpeechGPT2,
    Fish-Agent) lack published papers and are excluded from the timeline figure, which may leave gaps for practitioners
    interested in the deployed-systems landscape. The SuperCLUE benchmark, one of the more comprehensive interaction
    evaluations listed, is not open-source and focuses on Mandarin, limiting its utility for the broader research
    community.
  - 'Open questions surfaced include: whether speech tokenisers can be designed to enforce text-space alignment
    during encoding, eliminating the need for large paired corpora; what granularity of temporal alignment priors
    (sentence, word, phoneme level) is optimal for spoken dialogue training; and how preference optimisation techniques
    can be adapted for the joint text-speech output space.'
  caveats: []
- id: '2411.19842'
  published_date: "2024-11-29"
  entry_date: '2026-07-29'
  year: 2024
  venue: arXiv
  task:
  - codec
  architecture:
  - transformer-enc-dec
  - VAE
  relevance: medium
  evidence_role:
  - architecture_variant
  - evaluation_caution
  current_role: influential
  method_family:
  - transformer_ssl_encoders_and_adapters
  - vae_and_quantized_ssl_representations
  claims:
  - claim_id: scaling_transformer_architecture_parameter_count_in_neural_audio_codecs_produces
    role: supports
    claim: Scaling transformer architecture parameter count in neural audio codecs produces consistent quality improvements
      across objective and subjective metrics.
    source: §4.6, Table 4
    evidence: Scaling transformer architecture parameter count in neural audio codecs produces consistent quality
      improvements across objective and subjective metrics.
    confidence: high
    relevance: medium
  - claim_id: finite_scalar_quantization_achieves_near_perfect_codebook_utilization_without_explicit
    role: supports
    claim: Finite scalar quantization achieves near-perfect codebook utilization without explicit utilization regularization,
      simplifying downstream generative modeling compared to RVQ.
    source: §3.2, §A.8, Table 9
    evidence: Finite scalar quantization achieves near-perfect codebook utilization without explicit utilization
      regularization, simplifying downstream generative modeling compared to RVQ.
    confidence: high
    relevance: medium
  - claim_id: a_neural_codec_trained_exclusively_on_english_speech_can_generalize
    role: supports
    claim: A neural codec trained exclusively on English speech can generalize effectively to unseen languages,
      outperforming multilingual-trained baselines of similar scale on most objective metrics.
    source: §A.5, Table 7
    evidence: A neural codec trained exclusively on English speech can generalize effectively to unseen languages,
      outperforming multilingual-trained baselines of similar scale on most objective metrics.
    confidence: high
    relevance: low
  - claim_id: perceptual_losses_derived_from_self_supervised_speech_models_wavlm_large
    role: supports
    claim: Perceptual losses derived from self-supervised speech models (WavLM-Large features) are critical for
      achieving intelligible reconstruction at low bitrates, beyond what adversarial and spectral reconstruction
      losses alone provide.
    source: §3.4, §A.1, Table 3
    evidence: Perceptual losses derived from self-supervised speech models (WavLM-Large features) are critical for
      achieving intelligible reconstruction at low bitrates, beyond what adversarial and spectral reconstruction
      losses alone provide.
    confidence: high
    relevance: high
  - claim_id: systematic_spectral_bias_in_multi_resolution_stft_discriminators_arising_from
    role: supports
    claim: Systematic spectral bias in multi-resolution STFT discriminators, arising from power-of-two FFT configurations,
      causes periodic reconstruction artifacts that disproportionately affect large-capacity codec architectures.
    source: §3.3, §B.5
    evidence: Systematic spectral bias in multi-resolution STFT discriminators, arising from power-of-two FFT configurations,
      causes periodic reconstruction artifacts that disproportionately affect large-capacity codec architectures.
    confidence: high
    relevance: low
  limitations:
  - Training data is 16 kHz English audiobook speech only (105k hours). Multilingual generalization results are
    promising but the model was not trained or optimized for non-English data; claims about multilingual capability
    should be interpreted cautiously relative to models with dedicated multilingual training at scale.
  - The model has not been evaluated on noisy speech, overlapping speakers, or environmental audio, which are common
    real-world conditions. The large parameter count (950M) requires substantially more compute than lighter baselines
    (DAC at 76M, Mimi at 80M); while the RTF is acceptable on H100 GPUs for longer utterances, latency for short
    clips is roughly 3x that of smaller models, which matters for streaming applications. The post-hoc Residual
    FSQ decomposition is restricted to specific level configurations (L = 2^n + 1); arbitrary bitrate targets are
    not directly achievable without retraining. The systematic bias analysis in the discriminator (§B.5) raises
    open questions about whether similar biases appear in other convolutional discriminator architectures and how
    to address them in the general case.
  caveats: []
- id: '2412.04724'
  published_date: "2024-12-06"
  entry_date: '2026-07-29'
  year: 2024
  venue: arXiv
  task:
  - VC
  architecture:
  - flow-matching
  relevance: medium
  evidence_role:
  - architecture_variant
  - acceleration_evidence
  - control_evidence
  current_role: influential
  method_family:
  - flow_matching_with_ssl_conditioning
  claims:
  - claim_id: independent_timbre_and_style_transfer_from_distinct_unseen_speakers_can
    role: supports
    claim: Independent timbre and style transfer from distinct unseen speakers can be achieved without degrading
      either attribute when conditioning streams are separated via parallel cross-attention with adaptive gating.
    source: §DualAGC, Table 2
    evidence: Independent timbre and style transfer from distinct unseen speakers can be achieved without degrading
      either attribute when conditioning streams are separated via parallel cross-attention with adaptive gating.
    confidence: high
    relevance: medium
  - claim_id: flow_matching_enables_zero_shot_voice_conversion_that_surpasses_diffusion
    role: supports
    claim: Flow matching enables zero-shot voice conversion that surpasses diffusion baselines in both sample quality
      and inference speed, with quality stabilising in as few as 10 ODE steps.
    source: §Conditional Flow Matching, Table 1, Table 3
    evidence: Flow matching enables zero-shot voice conversion that surpasses diffusion baselines in both sample
      quality and inference speed, with quality stabilising in as few as 10 ODE steps.
    confidence: high
    relevance: low
  - claim_id: timbre_leakage_into_style_representations_is_a_failure_mode_in
    role: supports
    claim: Timbre leakage into style representations is a failure mode in jointly trained VC systems, and adversarial
      disentanglement via gradient reversal measurably reduces this cross-contamination.
    source: §Training Objectives, Table 4
    evidence: Timbre leakage into style representations is a failure mode in jointly trained VC systems, and adversarial
      disentanglement via gradient reversal measurably reduces this cross-contamination.
    confidence: high
    relevance: medium
  - claim_id: using_multiple_reference_utterances_alongside_a_pre_trained_speaker_verification
    role: supports
    claim: Using multiple reference utterances alongside a pre-trained speaker verification prior as the timbre
      attention key dramatically improves intelligibility in cross-attention-based timbre modeling, as its removal
      causes WER to collapse from 2% to over 22%.
    source: §DualAGC, Table 4
    evidence: Using multiple reference utterances alongside a pre-trained speaker verification prior as the timbre
      attention key dramatically improves intelligibility in cross-attention-based timbre modeling, as its removal
      causes WER to collapse from 2% to over 22%.
    confidence: high
    relevance: medium
  - claim_id: non_autoregressive_vc_generation_maintains_near_constant_latency_regardless_of
    role: supports
    claim: Non-autoregressive VC generation maintains near-constant latency regardless of utterance length, a qualitative
      advantage over token-by-token autoregressive decoding that is not captured by fixed-length RTF comparisons.
    source: §Experimental Results on Zero-shot VC
    evidence: Non-autoregressive VC generation maintains near-constant latency regardless of utterance length, a
      qualitative advantage over token-by-token autoregressive decoding that is not captured by fixed-length RTF
      comparisons.
    confidence: high
    relevance: low
  limitations:
  - The factorized codec used as the style extractor is referenced as a public tool but not described in detail
    in the paper, which restricts exact reproducibility of the style extraction stage.
  - The evaluation covers English speakers only (VCTK for timbre, ESD for style), leaving cross-lingual style transfer
    untested. The style space is effectively bounded by the factorized codec's subspace representation, and it is
    unclear how fine-grained or compositional the style control is in practice beyond the five ESD emotion categories.
    The system uses HiFi-GAN as vocoder rather than a codec-based decoder, which may limit audio bandwidth compared
    to neural codec approaches. Results on real-world noisy conditions are not reported; training was filtered by
    DNSMOS quality, so robustness to in-the-wild speech is assumed but unverified.
  caveats: []
- id: '2409.09098'
  published_date: "2025-01-09"
  entry_date: '2026-07-29'
  year: 2025
  venue: arXiv
  task:
  - TTS
  architecture:
  - transformer-enc-dec
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  - evaluation_caution
  current_role: active_evidence
  method_family:
  - transformer_ssl_encoders_and_adapters
  claims:
  - claim_id: speaker_accent_entanglement_in_accent_identification_models_causes_poor_generalisation
    role: supports
    claim: Speaker-accent entanglement in accent identification models causes poor generalisation to unseen speakers
      and limits the utility of accent embeddings for conditioning TTS systems.
    source: §III.A, §IV.A, Table IV
    evidence: The CommonAccent baseline achieves 0.96 accuracy on seen speakers but only 0.43 on unseen speakers
      (gap 0.53), and has a high SCSC of 0.236, indicating that it memorises speaker-to-accent mappings. GenAID
      with information bottleneck and adversarial training reduces the gap to 0.06 and SCSC to 0.079.
    confidence: high
    relevance: medium
  - claim_id: continuous_accent_embeddings_extracted_from_a_speaker_agnostic_model_provide
    role: supports
    claim: Continuous accent embeddings extracted from a speaker-agnostic model provide stronger accent conditioning
      for zero-shot TTS than discrete one-hot accent labels.
    source: §IV.B, Tables V–VII
    evidence: AccentBox conditioned on continuous GenAID embeddings achieves higher accent cosine similarity than
      the Accent_ID system using one-hot accent embeddings in both inherent and cross accent generation, and is
      preferred by listeners in subjective accent similarity tests.
    confidence: high
    relevance: medium
  - claim_id: zero_shot_accent_generation_introduces_a_trade_off_between_accent
    role: complicates
    claim: Zero-shot accent generation introduces a trade-off between accent fidelity and naturalness that varies
      with accent data coverage.
    source: §IV.B, Table VI
    evidence: AccentBox shows higher naturalness preference for American accent (60.0% preferred over Baseline,
      p=0.011) but lower preference for Irish accent (33.9%), where limited training data and sensitivity to monotonic
      prosody in reference speech cause degradation.
    confidence: high
    relevance: medium
  - claim_id: objective_speaker_similarity_metrics_may_not_align_with_subjective_listener
    role: complicates
    claim: Objective speaker similarity metrics may not align with subjective listener perception when accent and
      speaker identity are jointly manipulated.
    source: §IV.B, Tables V–VI
    evidence: For inherent accent generation, AccentBox achieves lower SpkCos (0.8293) than Baseline (0.8413) and
      VALL-E X (0.8605) in objective evaluation, yet listeners subjectively prefer AccentBox for speaker similarity
      (70.0%, p=0.002 for American accent), suggesting the speaker verification model is biased toward common accent
      patterns.
    confidence: high
    relevance: low
  - claim_id: standard_zs_tts_evaluations_based_on_naturalness_and_speaker_similarity
    role: refines
    claim: Standard ZS-TTS evaluations based on naturalness and speaker similarity fail to detect accent hallucination,
      underrepresenting the accent fidelity gap between TTS systems trained predominantly on American English and
      target accented speakers.
    source: §I, §III.B
    evidence: The paper demonstrates that current SOTA ZS-TTS systems (including VALL-E X) generate a default American-English
      accent regardless of the target speaker's accent, an artefact not captured by conventional MOS or speaker-similarity
      metrics. Accent cosine similarity and subjective accent preference tests are introduced as complementary metrics.
    confidence: high
    relevance: low
  limitations:
  - Subjective listening tests are restricted to two accents (American and Irish) due to budget constraints; the
    cross-accent and unseen-accent generation results lack systematic subjective evaluation. The Irish accent results
    show degraded naturalness, indicating the system is sensitive to data volume and reference speech quality in
    ways that may not generalise across all 13 accents.
  - The TTS backbone (YourTTS, VITS-based) is several generations behind current LLM-based and flow-matching ZS-TTS
    systems. The authors chose YourTTS for stability and compute efficiency, but the quality ceiling limits competitiveness
    with current state-of-the-art naturalness. The paper does not evaluate WER, arguing that ASR models are biased
    against accented speech; this is a reasonable methodological choice but limits comparability with other work.
    Unseen accent generation is demonstrated only via audio samples on the demo page, with no quantitative evaluation.
  caveats: []
- id: 2025.coling-main.518
  published_date: "2025-01-19"
  entry_date: '2026-07-29'
  year: 2025
  venue: COLING
  task:
  - TTS
  architecture:
  - flow-matching
  - hybrid
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  - evaluation_caution
  current_role: active_evidence
  method_family:
  - flow_matching_with_ssl_conditioning
  - hybrid_semantic_acoustic_ssl_systems
  claims:
  - claim_id: applying_conditional_flow_matching_to_a_self_supervised_prosody_latent
    role: supports
    claim: Applying conditional flow matching to a self-supervised prosody latent space rather than to full acoustic
      features enables diverse prosody generation at inference without requiring a reference utterance.
    source: §2.3, §3.3
    evidence: Applying conditional flow matching to a self-supervised prosody latent space rather than to full acoustic
      features enables diverse prosody generation at inference without requiring a reference utterance.
    confidence: high
    relevance: high
  - claim_id: self_supervised_speech_representations_wavlm_provide_a_more_effective_prosody
    role: supports
    claim: Self-supervised speech representations (WavLM) provide a more effective prosody conditioning signal for
      TTS than conventional pitch and energy predictors, as shown by ablation.
    source: §3.4, Table 3
    evidence: Self-supervised speech representations (WavLM) provide a more effective prosody conditioning signal
      for TTS than conventional pitch and energy predictors, as shown by ablation.
    confidence: high
    relevance: high
  - claim_id: flow_matching_in_a_prosody_latent_space_achieves_comparable_quality
    role: supports
    claim: Flow matching in a prosody latent space achieves comparable quality to diffusion-based prosody modeling
      with substantially lower computational cost.
    source: §3.4, Table 3
    evidence: Flow matching in a prosody latent space achieves comparable quality to diffusion-based prosody modeling
      with substantially lower computational cost.
    confidence: high
    relevance: medium
  - claim_id: a_small_number_of_flow_matching_function_evaluations_n_1
    role: supports
    claim: A small number of flow matching function evaluations (n=1) is sufficient to match or exceed legacy TTS
      baselines on MOS and WER, confirming the sample efficiency of flow matching for prosody.
    source: §3.3, Table 2
    evidence: A small number of flow matching function evaluations (n=1) is sufficient to match or exceed legacy
      TTS baselines on MOS and WER, confirming the sample efficiency of flow matching for prosody.
    confidence: high
    relevance: medium
  limitations:
  - The model is validated only on single-speaker LJSpeech; extension to multi-speaker and zero-shot settings is
    explicitly acknowledged as future work. The model architecture size is not reported. WavLM parameters are frozen
    throughout training, which may limit adaptation to unusual prosody distributions. The comparison set does not
    include the most recent flow matching TTS systems (Voicebox, Matcha-TTS, E2 TTS) — only legacy baselines (FastSpeech
    2, VITS, StyleTTS 2) and one diffusion model (DiffProsody). The absence of speaker diversity means the prosody
    variability shown may primarily reflect intra-speaker diversity of LJSpeech rather than generalizable prosodic
    modeling.
  caveats: []
- id: '2502.04128'
  published_date: "2025-02-06"
  entry_date: '2026-07-29'
  year: 2025
  venue: arXiv
  task:
  - TTS
  architecture:
  - autoregressive-LM
  relevance: medium
  evidence_role:
  - architecture_variant
  - acceleration_evidence
  - control_evidence
  current_role: influential
  method_family:
  - autoregressive_ssl_conditioned_models
  claims:
  - claim_id: single_stage_autoregressive_tts_trained_with_next_token_prediction_over
    role: supports
    claim: Single-stage autoregressive TTS trained with next-token prediction over discrete speech tokens is competitive
      with multi-stage AR+NAR pipelines on intelligibility and speaker similarity in continuation mode, though SIM-o
      gaps remain due to codec acoustic reconstruction limits.
    source: §3.2.4, Table 3
    evidence: Single-stage autoregressive TTS trained with next-token prediction over discrete speech tokens is
      competitive with multi-stage AR+NAR pipelines on intelligibility and speaker similarity in continuation mode,
      though SIM-o gaps remain due to codec acoustic reconstruction limits.
    confidence: high
    relevance: low
  - claim_id: both_model_scale_and_training_data_volume_independently_improve_tts
    role: supports
    claim: Both model scale and training data volume independently improve TTS quality across naturalness, prosody,
      and text comprehension, consistent with scaling laws observed in text LLMs.
    source: §2.3, §3.2.2, Tables 2, 4
    evidence: Both model scale and training data volume independently improve TTS quality across naturalness, prosody,
      and text comprehension, consistent with scaling laws observed in text LLMs.
    confidence: high
    relevance: medium
  - claim_id: inference_time_compute_scaling_via_speech_understanding_verifiers_can_substantially
    role: complicates
    claim: Inference-time compute scaling via speech understanding verifiers can substantially improve speaker similarity
      and emotional expressiveness beyond what train-time scaling alone achieves, at the cost of additional inference
      compute.
    source: §2.4, §3.2.3, Figure 2, Table 2
    evidence: Inference-time compute scaling via speech understanding verifiers can substantially improve speaker
      similarity and emotional expressiveness beyond what train-time scaling alone achieves, at the cost of additional
      inference compute.
    confidence: high
    relevance: low
  - claim_id: pure_process_reward_model_beam_search_for_tts_is_prone
    role: supports
    claim: Pure process reward model beam search for TTS is prone to mode collapse that degrades content accuracy
      (WER), and a hybrid partial-PRM strategy is needed to preserve both speaker similarity and intelligibility.
    source: §3.2.3, Figure 2
    evidence: Pure process reward model beam search for TTS is prone to mode collapse that degrades content accuracy
      (WER), and a hybrid partial-PRM strategy is needed to preserve both speaker similarity and intelligibility.
    confidence: high
    relevance: low
  - claim_id: single_vq_codecs_can_achieve_intelligibility_and_naturalness_competitive_with
    role: supports
    claim: Single-VQ codecs can achieve intelligibility and naturalness competitive with multi-layer RVQ codecs
      at the same token rate, but acoustic fidelity (speaker similarity) remains the limiting factor for single-VQ
      reconstruction.
    source: §3.1.2, Table 1
    evidence: Single-VQ codecs can achieve intelligibility and naturalness competitive with multi-layer RVQ codecs
      at the same token rate, but acoustic fidelity (speaker similarity) remains the limiting factor for single-VQ
      reconstruction.
    confidence: high
    relevance: low
  limitations:
  - 'The SIM-o gap between Llasa and RVQ-based baselines is intrinsic to the single-VQ design: acoustic reconstruction
    from a 65,536-entry single codebook at 50 Hz is weaker than 8-layer RVQ codecs, and this gap is only partially
    recovered by inference-time search. Systems requiring high timbre fidelity in a single inference pass would
    need a different codec design.'
  - Inference-time compute scaling requires running multiple candidates (beam search or Best-of-N) with auxiliary
    verifier models, which increases latency and compute cost substantially and makes the approach unsuitable for
    real-time or low-resource applications. The paper does not characterize latency or wall-clock overhead of the
    search strategies.
  - The text understanding evaluation uses expert-rated 3-point discrete scores, which are not directly comparable
    across papers and rely on a small number of sentences per category. The evaluation methodology is adapted from
    BASE TTS but the inter-rater reliability is not reported.
  - Models are trained on mixed Mandarin/English data, but the language coverage and balance are not fully documented.
    The internal data component of the 250k-hour corpus is not described, limiting reproducibility.
  caveats: []
- id: '2502.06490'
  published_date: "2025-02-10"
  entry_date: '2026-07-29'
  year: 2025
  venue: arXiv
  task:
  - TTS
  - VC
  - codec
  architecture: []
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: influential
  method_family: []
  claims:
  - claim_id: acoustic_tokens_and_semantic_tokens_occupy_fundamentally_distinct_points_in
    role: complicates
    claim: Acoustic tokens and semantic tokens occupy fundamentally distinct points in a reconstruction-versus-semantics
      trade-off space, and no single tokenization strategy currently achieves strong performance on both axes simultaneously.
    source: §VI-C, Table I
    evidence: Acoustic tokens and semantic tokens occupy fundamentally distinct points in a reconstruction-versus-semantics
      trade-off space, and no single tokenization strategy currently achieves strong performance on both axes simultaneously.
    confidence: high
    relevance: medium
  - claim_id: speaker_disentanglement_in_acoustic_tokens_enables_voice_conversion_capability_but
    role: supports
    claim: Speaker disentanglement in acoustic tokens enables voice conversion capability but consistently reduces
      reconstruction quality metrics at equivalent bitrates.
    source: §VI-D, Table I
    evidence: Speaker disentanglement in acoustic tokens enables voice conversion capability but consistently reduces
      reconstruction quality metrics at equivalent bitrates.
    confidence: high
    relevance: low
  - claim_id: k_means_clustering_on_ssl_model_embeddings_discards_prosody_information
    role: supports
    claim: K-means clustering on SSL model embeddings discards prosody information more severely than supervised
      or internally-quantized semantic tokenizers, making offline clustering ill-suited for tasks requiring prosody
      fidelity.
    source: §VI-C, §VI-D, Table I
    evidence: K-means clustering on SSL model embeddings discards prosody information more severely than supervised
      or internally-quantized semantic tokenizers, making offline clustering ill-suited for tasks requiring prosody
      fidelity.
    confidence: high
    relevance: high
  - claim_id: acoustic_byte_pair_encoding_achieves_greater_length_reduction_on_tokens
    role: supports
    claim: Acoustic byte-pair encoding achieves greater length reduction on tokens with lower information density,
      such as speaker-decoupled and semantic tokens, than on general-purpose acoustic tokens.
    source: §V-A, Figure 8
    evidence: Acoustic byte-pair encoding achieves greater length reduction on tokens with lower information density,
      such as speaker-decoupled and semantic tokens, than on general-purpose acoustic tokens.
    confidence: high
    relevance: medium
  - claim_id: single_codebook_tokens_at_very_low_frame_rates_improve_compatibility
    role: supports
    claim: Single-codebook tokens at very low frame rates improve compatibility with language model generation but
      currently exhibit measurable quality and intelligibility degradation relative to multi-codebook or higher
      frame-rate alternatives.
    source: §VIII.1, §VI-C
    evidence: Single-codebook tokens at very low frame rates improve compatibility with language model generation
      but currently exhibit measurable quality and intelligibility degradation relative to multi-codebook or higher
      frame-rate alternatives.
    confidence: high
    relevance: medium
  limitations:
  - The experimental comparisons are conducted on English data only (LibriTTS, LibriSpeech), leaving multilingual
    tokenization trade-offs unexplored. The unified vocoder (CTX-vec2wav) is specifically designed for semantic
    tokens, which may introduce a systematic advantage for semantic token types in reconstruction experiments. Not
    all acoustic tokens support voice conversion in the paper's framework, so the VC comparison covers only a subset
    of systems.
  - 'Open questions identified by the survey include: the bitrate lower bound for single-codebook tokens with acceptable
    intelligibility; whether causal SSL architectures can match non-causal models for semantic token quality; how
    VFR tokens perform on generative tasks beyond ASR; and whether token vocoders trained at scale can match flow
    matching-based alternatives for timbre controllability.'
  caveats: []
- id: '2502.07243'
  published_date: "2025-02-11"
  entry_date: '2026-07-29'
  year: 2025
  venue: ICLR
  task:
  - TTS
  - VC
  architecture:
  - autoregressive-LM
  - flow-matching
  - hybrid
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: influential
  method_family:
  - autoregressive_ssl_conditioned_models
  - flow_matching_with_ssl_conditioning
  - hybrid_semantic_acoustic_ssl_systems
  claims:
  - claim_id: the_vq_vae_codebook_vocabulary_size_can_function_as_a
    role: supports
    claim: The VQ-VAE codebook vocabulary size can function as a self-supervised information bottleneck for progressive
      disentanglement of timbre, style, and linguistic content in self-supervised speech representations.
    source: §3.1, Table 2
    evidence: The VQ-VAE codebook vocabulary size can function as a self-supervised information bottleneck for progressive
      disentanglement of timbre, style, and linguistic content in self-supervised speech representations.
    confidence: high
    relevance: high
  - claim_id: zero_shot_style_imitation_accent_and_emotion_conversion_without_annotation
    role: supports
    claim: Zero-shot style imitation (accent and emotion conversion) without annotation can match or exceed supervised
      baselines that rely on parallel corpora and style labels.
    source: §4.3, Table 4
    evidence: Zero-shot style imitation (accent and emotion conversion) without annotation can match or exceed supervised
      baselines that rely on parallel corpora and style labels.
    confidence: high
    relevance: medium
  - claim_id: hybrid_two_stage_pipelines_combining_autoregressive_style_modeling_with_flow
    role: supports
    claim: Hybrid two-stage pipelines combining autoregressive style modeling with flow-matching acoustic generation
      can decouple style and timbre control more effectively than single-stage approaches that use in-context learning
      to mimic all speech attributes jointly.
    source: §3.4, Tables 3–5
    evidence: Hybrid two-stage pipelines combining autoregressive style modeling with flow-matching acoustic generation
      can decouple style and timbre control more effectively than single-stage approaches that use in-context learning
      to mimic all speech attributes jointly.
    confidence: high
    relevance: medium
  - claim_id: autoregressive_models_in_zero_shot_tts_consistently_trade_intelligibility_higher
    role: supports
    claim: Autoregressive models in zero-shot TTS consistently trade intelligibility (higher WER) for stronger style
      imitation compared to non-autoregressive alternatives trained on the same data.
    source: §4.4, Tables 5, 9
    evidence: Autoregressive models in zero-shot TTS consistently trade intelligibility (higher WER) for stronger
      style imitation compared to non-autoregressive alternatives trained on the same data.
    confidence: high
    relevance: medium
  - claim_id: duration_reduction_on_content_tokens_improves_style_transfer_fidelity_by
    role: supports
    claim: Duration reduction on content tokens improves style transfer fidelity by removing unit-level duration
      patterns that encode source speaking style.
    source: §4.5, Table 6
    evidence: Duration reduction on content tokens improves style transfer fidelity by removing unit-level duration
      patterns that encode source speaking style.
    confidence: high
    relevance: medium
  limitations:
  - Style imitation evaluations (Table 4) use demo website samples from baseline systems as the test set, meaning
    evaluation conditions (recording environment, speaker demographics, utterance content) differ between Vevo and
    baselines. These comparisons are suggestive but not controlled, and the reported improvements should be treated
    as approximate.
  - Training is restricted to English audiobook speech (clean, single-domain), and no multilingual or expressive
    speech experiments are reported. The content-style token vocabulary size (K_s = 4096) and content token vocabulary
    size (K_c = 32) are empirically selected; the authors note these may not be globally optimal. The AR content-style
    model has 463M parameters and requires sequential decoding, introducing latency that could be problematic for
    streaming applications. The self-supervised disentanglement quality depends on HuBERT-Large features, requiring
    a large pre-trained SSL model as a prerequisite. Style controllability through a single reference utterance
    may be brittle for rare or highly expressive speaking styles not represented in the audiobook training distribution.
  caveats: []
- id: '2502.17239'
  published_date: "2025-02-24"
  entry_date: '2026-07-29'
  year: 2025
  venue: arXiv
  task:
  - SCA
  - TTS
  architecture:
  - autoregressive-LM
  - flow-matching
  - hybrid
  relevance: medium
  evidence_role:
  - architecture_variant
  - evaluation_caution
  current_role: influential
  method_family:
  - autoregressive_ssl_conditioned_models
  - flow_matching_with_ssl_conditioning
  - hybrid_semantic_acoustic_ssl_systems
  claims:
  - claim_id: multi_codebook_rvq_tokenizers_with_semantic_alignment_objectives_better_preserve
    role: supports
    claim: Multi-codebook RVQ tokenizers with semantic alignment objectives better preserve both acoustic and linguistic
      content than single-codebook designs, with each additional layer reducing ASR error substantially up to 8
      layers.
    source: §3.1, Table 1
    evidence: Multi-codebook RVQ tokenizers with semantic alignment objectives better preserve both acoustic and
      linguistic content than single-codebook designs, with each additional layer reducing ASR error substantially
      up to 8 layers.
    confidence: high
    relevance: medium
  - claim_id: staged_pretraining_frozen_llm_first_then_joint_training_measurably_reduces
    role: supports
    claim: Staged pretraining (frozen LLM first, then joint training) measurably reduces intelligence degradation
      in end-to-end speech LMs relative to single-stage joint training.
    source: §3.3.1, Table 5
    evidence: Staged pretraining (frozen LLM first, then joint training) measurably reduces intelligence degradation
      in end-to-end speech LMs relative to single-stage joint training.
    confidence: high
    relevance: medium
  - claim_id: text_guided_aligned_generation_where_the_model_completes_text_tokens
    role: supports
    claim: Text-guided aligned generation, where the model completes text tokens before emitting the corresponding
      audio tokens, mitigates semantic incoherence in speech LM outputs.
    source: §3.3, §4.1
    evidence: Text-guided aligned generation, where the model completes text tokens before emitting the corresponding
      audio tokens, mitigates semantic incoherence in speech LM outputs.
    confidence: high
    relevance: medium
  - claim_id: flow_matching_decoders_trained_as_post_vq_refinement_stages_recover
    role: supports
    claim: Flow-matching decoders trained as post-VQ refinement stages recover significant audio quality lost during
      quantisation, with UTMOS improvements of 0.6 points possible without retraining the LLM.
    source: §3.2, Table 3
    evidence: Flow-matching decoders trained as post-VQ refinement stages recover significant audio quality lost
      during quantisation, with UTMOS improvements of 0.6 points possible without retraining the LLM.
    confidence: high
    relevance: medium
  - claim_id: end_to_end_speech_lms_evaluated_in_s_s_mode
    role: supports
    claim: End-to-end speech LMs evaluated in S→S mode exhibit notably lower benchmark performance than the same
      model evaluated in S→T mode, indicating that audio token generation itself introduces a quality penalty beyond
      the comprehension step.
    source: §4.3, Table 8
    evidence: End-to-end speech LMs evaluated in S→S mode exhibit notably lower benchmark performance than the same
      model evaluated in S→T mode, indicating that audio token generation itself introduces a quality penalty beyond
      the comprehension step.
    confidence: high
    relevance: medium
  limitations:
  - The TTS quality evaluation is limited to an in-house test set (MED-TTS); no comparison against standard TTS
    benchmarks (VCTK, LJSpeech, LibriTTS test-clean) or against dedicated TTS systems is reported. The naturalness
    and speaker similarity of the generated speech relative to state-of-the-art TTS systems is therefore unknown.
  - The intelligence gap to GPT-4o-Audio remains large (roughly 15-20 percentage points on QA tasks), and the paper
    does not explain what architectural or data factors account for this difference. The OpenAudioBench evaluation
    uses GPT-4o as judge, which may introduce evaluation bias. The full-duplex and interruption-handling capabilities
    common in deployed spoken conversational agents are not evaluated. Training data mix decisions (e.g. dropping
    INTLV audio loss, ITTS loss design) are motivated empirically but without systematic ablation.
  caveats: []
- id: '2503.01710'
  published_date: "2025-03-03"
  entry_date: '2026-07-29'
  year: 2025
  venue: arXiv
  task:
  - TTS
  architecture:
  - autoregressive-LM
  relevance: medium
  evidence_role:
  - architecture_variant
  - acceleration_evidence
  - control_evidence
  current_role: influential
  method_family:
  - autoregressive_ssl_conditioned_models
  claims:
  - claim_id: disentangling_speech_tokens_into_linguistic_content_and_speaker_attributes_within
    role: supports
    claim: Disentangling speech tokens into linguistic content and speaker attributes within a single-stream codec
      enables a standard LLM to perform zero-shot TTS without a multi-stage pipeline.
    source: §3, §4.1
    evidence: Disentangling speech tokens into linguistic content and speaker attributes within a single-stream
      codec enables a standard LLM to perform zero-shot TTS without a multi-stage pipeline.
    confidence: high
    relevance: low
  - claim_id: small_llm_backbones_can_achieve_competitive_zero_shot_tts_intelligibility
    role: supports
    claim: Small LLM backbones can achieve competitive zero-shot TTS intelligibility when the codec reduces per-token
      modeling complexity through semantic alignment.
    source: §6.4, Table 4
    evidence: Small LLM backbones can achieve competitive zero-shot TTS intelligibility when the codec reduces per-token
      modeling complexity through semantic alignment.
    confidence: high
    relevance: low
  - claim_id: single_stage_autoregressive_tts_consistently_trails_multi_stage_or_non
    role: supports
    claim: Single-stage autoregressive TTS consistently trails multi-stage or non-autoregressive methods on speaker
      similarity metrics, even when intelligibility is comparable.
    source: §6.4, Table 4, Limitation
    evidence: Single-stage autoregressive TTS consistently trails multi-stage or non-autoregressive methods on speaker
      similarity metrics, even when intelligibility is comparable.
    confidence: high
    relevance: low
  - claim_id: fsq_based_global_token_quantization_with_learnable_cross_attention_queries
    role: supports
    claim: FSQ-based global token quantization with learnable cross-attention queries produces better speaker attribute
      representation than group-VQ at equivalent token lengths.
    source: §6.2, Table 2
    evidence: FSQ-based global token quantization with learnable cross-attention queries produces better speaker
      attribute representation than group-VQ at equivalent token lengths.
    confidence: high
    relevance: medium
  - claim_id: attribute_controllable_tts_benefits_from_hierarchical_coarse_to_fine_prediction
    role: supports
    claim: Attribute-controllable TTS benefits from hierarchical coarse-to-fine prediction within the LM inference
      loop rather than requiring separate conditioning modules.
    source: §4.1, §6.3
    evidence: Attribute-controllable TTS benefits from hierarchical coarse-to-fine prediction within the LM inference
      loop rather than requiring separate conditioning modules.
    confidence: high
    relevance: medium
  limitations:
  - Speaker similarity in zero-shot cloning is meaningfully lower than multi-stage methods (SIM 0.672 vs. 0.774
    for MaskGCT on test-zh). The paper attributes this to AR variability without explicit disentanglement constraints
    between semantic and global tokens, and no solution is evaluated in this work.
  - The VoxBox training data and BiCodec codec are trained on separate, relatively limited datasets (3k hours for
    BiCodec; 102.5k hours for the LM). The BiCodec training data is English-only (LibriSpeech + Emilia EN/CN), which
    may limit acoustic reconstruction quality for languages outside this distribution.
  - The attribution control is evaluated primarily for gender, pitch, and speed; there is no evaluation of emotion
    control despite VoxBox containing emotion annotations for many source datasets.
  - Fine-grained numerical pitch control is evaluated on figures (Figs. 4–5) but no quantitative accuracy metric
    against target pitch values is reported, making it difficult to assess precision.
  caveats: []
- id: '2504.08528'
  published_date: "2025-04-11"
  entry_date: '2026-07-29'
  year: 2025
  venue: arXiv
  task:
  - SCA
  architecture: []
  relevance: medium
  evidence_role:
  - architecture_variant
  - acceleration_evidence
  - control_evidence
  - infrastructure
  current_role: influential
  method_family: []
  claims:
  - claim_id: models_that_excel_at_content_level_spoken_language_understanding_tend
    role: supports
    claim: Models that excel at content-level spoken language understanding tend to underperform on tasks requiring
      reasoning about acoustic and paralinguistic details, indicating that these capabilities require separate optimisation.
    source: §7.2, Table discussion; VoiceBench vs. MMAU comparison
    evidence: Models that excel at content-level spoken language understanding tend to underperform on tasks requiring
      reasoning about acoustic and paralinguistic details, indicating that these capabilities require separate optimisation.
    confidence: high
    relevance: medium
  - claim_id: cascaded_speech_pipelines_and_end_to_end_spoken_language_models
    role: supports
    claim: 'Cascaded speech pipelines and end-to-end spoken language models have complementary strengths: cascades
      are stronger on semantic tasks, while end-to-end models better handle speaker and paralinguistic information.'
    source: §7.2 "Comparisons of cascaded vs. end-to-end SLMs"
    evidence: 'Cascaded speech pipelines and end-to-end spoken language models have complementary strengths: cascades
      are stronger on semantic tasks, while end-to-end models better handle speaker and paralinguistic information.'
    confidence: high
    relevance: medium
  - claim_id: jointly_training_a_speech_encoder_with_an_llm_backbone_improves
    role: supports
    claim: Jointly training a speech encoder with an LLM backbone improves instruction following and paralinguistic
      understanding but introduces risk of catastrophic forgetting of pre-trained text capabilities.
    source: §7.2, §4.3
    evidence: Jointly training a speech encoder with an LLM backbone improves instruction following and paralinguistic
      understanding but introduces risk of catastrophic forgetting of pre-trained text capabilities.
    confidence: high
    relevance: medium
  - claim_id: full_duplex_spoken_dialogue_requires_architectural_support_beyond_turn_taking
    role: supports
    claim: Full-duplex spoken dialogue requires architectural support beyond turn-taking assumptions, and current
      approaches (dual-channel and time-multiplexing) each involve significant trade-offs in latency, naturalness,
      and modelling complexity.
    source: §6
    evidence: Full-duplex spoken dialogue requires architectural support beyond turn-taking assumptions, and current
      approaches (dual-channel and time-multiplexing) each involve significant trade-offs in latency, naturalness,
      and modelling complexity.
    confidence: high
    relevance: low
  - claim_id: the_tokenisation_choice_phonetic_tokens_versus_audio_codec_tokens_determines
    role: complicates
    claim: The tokenisation choice (phonetic tokens versus audio codec tokens) determines the trade-off between
      linguistic coherence and speaker or acoustic fidelity in spoken language model outputs.
    source: §3.1.2 "Comparison of token types"
    evidence: The tokenisation choice (phonetic tokens versus audio codec tokens) determines the trade-off between
      linguistic coherence and speaker or acoustic fidelity in spoken language model outputs.
    confidence: high
    relevance: low
  limitations:
  - The survey is primarily a snapshot of English-centric, high-resource SLM research. Multilingual, low-resource,
    and accessibility-focused SLMs receive minimal coverage, and the authors acknowledge that SLM research has largely
    not addressed dialects, accents, or speech-related medical conditions.
  - 'Beyond that scope limitation, several structural gaps are identified:'
  - '- Scaling behaviour for speech+text LMs and speech-aware text LMs is entirely unknown; existing scaling studies
    apply only to pure speech LMs. - Most SLM evaluations use non-overlapping benchmarks, making direct performance
    comparisons across model families impossible. - Very few SLMs are fully open-source (code, weights, and data),
    which prevents controlled ablations of design choices. - The representation of spoken output in SLMs is unsettled:
    discrete tokens, continuous flows, and hybrid approaches have not been systematically compared under controlled
    conditions. - Safety and trustworthiness evaluation for SLMs is nascent; non-verbal toxicity and speaker-type
    bias remain largely unstudied.'
  caveats: []
- id: '2504.10344'
  published_date: "2025-04-14"
  entry_date: '2026-07-29'
  year: 2025
  venue: arXiv
  task:
  - codec
  - TTS
  architecture:
  - autoregressive-LM
  - VAE
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: active_evidence
  method_family:
  - autoregressive_ssl_conditioned_models
  - vae_and_quantized_ssl_representations
  claims:
  - claim_id: frame_level_quantization_without_cross_frame_context_limits_codec_semantic
    role: supports
    claim: Frame-level quantization without cross-frame context limits codec semantic richness and increases the
      difficulty of autoregressive LM training on the resulting tokens.
    source: §1, Figure 1
    evidence: Frame-level quantization without cross-frame context limits codec semantic richness and increases
      the difficulty of autoregressive LM training on the resulting tokens.
    confidence: high
    relevance: low
  - claim_id: initialising_vq_codebook_entries_from_semantic_priors_derived_from_self
    role: supports
    claim: Initialising VQ codebook entries from semantic priors derived from self-supervised models improves downstream
      recognition accuracy without additional distillation overhead.
    source: §3.2, Table 6
    evidence: Initialising VQ codebook entries from semantic priors derived from self-supervised models improves
      downstream recognition accuracy without additional distillation overhead.
    confidence: high
    relevance: high
  - claim_id: lower_token_sequence_length_lower_hz_frame_rate_improves_both
    role: supports
    claim: Lower token-sequence length (lower Hz frame rate) improves both training and inference efficiency for
      audio language models, independent of bitrate.
    source: Appendix B, Table 8
    evidence: Lower token-sequence length (lower Hz frame rate) improves both training and inference efficiency
      for audio language models, independent of bitrate.
    confidence: high
    relevance: medium
  - claim_id: reconstruction_quality_alone_is_an_insufficient_predictor_of_a_codec
    role: supports
    claim: 'Reconstruction quality alone is an insufficient predictor of a codec''s suitability for autoregressive
      language modelling: codecs with stronger semantic content produce more robust LM-based TTS even at higher
      bitrates.'
    source: §4.6, Table 4
    evidence: 'Reconstruction quality alone is an insufficient predictor of a codec''s suitability for autoregressive
      language modelling: codecs with stronger semantic content produce more robust LM-based TTS even at higher
      bitrates.'
    confidence: high
    relevance: low
  - claim_id: a_masked_autoencoder_auxiliary_loss_during_codec_training_increases_semantic
    role: supports
    claim: A masked autoencoder auxiliary loss during codec training increases semantic information in learned representations
      at a modest reconstruction cost.
    source: §3.3, Table 6
    evidence: A masked autoencoder auxiliary loss during codec training increases semantic information in learned
      representations at a modest reconstruction cost.
    confidence: high
    relevance: low
  limitations:
  - All downstream LM experiments use the same 1B-parameter LLaMA backbone with limited training data (2000 hours
    speech, ~500 hours each sound/music). Gains observed may not transfer to larger-scale audio LM systems, where
    baseline tokenizers may reach ceiling performance.
  - The two-stage training adds complexity and discards large sub-networks (MAE encoder/decoder, AR prediction transformer)
    after training, increasing resource cost without reuse. The authors note this explicitly and flag it as future
    work. Sound and music reconstruction remains challenging at 0.41 kbps, with VISQOL scores substantially lower
    than speech. Semantic retention still lags SSL models for ASR (18.3% WER vs. 6.2% for WavLM), indicating the
    codec does not fully substitute for purpose-built SSL representations. Code and model weights were not released
    at submission time, limiting reproducibility.
  caveats: []
- id: iclr-2025-dGSOn7sdWg
  published_date: "2025-04-24"
  entry_date: '2026-07-29'
  year: 2025
  venue: ICLR
  task:
  - SCA
  architecture:
  - autoregressive-LM
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: influential
  method_family:
  - autoregressive_ssl_conditioned_models
  claims:
  - claim_id: reducing_the_token_rate_of_speech_representations_below_10_hz
    role: supports
    claim: Reducing the token rate of speech representations below 10 Hz is sufficient to preserve semantic content
      adequate for spoken language modelling, while substantially improving training and inference efficiency.
    source: §5.6, Table 6
    evidence: SyllableLM at 6.25 Hz, 90M parameters, matches or exceeds TWIST models up to 13B parameters on sBLIMP
      semantic understanding benchmarks, with 30x less training compute and 4.5x faster inference than an equal-sized
      TWIST baseline.
    confidence: high
    relevance: medium
  - claim_id: the_loss_surface_of_a_masked_prediction_ssl_model_encodes
    role: supports
    claim: The loss surface of a masked-prediction SSL model encodes latent syllabic segmentation boundaries discoverable
      without additional supervised signal or cross-modal supervision.
    source: §3.1, Table 1
    evidence: LossPred, applied to a frozen HuBERT student-teacher pair without any training, achieves F1-50 of
      59.6 on syllabic boundary detection, outperforming the feature-similarity-based baseline (47.3) while requiring
      no fine-tuning.
    confidence: high
    relevance: high
  - claim_id: iterative_student_teacher_distillation_over_pseudo_syllabic_boundaries_progressively_sharpens
    role: supports
    claim: Iterative student-teacher distillation over pseudo-syllabic boundaries progressively sharpens SSL encoder
      representations toward syllable-level organisation.
    source: §5.3, Table 2
    evidence: SylBoost applied to HuBERT improves boundary detection F1 from 60.1 (LossPred initialisation) to 70.2
      after two iterations; applying it to Data2Vec2 reaches 73.2, each iteration producing a measurable gain over
      the previous.
    confidence: high
    relevance: high
  - claim_id: low_frequency_speech_units_that_improve_semantic_modelling_efficiency_may
    role: complicates
    claim: Low-frequency speech units that improve semantic modelling efficiency may sacrifice robustness to speaker
      rate variation.
    source: Appendix A.5, Table 11
    evidence: SylBoost unit counts collapse under audio speedups of 0.5x and 0.6x relative to original length, performing
      comparably to SD-HuBERT only at mild speedups (0.8x–0.9x range), while showing greater robustness to slowdowns.
    confidence: high
    relevance: medium
  - claim_id: units_optimised_for_semantic_modelling_in_audiobook_speech_may_lose
    role: complicates
    claim: Units optimised for semantic modelling in audiobook speech may lose paralinguistic information, limiting
      applicability to domains requiring prosodic or tonal fidelity.
    source: §6
    evidence: The authors note that low-frequency SylBoost units may lose paralinguistic features such as tone,
      and the entire evaluation is conducted on audiobook data (LibriSpeech / LibriLight); performance on spontaneous
      or multi-speaker speech is not reported.
    confidence: high
    relevance: medium
  limitations:
  - All training and evaluation uses English audiobook data (LibriSpeech, LibriLight). Generalisation to spontaneous
    conversational speech, other languages, or multi-speaker settings is not demonstrated.
  - The interleaved vocoder decoding pipeline introduces a dependency on the TWIST tokeniser and vocoder for resynthesis,
    meaning the system is not fully end-to-end and unit bitrate for the final waveform is partially bounded by TWIST's
    own quality ceiling (WER 6.3%). The efficiency gains in the SpeechLM do not fully apply to the decoding pipeline,
    which involves an additional language model. Scaling beyond 300M parameters was not attempted due to compute
    constraints, leaving open whether the efficiency advantage persists at very large model scales. The base encoder
    quality (Data2Vec2 vs newer models like w2v-BERT 2.0) is acknowledged as a confounding factor in cross-model
    comparisons.
  caveats: []
- id: 2025.findings-naacl.130
  published_date: "2025-04-29"
  entry_date: '2026-07-29'
  year: 2025
  venue: NAACL
  task:
  - TTS
  architecture:
  - transformer-enc-dec
  - GAN
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  - evaluation_caution
  current_role: active_evidence
  method_family:
  - transformer_ssl_encoders_and_adapters
  - gan_decoders_for_ssl_units
  claims:
  - claim_id: in_video_to_speech_synthesis_discrete_acoustic_unit_intermediate_representations
    role: supports
    claim: In video-to-speech synthesis, discrete acoustic-unit intermediate representations preserve speech content
      but discard speaker-identifying acoustic detail relative to continuous Mel-spectrogram representations.
    source: §3, Table 1, Figure 1
    evidence: With ground-truth input held fixed, Unit-HiFiGAN scores SECS 0.555 / EER 40.52 versus HiFi-GAN's SECS
      0.894 / EER 22.96, and a Mel-spectrogram visual comparison shows the unit-based vocoder output diverges from
      ground truth in the frequency domain.
    confidence: high
    relevance: medium
  - claim_id: audio_visual_pre_trained_visual_encoders_can_supply_enough_speaker
    role: supports
    claim: Audio-visual pre-trained visual encoders can supply enough speaker-identity information from silent video
      to make explicit speaker embeddings unnecessary in video-to-speech synthesis.
    source: §5.2, Table 4
    evidence: DiVISe attains the best or near-best SECS scores on LRS2 (0.609) and LRS3 (0.624) among all compared
      methods, including several that use audio speaker embeddings during training or inference (SVTS, Multi-Task).
    confidence: high
    relevance: medium
  - claim_id: gains_in_objective_intelligibility_metrics_from_architectural_changes_do_not
    role: complicates
    claim: Gains in objective intelligibility metrics from architectural changes do not necessarily transfer to
      subjective audio-quality preference, especially when speaker-identity cues are removed from the listening
      context.
    source: Appendix C, Table 16
    evidence: In an audio-only MOS test without speaker reference images, Unit-HiFiGAN scores higher (4.37±0.11)
      than the Mel-based HiFi-GAN (4.24±0.12), the inverse of the speaker-matching and intelligibility rankings
      observed elsewhere in the paper.
    confidence: high
    relevance: medium
  - claim_id: the_benefit_of_audio_visual_pre_training_for_video_to
    role: refines
    claim: The benefit of audio-visual pre-training for video-to-speech synthesis is not uniform across model components;
      it most strongly aids the component responsible for output representation choice rather than uniformly improving
      all downstream metrics.
    source: §5.4.2, Table 7
    evidence: Removing pre-training degrades DiVISe across all four reported metrics (SECS, EER, ESTOI, WER), but
      for ReVISE removing pre-training mainly hurts intelligibility (WER 36.03 to 77.24) while leaving speaker metrics
      comparatively unaffected (SECS 0.5384 to 0.5304).
    confidence: high
    relevance: medium
  limitations:
  - Reported WER remains high (35.68-36.24% in the full-resource setting), and the paper's own subjective MOS audio-quality
    test (without speaker context) ranks the proposed Mel-based vocoder below the unit-based baseline, indicating
    the speaker-preservation gains are not accompanied by a clear audio-quality win in isolation.
  - The paper requires the same heavy mouth-region preprocessing pipeline as AV-HuBERT, which the authors note limits
    real-time applicability (§8, Limitations). Evaluation is restricted to English-only corpora (LRS2, LRS3); the
    authors explicitly flag multilingual generalization as untested due to compute constraints. The model's WER
    substantially trails dedicated ASR or text-conditioned TTS systems, reflecting the inherent difficulty of inferring
    content purely from lip movements without textual or acoustic priors. Latency analysis (§6) shows the conformer
    module, while improving intelligibility, reduces throughput relative to a no-conformer variant, an explicit
    accuracy/speed trade-off the paper surfaces but does not resolve.
  caveats: []
- id: 2025.findings-naacl.471
  published_date: "2025-04-29"
  entry_date: '2026-07-29'
  year: 2025
  venue: NAACL
  task:
  - evaluation
  architecture: []
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: active_evidence
  method_family: []
  claims:
  - claim_id: prosodic_cues_in_natural_speech_carry_information_sufficient_to_partially
    role: supports
    claim: Prosodic cues in natural speech carry information sufficient to partially guide spoken comprehension
      models above chance level.
    source: §4.1, Table 2
    evidence: A WavLM-based SQA model trained and tested on 300 Hz low-pass filtered (prosodic condition) audio
      achieves FF1 18.49 on SLUE-SQA-5 test, substantially above the white-noise chance baseline of 6.03.
    confidence: high
    relevance: medium
  - claim_id: ssl_based_spoken_language_models_do_not_effectively_leverage_prosodic
    role: complicates
    claim: SSL-based spoken language models do not effectively leverage prosodic information when lexical cues are
      simultaneously available, even when lexical data is a small minority of training.
    source: §4.2, Figure 5
    evidence: When only 10% of training data carries lexical information and 90% is prosodic, models rapidly shift
      toward lexical strategies, with lexical and natural evaluation losses converging near prosodic-only levels
      within training.
    confidence: high
    relevance: high
  - claim_id: the_use_of_tts_synthesised_speech_in_sqa_training_data
    role: complicates
    claim: The use of TTS-synthesised speech in SQA training data may not capture the prosodic characteristics of
      natural speech, limiting the study of prosody in comprehension models.
    source: §1, §3.1
    evidence: The paper selects SLUE-SQA-5 specifically because prior SQA datasets relied on TTS synthesis, whose
      prosodic properties differ from natural speech; the prosodic condition experiments are only valid under the
      natural-speech assumption.
    confidence: high
    relevance: medium
  - claim_id: prosodic_cues_in_sqa_are_at_least_partially_question_sensitive
    role: refines
    claim: Prosodic cues in SQA are at least partially question-sensitive rather than globally highlighting salient
      passage regions.
    source: §4.1, Table 4
    evidence: Random pairing of questions and contexts in the prosodic condition reduces FF1 from 18.49 to 9.77,
      showing the model's prosodic utilization depends on question-context alignment, not only on passage-level
      salience signals.
    confidence: high
    relevance: medium
  limitations:
  - 'The prosodic and lexical conditions cannot achieve perfect information separation: the 300 Hz low-pass filter
    retains some residual lexical information (WER ~57% at 300 Hz cut-off), and flattening F0/intensity preserves
    rhythm. Results for the prosodic condition therefore reflect a lower bound on prosodic-only performance, and
    the lexical condition''s prosodic residue may slightly inflate cross-condition scores.'
  - The study uses extractive SQA where answers are timestamped spans in spoken passages. This frames prosody as
    a localization cue; its role in open-ended or inferential comprehension tasks is not addressed. The choice of
    deeper WavLM layers (stronger semantic encoding) likely underrepresents prosodic information relative to earlier
    layers, which may encode prosody more directly. The corpus (Spoken Wikipedia) is read-aloud speech, so findings
    may not generalize to conversational or spontaneous speaking styles where prosody is more variable and communicatively
    richer.
  caveats: []
- id: 2025.naacl-demo.12
  published_date: "2025-04-29"
  entry_date: '2026-07-29'
  year: 2025
  venue: NAACL
  task:
  - TTS
  - SCA
  architecture:
  - autoregressive-LM
  relevance: high
  evidence_role:
  - core_evidence
  - control_evidence
  - infrastructure
  current_role: active_evidence
  method_family:
  - autoregressive_ssl_conditioned_models
  claims:
  - claim_id: a_speech_language_model_can_be_initialized_from_a_pre
    role: supports
    claim: A speech language model can be initialized from a pre-trained text LLM and jointly trained on speech
      recognition, speech synthesis, text continuation, and audio continuation without substantially degrading the
      text-only capability of the base model.
    source: §4.3, Table 5
    evidence: The 1.7B multi-task model trained on ASR, TTS, TextLM, and AudioLM objectives scores MMLU 30.5, ARC-C
      41.3, and HellaSwag 50.4, close to the text-only LLaMA-3.2-1B baseline (32.2, 32.8, 41.2) despite carrying
      three additional speech tasks.
    confidence: high
    relevance: medium
  - claim_id: concatenating_neural_codec_tokens_with_self_supervised_speech_representations_frame
    role: supports
    claim: Concatenating neural codec tokens with self-supervised speech representations frame-by-frame is a viable
      tokenization strategy for both speech understanding and generation tasks within a single sequential model.
    source: §3.3
    evidence: The "Codec_SSL" scheme (ESPnet-Codec combined with XEUS SSL tokens) is used for the headline ASR and
      TTS experiments and the paper reports it "behaves well in both speech understanding and generation" *(§3.3)*.
    confidence: high
    relevance: high
  - claim_id: a_decoder_only_autoregressive_speech_language_model_can_match_or
    role: supports
    claim: A decoder-only autoregressive speech language model can match or exceed dedicated, larger ASR-only systems
      on English benchmarks while using substantially fewer parameters.
    source: §4.2, Table 3
    evidence: A 442M-parameter ESPnet-SpeechLM ASR model reaches average WER 5.4% across six English test sets,
      matching OWSM v3.1-medium (1.02B, 5.4%) and beating Whisper-small (244M, 6.4%) and Whisper-medium (769M, 5.7%).
    confidence: high
    relevance: medium
  - claim_id: cross_system_comparisons_of_speech_language_models_reported_in_the
    role: complicates
    claim: Cross-system comparisons of speech language models reported in the literature are frequently not run
      under matched conditions, limiting how much can be concluded from any single performance table.
    source: §4.3, Table 5
    evidence: In the multi-task comparison (Table 5), competitor numbers for Moshi, VITA, GLM-4-Voice, and others
      are taken directly from their own published reports rather than reproduced by the authors, and the paper explicitly
      flags this with footnote markers.
    confidence: high
    relevance: medium
  - claim_id: combining_a_codec_tokenizer_with_a_self_supervised_tokenizer_frame
    role: complicates
    claim: Combining a codec tokenizer with a self-supervised tokenizer frame-by-frame is reported as an effective
      design choice but is not validated against single-tokenizer ablations in the same controlled setting.
    source: §3.3
    evidence: The claim that Codec_SSL tokenization "behaves well" rests on a single line of justification without
      a paired ablation against codec-only or SSL-only tokenization on the same task and dataset.
    confidence: high
    relevance: high
  limitations:
  - The multi-task model's TTS quality (Proxy MOS 3.99, WER 6.0%) is noticeably weaker than the single-task TTS
    model trained on the same architecture (Proxy MOS 4.03, WER 3.1%), indicating a capacity or interference cost
    to joint multi-task training that the paper reports but does not analyze further. Most training and evaluation
    data is English-only (the multilingual text corpus is used only for the TextLM objective, not for speech tasks),
    so the demonstrated speech capabilities are not evidence of multilingual generalization. Competitor numbers
    in the multi-task comparison table are drawn from third-party reports under unmatched training data and conditions
    rather than reproduced by the authors, which the paper itself notes. As a system/demo paper, the contribution
    is the toolkit and its reference recipes rather than a novel architecture or training method; the headline numbers
    serve to validate functionality rather than push state of the art on any individual benchmark.
  caveats: []
- id: 2025.naacl-long.484
  published_date: "2025-04-29"
  entry_date: '2026-07-29'
  year: 2025
  venue: NAACL
  task:
  - SCA
  - TTS
  architecture:
  - autoregressive-LM
  relevance: medium
  evidence_role:
  - architecture_variant
  - acceleration_evidence
  - control_evidence
  current_role: active_evidence
  method_family:
  - autoregressive_ssl_conditioned_models
  claims:
  - claim_id: a_single_stream_control_token_representation_for_full_duplex_dialogue
    role: supports
    claim: A single-stream control-token representation for full-duplex dialogue can match or exceed a doubled-channel
      alternating representation in both naturalness and behavioral adherence, while using fewer tokens per unit
      of audio duration.
    source: §7.1, §7.2, Table 4, Table 5
    evidence: On the same Llama3.2-1B backbone and Behavior-SD training data, the streamlined-unit variant scores
      4.09 naturalness MOS and 0.58 interruption adherence vs. 3.90 MOS and 1.21 interruption adherence for the
      alternating-unit variant.
    confidence: high
    relevance: low
  - claim_id: conditioning_a_speech_prompted_tts_synthesizer_on_a_speaker_s
    role: supports
    claim: Conditioning a speech-prompted TTS synthesizer on a speaker's very first utterance, in addition to their
      most recent utterance, materially improves long-range speaker identity consistency in independently synthesized
      multi-turn dialogue.
    source: §7.3, Table 6
    evidence: WavLM-Base+ cosine similarity between the prompt and synthesized speech rises from 0.680 (no prompt)
      to 0.889 when conditioning on the first utterance alone, and combining first-utterance and previous-utterance
      prompts yields the best transition smoothness (0.872) without degrading global identity (0.885).
    confidence: high
    relevance: low
  - claim_id: pretraining_a_dialogue_generation_model_on_text_only_dialogue_before
    role: complicates
    claim: Pretraining a dialogue generation model on text-only dialogue before fine-tuning on speech-unit sequences
      improves adherence to narrative content but is not necessary for behavioral adherence.
    source: §7.2, Table 5
    evidence: Removing the text-dialogue pretraining stage drops narrative adherence (GPT-4o-rated) from 3.11 to
      2.80, while behavioral adherence scores for filler words, backchannels, and interruptions remain comparable
      (0.10/0.87/0.64 vs. 0.15/0.87/0.58 with pretraining).
    confidence: high
    relevance: low
  - claim_id: cascaded_llm_then_tts_pipelines_for_spoken_dialogue_generation_can
    role: complicates
    claim: Cascaded LLM-then-TTS pipelines for spoken dialogue generation can achieve strong narrative coherence
      and sound quality but struggle to maintain conversational meaningfulness because they lack any mechanism to
      track behavioral context (speaker turn identity, backchannel placement) across the full dialogue.
    source: §7.1, §6.2, Table 3
    evidence: GPT-4o and Llama3-70B cascaded baselines achieve the highest narrative adherence scores (4.58 and
      4.09) but their meaningfulness MOS (3.97 for both) is lower than the proposed model (4.04), attributed by
      the authors to speaker confusion and misplaced backchannels in cascaded outputs.
    confidence: high
    relevance: low
  limitations:
  - All quantitative evaluation (human MOS, behavioral adherence, speaker consistency) is conducted entirely on
    Behavior-SD's own synthetic test split, which is itself generated by the same LLM and TTS pipeline used for
    training data. No evaluation is reported on real, human-recorded spoken dialogue, so it remains untested whether
    behavioral adherence and naturalness gains transfer outside the paper's own synthetic data distribution.
  - The authors note occasional mispronunciation, limited control over complex emotions beyond the five discrete
    categories used for style captioning, and limited non-lexical vocalization beyond laughter. The word error rate
    on the Behavior-SD test split (3.55%, measured via Whisper-Large V3) is higher than the underlying CosyVoice
    TTS system's WER on LibriSpeech (2.89%), attributed to filler-word insertion and proper-name misrecognition
    rather than a fundamental synthesis quality gap. The dataset and model are limited to two-speaker dialogues;
    the authors flag multi-speaker extension as future work. Behavior-SD's narratives derive from SODA, which itself
    derives from a smaller set of social-commonsense scenarios, so diversity of conversational topics is bounded
    by that source. As with any LLM/TTS-synthesized dataset, the authors flag inherited biases from training data
    and the risk of misuse for voice impersonation or deepfake audio.
  caveats: []
- id: 2025.naacl-short.65
  published_date: "2025-04-29"
  entry_date: '2026-07-29'
  year: 2025
  venue: NAACL
  task:
  - TTS
  architecture:
  - hybrid
  - GAN
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: active_evidence
  method_family:
  - hybrid_semantic_acoustic_ssl_systems
  - gan_decoders_for_ssl_units
  claims:
  - claim_id: ssl_feature_spaces_from_pre_trained_models_encode_cross_speaker
    role: supports
    claim: SSL feature spaces from pre-trained models encode cross-speaker structure that enables zero-shot voice
      transfer through nearest-neighbor retrieval, without speaker-specific training data.
    source: §2.1, Table 1
    evidence: kNN-TTS uses WavLM-Large layer 6 features, where frames from different speakers that are linearly
      close share phonetic information while preserving speaker identity; kNN retrieval over these features achieves
      SECS 0.72 and competitive MOS scores trained only on 24h of single-speaker LJSpeech data.
    confidence: high
    relevance: high
  - claim_id: zero_shot_multi_speaker_tts_competitive_with_large_multi_speaker
    role: supports
    claim: Zero-shot multi-speaker TTS competitive with large multi-speaker end-to-end systems can be achieved with
      single-speaker transcribed training data by delegating speaker identity to inference-time retrieval.
    source: §4, Table 1
    evidence: GlowkNN-TTS (24h training, single speaker) achieves N-MOS and S-MOS within the confidence intervals
      of HierSpeech++ (2,796h, 7299 speakers) and XTTS (27,282h, multi-speaker) on LibriSpeech test-clean.
    confidence: high
    relevance: medium
  - claim_id: retrieval_based_zero_shot_tts_requires_substantially_more_reference_audio
    role: complicates
    claim: Retrieval-based zero-shot TTS requires substantially more reference audio from the target speaker than
      embedding-based approaches to achieve sufficient quality.
    source: §Limitations, Figure 3b
    evidence: kNN-TTS requires approximately 30 seconds of target speaker audio for suitable intelligibility and
      around 1 minute for speaker similarity to plateau, whereas competing embedding-based systems show diminishing
      returns beyond 10-30 seconds of reference audio.
    confidence: high
    relevance: medium
  - claim_id: frame_level_knn_speaker_transfer_does_not_address_speaker_specific
    role: complicates
    claim: Frame-level kNN speaker transfer does not address speaker-specific duration and rhythm, leaving prosodic
      timing patterns fixed to the training speaker.
    source: §Limitations "Rhythmic variations"
    evidence: In kNN-TTS, utterance duration is determined entirely by the single-speaker Text-to-SSL model; frame-by-frame
      retrieval substitutes voice quality but does not adapt speaking rate or rhythm to the target speaker.
    confidence: high
    relevance: medium
  limitations:
  - 'The reference audio requirement is a practical limitation: kNN-TTS needs approximately 30 seconds of target
    speaker audio for usable intelligibility, which is notably higher than embedding-based competitors that can
    function with shorter clips. This restricts applicability in truly few-shot or single-utterance zero-shot scenarios.'
  - Duration adaptation to the target speaker is not addressed; the speaking rate and rhythm of the output always
    reflect the training speaker (LJSpeech). The paper proposes Urhythmic-style rhythm modeling as future work.
    Evaluation is English-only, and while the authors note potential for cross-lingual transfer (via kNN-VC cross-lingual
    capabilities), this is not demonstrated. Using mel-spectrogram features as an alternative to SSL features was
    ablated and found completely ineffective, confirming the dependency on WavLM's particular representational structure.
  caveats: []
- id: 2025.iwsds-1.11
  published_date: "2025-05-01"
  entry_date: '2026-07-29'
  year: 2025
  venue: IWSDS
  task:
  - SCA
  architecture: []
  relevance: medium
  evidence_role:
  - architecture_variant
  current_role: active_evidence
  method_family: []
  claims:
  - claim_id: self_supervised_speech_representations_capture_sufficient_prosodic_information_to_classify
    role: supports
    claim: Self-supervised speech representations capture sufficient prosodic information to classify paralinguistic
      attitudes from acoustic input alone, without linguistic features.
    source: §4.1, Table 4, Table 5
    evidence: A frozen HuBERT-large model using layer 12 embeddings averaged over time achieves macro-F1 of 0.909
      on a four-class attitude task in Japanese reading speech, exceeding human listener performance of 0.829 on
      the same data.
    confidence: high
    relevance: high
  - claim_id: environmental_reverberation_poses_a_more_persistent_challenge_than_additive_noise
    role: complicates
    claim: Environmental reverberation poses a more persistent challenge than additive noise for paralinguistic
      recognition models, and is not adequately addressed by speech enhancement postprocessing.
    source: §4.3, Table 7
    evidence: Under noisy conditions, MP-SENet speech enhancement recovers macro-F1 from 0.625 to 0.844; under reverberant
      conditions, the same model raises F1 only from 0.449 to 0.492, attributed to the difficulty of recovering
      prosodic structure from reverberant speech.
    confidence: high
    relevance: medium
  - claim_id: data_augmentation_with_noise_and_reverberation_during_training_does_not
    role: complicates
    claim: Data augmentation with noise and reverberation during training does not close the performance gap when
      those conditions appear at inference time.
    source: §4, §4.3, Table 7
    evidence: Despite four-fold augmentation using noise (DEMAND, MUSAN, FSD50K) and room impulse responses (BIRD)
      at varying SNR levels, macro-F1 drops from 0.912 (clean) to 0.625 (noisy) and 0.449 (noisy-reverberant) on
      held-out test conditions.
    confidence: high
    relevance: medium
  - claim_id: paralinguistic_attitude_recognition_systems_trained_on_reading_speech_datasets_may
    role: complicates
    claim: Paralinguistic attitude recognition systems trained on reading-speech datasets may overstate real-world
      spoken dialogue performance, where speech production is more spontaneous and less controlled.
    source: §3, §5
    evidence: The dataset consists entirely of scripts read with intended attitudes by crowd workers and actors;
      the authors identify generalisation to naturally-occurring speech directed at dialogue systems as unresolved
      future work.
    confidence: high
    relevance: low
  limitations:
  - 'The corpus is a Japanese reading-speech dataset where speakers deliberately produce each attitude; it does
    not reflect the ambiguity or variability of spontaneous human-machine dialogue speech. The four-class taxonomy
    covers core boundary attitudes but may miss finer-grained paralinguistic distinctions. Reverberation robustness
    remains an open problem: speech enhancement postprocessing helps only marginally, and joint enhancement-recognition
    models are proposed but not explored. The paper does not address how an inferred attitude label should influence
    downstream response generation policy within the dialogue system.'
  caveats: []
- id: '2505.13000'
  published_date: "2025-05-19"
  entry_date: '2026-07-29'
  year: 2025
  venue: arXiv
  task:
  - codec
  - TTS
  architecture:
  - autoregressive-LM
  relevance: medium
  evidence_role:
  - architecture_variant
  - acceleration_evidence
  - control_evidence
  current_role: influential
  method_family:
  - autoregressive_ssl_conditioned_models
  claims:
  - claim_id: directly_encoding_ssl_features_into_the_first_rvq_layer_preserves
    role: supports
    claim: Directly encoding SSL features into the first RVQ layer preserves significantly more semantic content
      than distilling SSL representations into codec tokens, particularly for tonal languages where pitch information
      is phonemically critical.
    source: §4.2, Table 2
    evidence: Directly encoding SSL features into the first RVQ layer preserves significantly more semantic content
      than distilling SSL representations into codec tokens, particularly for tonal languages where pitch information
      is phonemically critical.
    confidence: high
    relevance: high
  - claim_id: neural_audio_codecs_operating_at_lower_frame_rates_with_more
    role: supports
    claim: Neural audio codecs operating at lower frame rates with more quantization layers achieve superior audio
      quality at equivalent bitrates compared to higher frame rate codecs with fewer quantization layers.
    source: §4.3, Table 3
    evidence: Neural audio codecs operating at lower frame rates with more quantization layers achieve superior
      audio quality at equivalent bitrates compared to higher frame rate codecs with fewer quantization layers.
    confidence: high
    relevance: medium
  - claim_id: semantic_enhancement_of_the_first_rvq_layer_improves_downstream_tts
    role: supports
    claim: Semantic enhancement of the first RVQ layer improves downstream TTS speaker similarity as well as intelligibility,
      because higher semantic fidelity in RVQ-1 enables the waveform stream's quantizers to focus on acoustic detail
      rather than recovering content information.
    source: §4.4, Table 4
    evidence: Semantic enhancement of the first RVQ layer improves downstream TTS speaker similarity as well as
      intelligibility, because higher semantic fidelity in RVQ-1 enables the waveform stream's quantizers to focus
      on acoustic detail rather than recovering content information.
    confidence: high
    relevance: low
  - claim_id: the_quality_gap_between_distillation_based_and_direct_encoding_semantic
    role: supports
    claim: The quality gap between distillation-based and direct-encoding semantic codecs is substantially larger
      in Mandarin than in English, revealing a systematic limitation of distillation approaches for tonal languages.
    source: §4.2, Table 2
    evidence: The quality gap between distillation-based and direct-encoding semantic codecs is substantially larger
      in Mandarin than in English, revealing a systematic limitation of distillation approaches for tonal languages.
    confidence: high
    relevance: medium
  limitations:
  - The 12.5Hz DualCodec-based TTS systems consistently underperform their 25Hz counterparts on both WER and speaker
    similarity (Table 4, Table 6), indicating that the quality upper bound of the 12.5Hz variant is not yet competitive
    with the best open-source systems at 50Hz despite the frame rate reduction improving inference speed.
  - The paper evaluates TTS only on Seed-TTS-Eval; no subjective TTS listening tests are reported, so the MUSHRA
    gains in codec reconstruction may not fully translate to perceived TTS naturalness. Speaker similarity scores
    with DualCodec-VALLE remain below those of MaskGCT baselines that use separate semantic and acoustic tokenizers,
    suggesting the unified approach has not yet matched the best-performing two-stage pipeline design. The DualCodec
    encoder is substantially heavier than baselines (628M vs 38M for Mimi) due to the frozen w2v-BERT-2.0 model,
    increasing training-time compute, though the decoder remains lightweight for inference.
  caveats: []
- id: '2506.10274'
  published_date: "2025-06-12"
  entry_date: '2026-07-29'
  year: 2025
  venue: arXiv
  task:
  - codec
  - TTS
  - evaluation
  architecture: []
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  - evaluation_caution
  - infrastructure
  current_role: influential
  method_family: []
  claims:
  - claim_id: no_discrete_audio_tokenizer_consistently_outperforms_others_across_reconstruction_downstream
    role: supports
    claim: No discrete audio tokenizer consistently outperforms others across reconstruction, downstream discriminative
      tasks, and acoustic language modeling; the optimal tokenizer is task- and domain-dependent.
    source: §3.4, Figure 4
    evidence: No discrete audio tokenizer consistently outperforms others across reconstruction, downstream discriminative
      tasks, and acoustic language modeling; the optimal tokenizer is task- and domain-dependent.
    confidence: high
    relevance: medium
  - claim_id: increasing_the_number_of_codebooks_improves_signal_reconstruction_quality_but
    role: supports
    claim: Increasing the number of codebooks improves signal reconstruction quality but often degrades downstream
      task performance by adding redundancy that burdens representation-level learning.
    source: §3.2, §3.2 "Impact of Codebook Size"
    evidence: Increasing the number of codebooks improves signal reconstruction quality but often degrades downstream
      task performance by adding redundancy that burdens representation-level learning.
    confidence: high
    relevance: medium
  - claim_id: semantic_distillation_aligning_early_rvq_layers_with_ssl_features_improves
    role: supports
    claim: Semantic distillation (aligning early RVQ layers with SSL features) improves phonetic content preservation
      in acoustic tokenizers but may reduce cross-domain generalization when the distillation source is speech-specific.
    source: §4.2 "Distillation Effect", Table 16
    evidence: Semantic distillation (aligning early RVQ layers with SSL features) improves phonetic content preservation
      in acoustic tokenizers but may reduce cross-domain generalization when the distillation source is speech-specific.
    confidence: high
    relevance: high
  - claim_id: domain_alignment_between_tokenizer_training_data_and_evaluation_domain_is
    role: supports
    claim: Domain alignment between tokenizer training data and evaluation domain is the dominant factor in discrete
      audio codec performance, outweighing quantization method or bitrate choices.
    source: §4.2 "Data Domains"
    evidence: Domain alignment between tokenizer training data and evaluation domain is the dominant factor in discrete
      audio codec performance, outweighing quantization method or bitrate choices.
    confidence: high
    relevance: low
  - claim_id: continuous_speech_representations_e_g_wavlm_large_consistently_outperform_all
    role: supports
    claim: Continuous speech representations (e.g., WavLM-large) consistently outperform all discrete tokenizers
      on discriminative tasks, with the gap widening in low-resource conditions.
    source: §3.2, Table 7
    evidence: Continuous speech representations (e.g., WavLM-large) consistently outperform all discrete tokenizers
      on discriminative tasks, with the gap widening in low-resource conditions.
    confidence: high
    relevance: high
  limitations:
  - The TTS and audio LM evaluations are conducted with constrained training budgets (VALL-E on LibriTTS only; 300M
    audio LM at half the original training compute), which limits the practical transferability of findings on TTS
    tokenizer ranking to large-scale production settings.
  - The acoustic LM evaluation uses a single architecture (Qwen-2.5 based, 357M) and a single dataset (LibriHeavy),
    so findings on which tokenizers support better SLMs may not generalise to other LM architectures or training
    scales. The ablation study (Section 4) is restricted to a DAC-backbone framework; FSQ and SVQ conclusions may
    not transfer to other encoder-decoder designs. Evaluation metrics for audio generation (FAD, KLD, CLAP) conflate
    vocoder quality with language model quality, making it difficult to attribute performance differences to the
    tokenizer's representational properties versus its decoder quality. The paper does not evaluate streaming tokenizers
    under actual latency constraints, limiting guidance for real-time deployment. Trustworthiness considerations
    (voice deepfakes, bias) are raised as open concerns but not empirically evaluated.
  caveats: []
- id: '2507.00808'
  published_date: "2025-07-02"
  entry_date: '2026-07-29'
  year: 2025
  venue: arXiv
  task:
  - TTS
  architecture:
  - transformer-enc-dec
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  - evaluation_caution
  current_role: active_evidence
  method_family:
  - transformer_ssl_encoders_and_adapters
  claims:
  - claim_id: iterative_natural_language_feedback_can_progressively_refine_the_speaking_style
    role: supports
    claim: Iterative natural language feedback can progressively refine the speaking style of synthesized speech
      without accumulating naturalness degradation.
    source: §4.2, §4.4, Figure 4, Figure 6
    evidence: Over three interaction sessions, Iterative (ours) significantly outperformed the Identical baseline
      on a 5-point style refinement MOS, and naturalness MOS showed no significant difference between Iterative
      and the oracle condition across all style groups.
    confidence: high
    relevance: medium
  - claim_id: global_speech_embedding_based_conditioning_cannot_accurately_reflect_fine_grained
    role: complicates
    claim: Global speech embedding-based conditioning cannot accurately reflect fine-grained positional or linguistic
      instructions in expressive TTS.
    source: §5.1, §5.2, Table 4
    evidence: Low-scoring examples in the test set contained directions targeting specific word positions ("at the
      beginning", "at the end", "for the part of...") or linguistic modifications ("hold your breath", "place just
      a slight pause between words"), which the speech embedding manipulation approach could not handle.
    confidence: high
    relevance: medium
  - claim_id: semantic_similarity_of_style_direction_text_not_exact_wording_governs
    role: supports
    claim: Semantic similarity of style direction text, not exact wording, governs how well listeners perceive style
      refinement as aligned with the instruction.
    source: §4.3, Figure 5
    evidence: In the style refinement accuracy evaluation, directions semantically similar to the one used for refinement
      (Random Similar) yielded scores comparable to the Matched condition, while semantically dissimilar directions
      scored significantly lower across all style groups.
    confidence: high
    relevance: medium
  - claim_id: holistic_subjective_evaluation_scales_may_not_adequately_capture_fine_grained
    role: complicates
    claim: Holistic subjective evaluation scales may not adequately capture fine-grained stylistic alignment in
      iterative TTS refinement tasks.
    source: §4.2, §5.3
    evidence: Even the Actor-Guided oracle condition scored around 3 out of 5 on the iterative style refinement
      MOS, which the authors attribute to the evaluation task not fully discriminating subtle style differences;
      similar evaluation difficulties have been noted in text-to-image/video generation research.
    confidence: high
    relevance: low
  limitations:
  - All training and evaluation data is proprietary in-house Japanese speech from two voice actors. No public dataset
    is used, and no results are reported outside this setup. Reproducibility and generalization are untested.
  - 'The style refiner is speaker-dependent; the authors plan to extend to speaker-independent operation as future
    work. The current model refines only paralinguistic information (speaking style via global embeddings), not
    linguistic content, so instructions involving pauses, stress, or pitch accent placement cannot be followed.
    The directions cover only two of four practical categories from actual recording sessions (paralinguistic and
    text-expressible linguistic instructions), omitting demonstrative and gestural instructions entirely. The evaluation
    task design is also noted as an open problem: the relatively low absolute scores even under oracle conditions
    suggest that existing MOS paradigms do not cleanly measure this type of fine-grained iterative alignment, and
    more sensitive evaluation methods are needed.'
  caveats: []
- id: '2507.02176'
  published_date: "2025-07-02"
  entry_date: '2026-07-29'
  year: 2025
  venue: arXiv
  task:
  - evaluation
  architecture: []
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  - evaluation_caution
  current_role: active_evidence
  method_family: []
  claims:
  - claim_id: asv_embeddings_encode_static_anatomical_speech_features_but_systematically_fail
    role: supports
    claim: ASV embeddings encode static anatomical speech features but systematically fail to represent dynamic
      behavioral identity markers such as rhythm and timing patterns.
    source: §3.2, Figure 1
    evidence: Lasso regression predicting handcrafted features from ASV embeddings across LibriSpeech, ARCTIC, and
      L2-ARCTIC shows high r² for mean pitch, HNR, shimmer, and α-ratio, but near-zero r² for speech rate, voiced
      and unvoiced segment lengths, and pitch standard deviation across all seven tested ASV models.
    confidence: high
    relevance: medium
  - claim_id: eer_based_speaker_similarity_measurements_in_speech_synthesis_evaluation_are
    role: complicates
    claim: EER-based speaker similarity measurements in speech synthesis evaluation are susceptible to confounding
      factors unrelated to voice identity, which can invalidate comparisons between synthesis systems.
    source: §3.3, Table 2
    evidence: Duration-sorting same-speaker utterances depresses EER from 50% to 30–39% across all ASV models; SNR
      20 dB noise reduces EER to 15–38%; equalization shifts cause near-total failure in GE2E. Re-equalization and
      duration matching restore correct EER.
    confidence: high
    relevance: low
  - claim_id: characterizing_speaker_rhythm_for_identity_assessment_requires_modeling_phoneme_duration
    role: supports
    claim: Characterizing speaker rhythm for identity assessment requires modeling phoneme-duration distributions
      rather than aggregate measures such as mean speech rate.
    source: §3.4, Figure 2, Table 3
    evidence: Many L2-ARCTIC speakers share similar syllable rates but show substantially different voiced segment
      duration patterns; U3D Wasserstein distances clearly separate same-speaker pairs (avg. 2.15) from nearest-by-speech-rate
      pairs (18.40) and random pairs (21.53).
    confidence: high
    relevance: medium
  - claim_id: self_supervised_speech_unit_representations_serve_as_language_agnostic_substitutes
    role: refines
    claim: Self-supervised speech unit representations serve as language-agnostic substitutes for phoneme labels
      in rhythm analysis, avoiding the need for forced alignment.
    source: §3.4, Table 3
    evidence: 'U3D using HuBERT-derived unsupervised clusters achieves Wasserstein distance separation between speaker
      conditions nearly identical to forced-alignment-based phoneme rhythm analysis (same: 2.15 vs. 2.48; nearest:
      18.40 vs. 18.37; random: 21.53 vs. 24.43).'
    confidence: high
    relevance: high
  limitations:
  - U3D is validated as a discriminative metric (separating speaker pairs) but has not been validated against human
    perceptual judgments of rhythm similarity. Whether Wasserstein distances correlate with listeners' perception
    of rhythmic difference between voices remains untested.
  - The paper is limited to neutral speech, with the authors explicitly noting extension to expressive or conversational
    speech as future work. Experiments use clean studio-quality recordings (ARCTIC, L2-ARCTIC), so the behavior
    of confounding factors in in-the-wild multi-condition data remains untested. Recommended mitigation strategies
    (duration matching, re-equalization) assume access to the same text prompts used in genuine recordings, which
    may not be feasible in all evaluation settings. U3D's language-agnosticism is argued theoretically but not empirically
    tested on non-English languages.
  caveats: []
- id: '2507.03887'
  published_date: "2025-07-05"
  entry_date: '2026-07-29'
  year: 2025
  venue: arXiv
  task:
  - TTS
  architecture:
  - flow-matching
  relevance: medium
  evidence_role:
  - architecture_variant
  current_role: active_evidence
  method_family:
  - flow_matching_with_ssl_conditioning
  claims:
  - claim_id: joint_training_of_a_tts_generator_and_a_speech_discriminator
    role: supports
    claim: Joint training of a TTS generator and a speech discriminator with aligned optimization objectives improves
      out-of-domain generalization of model-specific attribution compared to training the discriminator independently.
    source: §5.4, Table 3
    evidence: Jointly training F5-TTS and a wav2vec 2.0 + LCNN discriminator improves out-of-domain AUC from 0.8823
      to 0.9421 and EER from 18.99% to 11.50%, while independent discriminator training on the same data shows substantially
      weaker generalization.
    confidence: high
    relevance: medium
  - claim_id: watermark_free_tts_traceability_through_joint_training_can_preserve_or
    role: supports
    claim: Watermark-free TTS traceability through joint training can preserve or slightly improve synthesis quality
      relative to the base model.
    source: §5.6, Table 5
    evidence: The jointly trained F5-TTS achieves WER 2.033% (vs. 2.202% baseline), UTMOS 3.958 (vs. 3.926), and
      speaker similarity 0.661 (vs. 0.659) on LibriSpeech-PC test-clean, indicating no quality penalty from the
      added traceability objective.
    confidence: high
    relevance: medium
  - claim_id: tts_attribution_discriminators_based_on_audio_feature_patterns_are_vulnerable
    role: complicates
    claim: TTS attribution discriminators based on audio feature patterns are vulnerable to additive noise and pitch
      manipulation, but robust to common audio processing operations.
    source: §5.5, Table 4
    evidence: Out-of-domain AUC drops by 0.1233 under additive noise (MUSAN) and 0.0901 under pitch shifting, while
      remaining stable (within 0.025 AUC) under resampling, time-stretching, reverb, volume scaling, and MP3 compression.
    confidence: high
    relevance: medium
  - claim_id: differentiability_of_the_full_synthesis_pipeline_from_text_input_to
    role: complicates
    claim: Differentiability of the full synthesis pipeline from text input to output waveform is a necessary precondition
      for watermark-free traceability via joint generator-discriminator training, excluding discrete token-based
      TTS architectures.
    source: §6
    evidence: The paper explicitly notes that VALL-E-type models using non-differentiable discrete token generation
      are incompatible with this framework because the discriminator's loss cannot backpropagate through the discrete
      generation step.
    confidence: high
    relevance: medium
  limitations:
  - The method is evaluated on a single TTS backbone (F5-TTS) with a single discriminator architecture, and the
    out-of-domain test uses only three other TTS systems. Generalization of the approach to a broader range of TTS
    models and adversarial conditions is untested.
  - The framework is architecturally restricted to TTS models that support end-to-end differentiation from text
    input to waveform, ruling out discrete token-based systems (VALL-E and its descendants). Additive noise and
    pitch shifting substantially degrade discriminator performance, suggesting that an adversary with access to
    these transformations could evade attribution. The paper does not evaluate whether the discriminator's traceability
    signal remains effective after the fine-tuned model is further fine-tuned by a third party or after the vocoder
    is swapped at inference time (though vocoder-robustness is claimed by design). Code release is conditional on
    paper acceptance and not yet available.
  caveats: []
- id: '2507.03912'
  published_date: "2025-07-05"
  entry_date: '2026-07-29'
  year: 2025
  venue: arXiv
  task:
  - TTS
  architecture: []
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: minor
  method_family: []
  claims:
  - claim_id: combining_acoustic_foundation_model_features_with_phoneme_level_linguistic_bert
    role: supports
    claim: Combining acoustic foundation model features with phoneme-level linguistic BERT features improves automatic
      prosody label prediction over either modality alone.
    source: §5.4, Table 1
    evidence: On the CSJ corpus, HuBERT-base + PnG BERT achieves 89.8% ACC accuracy versus 89.0% with HuBERT alone
      and 82.5% with PnG BERT alone; consistent gains appear on HL and BI labels.
    confidence: high
    relevance: high
  - claim_id: ssl_speech_features_substantially_outperform_hand_crafted_acoustic_features_for
    role: supports
    claim: SSL speech features substantially outperform hand-crafted acoustic features for automatic prosody annotation.
    source: §5.4, Table 1
    evidence: HuBERT-base achieves 89.8% ACC accuracy versus 75.2% for melspectrogram and 62.9% for F0 alone on
      CSJ; the gap widens further when comparing macro F1 scores on rare prosodic classes.
    confidence: high
    relevance: high
  - claim_id: language_matched_pre_training_benefits_ssl_model_utility_for_prosody
    role: supports
    claim: Language-matched pre-training benefits SSL model utility for prosody prediction in a pitch-accent language.
    source: §5.6, Table 2
    evidence: Japanese-trained HuBERT-base and wav2vec2.0-base outperform English and multilingual SSL variants
      across all four prosodic label types on CSJ, though margins are modest (under 1 percentage point on most labels).
    confidence: high
    relevance: high
  - claim_id: prosodic_boundary_and_pause_signals_cannot_be_reliably_estimated_from
    role: complicates
    claim: Prosodic boundary and pause signals cannot be reliably estimated from text alone; explicit acoustic evidence
      is required.
    source: §5.5
    evidence: Without acoustic input, pause presence prediction fails to distinguish short pauses; boundary pitch
      movement symbols in ACC labels show higher error rates than when HuBERT features are included.
    confidence: high
    relevance: medium
  limitations:
  - The paper does not evaluate whether the predicted prosodic labels improve downstream TTS naturalness or prosody
    controllability. All results are prosody annotation accuracy numbers on CSJ; the claimed benefit to TTS training
    remains unvalidated.
  - The system is evaluated only on Japanese, a pitch-accent language with specific prosodic structure. Applicability
    to other languages with different prosodic hierarchies (e.g., stress-timed languages with ToBI conventions)
    is not tested. The annotation model depends on accurate phoneme alignments from a forced aligner; errors in
    alignment propagate directly into label prediction. Finally, CSJ is a monologue corpus of academic lectures;
    generalization to conversational or spontaneous speech styles with different prosodic distributions is an open
    question.
  caveats: []
- id: '2507.08012'
  published_date: "2025-07-05"
  entry_date: '2026-07-29'
  year: 2025
  venue: arXiv
  task:
  - TTS
  architecture:
  - transformer-enc-dec
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: minor
  method_family:
  - transformer_ssl_encoders_and_adapters
  claims:
  - claim_id: latent_controllable_features_absent_from_a_tts_model_s_training
    role: supports
    claim: Latent controllable features absent from a TTS model's training annotations can be discovered by applying
      PCA to embeddings of fixed-input generated samples and enrolling the identified dimensions as new description-prompt
      labels.
    source: §4.3, Tables 2–3
    evidence: For T3 (no emotion labels), PCA of 1,000 Wav2Vec2 embeddings revealed emotional intensity as the primary
      variance axis; iterative re-labelling and fine-tuning progressively separated neutral from emotive utterances,
      improving neutral cluster assignment from 59.3% to 89.3% across rounds.
    confidence: high
    relevance: medium
  - claim_id: variance_based_feature_discovery_in_prompt_based_tts_is_unreliable
    role: complicates
    claim: Variance-based feature discovery in prompt-based TTS is unreliable when the model's output distribution
      is highly constrained by existing conditioning labels.
    source: §4.4, Figure 6
    evidence: T3-emotion, already trained with explicit emotion and intensity labels, generated highly consistent
      F0 contours under neutral-emotion prompts (Figure 6), leaving insufficient variance for meaningful feature
      discovery in the fixed-input analysis set.
    confidence: high
    relevance: medium
  - claim_id: embedding_based_clustering_of_tts_output_variance_can_surface_acoustic
    role: complicates
    claim: Embedding-based clustering of TTS output variance can surface acoustic artefacts of the training corpus
      (such as recording-condition variation) rather than prosodic features of interest.
    source: §4.4, Figure 8
    evidence: For T3-emotion with a diverse analysis set, the principal component correlated strongly with GeMaps-v01b
      loudness features attributable to microphone distance differences in the Talromur-3 corpus, not to prosody.
    confidence: high
    relevance: medium
  - claim_id: self_supervised_speech_representations_suitable_for_prosody_analysis_are_inherently
    role: complicates
    claim: Self-supervised speech representations suitable for prosody analysis are inherently entangled with speaker
      identity and linguistic content, requiring fixed-input generation sets to isolate prosodic variation.
    source: §3.2, Figure 1
    evidence: Wav2Vec2 summary embeddings cluster by target text and speaker identity regardless of which network
      layer is used (Figure 1), motivating the fixed-input analysis design where text, speaker, and prompt are held
      constant.
    confidence: high
    relevance: high
  limitations:
  - No subjective listening tests are reported. All quality and controllability assessments use automatic metrics
    (ASR-based WER, speaker embedding cosine similarity, diversity score). Whether the discovered features correspond
    to perceptually meaningful and user-controllable dimensions is not established.
  - The method is evaluated on a single Icelandic speaker (Ingrid) for the fine-tuning stages, with no cross-speaker
    or cross-language generalisation experiments. The paper does not address why T3-emotion fails and T3 succeeds
    beyond noting the former's consistent output distribution; whether adding a more diverse analysis set would
    help, or whether an alternative embedding choice (not Wav2Vec2) would be less sensitive to recording conditions,
    remains open. The use of Icelandic as the sole test language, while noted as language-agnostic in principle,
    is unverified on any other language.
  caveats: []
- id: '2507.04349'
  published_date: "2025-07-06"
  entry_date: '2026-07-29'
  year: 2025
  venue: arXiv
  task:
  - TTS
  architecture:
  - flow-matching
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: active_evidence
  method_family:
  - flow_matching_with_ssl_conditioning
  claims:
  - claim_id: frozen_backbone_adapter_training_can_add_fine_grained_time_varying
    role: supports
    claim: Frozen-backbone adapter training can add fine-grained, time-varying conditioning to pretrained flow-matching
      TTS with substantially less data than full fine-tuning.
    source: §1, §4.4, Table 5
    evidence: ControlNet blocks trained on approximately 400 hours of public emotion data achieve higher Emo-SIM
      and Aro-Val SIM than EmoCtrl-TTS, which requires 87k hours of training including 27k hours of in-house emotion
      data, while preserving the backbone's zero-shot voice cloning capability.
    confidence: high
    relevance: medium
  - claim_id: stronger_emotion_conditioning_in_flow_matching_tts_introduces_a_trade
    role: complicates
    claim: Stronger emotion conditioning in flow-matching TTS introduces a trade-off with text intelligibility.
    source: §4.3.4, Table 4
    evidence: Increasing control scale lambda from 0 to 1 consistently improves AutoPCP, Emo-SIM, and Aro-Val SIM
      but raises WER from 2.9% to 9.58% on the EMO-Change benchmark, showing that stronger emotion modulation introduces
      acoustic variations that degrade phoneme-level precision.
    confidence: high
    relevance: medium
  - claim_id: transformer_blocks_in_dit_based_tts_models_contribute_unequally_to
    role: supports
    claim: Transformer blocks in DiT-based TTS models contribute unequally to speaker identity and text intelligibility,
      and block selection is important for conditional control.
    source: §4.3.1, Figure 2, Table 3
    evidence: Layer-wise skip analysis on F5-TTS shows that removing specific blocks dramatically increases WER
      and reduces speaker similarity; excluding those critical blocks from ControlNet connections yields WER 0%
      and SIM-o 0.684 versus 8.9% WER and 0.630 SIM-o when all blocks are connected.
    confidence: high
    relevance: medium
  - claim_id: emotion_in_a_flow_matching_trajectory_is_concentrated_at_early
    role: supports
    claim: Emotion in a flow-matching trajectory is concentrated at early denoising steps, and restricting conditioning
      to this interval improves both efficiency and intelligibility.
    source: §4.3.2, Table 1, Table 8
    evidence: Training with flow step interval [0, 0.1] achieves Emo-SIM 0.565 and Aro-Val SIM 0.876 with 1.9% WER,
      whereas training on the full [0, 1] interval degrades to Emo-SIM 0.389 and Aro-Val SIM 0.674 with 0% WER;
      applying ControlNet only in early steps reduces per-sample inference time from 5.4s to 4.2s.
    confidence: high
    relevance: medium
  - claim_id: frame_level_emotion_features_from_self_supervised_ser_models_require
    role: complicates
    claim: Frame-level emotion features from self-supervised SER models require temporal smoothing to serve as effective
      conditioning signals for TTS.
    source: §4.3.3, Table 2
    evidence: Using emotion window size of 1 (no smoothing) produces Emo-SIM 0.500 and WER 4.7%; a window size of
      30 achieves Emo-SIM 0.565 and WER 1.9%, demonstrating that token-level SER features without pooling lose emotional
      coherence.
    confidence: high
    relevance: high
  limitations:
  - The underlying SER model cannot reliably recognize non-verbal vocalizations (laughing, crying) since these are
    not well-captured by the arousal-valence-dominance regression framework trained on utterance-level labels. This
    restricts the expressiveness of the emotion conditioning relative to systems with dedicated non-verbal encoders
    (e.g., EmoCtrl-TTS with its NV encoder).
  - Most baseline comparisons use values reported in a prior paper (EmoCtrl-TTS) rather than independently reproduced,
    limiting the reliability of direct numerical comparisons for closed-source models. On the cross-lingual JVNV
    S2ST benchmark, TTS-CtrlNet's speaker similarity (0.464) remains behind several baselines trained on much larger
    data. The method inherits any failure modes of the backbone F5-TTS, including artifacts during high-pitch synthesis
    and language bias toward English and Chinese.
  caveats: []
- id: '2507.06116'
  published_date: "2025-07-08"
  entry_date: '2026-07-29'
  year: 2025
  venue: arXiv
  task:
  - evaluation
  architecture: []
  relevance: medium
  evidence_role:
  - architecture_variant
  - evaluation_caution
  current_role: minor
  method_family: []
  claims:
  - claim_id: multi_task_learning_with_an_auxiliary_synthesis_model_classification_objective
    role: supports
    claim: Multi-task learning with an auxiliary synthesis-model classification objective can improve absolute MOS
      prediction accuracy at the system level by encouraging the model to learn discriminative acoustic quality
      features.
    source: §III.C, Table 2
    evidence: The MoE system with auxiliary classification achieves system-level MSE 0.056 (rank 1 of seven teams),
      a 21% reduction in MSE over the strongest single-objective baseline (MSE 0.071).
    confidence: high
    relevance: medium
  - claim_id: improvements_in_system_level_mos_prediction_accuracy_do_not_reliably
    role: complicates
    claim: Improvements in system-level MOS prediction accuracy do not reliably transfer to utterance-level prediction
      performance.
    source: §IV, Table 1
    evidence: Despite ranking first on system-level MSE, the proposed model places 3rd-6th across utterance-level
      metrics; the authors identify disjoint train/test annotator sets as a primary cause, not merely architectural
      limitations.
    confidence: high
    relevance: medium
  - claim_id: the_performance_gap_between_system_level_and_utterance_level_mos
    role: refines
    claim: The performance gap between system-level and utterance-level MOS prediction reflects fundamentally different
      task requirements, not a data-scaling problem alone.
    source: §IV
    evidence: Doubling training data from 400 to 800 samples and introducing MoE architecture yields only marginal
      utterance-level gains, while system-level gains are large; the paper attributes residual difficulty to micro-feature
      variability and rater-specific bias.
    confidence: high
    relevance: medium
  - claim_id: high_correlation_between_predicted_and_actual_mos_at_the_system
    role: complicates
    claim: High correlation between predicted and actual MOS at the system level does not imply high correlation
      at the utterance level for the same model.
    source: §IV, Tables 1-2
    evidence: The proposed model achieves LCC 0.978 on system-level evaluation but only LCC 0.811 on utterance-level
      evaluation, with similarly large gaps in SRCC and KTAU, confirming that aggregation masks prediction errors
      at finer granularity.
    confidence: high
    relevance: medium
  limitations:
  - The training and test annotator pools are completely disjoint (raters 0-9 for training, raters 10-19 for testing).
    This means utterance-level results measure how well the model generalises to unseen rater subjectivity, an inherently
    harder problem than standard train/test splits assume. Utterance-level claims should be interpreted with this
    constraint in mind.
  - 'The dataset is small (800 samples across 12 systems, derived from 30 text entries), making it difficult to
    assess whether findings generalise to larger and more diverse evaluation sets. System and sample diversity are
    limited: all audio is generated from the same text prompts, which may suppress naturalness variation that appears
    in open-domain evaluation.'
  - The challenge setup is unnamed in the paper, making it impossible to compare directly against published baseline
    systems outside the competition tables. The paper proposes rater adaptation and fine-grained acoustic features
    as future directions but does not experiment with them.
  caveats: []
- id: '2412.18603'
  published_date: "2025-07-10"
  entry_date: '2026-07-29'
  year: 2025
  venue: ICML
  task:
  - SCA
  architecture:
  - autoregressive-LM
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  - evaluation_caution
  current_role: influential
  method_family:
  - autoregressive_ssl_conditioned_models
  claims:
  - claim_id: linear_time_state_space_sequence_models_enable_stable_long_form
    role: supports
    claim: Linear-time state-space sequence models enable stable long-form speech generation that Transformer spoken
      LMs cannot sustain at comparable scale.
    source: §7.1, §7.2, Tables 3-4
    evidence: SpeechSSM-2B maintains N-MOS of 4.12-4.16 across all time strata up to 4+ minutes and stable semantic
      coherence in 16-minute generation, while all Transformer-based baselines and windowed extensions of prior
      spoken LMs degrade rapidly beyond their training lengths.
    confidence: high
    relevance: medium
  - claim_id: separating_speaker_conditioning_into_a_dedicated_acoustic_synthesis_stage_preserves
    role: supports
    claim: Separating speaker conditioning into a dedicated acoustic synthesis stage preserves voice identity more
      effectively in long-form speech generation than encoding speaker information in semantic tokens.
    source: §6.1, §7.1, Tables 2-3
    evidence: SpeechSSM achieves SpkrSim of 0.85 in long-form evaluation (vs. 0.33-0.41 for GSLM and TWIST), attributed
      to SoundStorm's speaker-prompted acoustic stage handling identity independently of the semantic LM.
    confidence: high
    relevance: medium
  - claim_id: standard_evaluation_metrics_for_spoken_language_models_are_unreliable_or
    role: complicates
    claim: Standard evaluation metrics for spoken language models are unreliable or saturating for modern systems,
      particularly in long-form settings.
    source: §6.1, §4
    evidence: sWUGGY scores negatively correlate with generation quality at large token vocabularies (32k vs. hundreds);
      short-form holistic MOS approaches ground truth at 7 seconds while qualitative failures persist; transcript
      perplexity favors repetitive outputs at default temperatures. Reference-based embedding metrics and LLM-as-judge
      side-by-sides prove more discriminative.
    confidence: high
    relevance: low
  - claim_id: long_form_speech_generation_remains_far_below_human_quality_even
    role: complicates
    claim: Long-form speech generation remains far below human quality even for state-of-the-art spoken language
      models.
    source: §7.1, Table 3
    evidence: No model achieves any wins against LibriSpeech-Long ground truth in 4-minute side-by-side comparisons
      judged by an LLM on fluency, coherence, logicality, and interestingness, demonstrating a large remaining quality
      gap despite strong short-form performance.
    confidence: high
    relevance: medium
  limitations:
  - Model weights are not released and inference requires TPU infrastructure; independent reproducibility of the
    long-form results is not currently possible.
  - 'Comparison fairness across baselines is imperfect: systems differ in training data scale, semantic tokenizer
    choice, token vocabulary size, and whether text was seen during training. The authors partially address this
    with the matched SpeechTransformer-2B baseline, but cannot control all confounds across the full set of comparisons.'
  - The semantic coherence advantage may partly reflect USM-v2's speaker invariance (reducing token entropy from
    vocal variation) rather than the SSM architecture per se. SpeechSSM-X, the extemporaneous variant trained on
    216k hours of monologue speech, is evaluated qualitatively only without quantitative benchmarks. LibriSpeech-Long
    covers read audiobook speech; generalization to spontaneous, multi-speaker, or conversational long-form speech
    is not evaluated.
  caveats: []
- id: '2507.16632'
  published_date: "2025-07-22"
  entry_date: '2026-07-29'
  year: 2025
  venue: arXiv
  task:
  - SCA
  - TTS
  architecture:
  - autoregressive-LM
  - flow-matching
  - hybrid
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: influential
  method_family:
  - autoregressive_ssl_conditioned_models
  - flow_matching_with_ssl_conditioning
  - hybrid_semantic_acoustic_ssl_systems
  claims:
  - claim_id: interleaving_discrete_audio_tokens_with_text_tokens_in_a_shared
    role: supports
    claim: Interleaving discrete audio tokens with text tokens in a shared language model vocabulary enables end-to-end
      spoken dialogue systems that respond coherently to paralinguistic cues without requiring separate speech synthesis
      pipelines.
    source: §3.1, §4.6
    evidence: Interleaving discrete audio tokens with text tokens in a shared language model vocabulary enables
      end-to-end spoken dialogue systems that respond coherently to paralinguistic cues without requiring separate
      speech synthesis pipelines.
    confidence: high
    relevance: low
  - claim_id: reinforcement_learning_applied_to_audio_language_models_can_improve_reasoning
    role: supports
    claim: Reinforcement learning applied to audio language models can improve reasoning efficiency in complex acoustic
      scenarios while preserving generation quality.
    source: §3.4
    evidence: Reinforcement learning applied to audio language models can improve reasoning efficiency in complex
      acoustic scenarios while preserving generation quality.
    confidence: high
    relevance: medium
  - claim_id: retrieval_augmented_generation_and_external_tool_calling_can_substantially_reduce
    role: supports
    claim: Retrieval-augmented generation and external tool calling can substantially reduce hallucination and expand
      capability (timbre switching, web-grounded responses) in large audio language models without architectural
      redesign.
    source: §3.1, §4.5
    evidence: Retrieval-augmented generation and external tool calling can substantially reduce hallucination and
      expand capability (timbre switching, web-grounded responses) in large audio language models without architectural
      redesign.
    confidence: high
    relevance: medium
  - claim_id: comprehensive_multi_task_pre_training_across_asr_tts_translation_and
    role: supports
    claim: Comprehensive multi-task pre-training across ASR, TTS, translation, and conversation significantly improves
      spoken dialogue performance in low-resource languages and accented speech.
    source: §3.2, §4.1
    evidence: Comprehensive multi-task pre-training across ASR, TTS, translation, and conversation significantly
      improves spoken dialogue performance in low-resource languages and accented speech.
    confidence: high
    relevance: low
  - claim_id: existing_audio_language_model_benchmarks_fail_to_capture_fine_grained
    role: complicates
    claim: Existing audio language model benchmarks fail to capture fine-grained paralinguistic comprehension and
      voice-triggered tool invocation, leaving important capability dimensions systematically unmeasured.
    source: §4.2, §4.5
    evidence: Existing audio language model benchmarks fail to capture fine-grained paralinguistic comprehension
      and voice-triggered tool invocation, leaving important capability dimensions systematically unmeasured.
    confidence: high
    relevance: medium
  limitations:
  - Two of the main evaluation benchmarks (StepEval-Audio-Paralinguistic, StepEval-Audio-Toolcall) are introduced
    by the authors themselves and have not been validated independently. Results on these benchmarks may overstate
    absolute capability levels even if relative comparisons are informative.
  - Model size is not reported, which makes parameter-count comparisons with other open-source systems (Kimi-Audio,
    Qwen2.5-Omni) difficult to interpret fairly. The open-source mini variant uses Qwen2.5-7B as its backbone, but
    the full model's parameter count remains undisclosed.
  - The audio search tool relies on a proprietary library of hundreds of thousands of speech samples, limiting reproducibility
    of this capability. Latency characteristics are not reported; real-time performance is claimed via VAD and deployment
    infrastructure inherited from Step-Audio but not benchmarked here.
  - Evaluation is primarily Chinese-English bilingual. Performance on other language families, especially those
    with minimal pre-training coverage, is largely uncharacterized. The tool-calling evaluation uses synthesized
    speech conversations generated from text scripts rather than natural speech, which may not reflect real-world
    tool invocation patterns.
  caveats: []
- id: '2507.18897'
  published_date: "2025-07-25"
  entry_date: '2026-07-29'
  year: 2025
  venue: arXiv
  task:
  - codec
  architecture:
  - GAN
  relevance: medium
  evidence_role:
  - architecture_variant
  current_role: active_evidence
  method_family:
  - gan_decoders_for_ssl_units
  claims:
  - claim_id: single_quantizer_neural_codecs_can_match_multi_quantizer_systems_in
    role: supports
    claim: Single-quantizer neural codecs can match multi-quantizer systems in perceived speech quality at far lower
      bitrates when combined with stabilized VQ spaces and asymmetric decoder architectures.
    source: §4.3, Table 1
    evidence: HH-Codec achieves UTMOS 3.61 on LibriTTS test-clean at 0.3 kbps with a single quantizer, competitive
      with DAC's 4-quantizer configuration at 4 kbps (UTMOS 3.41) and markedly above SpeechTokenizer's single-quantizer
      variant at 0.75 kbps (UTMOS 1.26).
    confidence: high
    relevance: medium
  - claim_id: multi_layer_vq_training_with_single_layer_inference_where_auxiliary
    role: supports
    claim: Multi-layer VQ training with single-layer inference, where auxiliary quantizer layers act as regularizers,
      substantially improves codebook utilization in high-compression single-quantizer settings.
    source: §4.4, Tables 2-3
    evidence: SLM-VQ achieves 98% codebook utilization at 8192 entries versus 56% for Classic VQ and 92% for single-layer
      SLM-VQ, while improving UTMOS from 2.76 (Classic VQ) to 3.07 on LibriTTS test-other.
    confidence: high
    relevance: medium
  - claim_id: dual_domain_supervision_combining_intermediate_mel_spectrogram_and_final_audio
    role: supports
    claim: Dual-domain supervision combining intermediate mel-spectrogram and final audio reconstruction objectives
      is critical for stable high-compression neural codec training.
    source: §4.4, Table 2
    evidence: Reducing to single audio-domain supervision drops UTMOS from 3.07 to 1.85 and SPK-SIM from 0.64 to
      0.33 on LibriTTS test-other, the largest degradation across all ablation variants.
    confidence: high
    relevance: low
  - claim_id: standard_adversarial_codec_training_recipes_break_down_at_extreme_compression
    role: complicates
    claim: Standard adversarial codec training recipes break down at extreme compression ratios, requiring architectural
      and procedural modifications to avoid collapse.
    source: §1
    evidence: Below 0.3 kbps with existing methods, the paper documents adversarial training collapse, a 63% UTMOS
      drop below 30 tokens/s, 43% codebook utilization at 8192 entries, and minimal benefit from expanding training
      data, all addressed in HH-Codec through SLM-VQ and progressive training.
    confidence: high
    relevance: low
  limitations:
  - HuBERT-based semantic distillation is trained on English, so multilingual performance of SLM-VQ is unknown.
    The downstream spoken language modeling experiment measures only training loss reduction rather than end-to-end
    TTS quality, leaving it unclear how the 24-token-per-second compression affects downstream synthesis intelligibility
    and naturalness. Training data conditions differ across compared baselines (WavTokenizer and DAC use larger
    or different datasets), which limits direct attribution of gains to architecture versus data. The ablation study
    is conducted on a subset of training data only (LibriTTS train-100/360, not the full training set including
    Emilia), which may underestimate some component contributions.
  caveats: []
- id: 2025.acl-demo.37
  published_date: "2025-07-27"
  entry_date: '2026-07-29'
  year: 2025
  venue: ACL
  task:
  - VC
  architecture:
  - hybrid
  relevance: medium
  evidence_role:
  - architecture_variant
  - acceleration_evidence
  - control_evidence
  - evaluation_caution
  current_role: active_evidence
  method_family:
  - hybrid_semantic_acoustic_ssl_systems
  claims:
  - claim_id: articulatory_feature_spaces_enable_interpretable_content_speaker_disentanglement_in_voice
    role: supports
    claim: Articulatory feature spaces enable interpretable content-speaker disentanglement in voice conversion
      without sacrificing intelligibility relative to SSL-based approaches.
    source: §5.3, Table 1
    evidence: Articulatory feature spaces enable interpretable content-speaker disentanglement in voice conversion
      without sacrificing intelligibility relative to SSL-based approaches.
    confidence: high
    relevance: high
  - claim_id: real_time_zero_shot_voice_conversion_on_cpu_hardware_is
    role: supports
    claim: Real-time zero-shot voice conversion on CPU hardware is achievable below 70 ms end-to-end latency while
      maintaining naturalness MOS above 3.8.
    source: §3.6, Table 1
    evidence: Real-time zero-shot voice conversion on CPU hardware is achievable below 70 ms end-to-end latency
      while maintaining naturalness MOS above 3.8.
    confidence: high
    relevance: low
  - claim_id: causal_ddsp_vocoders_conditioned_on_articulatory_features_provide_competitive_synthesis
    role: supports
    claim: Causal DDSP vocoders conditioned on articulatory features provide competitive synthesis quality compared
      to GAN-based alternatives at substantially lower computational cost.
    source: §2.3, §3.5
    evidence: Causal DDSP vocoders conditioned on articulatory features provide competitive synthesis quality compared
      to GAN-based alternatives at substantially lower computational cost.
    confidence: high
    relevance: medium
  - claim_id: voice_conversion_systems_trained_with_static_noise_augmentation_degrade_gracefully
    role: complicates
    claim: Voice conversion systems trained with static noise augmentation degrade gracefully down to approximately
      20 dB SNR input but fail at 10 dB, suggesting a practical noise floor for real-time deployment.
    source: §5.4
    evidence: Voice conversion systems trained with static noise augmentation degrade gracefully down to approximately
      20 dB SNR input but fail at 10 dB, suggesting a practical noise floor for real-time deployment.
    confidence: high
    relevance: low
  limitations:
  - '- EMA representation omits nasal cavity and laryngeal dynamics, limiting modeling of nasal sounds and vocal
    fry. - Pseudo-EMA labels come from a WavLM model pretrained on English; cross-lingual performance is limited.
    - Sensitivity to input quality below 20 dB SNR, especially for white noise. - Model weights will not be open-sourced
    due to misuse concerns. - Future work: prompt-free conversion by offline target speaker design (gender, age,
    emotion, accent).'
  caveats: []
- id: 2025.acl-long.681
  published_date: "2025-07-27"
  entry_date: '2026-07-29'
  year: 2025
  venue: ACL
  task:
  - SCA
  architecture:
  - autoregressive-LM
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  - infrastructure
  current_role: active_evidence
  method_family:
  - autoregressive_ssl_conditioned_models
  claims:
  - claim_id: large_scale_instruction_datasets_that_span_diverse_acoustic_and_paralinguistic
    role: supports
    claim: Large-scale instruction datasets that span diverse acoustic and paralinguistic dimensions improve speech-text
      LLM generalization on instruction-following benchmarks beyond what smaller task-specific corpora achieve.
    source: §5.2, Table 3
    evidence: SIFT-LLM outperforms O-ASQA-LLM (same backbone, fine-tuned on 2.7M-example OpenASQA) across all benchmarks,
      including 46.1% vs. 22.9% on EvalSIFT closed-ended and 57.4% vs. 45.9% on Dynamic-Superb.
    confidence: high
    relevance: medium
  - claim_id: instruction_fine_tuning_on_diverse_speech_tasks_reduces_specialized_performance
    role: complicates
    claim: Instruction fine-tuning on diverse speech tasks reduces specialized performance on foundational tasks
      relative to the pre-trained model checkpoint.
    source: §5.3, Table 4
    evidence: SIFT-LLM WER on LibriSpeech test-clean rises from 2.5% (50K pre-training checkpoint) to 3.5% after
      instruction fine-tuning, consistent with behavior observed in Qwen2-Audio and its instruction fine-tuned variant.
    confidence: high
    relevance: medium
  - claim_id: codec_representations_that_jointly_encode_acoustic_and_semantic_content_enable
    role: supports
    claim: Codec representations that jointly encode acoustic and semantic content enable more accurate controllable
      speech generation than semantic-only representations.
    source: §5.4, Table 6
    evidence: X-codec2 (fused semantic and acoustic codes) achieves QWK of 0.69 for pitch variation versus 0.15
      for HuBERT codes (semantic-only) on SIFT-LLM GEN controllable generation evaluations.
    confidence: high
    relevance: low
  - claim_id: semantic_only_discrete_speech_representations_are_insufficient_for_reliable_speaker
    role: complicates
    claim: Semantic-only discrete speech representations are insufficient for reliable speaker-dependent controllable
      generation, as they primarily encode content rather than prosodic and acoustic identity.
    source: §5.4, §F.3
    evidence: The HuBERT-code vocoder setup achieves approximately 50% gender accuracy against instruction-specified
      gender, compared to 95.8% for X-codec2; pitch variation QWK from original audio degrades substantially in
      HuBERT re-synthesis (0.11 QWK, Table 18 Appendix F.3).
    confidence: high
    relevance: medium
  limitations:
  - SIFT-LLM does not achieve state-of-the-art results on foundational speech tasks; the instruction fine-tuning
    trades off specialized ASR and translation capability, limiting its use as a drop-in replacement for task-specialized
    models.
  - Controllable generation evaluation relies on automatically extracted acoustic features compared against instruction-specified
    targets, which indirectly assesses controllability but cannot measure perceptual quality or naturalness. The
    QWK and MAE metrics measure categorical agreement rather than human-perceived fidelity.
  - The LLM-as-a-judge evaluation methodology for open-ended and classification tasks introduces judge model dependency;
    results may shift with different judge LLMs, and the interaction between judge capability and task difficulty
    is not systematically studied.
  - Hallucination when queried about content unrelated to the input audio is noted as an open issue; attempts to
    mitigate it by adding off-topic instructions with higher weight degraded speech understanding performance, suggesting
    a robustness-capability tension.
  caveats: []
- id: 2025.acl-long.682
  published_date: "2025-07-27"
  entry_date: '2026-07-29'
  year: 2025
  venue: ACL
  task:
  - TTS
  - SCA
  - evaluation
  architecture:
  - autoregressive-LM
  - GAN
  - hybrid
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  - infrastructure
  current_role: active_evidence
  method_family:
  - autoregressive_ssl_conditioned_models
  - gan_decoders_for_ssl_units
  - hybrid_semantic_acoustic_ssl_systems
  claims:
  - claim_id: gan_based_vocoders_occupy_the_dominant_position_in_production_speech
    role: supports
    claim: GAN-based vocoders occupy the dominant position in production speech generation systems because their
      computational efficiency advantage over autoregressive and diffusion alternatives is orders of magnitude,
      at acceptable perceptual quality cost.
    source: §3.3.1, §F.3, Table 8
    evidence: GAN-based vocoders occupy the dominant position in production speech generation systems because their
      computational efficiency advantage over autoregressive and diffusion alternatives is orders of magnitude,
      at acceptable perceptual quality cost.
    confidence: high
    relevance: medium
  - claim_id: continued_pretraining_from_a_text_language_model_checkpoint_improves_speechlm
    role: supports
    claim: Continued pretraining from a text language model checkpoint improves SpeechLM convergence and downstream
      task performance compared to random initialization.
    source: §4.2.1
    evidence: Continued pretraining from a text language model checkpoint improves SpeechLM convergence and downstream
      task performance compared to random initialization.
    confidence: high
    relevance: medium
  - claim_id: semantic_tokenizers_and_acoustic_tokenizers_represent_complementary_capability_profiles_strong
    role: supports
    claim: Semantic tokenizers and acoustic tokenizers represent complementary capability profiles — strong semantic
      content fidelity versus strong acoustic reconstruction fidelity — and no single tokenizer type dominates both
      dimensions.
    source: §3.1, §F.2, Table 6
    evidence: Semantic tokenizers and acoustic tokenizers represent complementary capability profiles — strong semantic
      content fidelity versus strong acoustic reconstruction fidelity — and no single tokenizer type dominates both
      dimensions.
    confidence: high
    relevance: medium
  - claim_id: post_alignment_of_speechlms_via_preference_optimization_addresses_qualitatively_different
    role: supports
    claim: Post-alignment of SpeechLMs via preference optimization addresses qualitatively different failure modes
      (semantic inconsistency, token distribution mismatch) than post-alignment of text LLMs.
    source: §4.2.3
    evidence: Post-alignment of SpeechLMs via preference optimization addresses qualitatively different failure
      modes (semantic inconsistency, token distribution mismatch) than post-alignment of text LLMs.
    confidence: high
    relevance: medium
  - claim_id: full_duplex_spoken_interaction_simultaneous_bidirectional_speech_with_interruption_support
    role: supports
    claim: Full-duplex spoken interaction — simultaneous bidirectional speech with interruption support — requires
      joint modeling of both speaker streams and remains an open research challenge.
    source: §4.3
    evidence: Full-duplex spoken interaction — simultaneous bidirectional speech with interruption support — requires
      joint modeling of both speaker streams and remains an open research challenge.
    confidence: high
    relevance: medium
  limitations:
  - 'Coverage necessarily lags the field: systems published after mid-2024 receive limited treatment, and the overall
    corpus skews heavily toward English and Mandarin. The safety section identifies toxicity and speaker privacy
    risks but does not analyze them quantitatively. End-to-end training that backpropagates gradients from vocoder
    output to tokenizer input is flagged as a potentially high-value research direction but remains unexplored.
    The question of whether incorporating text modality fundamentally benefits or constrains speech intelligence
    is left open.'
  caveats: []
- id: 2025.acl-long.790
  published_date: "2025-07-27"
  entry_date: '2026-07-29'
  year: 2025
  venue: ACL
  task:
  - VC
  architecture:
  - flow-matching
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: active_evidence
  method_family:
  - flow_matching_with_ssl_conditioning
  claims:
  - claim_id: self_consistency_training_enables_shortcut_flow_matching_to_match_full
    role: supports
    claim: Self-consistency training enables shortcut flow matching to match full-step quality in voice conversion
      with as few as two inference steps.
    source: §4.2, Table 1
    evidence: R-VC at NFE=2 matches NFE=10 across all quality metrics (SECS 0.930 vs 0.931, UTMOS 4.1 vs 4.1, QMOS
      4.03 vs 4.05, SMOS 4.11 vs 4.12) while reducing inference time by 2.83x; vanilla CFM degrades sharply below
      10 steps.
    confidence: high
    relevance: low
  - claim_id: explicit_rhythm_modeling_via_a_target_conditioned_duration_model_substantially
    role: supports
    claim: Explicit rhythm modeling via a target-conditioned duration model substantially improves emotion style
      transfer in zero-shot VC.
    source: §4.3, §4.5, Table 2, Table 4
    evidence: Removing the duration module from R-VC drops the emotion score from 0.59 to 0.425 on the ESD dataset,
      while also increasing WER slightly; baselines that preserve source rhythm score 0.395-0.489.
    confidence: high
    relevance: medium
  - claim_id: fine_grained_duration_prediction_in_non_autoregressive_models_introduces_instability
    role: complicates
    claim: Fine-grained duration prediction in non-autoregressive models introduces instability in voice conversion
      that coarser duration strategies do not fully resolve.
    source: §7, Table 4
    evidence: R-VC's masked transformer duration model produces occasional over-extended pronunciations; sentence-level
      duration as a fallback yields worse WER (9.86 vs 6.95) and UTMOS (3.58 vs 3.85), offering no stability improvement
      in practice.
    confidence: high
    relevance: low
  - claim_id: data_perturbation_before_discrete_content_tokenisation_reduces_timbre_leakage_more
    role: supports
    claim: Data perturbation before discrete content tokenisation reduces timbre leakage more effectively than relying
      on the self-supervised representation alone.
    source: §4.5, Table 4
    evidence: Removing pitch perturbation before HuBERT token extraction degrades WER from 3.51 to 7.28 and speaker
      similarity from 0.930 to 0.869, confirming that perturbation actively suppresses content-irrelevant speaker
      information.
    confidence: high
    relevance: high
  limitations:
  - 'The masked transformer duration model has a known instability: inaccurate predictions cause over-extended pronunciations.
    Sentence-level duration as an alternative proved worse in both stability and quality, leaving robust duration
    modeling as an unresolved challenge.'
  - The system is evaluated only on English (MLS, LibriSpeech, ESD) and English Seed-TTS subsets; generalisation
    to cross-lingual or multilingual VC is untested. Training data (20k hours) is smaller than top competitors such
    as CosyVoice-VC (171k hours), which makes speaker similarity comparisons somewhat favourable to R-VC but also
    means that high-similarity performance on out-of-distribution accents or recording conditions is unknown. The
    RTF of 0.12 using 2-step inference is faster than most flow-matching competitors but still 20% slower than non-diffusion
    methods (FACodec-VC RTF 0.10).
  caveats: []
- id: 2025.acl-long.817
  published_date: "2025-07-27"
  entry_date: '2026-07-29'
  year: 2025
  venue: ACL
  task:
  - SCA
  - TTS
  architecture:
  - autoregressive-LM
  relevance: medium
  evidence_role:
  - architecture_variant
  - acceleration_evidence
  - control_evidence
  current_role: active_evidence
  method_family:
  - autoregressive_ssl_conditioned_models
  claims:
  - claim_id: boundary_aware_speech_representations_are_critical_for_enabling_offline_trained
    role: supports
    claim: Boundary-aware speech representations are critical for enabling offline-trained speech LLMs to perform
      simultaneous inference via wait-k strategies.
    source: §5.1, §5.2, Tables 4, 11
    evidence: CIF-based boundary-aware prompts outperform fixed downsampling by approximately 4 ASR-BLEU points
      at equivalent latency on Es-En, Fr-En, and De-En CVSS-C test sets, with a parallel ~4 BLEU gap on text output
      confirming the bottleneck is LLM prediction rather than speech synthesis.
    confidence: high
    relevance: medium
  - claim_id: offline_training_combined_with_test_time_simultaneous_inference_policies_can
    role: supports
    claim: Offline training combined with test-time simultaneous inference policies can match or outperform systems
      trained specifically for streaming in speech-to-speech translation.
    source: §5.1, Fig. 4, Table 1
    evidence: SimulS2S-LLM, trained offline, consistently outperforms the streaming-trained StreamSpeech model across
      all three language pairs at comparable latency, achieving up to 4 ASR-BLEU improvement on Es-En.
    confidence: high
    relevance: low
  - claim_id: aggregating_llm_hidden_states_across_multiple_layers_improves_discrete_speech
    role: supports
    claim: Aggregating LLM hidden states across multiple layers improves discrete speech token prediction compared
      to using only the final layer.
    source: §5.3, Fig. 6
    evidence: Multi-layer hidden state weighting yields approximately 1 ASR-BLEU improvement over last-layer-only
      decoding on CVSS-C Es-En Simul-S2ST, attributed to the final layer's focus on semantic text information at
      the expense of acoustic richness needed for speech token generation.
    confidence: high
    relevance: medium
  - claim_id: llm_based_approaches_to_simultaneous_speech_generation_face_a_latency
    role: complicates
    claim: LLM-based approaches to simultaneous speech generation face a latency penalty from LLM inference overhead
      that narrows the practical quality-latency advantage over non-LLM methods.
    source: §D, Tables 8-10, Limitations
    evidence: Computation-aware ATD for SimulS2S-LLM is substantially higher than standard ATD (e.g., 4239ms vs.
      3440ms at k=8 on Es-En), and the system is not evaluated at very low latency regimes (AL < 1s) where reordering
      requirements make offline-trained models unsuitable.
    confidence: high
    relevance: low
  - claim_id: shallow_fusion_of_n_gram_language_models_with_ctc_decoding
    role: supports
    claim: Shallow fusion of n-gram language models with CTC decoding of discrete speech tokens improves simultaneous
      speech translation quality without increasing latency.
    source: §5.4, Table 2
    evidence: n-gram LM fusion over greedy CTC search improves ASR-BLEU from 24.7 to 26.3 on CVSS-C Es-En at identical
      ATD of 3439ms.
    confidence: high
    relevance: low
  limitations:
  - The system is not evaluated at very low latency (AL < 1s), a regime the authors identify as unsuitable for offline-trained
    models due to reordering requirements. This excludes SimulS2S-LLM from the most latency-critical applications.
    Computation-aware latency is substantially higher than the reported ATD, and all experiments use 7B/8B open-source
    LLMs on small datasets (70-174 hours per language pair), leaving scalability to larger models and data unverified.
  - The evaluation is limited to three European language pairs in a single translation direction each. Language
    pairs with greater structural divergence or more extensive reordering would stress the wait-k assumption more
    severely. The system has not been evaluated on offline inference or tasks other than S2ST and S2TT, despite
    the claim that offline training preserves such capabilities. Long-form simultaneous speech translation is also
    untested due to lack of suitable data.
  caveats: []
- id: 2025.acl-long.87
  published_date: "2025-07-27"
  entry_date: '2026-07-29'
  year: 2025
  venue: ACL
  task:
  - VC
  architecture:
  - flow-matching
  relevance: medium
  evidence_role:
  - architecture_variant
  - acceleration_evidence
  - control_evidence
  current_role: active_evidence
  method_family:
  - flow_matching_with_ssl_conditioning
  claims:
  - claim_id: combining_asr_derived_phonetic_features_and_quantized_self_supervised_representations
    role: supports
    claim: Combining ASR-derived phonetic features and quantized self-supervised representations via adaptive fusion
      reduces timbre leakage while preserving paralinguistic content in zero-shot voice conversion.
    source: §5.3, Table 3
    evidence: Removing the PPG branch (WavLM-only) causes SMOS to drop from 4.11 to 3.07 and SECS from 0.71 to 0.45
      on LibriTTS, indicating that SSL features alone carry substantial timbre leakage; removing the SSL branch
      degrades NMOS and WER, confirming PPGs alone lose paralinguistic richness.
    confidence: high
    relevance: high
  - claim_id: flow_matching_provides_faster_inference_than_diffusion_based_voice_conversion
    role: supports
    claim: Flow matching provides faster inference than diffusion-based voice conversion systems without sacrificing
      speaker similarity or naturalness.
    source: §5.1, Table 1
    evidence: Takin-VC achieves RTF 0.154, lower than DiffVC (0.294), NS2VC (0.347), and SeedVC (0.341), while simultaneously
      outperforming these baselines on NMOS, SMOS, and SECS.
    confidence: high
    relevance: low
  - claim_id: global_time_invariant_speaker_embeddings_are_insufficient_for_robust_timbre
    role: complicates
    claim: Global, time-invariant speaker embeddings are insufficient for robust timbre modeling in expressive zero-shot
      voice conversion.
    source: §5.3, Table 4
    evidence: Removing the context-aware cross-attention module (which aligns source content with target timbre
      dynamically) drops SMOS from 4.11 to 3.61 and SECS from 0.71 to 0.58, while the memory-augmented module removal
      drops SECS to 0.52. Both modules provide content-sensitive timbre conditioning beyond a static speaker embedding
      alone.
    confidence: high
    relevance: low
  - claim_id: cross_gender_voice_conversion_consistently_yields_lower_speaker_similarity_than
    role: complicates
    claim: Cross-gender voice conversion consistently yields lower speaker similarity than same-gender conversion
      even in well-trained systems.
    source: §5.2, Table 2
    evidence: 'On the large-scale multilingual dataset, same-gender pairs (F2F: SECS 0.74; M2M: 0.73) outperform
      cross-gender pairs (F2M: 0.71; M2F: 0.70) in speaker embedding cosine similarity, a gap that persists across
      all conversion directions.'
    confidence: high
    relevance: low
  - claim_id: quantizing_self_supervised_speech_features_before_content_encoding_reduces_timbre
    role: refines
    claim: Quantizing self-supervised speech features before content encoding reduces timbre leakage more effectively
      than using continuous SSL representations directly.
    source: §3.2, §5.3, Table 3
    evidence: The RVQ quantizer (codebook size 8,200) applied to WavLM features is the key mechanism for timbre
      suppression in the hybrid encoder; ablation with WavLM-only (continuous features without adaptive fusion)
      shows SECS drops to 0.45 compared to 0.71 for the full model, consistent with timbre leakage from unquantized
      SSL features.
    confidence: high
    relevance: high
  limitations:
  - All large-scale training data (500k hours) and the 100-speaker evaluation set are proprietary and not publicly
    available. The large-scale results cannot be reproduced by external researchers, and it is unclear how much
    of the gain over competitive baselines is attributable to data scale rather than the proposed modules.
  - The paper does not include targeted evaluation of paralinguistic preservation (breathing, crying, emotion transfer),
    despite listing this as a primary contribution. NMOS and SMOS measure general naturalness and speaker similarity
    but are not designed to capture expressive fidelity specifically. Speech editing under zero-shot conditions
    is acknowledged as out of scope and a direction for future work. Ethical risks from voice impersonation are
    noted but no technical mitigations are proposed.
  caveats: []
- id: 2025.acl-long.997
  published_date: "2025-07-27"
  entry_date: '2026-07-29'
  year: 2025
  venue: ACL
  task:
  - SCA
  architecture:
  - autoregressive-LM
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: active_evidence
  method_family:
  - autoregressive_ssl_conditioned_models
  claims:
  - claim_id: preference_optimization_with_llm_based_semantic_feedback_improves_long_range
    role: supports
    claim: Preference optimization with LLM-based semantic feedback improves long-range semantic coherence in textless
      spoken language models beyond next-token prediction training.
    source: §5.2, Table 2
    evidence: DPO with Mistral-score preference data raises T-StoryCloze from 69.7% to 74.2% (1.3B model) and from
      75.4% to 83.8% (7B model) over the pre-trained TWIST baseline; curriculum learning further improves both to
      76.1% and 85.6% respectively.
    confidence: high
    relevance: medium
  - claim_id: llm_based_semantic_scoring_provides_more_effective_alignment_signal_for
    role: supports
    claim: LLM-based semantic scoring provides more effective alignment signal for speech continuation than perplexity-based
      preference selection.
    source: §5.1, Table 1
    evidence: On the 1.3B model, Align-SLM with Mistral score improves T-StoryCloze by +4.5 and GPT4-o score by
      +0.24 over the pre-trained baseline, whereas the PPL-based variant degrades T-StoryCloze by 2.0 and improves
      GPT4-o by only +0.03.
    confidence: high
    relevance: medium
  - claim_id: curriculum_learning_with_iteratively_tightened_preference_thresholds_yields_progressive_semantic
    role: supports
    claim: Curriculum learning with iteratively tightened preference thresholds yields progressive semantic improvements
      in preference-optimized spoken language models.
    source: §5.3, Appendix E, Table 2
    evidence: Two curriculum learning iterations on the 7B model with LibriSpeech data improve T-StoryCloze from
      83.8% (Align-SLM) to 85.6% (Align-SLM+CL) and GPT4-o from 3.50 to 3.56; further iterations continue to improve
      most metrics.
    confidence: high
    relevance: medium
  - claim_id: semantic_gains_from_preference_optimization_do_not_transfer_uniformly_to
    role: complicates
    claim: Semantic gains from preference optimization do not transfer uniformly to lexical and grammatical aspects
      of spoken language modeling.
    source: §5.5, Table 2
    evidence: Align-SLM shows only marginal sBLIMP improvement (+1.3 for the 1.3B Mistral variant) while AudioLM
      and SyllableLM, which use no preference training but have better speech token designs, outperform on grammatical
      correctness.
    confidence: high
    relevance: medium
  - claim_id: textless_end_to_end_spoken_language_models_remain_substantially_behind
    role: complicates
    claim: Textless end-to-end spoken language models remain substantially behind cascaded ASR+LLM pipelines in
      semantic understanding even with preference optimization.
    source: §5.5, Table 2
    evidence: The best Align-SLM variant achieves 77.9% sWUGGY and 86.8% T-StoryCloze, compared to 79.2% and 94.8%
      for the ASR+LLM cascade topline, a gap that preference optimization alone does not close.
    confidence: high
    relevance: medium
  limitations:
  - Training and evaluation are restricted to English audiobook speech (LibriSpeech, MLS), a clean and structured
    domain. Generalisation to spontaneous, noisy, or multi-speaker speech is untested, and the automatic preference
    pipeline's reliance on ASR transcription quality may degrade under real-world acoustic conditions.
  - The paper does not compare against text-injecting SLMs (SPIRITLM, Moshi, VoxtLM) under matched data and compute
    budgets, so the relative cost of maintaining the textless constraint is not quantified. Only the semantic dimension
    of SLM quality is addressed; paralinguistic aspects such as prosody, speaking style, and emotion are explicitly
    out of scope and represent open directions for extending the framework.
  caveats: []
- id: 2025.findings-acl.631
  published_date: "2025-07-27"
  entry_date: '2026-07-29'
  year: 2025
  venue: ACL
  task:
  - SCA
  architecture:
  - autoregressive-LM
  relevance: medium
  evidence_role:
  - architecture_variant
  - acceleration_evidence
  current_role: active_evidence
  method_family:
  - autoregressive_ssl_conditioned_models
  claims:
  - claim_id: text_lm_initialization_substantially_accelerates_slm_convergence_and_improves_final
    role: supports
    claim: Text LM initialization substantially accelerates SLM convergence and improves final model quality within
      fixed compute budgets.
    source: §4.1, Figure 2, Figure 3
    evidence: TWIST-style initialization from Qwen2.5-0.5B outperforms uninitialised variants across all evaluated
      model families; by end of training, Qwen2.5 with initialization surpasses alternatives even with fewer training
      tokens.
    confidence: high
    relevance: medium
  - claim_id: synthetic_tts_data_improves_semantic_modeling_in_slms_when_combined
    role: supports
    claim: Synthetic TTS data improves semantic modeling in SLMs when combined with real speech, but replacing real
      data with synthetic data alone degrades performance.
    source: §4.2, Table 1
    evidence: Adding sTinyStories to the training mix boosts TSC from 71.14 to 78.01 and GenPPL from 145.4 to 88.3
      for Qwen-0.5B; training on synthetic data exclusively yields sBLIMP 52.35 versus 56.45 for the mixed baseline.
    confidence: high
    relevance: medium
  - claim_id: preference_optimization_with_synthetically_generated_preference_data_substantially_improves_slm
    role: supports
    claim: Preference optimization with synthetically generated preference data substantially improves SLM generation
      quality even under tight compute constraints.
    source: §4.4, Figure 5, Table 2
    evidence: 30 minutes of DPO training on SpokenSwag improves TSC from 78.01 to 82.04 and GenPPL from 88.3 to
      62.8; the gain saturates quickly and additional DPO budget can degrade results.
    confidence: high
    relevance: medium
  - claim_id: slm_scaling_laws_accurately_predict_compute_optimal_performance_boundaries_for
    role: contradicts
    claim: SLM scaling laws accurately predict compute-optimal performance boundaries for academic-scale training.
    source: §5, Table 2
    evidence: Slam on one A5000 GPU-day achieves TSC 82.04 against the scaling-law predicted optimal of TSC 70.49
      for the same compute, and sBLIMP 58.86 against the predicted 56.85; the authors attribute the gap to text
      LM initialization and synthetic data not being factored into scaling law derivations.
    confidence: high
    relevance: medium
  - claim_id: diverse_multi_domain_speech_training_data_improves_slm_generalization
    role: complicates
    claim: Diverse multi-domain speech training data improves SLM generalization.
    source: §4.2, Table 1
    evidence: Augmenting LibriLight and LibriSpeech with VoxPopuli, TED-LIUM, PeopleSpeech, and SWC consistently
      degrades TSC and SSC for both OPT-125M and Qwen-0.5B under the single-GPU budget; the authors hypothesize
      that modelling acoustic diversity requires more compute than the Slamming budget allows.
    confidence: high
    relevance: medium
  limitations:
  - The evaluation uses only semantic modeling benchmarks (TSC, SSC, sBLIMP) and generation perplexity. Acoustic
    and prosodic quality are not evaluated; the authors note that benchmarks such as SALMon may reveal further limitations
    of low-resource SLM training in these dimensions. *(§7)*
  - The recipe relies exclusively on HuBERT semantic tokens, leaving open whether the same efficiency gains apply
    with newer tokenizers such as Mimi or SylBoost. Text-speech interleaving underperforms speech-only training
    at the single-GPU budget due to vocabulary expansion and fewer speech tokens per step; whether this changes
    at higher compute is unresolved. Out-of-domain generalization (People Speech) shows that Slam scaled matches
    but does not clearly outperform TWIST-7B, despite Slam using only audiobook-domain training data that TWIST
    also trained on.
  caveats: []
- id: 2025.findings-acl.71
  published_date: "2025-07-27"
  entry_date: '2026-07-29'
  year: 2025
  venue: ACL
  task:
  - SCA
  architecture:
  - autoregressive-LM
  relevance: medium
  evidence_role:
  - architecture_variant
  current_role: active_evidence
  method_family:
  - autoregressive_ssl_conditioned_models
  claims:
  - claim_id: explicit_listening_comprehension_objectives_improve_cross_modal_understanding_in_multimodal
    role: supports
    claim: Explicit listening comprehension objectives improve cross-modal understanding in multimodal LLMs fine-tuned
      for spoken QA.
    source: §3.4, Table 3
    evidence: Removing the story listening comprehension sub-task from DAMSEL causes the largest single-task performance
      degradation (Table 3, ablations on ASK-QA with Speech-Qwen), confirming that transcription supervision is
      the most critical auxiliary signal.
    confidence: high
    relevance: medium
  - claim_id: multi_task_auxiliary_task_design_can_substitute_for_large_scale
    role: supports
    claim: Multi-task auxiliary task design can substitute for large-scale data collection in adapting MLLMs to
      spoken question answering.
    source: §3.2, Figure 4, Table A7
    evidence: DAMSEL with Speech-Qwen surpasses the prior state-of-the-art on Spoken-SQuAD using only 10% of training
      data (66.38 EM vs. DDNet 64.1 EM trained on 100% of data).
    confidence: high
    relevance: medium
  - claim_id: data_centric_multi_task_fine_tuning_yields_consistent_gains_even
    role: supports
    claim: Data-centric multi-task fine-tuning yields consistent gains even for frontier MLLMs with extensive prior
      multi-modal pre-training.
    source: §3.1
    evidence: Gemini Pro fine-tuned with DAMSEL on 1% of ASK-QA improves 5.7% relative over single-task tuning;
      improvement persists with full data (1.6% relative), despite Gemini having seen large-scale multi-modal training
      data and ASK-QA being newly synthesised.
    confidence: high
    relevance: medium
  - claim_id: tts_synthesised_speech_datasets_have_quality_ceilings_that_limit_model
    role: complicates
    claim: TTS-synthesised speech datasets have quality ceilings that limit model generalisation to naturalistic
      paralinguistic variation.
    source: §Limitations, §B.1
    evidence: The paper notes that WER-filtered ASK-QA transcriptions may not perfectly match synthesised speech,
      and that current TTS lacks perfect controllability for speaking rate and pitch. The dynamic evaluation pipeline
      also depends on TTS to regenerate spoken turns, compounding generation quality issues across multi-turn rollouts.
    confidence: high
    relevance: medium
  - claim_id: spoken_qa_performance_gains_from_multi_task_learning_are_largest
    role: refines
    claim: Spoken QA performance gains from multi-task learning are largest in the low-data regime and diminish
      as training data scales.
    source: §3.1, §3.3, Table 2
    evidence: Gemini Pro's DAMSEL advantage over single-task tuning narrows from 16.13% relative EM improvement
      (1% data) to 1.7% (full data) on SD-QA, and from 5.7% to 1.6% relative on ASK-QA.
    confidence: high
    relevance: medium
  limitations:
  - Performance relies on ground-truth transcriptions as listening comprehension targets. In practice, transcriptions
    from Spoken-SQuAD and SD-QA are derived from ASR systems with known errors. The paper acknowledges it is unclear
    whether slight transcription noise improves robustness or degrades performance.
  - A key open question is whether the DAMSEL data generation process scales to MLLM post-training. The paper demonstrates
    gains with fine-tuning but notes that verifying post-training effects is computationally infeasible in this
    work. The dynamic multi-turn evaluation for ASK-QA also depends on an LLM-based action classifier and user simulator,
    introducing additional sources of evaluation variance that are not thoroughly characterised.
  - The model currently targets auditory semantic understanding. The paper notes that more nuanced SCA tasks (such
    as monitoring user frustration in task guidance) require sensitivity to different paralinguistic dimensions
    not addressed by the current auxiliary tasks.
  caveats: []
- id: 2025.findings-acl.75
  published_date: "2025-07-27"
  entry_date: '2026-07-29'
  year: 2025
  venue: ACL
  task:
  - TTS
  architecture:
  - transformer-enc-dec
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: active_evidence
  method_family:
  - transformer_ssl_encoders_and_adapters
  claims:
  - claim_id: discrete_speech_unit_sequences_can_be_compressed_into_text_like
    role: supports
    claim: Discrete speech unit sequences can be compressed into text-like representations via n-gram language modelling,
      providing effective text-free guidance for both cross-modal and cross-lingual learning in speech translation.
    source: §2.1, §3.2, Table 2
    evidence: Unit language (2-gram merged mHubert units, K=3) achieves average ASR-BLEU of 21.5 on VoxPopuli across
      four language directions, matching recognised-text training and yielding +1.2 BLEU over the baseline that
      uses raw discrete units alone.
    confidence: high
    relevance: high
  - claim_id: combining_cross_modal_and_cross_lingual_auxiliary_objectives_in_multi
    role: complicates
    claim: Combining cross-modal and cross-lingual auxiliary objectives in multi-task sequence-to-sequence speech
      training can produce gradient interference that limits the gains of each individual loss.
    source: §2.3, §3.2, §4.3
    evidence: Simultaneously applying source (L_CM) and target (L_CL) unit language losses without task prompts
      yields combined gains no greater than either loss alone, due to CM guidance at middle encoder layers disrupting
      CL representation learning in upper layers, as confirmed by attention localness analysis.
    confidence: high
    relevance: medium
  - claim_id: learnable_task_specific_prompt_vectors_can_separate_competing_auxiliary_objectives
    role: supports
    claim: Learnable task-specific prompt vectors can separate competing auxiliary objectives within a shared Transformer
      encoder by enforcing distinct intermediate representations for each task.
    source: §2.4, §3.2, Table 2
    evidence: Two prompt vectors (b_CM and b_CL) prepended at different encoder layer boundaries, with a negative
      MSE diversification loss, fully recover the +1.2 BLEU combined gain from unit language guidance that is otherwise
      lost to task interference.
    confidence: high
    relevance: medium
  - claim_id: sequence_length_compression_is_a_critical_enabler_of_cross_lingual
    role: refines
    claim: Sequence length compression is a critical enabler of cross-lingual alignment in textless speech translation,
      and n-gram-merged unit sequences provide substantially better compression than BPE pseudo-languages.
    source: §4.1, §4.6, Figure 4, Table 6
    evidence: Unit language lengths fall consistently between raw units and character-level text across all four
      VoxPopuli language pairs; unit language CL training outperforms BPE pseudo-language by +0.5 BLEU average,
      attributed to n-gram contextual information absent in BPE merges.
    confidence: high
    relevance: medium
  limitations:
  - 'The evaluation uses ASR-BLEU only: output speech is recognised by ASR and scored with BLEU against text references.
    Naturalness, fluency, and speaker consistency of the synthesised speech are not assessed. The paper''s own Limitations
    section acknowledges this gap.'
  - The method is validated on four European language directions (Spanish, French, English) from VoxPopuli. Generalisation
    to typologically distant or low-resource languages, where phonological unit distributions may differ substantially,
    is unexamined. Unit language construction with K=3 requires approximately 12 hours of preprocessing per 160M-unit
    corpus, which may limit applicability at scale. The task prompt mitigates but does not fully eliminate the CM/CL
    conflict; the remaining interaction between the two objectives is not resolved.
  caveats: []
- id: '2503.11026'
  published_date: "2025-07-30"
  entry_date: '2026-07-29'
  year: 2025
  venue: arXiv
  task:
  - TTS
  architecture:
  - flow-matching
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: active_evidence
  method_family:
  - flow_matching_with_ssl_conditioning
  claims:
  - claim_id: conditioning_a_flow_matching_mel_spectrogram_generator_on_rich_multimodal
    role: supports
    claim: Conditioning a flow matching mel-spectrogram generator on rich multimodal speaker representations produces
      more consistent speaker identity in zero-shot cross-lingual speech synthesis than injecting a single speaker
      embedding at the vocoder stage.
    source: §5.4, Table 1
    evidence: MAVFlow achieves an average 36% improvement in speaker similarity (SS) over AV2AV across four language
      pairs on MuAViC, using OT-CFM conditioned on x-vector speaker embeddings plus facial emotion embeddings, while
      AV2AV uses d-vector conditioning in the vocoder only.
    confidence: high
    relevance: medium
  - claim_id: higher_quality_intermediate_mel_spectrogram_synthesis_propagates_benefits_to_downstream
    role: supports
    claim: Higher-quality intermediate mel-spectrogram synthesis propagates benefits to downstream talking-face
      generation even when the face decoder itself is unchanged.
    source: §5.5, Table 5
    evidence: MAVFlow improves LSE-C by +0.87, LSE-D by -0.49, and FID by -0.61 relative to AV2AV on LRS3 visual
      evaluation, despite using the same Wav2Lip face decoder, suggesting that the mel quality bottleneck affects
      face sync accuracy.
    confidence: high
    relevance: medium
  - claim_id: visual_emotion_conditioning_is_insufficient_on_its_own_to_improve
    role: complicates
    claim: Visual emotion conditioning is insufficient on its own to improve emotion expression in synthesized speech
      and requires concurrent audio speaker conditioning to be effective.
    source: §5.6, Table 8
    evidence: Adding only visual guidance to the CFM model marginally maintains speaker similarity (SS 0.056 vs
      0.057 without guidance) but reduces emotion accuracy from 28.66% to 26.83% on CREMA-D; the combination of
      audio and visual guidance is needed to reach 36.46%.
    confidence: high
    relevance: medium
  - claim_id: paralinguistic_and_linguistic_generation_objectives_are_compatible_in_zero_shot
    role: supports
    claim: Paralinguistic and linguistic generation objectives are compatible in zero-shot cross-lingual speech
      synthesis; improving speaker fidelity does not require sacrificing translation accuracy.
    source: §5.4, Tables 1 and 3
    evidence: MAVFlow maintains competitive ASR-BLEU scores (26.97 Es-En vs 26.57 for AV2AV and 28.66–30.55 for
      cascaded systems) while substantially improving speaker similarity, using the same unit translation module
      as AV2AV.
    confidence: high
    relevance: medium
  - claim_id: emotion_recognition_accuracy_in_synthesized_cross_lingual_speech_remains_far
    role: complicates
    claim: Emotion recognition accuracy in synthesized cross-lingual speech remains far below ground-truth levels
      even with multimodal conditioning, indicating that paralinguistic preservation is an unsolved challenge.
    source: §5.4, Table 2; §5.6, Table 7
    evidence: MAVFlow achieves 36.46% audio emotion accuracy vs a GT upper bound of 81.95% on CREMA-D, even with
      dual audio-visual guidance; additional training on an emotion-rich dataset (CREMA-D) raises this to 51.46%
      but still far below GT.
    confidence: high
    relevance: medium
  limitations:
  - The system relies on emotional cues from facial video alone; audio-side paralinguistics such as prosody and
    timbre variation are not used as emotion conditioning signals, which may limit emotion transfer when visual
    input is unavailable or low quality. The unit extractor and unit-to-unit translation modules are adopted unchanged
    from AV2AV, so improvements in semantic translation quality require addressing those upstream components separately.
    The Duration Length Regulator interpolates to match source duration, which may introduce length-related artifacts
    when source and translated speech have substantially different natural lengths. Evaluation is limited to five
    European languages with English as the target; generalization to typologically distant language pairs (e.g.,
    tonal languages, right-to-left scripts) is untested.
  caveats: []
- id: 2025.iwslt-1.5
  published_date: "2025-07-31"
  entry_date: '2026-07-29'
  year: 2025
  venue: IWSLT
  task:
  - SCA
  architecture:
  - autoregressive-LM
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: active_evidence
  method_family:
  - autoregressive_ssl_conditioned_models
  claims:
  - claim_id: grounding_speech_compression_in_explicit_speech_text_alignment_boundaries_improves
    role: supports
    claim: Grounding speech compression in explicit speech-text alignment boundaries improves spoken language understanding
      and cross-modal reasoning in speech-augmented language models.
    source: §4.3, Tables 2–3
    evidence: SSR-CONNECTOR (UNITY2) achieves 64.2% / 68.6% (0/5-shot) on Speech-MMLU vs. 40.5% / 42.75% for SPIRITLM
      (LLAMA3), and 74.8% on StoryCloze (S→T) vs. 61.6%, while maintaining 65.3% MMLU text accuracy.
    confidence: high
    relevance: medium
  - claim_id: fine_tuning_a_pre_trained_llm_on_speech_data_degrades
    role: complicates
    claim: Fine-tuning a pre-trained LLM on speech data degrades cross-modal understanding even when catastrophic
      forgetting is mitigated through multitask training.
    source: §5.2, Table 5
    evidence: All Stage 2 fine-tuning methods improve speech-only task performance (sWUGGY, sBLIMP) but consistently
      reduce Speech-MMLU and StoryCloze (S→T) accuracy; multitask fine-tuning limits the drop but cannot eliminate
      it.
    confidence: high
    relevance: medium
  - claim_id: distillation_from_text_embeddings_is_an_effective_pre_training_objective
    role: supports
    claim: Distillation from text embeddings is an effective pre-training objective for speech-text modality connectors,
      enabling cross-lingual transfer to ASR without explicit transcription supervision.
    source: §4.3, Table 3
    evidence: After Stage 1 distillation alone (no ASR objective), SSR-CONNECTOR achieves 5.6% / 4.0% WER on LibriSpeech
      clean (0/5-shot) via in-context prompting, using only the alignment-supervised distillation loss.
    confidence: high
    relevance: medium
  - claim_id: alignment_aware_speech_segmentation_reduces_effectiveness_on_lexical_tasks_that
    role: complicates
    claim: Alignment-aware speech segmentation reduces effectiveness on lexical tasks that depend on synthesized
      anomalous phoneme sequences.
    source: §4.3
    evidence: SSR-CONNECTOR underperforms on sWUGGY because the aligner was not trained on incorrectly spoken words
      (non-words), causing segmentation errors; the paper notes this as a known scope limitation of alignment-based
      methods.
    confidence: high
    relevance: medium
  limitations:
  - 'All experiments use a single language (English), a single LLM backbone (LLAMA3), and a single self-supervised
    speech encoder (DinoSR). The paper does not ablate the speech encoder, which is a notable gap since encoder
    quality directly determines the input representations. The Speech-MMLU benchmark excludes domains with synthesis
    quality issues (e.g., algebra), limiting its breadth as a general evaluation tool. The trade-off between enhanced
    speech understanding (Stage 2 fine-tuning) and degraded cross-modal performance remains partially unresolved:
    even with multitask fine-tuning, cross-modal tasks suffer. Extending SSR-CONNECTOR to languages with looser
    speech-text alignment (e.g., tonal languages, code-switched speech) is left to future work.'
  caveats: []
- id: 2025.sigdial-1.21
  published_date: "2025-08-01"
  entry_date: '2026-07-29'
  year: 2025
  venue: workshop
  task:
  - SCA
  architecture:
  - hybrid
  relevance: medium
  evidence_role:
  - architecture_variant
  - acceleration_evidence
  current_role: active_evidence
  method_family:
  - hybrid_semantic_acoustic_ssl_systems
  claims:
  - claim_id: purely_acoustic_real_time_turn_taking_inference_is_feasible_without
    role: supports
    claim: Purely acoustic, real-time turn-taking inference is feasible without relying on ASR or linguistic features.
    source: §5.3, Figure 6
    evidence: TRPDformer achieves AUC = 0.83 on CEJC using only CPC audio representations and a causal transformer,
      outperforming a VAD baseline with no transcription step.
    confidence: high
    relevance: low
  - claim_id: voice_activity_projection_models_are_not_directly_transferable_to_transition
    role: complicates
    claim: Voice activity projection models are not directly transferable to transition relevance point detection,
      even though the tasks appear structurally similar.
    source: §4.2
    evidence: Statistical analysis of CEJC shows that a theoretically perfect VAP model would achieve at most recall
      = 0.60 and precision = 0.71 on TRP detection, below TRPDformer's empirical performance, because speaker shift
      and TRP are not equivalent events.
    confidence: high
    relevance: medium
  - claim_id: learned_acoustic_response_timing_models_produce_perceptibly_more_natural_dialogue
    role: supports
    claim: Learned acoustic response-timing models produce perceptibly more natural dialogue turn-taking than fixed-threshold
      silence detection.
    source: §5.5, Figure 8
    evidence: In a preference test with 40 raters per stimulus pair, subjects preferred TRPDformer's response timing
      over the VAD baseline, including in cases where both models misdetected TRPs.
    confidence: high
    relevance: low
  - claim_id: frame_level_trp_models_trained_on_spontaneous_conversation_may_not
    role: complicates
    claim: Frame-level TRP models trained on spontaneous conversation may not generalise across languages or domains,
      limiting their direct applicability in multilingual or task-oriented dialogue systems.
    source: §6
    evidence: TRPDformer is trained and evaluated solely on CEJC (Japanese everyday conversation); the paper identifies
      cross-lingual and cross-domain robustness as open future directions.
    confidence: high
    relevance: low
  limitations:
  - Evaluation is limited to a single language (Japanese) and a single corpus (CEJC). Generalisation to other languages,
    accents, or task-oriented dialogue domains is untested, and the paper does not report real-time processing latency
    or robustness to noise — both of which are critical for deployment.
  - The paper also does not evaluate the downstream effect of TRPDformer on a full spoken dialogue system end-to-end,
    only on the timing component in isolation via a preference test. Response timing is studied independently of
    response content quality, and the paper acknowledges that future work should integrate TRPDformer into a complete
    dialogue system to assess its impact on overall conversational quality.
  caveats: []
- id: 2025.sigdial-1.51
  published_date: "2025-08-01"
  entry_date: '2026-07-29'
  year: 2025
  venue: workshop
  task:
  - SCA
  architecture: []
  relevance: medium
  evidence_role:
  - architecture_variant
  - acceleration_evidence
  current_role: minor
  method_family: []
  claims:
  - claim_id: modular_incremental_architectures_built_on_the_iu_framework_enable_flexible
    role: supports
    claim: Modular, incremental architectures built on the IU framework enable flexible composition of heterogeneous
      speech and perception components for real-time spoken dialogue.
    source: §2, §3
    evidence: rrSDS 2.0 supports seamless swapping of ASR modules (Whisper vs. Wav2vec) and vision modules (YOLOv11,
      YOLOv12, RT-DETR, SAM, DINOv2) through a shared IU interface, demonstrated in three multi-component robot
      pipelines.
    confidence: high
    relevance: low
  - claim_id: real_time_multimodal_spoken_dialogue_for_physical_and_simulated_robotic
    role: complicates
    claim: Real-time multimodal spoken dialogue for physical and simulated robotic platforms introduces synchronization
      challenges not present in text-only or speech-only systems.
    source: §4
    evidence: The authors report facing challenges balancing real-time, multimodal interaction when integrating
      rrSDS 2.0 across physical and simulated robots, without providing quantitative results characterizing the
      trade-offs.
    confidence: high
    relevance: low
  - claim_id: open_source_framework_level_integration_of_state_of_the_art
    role: supports
    claim: Open-source, framework-level integration of state-of-the-art ASR and vision components reduces the engineering
      overhead for building multimodal robotic dialogue systems.
    source: §1, §2
    evidence: rrSDS 2.0 wraps Whisper, Wav2vec 2.0, YOLO variants, SAM, DINOv2, MediaPipe, RASA 3.0, and HuggingFace
      text generation behind a common IU interface, with pypi distribution and improved documentation to lower the
      setup barrier.
    confidence: high
    relevance: low
  limitations:
  - No quantitative evaluation is reported, making it impossible to compare rrSDS 2.0 against other dialogue system
    frameworks on latency, accuracy, or robustness. The demo paper format means claims about incremental processing
    quality and real-time performance rest on qualitative demonstration rather than systematic measurement. Integration
    with additional benchmarks beyond ALFRED is noted as future work, as is compatibility with the Remdis framework
    for LLM-driven incremental dialogue.
  caveats: []
- id: '2508.14049'
  published_date: "2025-08-05"
  entry_date: '2026-07-29'
  year: 2025
  venue: arXiv
  task:
  - TTS
  architecture:
  - autoregressive-LM
  - flow-matching
  relevance: medium
  evidence_role:
  - architecture_variant
  current_role: active_evidence
  method_family:
  - autoregressive_ssl_conditioned_models
  - flow_matching_with_ssl_conditioning
  claims:
  - claim_id: large_multilingual_tts_systems_built_on_semantic_token_intermediaries_transfer
    role: supports
    claim: Large multilingual TTS systems built on semantic token intermediaries transfer to low-resource languages
      more readily than end-to-end spectrogram models.
    source: §2.1, §5.1
    evidence: Large multilingual TTS systems built on semantic token intermediaries transfer to low-resource languages
      more readily than end-to-end spectrogram models.
    confidence: high
    relevance: high
  - claim_id: decoupling_the_text_to_semantic_and_semantic_to_acoustic_stages
    role: supports
    claim: Decoupling the text-to-semantic and semantic-to-acoustic stages enables independent training and simplifies
      the addition of new languages without full system retraining.
    source: §2, §4.1
    evidence: Decoupling the text-to-semantic and semantic-to-acoustic stages enables independent training and simplifies
      the addition of new languages without full system retraining.
    confidence: high
    relevance: medium
  - claim_id: flow_matching_is_a_viable_replacement_for_diffusion_in_the
    role: supports
    claim: Flow matching is a viable replacement for diffusion in the acoustic generation stage of two-stage TTS
      pipelines, maintaining competitive quality at lower training complexity.
    source: §2.3, §5.1
    evidence: Flow matching is a viable replacement for diffusion in the acoustic generation stage of two-stage
      TTS pipelines, maintaining competitive quality at lower training complexity.
    confidence: high
    relevance: medium
  - claim_id: intelligibility_in_low_resource_languages_with_limited_training_data_remains
    role: complicates
    claim: Intelligibility in low-resource languages with limited training data remains markedly worse than high-resource
      languages within the same multilingual system.
    source: §5.1, Table 2
    evidence: Intelligibility in low-resource languages with limited training data remains markedly worse than high-resource
      languages within the same multilingual system.
    confidence: high
    relevance: medium
  limitations:
  - The evaluation relies exclusively on WER over 10 sentences per language with no MOS, SMOS, or naturalness scores.
    This makes it impossible to assess audio quality, expressiveness, or speaker similarity relative to baselines
    — core dimensions for a TTS system.
  - 'English dominates the training set at 58%, which may explain strong English results but raises questions about
    whether true cross-lingual transfer or data dominance is responsible. The system lacks prosody and pace control
    conditioning in M1, which the authors flag as future work. Zero-shot speaker fidelity for M2 is acknowledged
    as limited compared to infilling-based approaches like Seamless. Fine-tuning introduces hallucination that requires
    careful intervention (freezing classification heads only), suggesting the LM component is sensitive to distribution
    shift. Languages with fewer training hours (Assamese: 48h, Dogri: 8h, Rajasthani: 20h) show substantially weaker
    results.'
  caveats: []
- id: '2508.04141'
  published_date: "2025-08-06"
  entry_date: '2026-07-29'
  year: 2025
  venue: arXiv
  task:
  - TTS
  architecture:
  - autoregressive-LM
  - hybrid
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: active_evidence
  method_family:
  - autoregressive_ssl_conditioned_models
  - hybrid_semantic_acoustic_ssl_systems
  claims:
  - claim_id: generating_semantic_and_acoustic_tokens_simultaneously_in_a_single_autoregressive
    role: supports
    claim: Generating semantic and acoustic tokens simultaneously in a single autoregressive forward pass, rather
      than cascading semantic prediction before acoustic prediction, reduces word error rate and improves naturalness
      in zero-shot TTS.
    source: §V.A, Table I
    evidence: Generating semantic and acoustic tokens simultaneously in a single autoregressive forward pass, rather
      than cascading semantic prediction before acoustic prediction, reduces word error rate and improves naturalness
      in zero-shot TTS.
    confidence: high
    relevance: medium
  - claim_id: combining_specialist_ssl_models_for_distinct_speech_attributes_semantic_content
    role: supports
    claim: Combining specialist SSL models for distinct speech attributes (semantic content, acoustic texture, speaker
      identity) as frozen feature extractors enables more effective token-level disentanglement than using a single
      encoder for all attributes.
    source: §III.A, Tables III–IV
    evidence: Combining specialist SSL models for distinct speech attributes (semantic content, acoustic texture,
      speaker identity) as frozen feature extractors enables more effective token-level disentanglement than using
      a single encoder for all attributes.
    confidence: high
    relevance: high
  - claim_id: a_hybrid_ar_nar_design_that_enforces_independence_at_the
    role: supports
    claim: A hybrid AR+NAR design that enforces independence at the coarse token level and interdependence at the
      fine-grained level distributes modeling complexity more effectively than an all-AR or all-NAR approach.
    source: §V.B, Tables III–IV
    evidence: A hybrid AR+NAR design that enforces independence at the coarse token level and interdependence at
      the fine-grained level distributes modeling complexity more effectively than an all-AR or all-NAR approach.
    confidence: high
    relevance: medium
  - claim_id: parallel_semantic_acoustic_modeling_improves_naturalness_and_intelligibility_without_fully
    role: supports
    claim: Parallel semantic-acoustic modeling improves naturalness and intelligibility without fully closing the
      speaker similarity gap relative to systems with dedicated speaker embedding refinement.
    source: §V.A, Table I
    evidence: Parallel semantic-acoustic modeling improves naturalness and intelligibility without fully closing
      the speaker similarity gap relative to systems with dedicated speaker embedding refinement.
    confidence: high
    relevance: low
  limitations:
  - Speaker similarity lags behind CosyVoice (SMOS gap ~0.15–0.2 on English), suggesting the parallel architecture
    does not yet fully leverage speaker conditioning. UTMOS scores, while competitive, do not reach ground-truth
    levels. Model size is not reported, making compute comparisons difficult. The internal Chinese dataset and preprocessing
    pipeline (Emilia + NCSSD) are not publicly released, limiting reproducibility on that front. Extending the framework
    to prosody control, emotion conditioning, or cross-lingual voice conversion is not explored. The subjective
    decoupling evaluation (Section V.C) relies on 90% evaluator agreement rather than a standardized metric, leaving
    quantitative disentanglement assessment as an open question.
  caveats: []
- id: '2508.04996'
  published_date: "2025-08-07"
  entry_date: '2026-07-29'
  year: 2025
  venue: arXiv
  task:
  - VC
  architecture:
  - flow-matching
  relevance: medium
  evidence_role:
  - architecture_variant
  current_role: active_evidence
  method_family:
  - flow_matching_with_ssl_conditioning
  claims:
  - claim_id: ssl_features_improve_paralinguistic_expressiveness_in_voice_conversion_but_introduce
    role: supports
    claim: SSL features improve paralinguistic expressiveness in voice conversion but introduce timbre leakage and
      noise sensitivity that require explicit mitigation.
    source: §I, §II.B
    evidence: SSL features improve paralinguistic expressiveness in voice conversion but introduce timbre leakage
      and noise sensitivity that require explicit mitigation.
    confidence: high
    relevance: high
  - claim_id: random_feature_erasure_at_training_time_can_reduce_a_model
    role: supports
    claim: Random feature erasure at training time can reduce a model's over-reliance on information-rich but noise-sensitive
      representations without information bottleneck machinery.
    source: §II.B
    evidence: Random feature erasure at training time can reduce a model's over-reliance on information-rich but
      noise-sensitive representations without information bottleneck machinery.
    confidence: high
    relevance: medium
  - claim_id: implicit_alignment_borrowed_from_non_autoregressive_tts_can_improve_noise
    role: supports
    claim: Implicit alignment borrowed from non-autoregressive TTS can improve noise robustness in voice conversion
      by preventing the model from over-reconstructing noise-carrying source frames.
    source: §II.C
    evidence: Implicit alignment borrowed from non-autoregressive TTS can improve noise robustness in voice conversion
      by preventing the model from over-reconstructing noise-carrying source frames.
    confidence: high
    relevance: low
  - claim_id: shortcut_models_reduce_flow_matching_inference_steps_by_an_order
    role: supports
    claim: Shortcut Models reduce flow-matching inference steps by an order of magnitude with only marginal quality
      loss in voice conversion.
    source: §II.D, Table I
    evidence: Shortcut Models reduce flow-matching inference steps by an order of magnitude with only marginal quality
      loss in voice conversion.
    confidence: high
    relevance: low
  - claim_id: asr_based_bottleneck_features_and_ssl_representations_are_complementary_the
    role: supports
    claim: 'ASR-based bottleneck features and SSL representations are complementary: the former provides noise-robust
      linguistic content, the latter contributes paralinguistic fidelity that ASR training suppresses.'
    source: §I, §II.A
    evidence: 'ASR-based bottleneck features and SSL representations are complementary: the former provides noise-robust
      linguistic content, the latter contributes paralinguistic fidelity that ASR training suppresses.'
    confidence: high
    relevance: high
  limitations:
  - The model cannot synthesise arbitrarily long utterances. The implicit alignment mechanism introduces a maximum-length
    constraint analogous to that in E2TTS-style TTS systems. The authors flag this as a known limitation without
    providing an upper bound or workaround.
  - 'Evaluation scale is small: 100 clean and 50 noisy test utterances is insufficient to draw strong conclusions
    about generalisation across noise types or speaking styles. The noisy set recording conditions are not fully
    documented. Comparison to other noise-robust VC systems such as NORO ([[2411.19770|NORO]]) is absent — only
    Seed-VC and a VITS-VC internal baseline are used.'
  - 'The prosody preservation trade-off is acknowledged: REF-VC preserves source prosody well, but users may prefer
    target speaker style transfer instead. Future work is needed to support simultaneous prosody preservation and
    style transfer.'
  - Singing voice conversion is mentioned as a capability but receives no quantitative evaluation.
  caveats: []
- id: '2508.05385'
  published_date: "2025-08-07"
  entry_date: '2026-07-29'
  year: 2025
  venue: arXiv
  task:
  - TTS
  architecture:
  - transformer-enc-dec
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: active_evidence
  method_family:
  - transformer_ssl_encoders_and_adapters
  claims:
  - claim_id: automated_annotation_pipelines_for_non_verbal_vocalizations_can_match_or
    role: supports
    claim: Automated annotation pipelines for non-verbal vocalizations can match or exceed manually-annotated datasets
      in downstream NV generation and understanding tasks, while scaling at substantially lower cost.
    source: §5.1, Table 2; §5.2, Tables 4–5
    evidence: Automated annotation pipelines for non-verbal vocalizations can match or exceed manually-annotated
      datasets in downstream NV generation and understanding tasks, while scaling at substantially lower cost.
    confidence: high
    relevance: medium
  - claim_id: positional_accuracy_of_non_verbal_tag_annotations_is_a_more
    role: supports
    claim: Positional accuracy of non-verbal tag annotations is a more important determinant of NV controllability
      in TTS than dataset size alone.
    source: §5.1, Table 2
    evidence: Positional accuracy of non-verbal tag annotations is a more important determinant of NV controllability
      in TTS than dataset size alone.
    confidence: high
    relevance: medium
  - claim_id: frame_level_nv_detection_models_trained_exclusively_on_one_language
    role: supports
    claim: Frame-level NV detection models trained exclusively on one language can generalise to structurally different
      languages without retraining, suggesting that acoustic features of non-verbal sounds are largely language-agnostic.
    source: §2.2; §3
    evidence: Frame-level NV detection models trained exclusively on one language can generalise to structurally
      different languages without retraining, suggesting that acoustic features of non-verbal sounds are largely
      language-agnostic.
    confidence: high
    relevance: medium
  - claim_id: rule_based_nv_data_augmentation_produces_measurably_worse_nv_controllability
    role: supports
    claim: Rule-based NV data augmentation produces measurably worse NV controllability in fine-tuned TTS systems
      compared to models trained on naturally-occurring vocalizations.
    source: §5.1, Table 2
    evidence: Rule-based NV data augmentation produces measurably worse NV controllability in fine-tuned TTS systems
      compared to models trained on naturally-occurring vocalizations.
    confidence: high
    relevance: medium
  limitations:
  - The dataset is heavily skewed toward Chinese (over two-thirds of samples), not because the pipeline is language-dependent
    but because the crawled source material is predominantly Chinese. This imbalance means English NV performance,
    while competitive, may underperform on broader English benchmarks, and the paper does not provide results on
    out-of-domain English test sets.
  - The detection model is trained and evaluated on simulated test data (augmented speech + NV clips), not on naturally-occurring
    NV events in the wild. Whether the ~91% F1 reflects real-world performance is untested. The evaluation covers
    only six NV categories despite the dataset spanning ten. NVS is also relatively small (38K samples, ~131 hours)
    compared to comparable speech datasets used for general TTS; the observed SSIM and WER gaps relative to the
    base F5-TTS suggest this is a binding constraint.
  caveats: []
- id: '2508.06890'
  published_date: "2025-08-09"
  entry_date: '2026-07-29'
  year: 2025
  venue: ASRU
  task:
  - VC
  architecture:
  - GAN
  - transformer-enc-dec
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: active_evidence
  method_family:
  - gan_decoders_for_ssl_units
  - transformer_ssl_encoders_and_adapters
  claims:
  - claim_id: frame_level_emotion_representations_improve_speaker_emotion_classification_accuracy_and
    role: supports
    claim: Frame-level emotion representations improve speaker emotion classification accuracy and prosody transfer
      fidelity over utterance-level representations in voice conversion.
    source: §IV.C, Table I
    evidence: Frame-level emotion representations improve speaker emotion classification accuracy and prosody transfer
      fidelity over utterance-level representations in voice conversion.
    confidence: high
    relevance: low
  - claim_id: adversarial_disentanglement_via_gradient_reversal_layers_reduces_phonetic_leakage_into
    role: supports
    claim: Adversarial disentanglement via gradient reversal layers reduces phonetic leakage into emotion embeddings,
      improving intelligibility under cross-linguistic-content conversion.
    source: §IV.C, Table I
    evidence: Adversarial disentanglement via gradient reversal layers reduces phonetic leakage into emotion embeddings,
      improving intelligibility under cross-linguistic-content conversion.
    confidence: high
    relevance: medium
  - claim_id: explicit_conditioning_on_extracted_prosodic_features_f0_and_energy_from
    role: supports
    claim: Explicit conditioning on extracted prosodic features (F0 and energy) from an emotion reference transfers
      temporal dynamics more faithfully than implicit prediction from latent codes.
    source: §IV.A, Table I
    evidence: Explicit conditioning on extracted prosodic features (F0 and energy) from an emotion reference transfers
      temporal dynamics more faithfully than implicit prediction from latent codes.
    confidence: high
    relevance: medium
  - claim_id: training_time_prosody_augmentation_through_temporal_shifting_and_warping_improves
    role: supports
    claim: Training-time prosody augmentation through temporal shifting and warping improves robustness of prosody
      transfer under mismatched reference conditions without sacrificing naturalness.
    source: §IV.C, Table I
    evidence: Training-time prosody augmentation through temporal shifting and warping improves robustness of prosody
      transfer under mismatched reference conditions without sacrificing naturalness.
    confidence: high
    relevance: medium
  limitations:
  - Training and primary evaluation use only the English ESD corpus — 350 parallel utterances across 10 speakers
    and 5 emotion categories. This is a narrow domain; generalisation to spontaneous, noisy, or multilingual emotional
    speech is entirely untested.
  - The small, parallel ESD corpus makes it difficult to assess whether the disentanglement holds under more naturalistic
    or non-parallel conditions. The ablation study evaluates the seen scenario only; it is not clear whether the
    ablated variants degrade similarly on unseen speakers and emotions. Speaker classification accuracy (SCA) is
    reported as a zero-shot metric for the seen-speaker scenario but becomes undefined for unseen speakers, so that
    dimension of the zero-shot evaluation lacks a corresponding metric. Model size and inference speed are not reported,
    which matters for the real-time dubbing applications the paper motivates.
  caveats: []
- id: '2508.07273'
  published_date: "2025-08-10"
  entry_date: '2026-07-29'
  year: 2025
  venue: arXiv
  task:
  - SCA
  architecture:
  - transformer-enc-dec
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  - evaluation_caution
  current_role: active_evidence
  method_family:
  - transformer_ssl_encoders_and_adapters
  claims:
  - claim_id: training_speech_llms_on_data_that_jointly_encodes_contextual_and
    role: supports
    claim: Training Speech-LLMs on data that jointly encodes contextual and paralinguistic reasoning substantially
      outperforms training on isolated paralinguistic QA templates, even when the underlying speech encoder is frozen.
    source: §V.B.1, Table III
    evidence: Training Speech-LLMs on data that jointly encodes contextual and paralinguistic reasoning substantially
      outperforms training on isolated paralinguistic QA templates, even when the underlying speech encoder is frozen.
    confidence: high
    relevance: medium
  - claim_id: incorporating_dimensional_emotion_annotations_valence_arousal_dominance_alongside_categorical_labels
    role: supports
    claim: Incorporating dimensional emotion annotations (valence, arousal, dominance) alongside categorical labels
      into LLM-generated QA data diversifies training supervision and improves generalisation to complex emotional
      states.
    source: §II.D, §V.B.1
    evidence: Incorporating dimensional emotion annotations (valence, arousal, dominance) alongside categorical
      labels into LLM-generated QA data diversifies training supervision and improves generalisation to complex
      emotional states.
    confidence: high
    relevance: medium
  - claim_id: explicitly_injecting_emotion_metadata_into_inference_prompts_can_partially_compensate
    role: complicates
    claim: Explicitly injecting emotion metadata into inference prompts can partially compensate for a model's limited
      intrinsic paralinguistic understanding, but training on contextual-paralinguistic data yields more robust
      generalisation across question types.
    source: §V.B.3, Fig. 2
    evidence: Explicitly injecting emotion metadata into inference prompts can partially compensate for a model's
      limited intrinsic paralinguistic understanding, but training on contextual-paralinguistic data yields more
      robust generalisation across question types.
    confidence: high
    relevance: medium
  - claim_id: llm_judge_scores_for_open_ended_speech_language_model_evaluation
    role: supports
    claim: LLM judge scores for open-ended speech-language model evaluation correlate reliably with classification-based
      accuracy and F1 metrics on questions with deterministic answers, supporting their use as a proxy metric.
    source: §IV, §V.B.4, Table V
    evidence: LLM judge scores for open-ended speech-language model evaluation correlate reliably with classification-based
      accuracy and F1 metrics on questions with deterministic answers, supporting their use as a proxy metric.
    confidence: high
    relevance: low
  limitations:
  - The CPQA training data is derived from a proprietary in-house movie and TV dataset that is not publicly released,
    making exact replication of the training setup impossible for external researchers.
  - Emotion labels used both in training and inference prompts come from SER models rather than ground-truth annotations,
    introducing noise that may suppress performance on direct classification tasks while still benefiting contextual
    reasoning. The LLM-generated CPQA evaluation set contains evaluation confounds — direct emotion questions that
    benefit disproportionately from explicit metadata injection — which the authors flag but do not resolve in the
    current work, requiring stricter QA generation controls in follow-up.
  - 'The evaluation scope is narrow: CPQA and emotion-PQA benchmarks only. Performance on other SCA capabilities
    (dialogue coherence, turn-taking, ASR) is not assessed, so it is unclear whether the CPQA training data causes
    regression elsewhere. All experiments use a single base architecture (MERaLiON); generalisability to other Speech-LLM
    frameworks remains untested.'
  caveats: []
- id: '2508.07375'
  published_date: "2025-08-10"
  entry_date: '2026-07-29'
  year: 2025
  venue: arXiv
  task:
  - SCA
  architecture:
  - autoregressive-LM
  relevance: medium
  evidence_role:
  - architecture_variant
  - evaluation_caution
  current_role: active_evidence
  method_family:
  - autoregressive_ssl_conditioned_models
  claims:
  - claim_id: aggregating_text_tokens_at_the_dialogue_turn_level_provides_stronger
    role: supports
    claim: Aggregating text tokens at the dialogue-turn level provides stronger semantic guidance for full-duplex
      speech generation than word-level token-by-token text insertion.
    source: §2.2, §4.2.1, Table 2
    evidence: Aggregating text tokens at the dialogue-turn level provides stronger semantic guidance for full-duplex
      speech generation than word-level token-by-token text insertion.
    confidence: high
    relevance: low
  - claim_id: the_timing_of_text_insertion_into_full_duplex_dialogue_sequences
    role: supports
    claim: The timing of text insertion into full-duplex dialogue sequences is more sensitive to error than the
      length of inserted text, with mistimed insertions causing larger semantic degradation than incorrectly sized
      text chunks.
    source: §4.2.3, Table 4
    evidence: The timing of text insertion into full-duplex dialogue sequences is more sensitive to error than the
      length of inserted text, with mistimed insertions causing larger semantic degradation than incorrectly sized
      text chunks.
    confidence: high
    relevance: low
  - claim_id: increasing_the_training_loss_weight_of_text_tokens_relative_to
    role: supports
    claim: Increasing the training loss weight of text tokens relative to speech tokens in a text-speech interleaved
      model improves semantic quality without sacrificing turn-taking naturalness.
    source: §4.2.1, Table 2
    evidence: Increasing the training loss weight of text tokens relative to speech tokens in a text-speech interleaved
      model improves semantic quality without sacrificing turn-taking naturalness.
    confidence: high
    relevance: low
  - claim_id: fine_grained_turn_taking_benchmarks_are_more_informative_than_corpus
    role: supports
    claim: Fine-grained turn-taking benchmarks are more informative than corpus-level statistical correlations for
      evaluating full-duplex spoken dialogue models.
    source: §4.2.2
    evidence: Fine-grained turn-taking benchmarks are more informative than corpus-level statistical correlations
      for evaluating full-duplex spoken dialogue models.
    confidence: high
    relevance: low
  - claim_id: gpt_based_automated_semantic_evaluation_of_spoken_dialogue_aligns_closely
    role: supports
    claim: GPT-based automated semantic evaluation of spoken dialogue aligns closely with human preference judgements
      when score differences exceed one point, but reliability degrades for near-tied comparisons.
    source: Appendix A, Table 5
    evidence: GPT-based automated semantic evaluation of spoken dialogue aligns closely with human preference judgements
      when score differences exceed one point, but reliability degrades for near-tied comparisons.
    confidence: high
    relevance: low
  limitations:
  - All results are on the Fisher telephone conversation corpus only. Fisher's conversational style and acoustic
    conditions (telephone, English, spontaneous) are narrow, and the paper provides no evidence that TurnGuide generalises
    to other languages, speaking styles, or domains.
  - The model has not undergone safety alignment (RLHF or equivalent), which the authors acknowledge is a prerequisite
    for deployment. The theoretical first-package latency is 1.05 seconds — workable for some interactive applications
    but not near real-time; the authors note the vocoder imposes the dominant cost and that an optimised decoder
    would reduce this.
  - The evaluation uses GPT-4o as both the semantic evaluator and implicitly as an oracle for dialogue quality,
    which introduces a potential circularity if the trained model's outputs are biased toward patterns GPT-4o scores
    favourably. The corpus-level Pearson correlation analysis (Table 7) shows TurnGuide is at parity with baselines
    on turn-taking statistics but does not establish whether the improvement in GPT-score translates to user preference
    in live interaction.
  caveats: []
- id: '2508.08399'
  published_date: "2025-08-11"
  entry_date: '2026-07-29'
  year: 2025
  venue: arXiv
  task:
  - codec
  architecture:
  - VAE
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: active_evidence
  method_family:
  - vae_and_quantized_ssl_representations
  claims:
  - claim_id: fully_discrete_disentanglement_of_phonetic_prosodic_and_speaker_information_in
    role: complicates
    claim: Fully discrete disentanglement of phonetic, prosodic, and speaker information in a speech codec is achievable
      without phoneme labels or F0 supervision, at the cost of a small reconstruction quality degradation.
    source: §III, §IV.B, Table II
    evidence: Fully discrete disentanglement of phonetic, prosodic, and speaker information in a speech codec is
      achievable without phoneme labels or F0 supervision, at the cost of a small reconstruction quality degradation.
    confidence: high
    relevance: low
  - claim_id: quantizing_speaker_vectors_into_discrete_codes_reduces_speaker_identity_fidelity
    role: complicates
    claim: Quantizing speaker vectors into discrete codes reduces speaker identity fidelity compared to continuous
      speaker representations, posing a fundamental trade-off between LLM compatibility and speaker retention.
    source: §IV.B, Table III
    evidence: Quantizing speaker vectors into discrete codes reduces speaker identity fidelity compared to continuous
      speaker representations, posing a fundamental trade-off between LLM compatibility and speaker retention.
    confidence: high
    relevance: medium
  - claim_id: instance_normalization_of_ssl_residual_features_provides_a_label_free
    role: supports
    claim: Instance normalization of SSL residual features provides a label-free mechanism to separate time-invariant
      speaker statistics from time-variant prosodic content.
    source: §III.B
    evidence: Instance normalization of SSL residual features provides a label-free mechanism to separate time-invariant
      speaker statistics from time-variant prosodic content.
    confidence: high
    relevance: high
  - claim_id: fully_discrete_speech_codecs_can_match_conventional_voice_conversion_methods
    role: supports
    claim: Fully discrete speech codecs can match conventional voice conversion methods on intelligibility and naturalness
      while enabling attribute manipulation through codebook-level operations.
    source: §IV.B, Table III
    evidence: Fully discrete speech codecs can match conventional voice conversion methods on intelligibility and
      naturalness while enabling attribute manipulation through codebook-level operations.
    confidence: high
    relevance: low
  limitations:
  - All experiments use LibriSpeech clean speech (16 kHz, studio conditions); performance on noisy, spontaneous,
    or out-of-domain speech is untested. The one-shot VC evaluation uses only two reference speakers (one male,
    one female), limiting statistical confidence in the speaker similarity results.
  - The model is not tested on any downstream application (TTS, ASR, speech LM), despite this being the stated motivation.
    Whether the disentangled discrete tokens actually improve over non-disentangled tokens on downstream tasks remains
    an open question — the paper acknowledges this as future work. The GRVQ codebook dimensionality analysis shows
    a clear trade-off between bitrate and speaker identity, but optimal bitrate allocation across the three streams
    is not systematically explored. Prosody quantization codebook interpretability beyond F0 correlation (e.g.,
    energy, duration, speaking rate) is not investigated.
  caveats: []
- id: '2508.09600'
  published_date: "2025-08-13"
  entry_date: '2026-07-29'
  year: 2025
  venue: arXiv
  task:
  - SCA
  architecture:
  - autoregressive-LM
  - flow-matching
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  - evaluation_caution
  current_role: active_evidence
  method_family:
  - autoregressive_ssl_conditioned_models
  - flow_matching_with_ssl_conditioning
  claims:
  - claim_id: explicit_chain_of_thought_reasoning_over_paralinguistic_cues_emotion_age
    role: supports
    claim: Explicit chain-of-thought reasoning over paralinguistic cues (emotion, age, gender, sound events) improves
      empathetic response generation in speech-to-speech dialogue systems.
    source: §OSUM-EChat "Training", Stage 3 Empathy; Table 2
    evidence: Explicit chain-of-thought reasoning over paralinguistic cues (emotion, age, gender, sound events)
      improves empathetic response generation in speech-to-speech dialogue systems.
    confidence: high
    relevance: low
  - claim_id: pretraining_on_multitask_speech_understanding_before_speech_to_speech_dialogue
    role: supports
    claim: Pretraining on multitask speech understanding before speech-to-speech dialogue training reduces dependence
      on large-scale paired dialogue datasets while maintaining paralinguistic modelling quality.
    source: §OSUM-EChat "Training", Stage 1; Table 2
    evidence: Pretraining on multitask speech understanding before speech-to-speech dialogue training reduces dependence
      on large-scale paired dialogue datasets while maintaining paralinguistic modelling quality.
    confidence: high
    relevance: low
  - claim_id: synthetic_speech_to_speech_data_derived_from_tts_systems_exhibits
    role: supports
    claim: Synthetic speech-to-speech data derived from TTS systems exhibits reduced emotional expressiveness compared
      to real human speech, creating a systematic domain gap that degrades empathetic dialogue evaluation.
    source: §EChat-200K Dataset
    evidence: Synthetic speech-to-speech data derived from TTS systems exhibits reduced emotional expressiveness
      compared to real human speech, creating a systematic domain gap that degrades empathetic dialogue evaluation.
    confidence: high
    relevance: low
  - claim_id: automatic_empathy_evaluation_pipelines_using_llm_scoring_and_automatic_emotion
    role: supports
    claim: Automatic empathy evaluation pipelines using LLM scoring and automatic emotion classifiers diverge measurably
      from human judgements, primarily due to emotion classifier errors and LLM hallucinations.
    source: §Main Results "Results of Empathetic Intelligence"; Table 3
    evidence: Automatic empathy evaluation pipelines using LLM scoring and automatic emotion classifiers diverge
      measurably from human judgements, primarily due to emotion classifier errors and LLM hallucinations.
    confidence: high
    relevance: low
  - claim_id: native_multimodal_models_that_integrate_speech_token_prediction_directly_into
    role: supports
    claim: Native multimodal models that integrate speech token prediction directly into the LLM are better suited
      to capturing and generating paralinguistic nuance than modularly aligned architectures that treat speech decoding
      separately.
    source: §Introduction; §Related Work "End-to-End Spoken Dialogue System"
    evidence: Native multimodal models that integrate speech token prediction directly into the LLM are better suited
      to capturing and generating paralinguistic nuance than modularly aligned architectures that treat speech decoding
      separately.
    confidence: high
    relevance: medium
  limitations:
  - 'The EChat-200K dataset is almost entirely synthetic: query audio is generated by CosyVoice2 and response audio
    likewise, with real recordings comprising only a minority of the data. The authors acknowledge that emotional
    expressiveness of synthesised audio lags real human speech, and the training corpus does not include dynamic
    paralinguistic scenarios (e.g. emotional transitions, multi-speaker interactions). Generalisability to natural
    in-the-wild speech remains unvalidated.'
  - 'The EChat-eval automatic scoring pipeline — combining GPT-4o and emotion2vec-Large — produces rankings consistent
    with human evaluation but absolute scores that diverge meaningfully. The benchmark is therefore more reliable
    for ranking systems than for measuring absolute empathy. Sound event capability is evaluated only within the
    categories present in EChat-200K, which may not reflect the diversity of real conversational contexts. General
    linguistic intelligence (UltraEval-Audio, Table 4) regresses compared to the Qwen2.5-3B base: GSM8K drops from
    85 to 34 after the multi-stage training, indicating that the speech capability gain comes at a significant cost
    to LLM reasoning.'
  caveats: []
- id: '2508.11224'
  published_date: "2025-08-15"
  entry_date: '2026-07-29'
  year: 2025
  venue: ASRU
  task:
  - evaluation
  - TTS
  architecture: []
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: active_evidence
  method_family: []
  claims:
  - claim_id: ssl_models_using_frame_wise_masked_prediction_capture_relative_prosodic
    role: supports
    claim: SSL models using frame-wise masked prediction capture relative prosodic contours within an utterance
      rather than absolute acoustic magnitudes, making them insensitive to global intensity rescaling.
    source: §V-A, Fig. 1
    evidence: SSL models using frame-wise masked prediction capture relative prosodic contours within an utterance
      rather than absolute acoustic magnitudes, making them insensitive to global intensity rescaling.
    confidence: high
    relevance: high
  - claim_id: models_pretrained_to_predict_discrete_targets_encode_phoneme_like_structure
    role: supports
    claim: Models pretrained to predict discrete targets encode phoneme-like structure effectively at small cluster
      sizes, while models pretrained on continuous targets require larger cluster sizes to approach the same phonemic
      alignment.
    source: §V-A, Fig. 3
    evidence: Models pretrained to predict discrete targets encode phoneme-like structure effectively at small cluster
      sizes, while models pretrained on continuous targets require larger cluster sizes to approach the same phonemic
      alignment.
    confidence: high
    relevance: medium
  - claim_id: training_k_means_clustering_on_emotionally_expressive_speech_increases_the
    role: supports
    claim: Training k-means clustering on emotionally expressive speech increases the prosodic sensitivity of resulting
      tokens for most SSL model and layer combinations.
    source: §V-B, Table I
    evidence: Training k-means clustering on emotionally expressive speech increases the prosodic sensitivity of
      resulting tokens for most SSL model and layer combinations.
    confidence: high
    relevance: high
  - claim_id: applying_a_temporal_moving_average_to_ssl_features_before_k
    role: complicates
    claim: Applying a temporal moving average to SSL features before k-means clustering provides an adjustable trade-off
      between prosodic sensitivity and speaker invariance, with intermediate window sizes improving both simultaneously.
    source: §V-C, Fig. 5
    evidence: Applying a temporal moving average to SSL features before k-means clustering provides an adjustable
      trade-off between prosodic sensitivity and speaker invariance, with intermediate window sizes improving both
      simultaneously.
    confidence: high
    relevance: high
  - claim_id: differences_between_ssl_pretraining_objectives_in_their_token_level_linguistic
    role: supports
    claim: Differences between SSL pretraining objectives in their token-level linguistic and prosodic encoding
      are concentrated in the final transformer layers, while intermediate layers exhibit largely similar behaviour
      across model families.
    source: §V-A, §V-B
    evidence: Differences between SSL pretraining objectives in their token-level linguistic and prosodic encoding
      are concentrated in the final transformer layers, while intermediate layers exhibit largely similar behaviour
      across model families.
    confidence: high
    relevance: high
  limitations:
  - The analysis measures sensitivity via TER — a proxy for how much token sequences change in response to acoustic
    manipulation — rather than directly probing what information is decodable from the tokens. Whether the observed
    sensitivity differences translate to actual gains in downstream prosody-related tasks (prosody-conditional TTS,
    emphasis transfer, emotion recognition from discrete tokens) remains untested. The evaluation uses a single
    corpus (TIMIT) of read speech by native English speakers, which may limit generalisability to spontaneous, conversational,
    or multilingual speech. Duration was excluded from prosody analysis because deduplication collapses durational
    information; this omission means the benchmark does not cover the full prosody space. The study does not test
    acoustic tokens (neural codec outputs), focusing exclusively on semantic tokens.
  caveats: []
- id: '2508.11273'
  published_date: "2025-08-15"
  entry_date: '2026-07-29'
  year: 2025
  venue: arXiv
  task:
  - TTS
  architecture:
  - transformer-enc-dec
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: active_evidence
  method_family:
  - transformer_ssl_encoders_and_adapters
  claims:
  - claim_id: discretizing_self_supervised_speech_features_via_k_means_produces_more
    role: supports
    claim: Discretizing self-supervised speech features via k-means produces more stable prosodic control signals
      than retaining continuous SSL representations in encoder-decoder TTS.
    source: §5.6, Table 1
    evidence: Discretizing self-supervised speech features via k-means produces more stable prosodic control signals
      than retaining continuous SSL representations in encoder-decoder TTS.
    confidence: high
    relevance: high
  - claim_id: combining_continuous_spherical_emotion_vectors_with_discrete_ssl_prosody_tokens
    role: supports
    claim: Combining continuous spherical emotion vectors with discrete SSL prosody tokens improves both naturalness
      and intelligibility compared to spherical emotion vectors alone.
    source: §5.1, §5.3, Tables 1–2
    evidence: Combining continuous spherical emotion vectors with discrete SSL prosody tokens improves both naturalness
      and intelligibility compared to spherical emotion vectors alone.
    confidence: high
    relevance: high
  - claim_id: speaker_independent_prosody_conditioning_via_ssl_tokens_can_generalize_across
    role: supports
    claim: Speaker-independent prosody conditioning via SSL tokens can generalize across reference speakers with
      minimal degradation, enabling robust cross-speaker emotion transfer.
    source: §5.6, Table 1
    evidence: Speaker-independent prosody conditioning via SSL tokens can generalize across reference speakers with
      minimal degradation, enabling robust cross-speaker emotion transfer.
    confidence: high
    relevance: high
  - claim_id: semantic_text_encoders_contribute_to_emotional_and_prosodic_consistency_in
    role: supports
    claim: Semantic text encoders contribute to emotional and prosodic consistency in multilingual TTS even when
      they do not serve as direct decoder inputs.
    source: §5.6, Table 1
    evidence: Semantic text encoders contribute to emotional and prosodic consistency in multilingual TTS even when
      they do not serve as direct decoder inputs.
    confidence: high
    relevance: medium
  limitations:
  - '- Evaluations are limited to single-speaker datasets in two languages, making it unclear whether EmoSSLSphere
    generalises to multi-speaker, low-resource, or unseen-language scenarios. - Subjective listener panels are small,
    and emotional authenticity is assessed primarily via AVD RMSE as a proxy rather than direct perceptual emotion
    ratings. - Speaker similarity (SPK-SIM) is not evaluated, making it hard to quantify speaker fidelity claims.
    - Cross-speaker emotion transfer is described but not formally evaluated; inference always uses same-speaker
    reference audio. - Separate per-language encoder instances do not scale to many-language settings without significant
    parameter overhead. - Integration with semi-supervised training (EmoSphere++) and extension to zero-shot speaker
    scenarios are listed as future work.'
  caveats: []
- id: interspeech-2025-0115
  published_date: "2025-08-17"
  entry_date: '2026-07-29'
  year: 2025
  venue: Interspeech
  task:
  - codec
  architecture:
  - transformer-enc-dec
  relevance: medium
  evidence_role:
  - architecture_variant
  current_role: active_evidence
  method_family:
  - transformer_ssl_encoders_and_adapters
  claims:
  - claim_id: even_codecs_explicitly_trained_with_disentanglement_objectives_fail_to_cleanly
    role: complicates
    claim: Even codecs explicitly trained with disentanglement objectives fail to cleanly separate pitch from other
      speech attributes in their token embeddings.
    source: §2.3, §2.4
    evidence: Even codecs explicitly trained with disentanglement objectives fail to cleanly separate pitch from
      other speech attributes in their token embeddings.
    confidence: high
    relevance: medium
  - claim_id: linguistic_content_in_neural_audio_codec_representations_concentrates_in_the
    role: supports
    claim: Linguistic content in neural audio codec representations concentrates in the lowest RVQ scales regardless
      of whether distillation was used, but leaks into higher scales when the frame rate is very low.
    source: §2.1
    evidence: Linguistic content in neural audio codec representations concentrates in the lowest RVQ scales regardless
      of whether distillation was used, but leaks into higher scales when the frame rate is very low.
    confidence: high
    relevance: low
  - claim_id: a_masked_autoencoder_framework_can_bridge_codec_tokens_and_perceptual
    role: supports
    claim: A masked-autoencoder framework can bridge codec tokens and perceptual speech attributes bidirectionally,
      enabling voice conversion at dramatically lower bitrates than spectrogram-based equivalents.
    source: §3.1, §3.2
    evidence: A masked-autoencoder framework can bridge codec tokens and perceptual speech attributes bidirectionally,
      enabling voice conversion at dramatically lower bitrates than spectrogram-based equivalents.
    confidence: high
    relevance: low
  - claim_id: post_hoc_interpretability_tools_reveal_systematic_trade_offs_between_content
    role: supports
    claim: Post-hoc interpretability tools reveal systematic trade-offs between content accuracy and synthesis quality
      that differ by codec design, complicating the choice of codec for controllable speech generation.
    source: §3.2, Table 1
    evidence: Post-hoc interpretability tools reveal systematic trade-offs between content accuracy and synthesis
      quality that differ by codec design, complicating the choice of codec for controllable speech generation.
    confidence: high
    relevance: low
  limitations:
  - All experiments use LibriSpeech, a clean read-speech corpus with limited acoustic diversity. Whether the observed
    encoding patterns hold for spontaneous speech, expressive data, or noise-conditioned codecs is untested.
  - The study covers four specific codecs; the broader generalisation across the growing landscape of codec designs
    (including future multi-scale or end-to-end codec-LM systems) is an open question. Pitch estimation from codec
    tokens remains poor, and no remedy is proposed — it is unclear whether this is a fundamental limitation of RVQ-based
    representations or an artefact of the particular codecs studied. The bidirectional AnCoGen-Codec is compared
    only against the Melspectrogram baseline and not against dedicated disentanglement-oriented codec frameworks
    such as FreeCodec or SpeechFlow, which would provide stronger context for the synthesis results.
  caveats: []
- id: interspeech-2025-0143
  published_date: "2025-08-17"
  entry_date: '2026-07-29'
  year: 2025
  venue: Interspeech
  task:
  - TTS
  - evaluation
  architecture: []
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: active_evidence
  method_family: []
  claims:
  - claim_id: multimodal_fusion_of_data_driven_acoustic_and_word_level_linguistic
    role: supports
    claim: Multimodal fusion of data-driven acoustic and word-level linguistic embeddings outperforms unimodal and
      knowledge-based features for multilingual sentence mode classification.
    source: §4, Table 2
    evidence: Multimodal fusion of data-driven acoustic and word-level linguistic embeddings outperforms unimodal
      and knowledge-based features for multilingual sentence mode classification.
    confidence: high
    relevance: medium
  - claim_id: multilingual_ssl_representations_trained_on_large_corpora_can_transfer_sentence
    role: supports
    claim: Multilingual SSL representations trained on large corpora can transfer sentence-mode discriminative information
      across languages without language-specific training.
    source: §5
    evidence: Multilingual SSL representations trained on large corpora can transfer sentence-mode discriminative
      information across languages without language-specific training.
    confidence: high
    relevance: high
  - claim_id: state_of_the_art_asr_systems_are_unable_to_reliably
    role: supports
    claim: State-of-the-art ASR systems are unable to reliably detect exclamatory sentence mode from speech, producing
      recall rates below chance level for that class.
    source: §4, Table 3
    evidence: State-of-the-art ASR systems are unable to reliably detect exclamatory sentence mode from speech,
      producing recall rates below chance level for that class.
    confidence: high
    relevance: medium
  - claim_id: sentence_mode_prediction_performance_degrades_substantially_on_emotional_speech_due
    role: supports
    claim: Sentence mode prediction performance degrades substantially on emotional speech due to definitional overlap
      between exclamatory sentence mode and emotional expressiveness.
    source: §5, Table 4
    evidence: Sentence mode prediction performance degrades substantially on emotional speech due to definitional
      overlap between exclamatory sentence mode and emotional expressiveness.
    confidence: high
    relevance: medium
  limitations:
  - '- No end-to-end TTS experiment; the study measures sentence mode prediction accuracy, not synthesized prosody
    quality. - Labels derived from punctuation marks are a proxy for sentence mode; they may not reflect actual
    prosodic realization (especially for audiobooks where speakers may monotonize exclamatory passages). - Only
    three languages evaluated; coverage of typologically diverse languages is absent. - MLP classifier is simple;
    more powerful sequence models may perform better. - UAR is used rather than accuracy due to class imbalance,
    but class sizes differ substantially across languages.'
  caveats: []
- id: interspeech-2025-0166
  published_date: "2025-08-17"
  entry_date: '2026-07-29'
  year: 2025
  venue: Interspeech
  task:
  - SCA
  architecture:
  - transformer-enc-dec
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  - evaluation_caution
  current_role: active_evidence
  method_family:
  - transformer_ssl_encoders_and_adapters
  claims:
  - claim_id: training_a_speech_encoder_to_match_llm_responses_conditioned_on
    role: supports
    claim: Training a speech encoder to match LLM responses conditioned on emotion labels can transfer paralinguistic
      understanding to a frozen LLM without any LLM weight modification.
    source: §2.3, §4.1, §4.2, Table 1
    evidence: SpeechEmotionLlama achieves an LLM-annotated emotion understanding score of 7.54 versus 5.83 for the
      best cascaded baseline, and 81.51% SER accuracy versus 80.81% for the standalone SER model, while keeping
      Llama 3 8B entirely frozen throughout training.
    confidence: high
    relevance: medium
  - claim_id: cascading_a_high_accuracy_standalone_emotion_classifier_onto_a_speech
    role: complicates
    claim: Cascading a high-accuracy standalone emotion classifier onto a speech-LLM system provides only marginal
      improvement in emotional response quality.
    source: §3.2, §4.1, Table 1
    evidence: Prepending SER model predictions (80.81% accuracy) to Baseline 2 raises the LLM-annotated emotion
      score from 5.59 to only 5.83, while encoder-level alignment raises it to 7.54, suggesting that discretized
      emotion tags do not capture the richness of paralinguistic information as effectively as continuous encoder
      embeddings.
    confidence: high
    relevance: medium
  - claim_id: evaluating_paralinguistic_speech_llm_systems_is_inherently_difficult_when_training
    role: complicates
    claim: Evaluating paralinguistic speech-LLM systems is inherently difficult when training and evaluation data
      are fully proprietary.
    source: §3.1, §3.3
    evidence: All datasets (pre-training, fine-tuning, and test sets) are proprietary, precluding external replication;
      the LLM-based evaluation metric (Llama 3 70B as judge) also introduces a dependency on the evaluator model's
      own behavior, which is not systematically validated against human listeners.
    confidence: high
    relevance: low
  - claim_id: self_supervised_speech_pre_training_on_large_scale_multilingual_data
    role: supports
    claim: Self-supervised speech pre-training on large-scale multilingual data provides a strong initialization
      for downstream paralinguistic understanding in speech encoders.
    source: §2.2, §3.2, §4.2
    evidence: The Conformer encoder is pre-trained with BEST-RQ on approximately 15M hours before fine-tuning on
      emotion tasks; the system substantially outperforms baselines that use a speech encoder that is not further
      fine-tuned on emotion-related tasks, indicating the pre-training provides useful representations that emotion-specific
      fine-tuning can build upon.
    confidence: high
    relevance: high
  limitations:
  - All training and evaluation data are proprietary; results cannot be replicated externally and comparisons are
    restricted to in-house baselines. The role of data scale (1B-parameter encoder, 15M hours pre-training) versus
    the training paradigm itself is not ablated.
  - The system is evaluated only on English expressive speech from voice actors and a large but single-language
    corpus, leaving generalization to other languages, accents, and naturalistic (non-acted) expressive speech open.
    The test set is intentionally constructed to control for linguistic content by using utterances where the same
    sentence is spoken in multiple styles, which may not reflect realistic distribution. LLM-judged metrics (Llama
    3 70B as scorer) have not been validated against human listener judgments for empathy or response quality in
    this setup.
  caveats: []
- id: interspeech-2025-0203
  published_date: "2025-08-17"
  entry_date: '2026-07-29'
  year: 2025
  venue: Interspeech
  task:
  - VC
  architecture:
  - flow-matching
  - transformer-enc-dec
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: active_evidence
  method_family:
  - flow_matching_with_ssl_conditioning
  - transformer_ssl_encoders_and_adapters
  claims:
  - claim_id: natural_language_prompts_can_control_emotional_voice_conversion_at_parity
    role: supports
    claim: Natural language prompts can control emotional voice conversion at parity with reference speech for the
      majority of listeners, reducing reliance on hard-to-source reference audio.
    source: §3.2.2
    evidence: Natural language prompts can control emotional voice conversion at parity with reference speech for
      the majority of listeners, reducing reliance on hard-to-source reference audio.
    confidence: high
    relevance: low
  - claim_id: flow_matching_produces_noticeably_higher_speech_naturalness_and_audio_quality
    role: supports
    claim: Flow matching produces noticeably higher speech naturalness and audio quality in emotional voice conversion
      than GAN and autoencoder baselines.
    source: §3.2.1, Table 1
    evidence: Flow matching produces noticeably higher speech naturalness and audio quality in emotional voice conversion
      than GAN and autoencoder baselines.
    confidence: high
    relevance: low
  - claim_id: combining_categorical_emotion_labels_with_free_form_prompt_labels_through
    role: supports
    claim: Combining categorical emotion labels with free-form prompt labels through soft-label contrastive training
      improves emotion embedding quality over prompt-only or label-only training.
    source: §3.3, Table 2
    evidence: Combining categorical emotion labels with free-form prompt labels through soft-label contrastive training
      improves emotion embedding quality over prompt-only or label-only training.
    confidence: high
    relevance: medium
  - claim_id: an_explicit_scalar_intensity_gate_applied_to_emotional_embeddings_before
    role: supports
    claim: An explicit scalar intensity gate applied to emotional embeddings before content-emotion fusion improves
      both naturalness and emotion similarity in converted speech.
    source: §3.3, Table 2
    evidence: An explicit scalar intensity gate applied to emotional embeddings before content-emotion fusion improves
      both naturalness and emotion similarity in converted speech.
    confidence: high
    relevance: medium
  limitations:
  - The system is trained and evaluated entirely on a proprietary internal Mandarin corpus. No open-source data
    or model weights are released, and no cross-lingual or multi-speaker generalisation is tested.
  - Comparisons are restricted to older GAN and autoencoder baselines (StarGAN-EVC, Seq2seq-EVC, MixEmo); no diffusion-based
    or recent flow-matching EVC systems are included, so the claimed state-of-the-art position cannot be verified
    against the most competitive contemporaries. The evaluation is any-to-one (fixed target speaker identity), leaving
    any-to-any EVC performance unaddressed. Emotion coverage is limited to seven categorical classes; whether the
    natural language conditioning generalises to subtler or blended emotional states is untested.
  caveats: []
- id: interspeech-2025-0246
  published_date: "2025-08-17"
  entry_date: '2026-07-29'
  year: 2025
  venue: Interspeech
  task:
  - codec
  architecture:
  - transformer-enc-dec
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: active_evidence
  method_family:
  - transformer_ssl_encoders_and_adapters
  claims:
  - claim_id: speaker_invariant_discrete_speech_tokens_trained_with_a_dual_codebook
    role: supports
    claim: Speaker-invariant discrete speech tokens, trained with a dual-codebook objective, can improve both spoken
      language model performance and speech resynthesis intelligibility simultaneously.
    source: §3.2, §3.3, Tables 1–3
    evidence: Speaker-invariant discrete speech tokens, trained with a dual-codebook objective, can improve both
      spoken language model performance and speech resynthesis intelligibility simultaneously.
    confidence: high
    relevance: medium
  - claim_id: n_gram_predictability_and_phoneme_character_mutual_information_are_stronger
    role: supports
    claim: N-gram predictability and phoneme/character mutual information are stronger proxies for SLM downstream
      performance than the widely-used ABX error rate.
    source: §3.5, Figure 3b
    evidence: N-gram predictability and phoneme/character mutual information are stronger proxies for SLM downstream
      performance than the widely-used ABX error rate.
    confidence: high
    relevance: medium
  - claim_id: ssl_based_tokenizers_can_match_or_exceed_neural_codec_tokenizers
    role: supports
    claim: SSL-based tokenizers can match or exceed neural codec tokenizers on speech intelligibility (WER) at substantially
      lower bitrates when the encoder is fine-tuned to suppress speaker variation.
    source: §3.3, Table 3
    evidence: SSL-based tokenizers can match or exceed neural codec tokenizers on speech intelligibility (WER) at
      substantially lower bitrates when the encoder is fine-tuned to suppress speaker variation.
    confidence: high
    relevance: high
  - claim_id: supervised_fine_tuning_with_phoneme_recognition_targets_provides_consistent_gains
    role: supports
    claim: Supervised fine-tuning with phoneme recognition targets provides consistent gains over ASR targets for
      speech resynthesis intelligibility, suggesting phoneme alignment is more directly beneficial than word-level
      transcription for unit-based vocoders.
    source: §3.3, Table 3
    evidence: Supervised fine-tuning with phoneme recognition targets provides consistent gains over ASR targets
      for speech resynthesis intelligibility, suggesting phoneme alignment is more directly beneficial than word-level
      transcription for unit-based vocoders.
    confidence: high
    relevance: medium
  - claim_id: scaling_slm_model_size_has_diminishing_returns_on_tasks_that
    role: supports
    claim: Scaling SLM model size has diminishing returns on tasks that require sentence-level semantic coherence
      when the tokenizer quality is the primary bottleneck.
    source: §3.2, Table 2
    evidence: Scaling SLM model size has diminishing returns on tasks that require sentence-level semantic coherence
      when the tokenizer quality is the primary bottleneck.
    confidence: high
    relevance: medium
  limitations:
  - The unconstrained comparison in Table 2 pits a 150M SLM trained on 6k hours against models using 13B parameters
    and 150k+ hours. While the framing as "limited-resource" is accurate, direct comparisons against high-resource
    systems risk being misleading — the TSC gap (70.7 vs. 82.9 for SPIRIT LM) is large and likely attributable to
    scale rather than tokenizer quality.
  - The evaluation is English-only, and the authors acknowledge multilinguality as future work. The HiFi-GAN vocoder
    receives speaker and style IDs externally, so the speaker-invariance of the tokens is not tested in an open-vocabulary,
    zero-shot resynthesis scenario. The Expresso WER figures (>10%) are elevated by the expressive speaking styles
    (laughter, whisper), making absolute comparisons with non-expressive benchmarks impractical. The paper also
    does not report wall-clock training times or resource costs in sufficient detail for full reproducibility assessment.
  caveats: []
- id: interspeech-2025-0305
  published_date: "2025-08-17"
  entry_date: '2026-07-29'
  year: 2025
  venue: Interspeech
  task:
  - singing
  - VC
  architecture:
  - flow-matching
  - transformer-enc-dec
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: active_evidence
  method_family:
  - flow_matching_with_ssl_conditioning
  - transformer_ssl_encoders_and_adapters
  claims:
  - claim_id: flow_matching_decoders_produce_higher_audio_quality_than_gan_based
    role: supports
    claim: Flow matching decoders produce higher audio quality than GAN-based decoders in singing voice conversion
      when conditioning signal quality is held constant.
    source: §4.1, §4.2, Table 1, Table 2
    evidence: Replacing NeuCoSVC's GAN-based FastSVC decoder with a CFM module improves MCD from 8.634 to 7.220
      and MOS-Naturalness from 3.47 to 3.80 on OpenSinger, with the ablation confirming that even without the DCAM
      module the CFM-equipped model surpasses NeuCoSVC.
    confidence: high
    relevance: low
  - claim_id: ssl_feature_matching_prevents_timbre_leakage_in_singing_voice_conversion
    role: refines
    claim: SSL feature matching prevents timbre leakage in singing voice conversion but is insufficient on its own
      for high timbre similarity, because target timbre is distributed across the full reference utterance rather
      than captured by sparse nearest-neighbour retrieval.
    source: §4.2, Table 1, Table 2
    evidence: The -spk&att ablation (SSL matching only, no speaker embeddings or DCAM) achieves SSIM 0.709 vs. 0.692
      for NeuCoSVC (marginal improvement), while adding speaker embeddings with the full DCAM raises SSIM to 0.754
      — a larger relative gain than the matching step alone provides.
    confidence: high
    relevance: high
  - claim_id: cross_attention_fusion_of_speaker_embeddings_and_melody_features_with
    role: supports
    claim: Cross-attention fusion of speaker embeddings and melody features with shared content queries improves
      timbre similarity and audio coherence over simple feature concatenation in conditional singing voice conversion.
    source: §4.2, Table 2
    evidence: Removing the DCAM while retaining speaker embeddings (-att ablation) drops SSIM from 0.754 to 0.710
      and MCD from 7.220 to 8.129, demonstrating that the attention mechanism adds value beyond the conditioning
      signals themselves.
    confidence: high
    relevance: low
  - claim_id: one_shot_singing_voice_conversion_evaluations_remain_narrow_in_scope
    role: complicates
    claim: One-shot singing voice conversion evaluations remain narrow in scope, limiting the generalisability of
      reported gains.
    source: §3.1, §3.4, §5
    evidence: Experiments use a single Chinese singing dataset (OpenSinger), 20 samples for subjective evaluation
      with 15 listeners, and four unseen target speakers. Cross-language, multi-domain, or noisy-environment generalisation
      is explicitly deferred to future work.
    confidence: high
    relevance: low
  limitations:
  - The system is trained and evaluated exclusively on high-quality Chinese singing (OpenSinger, recorded in a professional
    studio), and the authors acknowledge that performance in noisy environments and cross-language settings is untested.
    The subjective evaluation is small (20 samples, 15 listeners), raising questions about statistical robustness.
    The pitch shifting strategy (scaling source pitch by the ratio of target-to-source median pitch) is a global
    heuristic that may not capture fine-grained vocal range adaptation. Model size, inference latency, and real-time
    factor are not reported. Whether the DCAM design generalises beyond the Chinese singing domain remains open.
  caveats: []
- id: interspeech-2025-0310
  published_date: "2025-08-17"
  entry_date: '2026-07-29'
  year: 2025
  venue: Interspeech
  task:
  - SCA
  - codec
  architecture:
  - autoregressive-LM
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: active_evidence
  method_family:
  - autoregressive_ssl_conditioned_models
  claims:
  - claim_id: moderately_coarse_fixed_width_segmentation_around_80_ms_combined_with
    role: supports
    claim: Moderately coarse fixed-width segmentation (around 80 ms) combined with a large K-means vocabulary outperforms
      fine-grained original-resolution tokenization on zero-shot spoken language understanding tasks without sacrificing
      accuracy.
    source: §4, Table 3
    evidence: Moderately coarse fixed-width segmentation (around 80 ms) combined with a large K-means vocabulary
      outperforms fine-grained original-resolution tokenization on zero-shot spoken language understanding tasks
      without sacrificing accuracy.
    confidence: high
    relevance: medium
  - claim_id: larger_segmentation_widths_require_proportionally_larger_vocabularies_to_preserve_phonetic
    role: supports
    claim: Larger segmentation widths require proportionally larger vocabularies to preserve phonetic discriminability,
      analogous to the phoneme-morpheme relationship in linguistics.
    source: §5.1
    evidence: Larger segmentation widths require proportionally larger vocabularies to preserve phonetic discriminability,
      analogous to the phoneme-morpheme relationship in linguistics.
    confidence: high
    relevance: medium
  - claim_id: variable_width_segmentation_based_on_linguistic_units_phoneme_syllable_word
    role: supports
    claim: Variable-width segmentation based on linguistic units (phoneme, syllable, word boundaries) does not consistently
      outperform fixed-width segmentation of matched median duration, and incurs additional computational cost.
    source: §5.3
    evidence: Variable-width segmentation based on linguistic units (phoneme, syllable, word boundaries) does not
      consistently outperform fixed-width segmentation of matched median duration, and incurs additional computational
      cost.
    confidence: high
    relevance: medium
  - claim_id: optimal_tokenization_settings_vary_across_spoken_language_understanding_benchmarks_suggesting
    role: supports
    claim: Optimal tokenization settings vary across spoken language understanding benchmarks, suggesting that ensembling
      multiple tokenization schemes may be necessary for broad SLU capability.
    source: §5.2
    evidence: Optimal tokenization settings vary across spoken language understanding benchmarks, suggesting that
      ensembling multiple tokenization schemes may be necessary for broad SLU capability.
    confidence: high
    relevance: medium
  limitations:
  - '- Evaluation is exclusively on SLU (understanding) tasks; no speech generation quality assessment. - Training
    data (LibriSpeech 960h) is small by current SLM standards; findings may differ at larger scale (LibriLight 60k).
    - Only HuBERT layer-9 as the SSL backbone; other models (WavLM, wav2vec 2.0) may yield different tradeoff curves.
    - Variable-width segmentation using predicted (rather than oracle) boundaries introduces inaccuracy; learned
    segment representations (as in Sylber) may be better than raw pooling. - Optimal tokenization varies per benchmark,
    motivating multi-token ensemble strategies not explored here.'
  caveats: []
- id: interspeech-2025-0383
  published_date: "2025-08-17"
  entry_date: '2026-07-29'
  year: 2025
  venue: Interspeech
  task:
  - VC
  architecture:
  - transformer-enc-dec
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  - evaluation_caution
  current_role: active_evidence
  method_family:
  - transformer_ssl_encoders_and_adapters
  claims:
  - claim_id: discrete_unit_voice_conversion_can_be_extended_to_control_subjective
    role: supports
    claim: Discrete-unit voice conversion can be extended to control subjective perceptual attributes beyond speaker
      identity by adding a scalar conditioning signal to the synthesis model.
    source: §3, §5.3, Figure 6
    evidence: A FastSpeech 2 model conditioned on HuBERT-based discrete units, ECAPA-TDNN speaker embeddings, and
      a target likability scalar successfully steered perceived likability for 3 of 4 speakers in pairwise preference
      tests.
    confidence: high
    relevance: low
  - claim_id: automatic_likability_predictors_based_on_tdnn_regression_on_crowd_sourced
    role: supports
    claim: Automatic likability predictors based on TDNN regression on crowd-sourced ratings can provide sufficient
      proxy labels to train large-scale perceptual attribute control systems.
    source: §4.2, Table 2
    evidence: The predictor achieved LCC 0.46 and SRCC 0.49 with human ratings (p < 3e-17) and 74% binary classification
      accuracy; its outputs were used to automatically annotate the JVS and JTES corpora for VC training.
    confidence: high
    relevance: medium
  - claim_id: strong_likability_control_and_speaker_identity_preservation_are_in_tension
    role: complicates
    claim: Strong likability control and speaker identity preservation are in tension in discrete-unit VC systems,
      particularly at extreme target values.
    source: §5.2, §5.3, Figures 4–6
    evidence: At target likability = 2 (outside the training range), CER increased substantially for female speakers
      and speaker m49 showed degraded speaker similarity and unexpected subjective likability drop; the inference-time
      scalar multiplier partially mitigates but does not eliminate this trade-off.
    confidence: high
    relevance: medium
  - claim_id: voice_likability_control_demonstrates_effective_behaviour_for_majority_speaker_groups
    role: complicates
    claim: Voice likability control demonstrates effective behaviour for majority speaker groups but can fail for
      individual speakers due to identity-likability interaction effects.
    source: §5.3, Figure 6
    evidence: Three of four speakers showed significant preference differences between target -1 and 1; speaker
      m49 exhibited the opposite trend, attributed to failure to preserve speaker identity, indicating that per-speaker
      variation is a real limitation.
    confidence: high
    relevance: medium
  - claim_id: perceived_voice_likability_is_a_multi_factorial_attribute_requiring_demographic
    role: refines
    claim: Perceived voice likability is a multi-factorial attribute requiring demographic-stratified modelling,
      not a single group-level score.
    source: §2, §4.2, Table 2
    evidence: The predictor uses four separate listener-group outputs (by gender and age); per-group LCC ranges
      from 0.36 to 0.41 while the aggregate LCC is 0.46, confirming systematic variation across listener demographics.
    confidence: high
    relevance: medium
  limitations:
  - The evaluation is limited to four speakers (two female, two male) and one language (Japanese), with training
    on Japanese speech corpora. Cross-lingual and multi-lingual generalisability is entirely untested.
  - 'The practical control range of the system is narrow: despite targeting values from -2 to 2, predicted likability
    shifts only from approximately -0.51 to -0.23, suggesting the model substantially operates within each speaker''s
    inherent likability range rather than achieving wide stylistic transfer. This is acknowledged by the authors
    but not resolved.'
  - The likability predictor's correlation with human ratings is statistically significant but moderate (LCC 0.46),
    meaning a substantial portion of subjective likability variance is not captured. Since this predictor is the
    primary source of training supervision, any systematic biases in its predictions will propagate to the VC model.
  - The scalar multiplier (s = 2.5 at inference vs. 1 at training) is manually tuned and distribution-shifted, which
    may limit out-of-distribution robustness. Future work noted by the authors includes extending to natural language
    specification of desired voice characteristics.
  caveats: []
- id: interspeech-2025-0433
  published_date: "2025-08-17"
  entry_date: '2026-07-29'
  year: 2025
  venue: Interspeech
  task:
  - VC
  architecture:
  - VAE
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: active_evidence
  method_family:
  - vae_and_quantized_ssl_representations
  claims:
  - claim_id: voice_conversion_architectures_designed_for_human_speech_require_non_trivial
    role: supports
    claim: Voice conversion architectures designed for human speech require non-trivial adaptation to generalise
      to non-human vocalizations with broad frequency ranges and transient-rich characteristics.
    source: §4.2.2, Table 1
    evidence: Replacing the proposed preprocessing pipeline with a conventional speech-focused one degraded WER
      from 24.02% to 62.29% and MOS-S from 3.78 to 3.53, indicating that human-speech assumptions about frame resolution
      and frequency range materially impair non-human sound conversion.
    confidence: high
    relevance: low
  - claim_id: isolating_style_conditioning_to_the_prior_network_and_normalizing_flow
    role: supports
    claim: Isolating style conditioning to the prior network and normalizing flow, and excluding it from the posterior
      encoder and decoder, reduces style leakage and improves speaker similarity in CVAE-based voice conversion.
    source: §4.2.3, Table 1
    evidence: Adding the style embedding to the audio encoder and decoder (w/ SEED ablation) reduced MOS-S from
      3.78 to 3.61, attributed to style overlap between the reference encoder output and latent acoustic tokens.
    confidence: high
    relevance: low
  - claim_id: kl_annealing_mitigates_posterior_collapse_in_vae_based_voice_conversion
    role: supports
    claim: KL annealing mitigates posterior collapse in VAE-based voice conversion and improves linguistic content
      preservation, particularly for complex non-human vocalizations.
    source: §4.2.3, Table 1
    evidence: Removing KL annealing increased CER from 15.48% to 28.89% and WER from 24.02% to 44.69%, while MOS
      scores changed minimally, indicating that linguistic clarity is the primary casualty of over-regularization
      in early training.
    confidence: high
    relevance: low
  - claim_id: existing_fundamental_frequency_f0_estimation_methods_are_not_reliable_for
    role: complicates
    claim: Existing fundamental frequency (F0) estimation methods are not reliable for sounds lacking a well-defined
      harmonic structure, constraining prosodic feature extraction in non-human voice conversion systems.
    source: §3.1
    evidence: The authors tested frame-level F0 from non-human sounds but found existing estimators (Praat, CREPE,
      SPICE, PESTO) exhibited limitations due to absent harmonic structure; the system falls back to energy-only
      prosodic features, leaving robust F0 extraction as an open problem.
    confidence: high
    relevance: low
  limitations:
  - The dataset is entirely internal and the evaluation uses only 9 human raters on an unspecified number of test
    samples, limiting reproducibility and the statistical reliability of MOS scores.
  - No publicly available data or code is confirmed, making direct comparison and reproduction difficult. The evaluation
    benchmarks non-human VC against baselines that were not adapted for non-human sounds, which is the correct setup
    for the paper's argument but means absolute MOS values are not comparable to human-speech VC literature.
  - 'F0 estimation for non-human sounds remains unsolved: the system uses energy-only prosodic features and omits
    pitch conditioning, which may limit prosodic expressiveness for vocalizations where pitch contour is perceptually
    salient (e.g., melodic birdsong). The evaluation scope is restricted to a set of internally defined sound categories
    (exclamations, designed voices, animal sounds from a commercial library), and generalisation to out-of-distribution
    non-human sounds is not tested.'
  caveats: []
- id: interspeech-2025-0438
  published_date: "2025-08-17"
  entry_date: '2026-07-29'
  year: 2025
  venue: Interspeech
  task:
  - VC
  architecture:
  - hybrid
  relevance: high
  evidence_role:
  - core_evidence
  - control_evidence
  - evaluation_caution
  current_role: active_evidence
  method_family:
  - hybrid_semantic_acoustic_ssl_systems
  claims:
  - claim_id: linear_transformations_of_self_supervised_speech_features_are_sufficient_for
    role: supports
    claim: Linear transformations of self-supervised speech features are sufficient for competitive voice conversion,
      without complex nonlinear decoders or model fine-tuning.
    source: §3.2, Table 1
    evidence: LinearVC's single learned projection matrix W on WavLM-Large layer 6 achieves WER 4.9%, EER 33.6%,
      and MUSHRA naturalness 62.5 on LibriSpeech test-clean, statistically indistinguishable from kNN-VC and SoundStorm
      in both naturalness and speaker similarity.
    confidence: high
    relevance: high
  - claim_id: content_and_speaker_identity_information_reside_in_orthogonal_low_dimensional
    role: supports
    claim: Content and speaker identity information reside in orthogonal low-dimensional subspaces within the same
      SSL layer, enabling voice style transfer through geometric manipulation.
    source: §4, Table 2; §5.2, Figure 4
    evidence: Constraining the linear transformation to rotation and reflection only achieves EER 27.7% vs. 31.8%
      for unconstrained; adding translation alone gives only EER 7.7% with intelligibility maintained (CER 2.9%),
      confirming a shared phonetic subspace across speakers. SVD factorization at rank 16 achieves CER < 4%, while
      speaker similarity requires rank ~100.
    confidence: high
    relevance: high
  - claim_id: high_naturalness_in_voice_conversion_does_not_imply_high_speaker
    role: complicates
    claim: High naturalness in voice conversion does not imply high speaker similarity — these objectives can trade
      off sharply depending on the system design.
    source: §3.2, Table 1
    evidence: FreeVC achieves the highest naturalness (71.1 MUSHRA) among all systems but the lowest speaker similarity
      (EER 10.5% vs. 33.6% for LinearVC and 38.9% for kNN-VC), indicating that perceptual smoothness and target-speaker
      fidelity are partly in tension.
    confidence: high
    relevance: low
  - claim_id: eer_as_an_objective_speaker_similarity_metric_provides_only_coarse
    role: refines
    claim: EER as an objective speaker similarity metric provides only coarse correspondence with perceived speaker
      similarity in voice conversion evaluation.
    source: §3.2
    evidence: The paper notes that small EER differences do not reliably track subjective similarity ratings, and
      that EER gives a coarse correspondence with perceptual judgements. LinearVC and kNN-VC have comparable subjective
      similarity (67.5 vs. 67.2) but different EERs (33.6 vs. 38.9).
    confidence: high
    relevance: low
  limitations:
  - The analysis is restricted to WavLM-Large (layer 6) on English LibriSpeech. Whether the orthogonal subspace
    structure holds for other SSL models (HuBERT, wav2vec 2.0), other layers, other languages, or noisy/spontaneous
    speech settings is left as future work. The vocoder is shared across all systems, making it difficult to isolate
    conversion quality from synthesis quality. The training data requirement (2.7 minutes per target speaker) means
    LinearVC is not a zero-shot system in the strictest sense; it requires a small per-speaker adaptation step.
    Performance at very low reference amounts is untested. Future applications the authors suggest include speech
    anonymization and phonetic content extraction.
  caveats: []
- id: interspeech-2025-0468
  published_date: "2025-08-17"
  entry_date: '2026-07-29'
  year: 2025
  venue: Interspeech
  task:
  - codec
  architecture:
  - VAE
  - GAN
  relevance: medium
  evidence_role:
  - architecture_variant
  - acceleration_evidence
  current_role: active_evidence
  method_family:
  - vae_and_quantized_ssl_representations
  - gan_decoders_for_ssl_units
  claims:
  - claim_id: directly_encoding_ssl_features_as_a_first_class_codec_stream
    role: supports
    claim: Directly encoding SSL features as a first-class codec stream produces stronger semantic preservation
      in RVQ-1 tokens than distillation from an SSL model, particularly for tonal languages where pitch fidelity
      is critical.
    source: §4.2, Table 2
    evidence: Directly encoding SSL features as a first-class codec stream produces stronger semantic preservation
      in RVQ-1 tokens than distillation from an SSL model, particularly for tonal languages where pitch fidelity
      is critical.
    confidence: high
    relevance: high
  - claim_id: operating_a_neural_codec_at_lower_frame_rates_with_more
    role: supports
    claim: Operating a neural codec at lower frame rates with more RVQ layers at fixed token rate improves audio
      quality over higher-frame-rate codecs with fewer layers at the same bitrate.
    source: §4.3, Table 3
    evidence: Operating a neural codec at lower frame rates with more RVQ layers at fixed token rate improves audio
      quality over higher-frame-rate codecs with fewer layers at the same bitrate.
    confidence: high
    relevance: low
  - claim_id: semantic_quality_of_rvq_1_tokens_is_a_primary_determinant
    role: supports
    claim: Semantic quality of RVQ-1 tokens is a primary determinant of downstream TTS intelligibility in autoregressive
      codec-based systems, independent of codec audio reconstruction quality.
    source: §4.4, Table 4
    evidence: Semantic quality of RVQ-1 tokens is a primary determinant of downstream TTS intelligibility in autoregressive
      codec-based systems, independent of codec audio reconstruction quality.
    confidence: high
    relevance: low
  - claim_id: an_ssl_based_semantic_stream_in_a_codec_encoder_can
    role: supports
    claim: An SSL-based semantic stream in a codec encoder can improve perceptual audio quality beyond what waveform-only
      codecs achieve, even when using the same decoder architecture.
    source: §4.3, Table 3
    evidence: An SSL-based semantic stream in a codec encoder can improve perceptual audio quality beyond what waveform-only
      codecs achieve, even when using the same decoder architecture.
    confidence: high
    relevance: high
  limitations:
  - The 12.5 Hz DualCodec-based TTS lags behind the 25 Hz variant in both WER and speaker similarity, indicating
    that the more aggressive downsampling introduces a ceiling on semantic accuracy that affects TTS quality. The
    paper acknowledges this gap as the primary remaining challenge.
  - The SSL model (w2v-BERT-2.0, 600M parameters, frozen) is required at TTS training time but not inference. This
    makes the training pipeline heavier than pure waveform codec approaches. It is also unclear whether the approach
    generalises to SSL models other than w2v-BERT-2.0, or whether the chosen 16th layer feature is optimal across
    languages beyond English and Mandarin.
  - Codec evaluation uses a single controlled bitrate band (~0.75 kbps); performance at higher bitrates typical
    of studio-quality TTS is not evaluated. The subjective test has only 8 participants, limiting statistical confidence
    for MUSHRA comparisons between closely-scoring systems.
  caveats: []
- id: interspeech-2025-0506
  published_date: "2025-08-17"
  entry_date: '2026-07-29'
  year: 2025
  venue: Interspeech
  task:
  - codec
  architecture:
  - transformer-enc-dec
  relevance: medium
  evidence_role:
  - architecture_variant
  current_role: active_evidence
  method_family:
  - transformer_ssl_encoders_and_adapters
  claims:
  - claim_id: neural_codec_token_spaces_capture_semantically_useful_information_beyond_audio
    role: supports
    claim: Neural codec token spaces capture semantically useful information beyond audio reconstruction, enabling
      effective use as pretraining targets for general audio representation models.
    source: §4, Table 1
    evidence: EnCodecMAE, which predicts frozen EnCodec RVQ tokens, outperforms BEATs (which uses random quantizer
      targets) and AudioMAE on global HEAREval score, with large+ST models reaching 99.2 on Audioset-only pretraining
      versus 96.0 for BEATs Iter 3.
    confidence: high
    relevance: low
  - claim_id: frame_level_temporal_representations_outperform_patch_based_representations_for_speech
    role: supports
    claim: Frame-level temporal representations outperform patch-based representations for speech tasks in universal
      audio models, while both approaches perform comparably on environmental sound classification.
    source: §4, Table 1
    evidence: EnCodecMAE (frame-level) achieves 96.3% speech command accuracy and 75.5% emotion recognition versus
      patch-based BEATs Iter 1's 91.0% and 68.7%, while matching or being slightly behind on ESC-50 (79.8 vs. 82.3).
    confidence: high
    relevance: medium
  - claim_id: self_supervised_universal_audio_representations_do_not_close_the_gap
    role: complicates
    claim: Self-supervised universal audio representations do not close the gap with dedicated speech SSL models
      on phoneme-level tasks such as ASR.
    source: §4, Table 2
    evidence: EnCodecMAE Large+ST achieves 8.59% WER on LibriSpeech with a language model, versus 2.94% for HuBERT
      Large trained solely on speech data. The paper attributes this to differences in target definition (MFCC clusters
      correlating with phonemes) and input representation (raw waveform).
    confidence: high
    relevance: high
  - claim_id: the_optimal_input_representation_for_self_supervised_audio_models_is
    role: refines
    claim: The optimal input representation for self-supervised audio models is task-dependent rather than universally
      preferable across audio domains.
    source: §4, Table 1
    evidence: Melspectrograms yield higher global HEAREval scores than EnCodec encoder outputs (95.9 vs. 83.4 base
      model), but EnCodec input outperforms melspectrograms on pitch prediction (NSynth), where codec features capture
      finer harmonic structure.
    confidence: high
    relevance: high
  - claim_id: pretraining_data_diversity_is_essential_for_universal_audio_representation_models
    role: supports
    claim: Pretraining data diversity is essential for universal audio representation; models pretrained on a single
      domain generalise poorly outside that domain.
    source: §4, Table 1
    evidence: Speech-only pretraining (LL6K only) gives a global score of 75.3, music-only (FMA only) gives 89.1,
      while the full mixture (AS+FMA+LL6K) reaches 97.1 global, with gains concentrated in the non-primary domain
      for each restricted setting.
    confidence: high
    relevance: medium
  limitations:
  - The ASR evaluation uses the SUPERB protocol (frozen encoder, BiLSTM probe, 100h fine-tune), which is designed
    to measure representation quality, not end-to-end ASR. The resulting WERs cannot be directly compared against
    fine-tuned ASR systems trained end-to-end and should be treated as a proxy for phoneme-level representation
    quality only.
  - Evaluation on HEAREval uses instance-level embeddings derived by averaging frame-level features, which discards
    temporal structure useful for tasks requiring sequence-level reasoning. The self-training stage cluster targets
    (k-means over 10k samples with k=1024) may be too coarse for high-resolution phonetic discrimination, which
    the authors identify as a direction for future work alongside exploring alternative training targets to improve
    ASR without sacrificing music and environmental performance. All models are evaluated with a frozen upstream
    encoder; fine-tuning experiments are not reported, leaving open whether the representations would improve further
    under task-specific adaptation.
  caveats: []
- id: interspeech-2025-0656
  published_date: "2025-08-17"
  entry_date: '2026-07-29'
  year: 2025
  venue: Interspeech
  task:
  - VC
  architecture:
  - hybrid
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  - evaluation_caution
  current_role: active_evidence
  method_family:
  - hybrid_semantic_acoustic_ssl_systems
  claims:
  - claim_id: cross_modal_feature_alignment_between_neural_signals_and_speaker_embeddings
    role: supports
    claim: Cross-modal feature alignment between neural signals and speaker embeddings can enable voice conversion
      without any target-speaker voice data.
    source: §3.1.2, §4.4
    evidence: The EEG-voice feature alignment module, trained with embedding MSE and speaker classification losses,
      produces speaker embeddings from EEG that drive FreeVC-based conversion to zero-shot quality (Naturalness
      MOS 4.00, Consistency obj 0.8026 for unseen speakers).
    confidence: high
    relevance: low
  - claim_id: non_speech_biometric_signals_can_encode_speaker_identity_information_sufficient
    role: supports
    claim: Non-speech biometric signals can encode speaker-identity information sufficient to guide voice timbre
      conversion.
    source: §4.4.1, Figure 2
    evidence: t-SNE visualisation shows synthesised speech clusters align with reference audio per speaker, and
      Homogeneity scores (0.9437–0.9465) exceed the FreeVC baseline (0.9371), indicating EEG features encode timbre-discriminative
      information.
    confidence: high
    relevance: medium
  - claim_id: zero_shot_voice_conversion_from_eeg_signals_requires_a_large
    role: complicates
    claim: Zero-shot voice conversion from EEG signals requires a large-scale speech-only pre-training stage to
      compensate for the scarcity and noise of paired EEG-speech data.
    source: §3.2
    evidence: The three-stage curriculum first pre-trains on VCTK (Stage I, speech only), then aligns EEG to pre-trained
      speaker embeddings (Stage II), before joint fine-tuning (Stage III). The authors explicitly attribute feasibility
      to leveraging the pre-trained VC model's representations.
    confidence: high
    relevance: low
  - claim_id: evaluation_of_eeg_driven_voice_conversion_is_fundamentally_limited_by
    role: complicates
    claim: Evaluation of EEG-driven voice conversion is fundamentally limited by the availability of paired EEG-speech
      corpora at the scale needed for generalisation.
    source: §4.1, §5
    evidence: The entire EEG evaluation uses the Single-Word-Production Dutch-iBIDS dataset (10 speakers, single-word
      utterances). Seen-speaker training uses 80% of this data; unseen-speaker evaluation uses k-fold over 10 speakers.
      The authors note that "more data will improve performance."
    confidence: high
    relevance: low
  limitations:
  - The evaluation dataset contains only 10 speakers producing single words in Dutch. Results on connected speech,
    diverse languages, and larger speaker populations are entirely untested. Generalisability claims should be treated
    as preliminary.
  - 'The baseline comparison (FreeVC) is not an equivalent system: FreeVC uses a target-speaker voice prompt while
    the proposed system uses EEG, so observed differences in quality scores may partly reflect the inherent difficulty
    of the EEG conditioning signal rather than architectural superiority. The paper does not report intelligibility
    metrics (WER/CER), making it impossible to assess whether semantic content is preserved faithfully. The EEG
    signals in the dataset come from intracranial recordings (high SNR), which may not transfer to consumer-grade
    scalp EEG equipment used in practical BCI deployments.'
  caveats: []
- id: interspeech-2025-0706
  published_date: "2025-08-17"
  entry_date: '2026-07-29'
  year: 2025
  venue: Interspeech
  task:
  - SCA
  architecture: []
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  - evaluation_caution
  current_role: active_evidence
  method_family: []
  claims:
  - claim_id: automated_llm_based_qa_generation_can_produce_paralinguistic_evaluation_sets
    role: supports
    claim: Automated LLM-based QA generation can produce paralinguistic evaluation sets that correlate closely with
      human-authored ones for assessing speech LLMs.
    source: §4, Table 2
    evidence: Qwen2-Audio-7B-Instruct achieves 60.28 (LLM QA) vs. 59.46 (Human QA) under the ChatGPT judge with
      Prompt 2, and 56.82 vs. 54.33 under the Llama-70B judge; both differences are small and directionally consistent.
    confidence: high
    relevance: low
  - claim_id: combining_categorical_and_dimensional_emotion_recognition_with_consistency_filtering_reduces
    role: supports
    claim: Combining categorical and dimensional emotion recognition with consistency filtering reduces annotation
      noise in in-the-wild speech data condensation.
    source: §3.1, Figure 3
    evidence: The dual SER consistency condition (categorical sentiment class aligned with valence score) applied
      on SG TV/Movie data achieved the highest UWA of 33.65% at valence thresholds x=0.5, y=0.4, outperforming single-paradigm
      labelling and improving class balance.
    confidence: high
    relevance: medium
  - claim_id: current_speech_llms_exhibit_weak_performance_on_contextual_empathetic_reasoning
    role: complicates
    claim: Current speech LLMs exhibit weak performance on contextual empathetic reasoning even when evaluated against
      well-formed paralinguistic QA.
    source: §4, §5
    evidence: Evaluation of Qwen2-Audio-7B-Instruct on the CPQA set reveals limitations in handling empathetic reasoning
      tasks; the paper identifies this as a motivating gap requiring both better data and more robust models.
    confidence: high
    relevance: medium
  - claim_id: scalable_automatic_qa_generation_from_speech_introduces_systematic_biases_that
    role: complicates
    claim: Scalable automatic QA generation from speech introduces systematic biases that require post-filtering
      to produce usable evaluation data.
    source: §3.2
    evidence: ChatGPT-generated QA contained repetitive question variants (multiple paraphrases asking about reasons
      behind emotion) and irrelevant questions assuming a text transcript was available; keyword-based post-filtering
      was required to remove these.
    confidence: high
    relevance: low
  limitations:
  - The evaluation benchmark is small (480 samples, 6.5 hours) and drawn from a single domain (Singaporean YouTube
    channels) in a mostly English-Mandarin context, so generalisation to other languages and speaking styles is
    unknown. The entire framework is validated against a single speech LLM (Qwen2-Audio-7B-Instruct), and the LLM
    judge scores are relatively similar across generated and human QA, which leaves open the question of whether
    the difference would matter more with a weaker model or a harder task. Speaker diarization is absent, causing
    the LLM to under-generate questions about multi-speaker interactions when multiple speakers share a gender.
    The internal SG TV/Movie dataset used for parameter tuning is not publicly available, making exact replication
    of the condensation configuration difficult.
  caveats: []
- id: interspeech-2025-0723
  published_date: "2025-08-17"
  entry_date: '2026-07-29'
  year: 2025
  venue: Interspeech
  task:
  - TTS
  architecture:
  - transformer-enc-dec
  - VAE
  relevance: medium
  evidence_role:
  - architecture_variant
  current_role: active_evidence
  method_family:
  - transformer_ssl_encoders_and_adapters
  - vae_and_quantized_ssl_representations
  claims:
  - claim_id: encoder_representations_in_trained_encoder_decoder_tts_models_encode_prosodic
    role: supports
    claim: Encoder representations in trained encoder-decoder TTS models encode prosodic and phonetic properties
      in a distributed, neuron-level format that is recoverable by linear classifiers.
    source: §3.1, §5.1
    evidence: Encoder representations in trained encoder-decoder TTS models encode prosodic and phonetic properties
      in a distributed, neuron-level format that is recoverable by linear classifiers.
    confidence: high
    relevance: medium
  - claim_id: gradient_ascent_in_a_compressed_vae_latent_space_produces_more
    role: supports
    claim: Gradient ascent in a compressed VAE latent space produces more controlled prosodic edits than direct
      activation manipulation in the ambient space, reducing off-manifold artifacts.
    source: §3.2.1, §5.2
    evidence: Gradient ascent in a compressed VAE latent space produces more controlled prosodic edits than direct
      activation manipulation in the ambient space, reducing off-manifold artifacts.
    confidence: high
    relevance: medium
  - claim_id: anchoring_latent_edits_to_a_vq_vae_prototype_codebook_is
    role: supports
    claim: Anchoring latent edits to a VQ-VAE prototype codebook is necessary to preserve phoneme identity under
      large prosodic shifts at inference time.
    source: §3.2.2, §5.2
    evidence: Anchoring latent edits to a VQ-VAE prototype codebook is necessary to preserve phoneme identity under
      large prosodic shifts at inference time.
    confidence: high
    relevance: medium
  - claim_id: post_hoc_activation_editing_can_correct_mispronunciations_in_grapheme_input
    role: supports
    claim: Post-hoc activation editing can correct mispronunciations in grapheme-input TTS without a pronunciation
      dictionary, using a speech-only correction query as the sole supervision signal.
    source: §5.3, Table 1
    evidence: Post-hoc activation editing can correct mispronunciations in grapheme-input TTS without a pronunciation
      dictionary, using a speech-only correction query as the sole supervision signal.
    confidence: high
    relevance: medium
  limitations:
  - '- Validated only on Tacotron 2; applicability to flow-matching or codec-based LM TTS is untested. - Feature
    entanglement remains: controlling duration slightly shifts pitch (acknowledged in the paper), suggesting that
    targeting specific neurons rather than full activation vectors is needed. - Mispronunciation correction requires
    a speech-only query word as supervision — not purely zero-shot. - LJSpeech is single-speaker; multi-speaker
    generalization is not evaluated.'
  caveats: []
- id: interspeech-2025-0756
  published_date: "2025-08-17"
  entry_date: '2026-07-29'
  year: 2025
  venue: Interspeech
  task:
  - SCA
  architecture:
  - hybrid
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: active_evidence
  method_family:
  - hybrid_semantic_acoustic_ssl_systems
  claims:
  - claim_id: structuring_speech_adapter_routing_around_distinct_affective_dimensions_improves_emotion
    role: supports
    claim: Structuring speech adapter routing around distinct affective dimensions improves emotion prediction and
      response quality in spoken dialogue systems compared to single-path adapters.
    source: §3.8, Table 4
    evidence: A-SMiLE's task-aware top-1/top-4 expert routing for VAD vs. response generation outperforms conv-based,
      Q-Former, and standard Transformer adapters across METEOR, ROUGE-L, and GPT-4o empathy scores on DailyTalk.
    confidence: high
    relevance: low
  - claim_id: multi_task_joint_training_of_emotion_regression_and_response_generation
    role: supports
    claim: Multi-task joint training of emotion regression and response generation improves both emotional alignment
      and text quality over response-only training in spoken dialogue.
    source: §3.6, Table 2
    evidence: EAML fine-tuning (Stage 2) consistently raises CCC scores and METEOR/ROUGE-L for both 0.5B and 7B
      backbones over text-only and cascaded baselines that receive VAD labels as auxiliary text rather than learned
      representations.
    confidence: high
    relevance: low
  - claim_id: discrete_categorical_emotion_labels_are_insufficient_for_capturing_fine_grained
    role: complicates
    claim: Discrete categorical emotion labels are insufficient for capturing fine-grained paralinguistic states
      such as sarcasm and depression in spoken dialogue.
    source: §3.1, §3.5, Table 1
    evidence: The hard-case benchmark (sourced from IEMOCAP and MUStARD) specifically targets emotionally complex
      states where existing systems trained on categorical emotion corpora fail to capture nuance; A-SMiLE's CCC-based
      VAD modeling shows substantial gains on exactly these cases.
    confidence: high
    relevance: low
  - claim_id: automatic_text_overlap_metrics_meteor_rouge_l_and_llm_based
    role: complicates
    claim: Automatic text-overlap metrics (METEOR, ROUGE-L) and LLM-based empathy scoring may not fully reflect
      human perceptual judgments of emotional appropriateness in dialogue responses.
    source: §3.2
    evidence: The paper evaluates response quality exclusively with METEOR, ROUGE-L, and GPT-4o-1120 empathy scores;
      no human listening evaluation is reported, leaving open the question of whether score gains translate to perceived
      empathy.
    confidence: high
    relevance: low
  limitations:
  - The evaluation relies entirely on reference-text metrics (METEOR, ROUGE-L) and GPT-4o-based empathy scores,
    with no human perceptual study. The DailyTalk corpus is 20 hours of scripted dialogue between two fixed speakers
    and may not represent the diversity of real-world conversations. The hard-case evaluation set is small (0.8
    hours). The system generates text responses rather than speech, leaving the downstream TTS step and its interaction
    with emotional conditioning unaddressed. Whether the cognitive specialisation motivation (routing to distinct
    brain-analogue experts) genuinely produces disentangled representations or is primarily a useful inductive bias
    remains untested analytically.
  caveats: []
- id: interspeech-2025-0815
  published_date: "2025-08-17"
  entry_date: '2026-07-29'
  year: 2025
  venue: Interspeech
  task:
  - VC
  architecture:
  - VAE
  - GAN
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: active_evidence
  method_family:
  - vae_and_quantized_ssl_representations
  - gan_decoders_for_ssl_units
  claims:
  - claim_id: discrete_speech_unit_representations_reduce_source_speaker_leakage_in_voice
    role: supports
    claim: Discrete speech unit representations reduce source speaker leakage in voice conversion but introduce
      pronunciation artefacts that degrade intelligibility compared to continuous feature counterparts.
    source: §4.2, Table 1
    evidence: Discrete speech unit representations reduce source speaker leakage in voice conversion but introduce
      pronunciation artefacts that degrade intelligibility compared to continuous feature counterparts.
    confidence: high
    relevance: high
  - claim_id: mix_style_layer_normalisation_mitigates_the_train_inference_mismatch_caused
    role: supports
    claim: Mix-style layer normalisation mitigates the train-inference mismatch caused by content-style dependence
      in style encoders, improving zero-shot generalisation on unseen speakers.
    source: §4.3, Table 2
    evidence: Mix-style layer normalisation mitigates the train-inference mismatch caused by content-style dependence
      in style encoders, improving zero-shot generalisation on unseen speakers.
    confidence: high
    relevance: medium
  - claim_id: enriching_global_style_embeddings_with_explicit_pitch_and_energy_features
    role: supports
    claim: Enriching global style embeddings with explicit pitch and energy features improves emotion transfer fidelity
      in expressive voice conversion beyond mel-spectrogram-only style encoding.
    source: §3.5, §4.3, Table 2
    evidence: Enriching global style embeddings with explicit pitch and energy features improves emotion transfer
      fidelity in expressive voice conversion beyond mel-spectrogram-only style encoding.
    confidence: high
    relevance: low
  - claim_id: cross_attention_fusion_of_local_f0_contours_with_content_embeddings
    role: supports
    claim: Cross-attention fusion of local F0 contours with content embeddings produces stronger prosodic alignment
      to the target than additive F0 injection in non-autoregressive voice conversion.
    source: §3.1, §4.3, Table 2
    evidence: Cross-attention fusion of local F0 contours with content embeddings produces stronger prosodic alignment
      to the target than additive F0 injection in non-autoregressive voice conversion.
    confidence: high
    relevance: low
  - claim_id: zero_shot_cross_lingual_voice_conversion_is_achievable_with_a
    role: supports
    claim: Zero-shot cross-lingual voice conversion is achievable with a monolingual training corpus when content
      representations are extracted from a multilingual speech model, though intelligibility degrades for unseen
      source languages.
    source: §4.4, Table 4
    evidence: Zero-shot cross-lingual voice conversion is achievable with a monolingual training corpus when content
      representations are extracted from a multilingual speech model, though intelligibility degrades for unseen
      source languages.
    confidence: high
    relevance: low
  limitations:
  - 'The proposed system incurs a substantial WER penalty relative to baselines: 7.98% vs. 5.01% (ESD) and 8.84%
    vs. 3.48% (LibriTTS) for the full model, with discrete units identified as the cause. This intelligibility regression
    is acknowledged but not resolved; future work is deferred.'
  - The evaluation uses a small subjective panel (15 listeners, 10–15 samples per model), limiting the statistical
    power of MOS comparisons. The cross-lingual results are restricted to English and German; how performance degrades
    for more distant language pairs is untested. The model is trained on English-only data, and German-to-English
    conversion shows a 30.84% WER, suggesting significant cross-lingual generalisation limits. The model size and
    computational cost are not reported, making it difficult to assess deployment feasibility.
  caveats: []
- id: interspeech-2025-0816
  published_date: "2025-08-17"
  entry_date: '2026-07-29'
  year: 2025
  venue: Interspeech
  task:
  - singing
  - VC
  architecture:
  - diffusion
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  - evaluation_caution
  current_role: active_evidence
  method_family:
  - diffusion_with_ssl_conditioning
  claims:
  - claim_id: converting_speech_timbre_to_singing_requires_cross_modal_speaker_embedding
    role: supports
    claim: Converting speech timbre to singing requires cross-modal speaker embedding alignment, and standard singer-identity
      conditioning generalises poorly across the speech-singing domain boundary.
    source: §1, §2.1
    evidence: Converting speech timbre to singing requires cross-modal speaker embedding alignment, and standard
      singer-identity conditioning generalises poorly across the speech-singing domain boundary.
    confidence: high
    relevance: medium
  - claim_id: cycle_training_strategies_that_simulate_paired_cross_domain_data_can
    role: supports
    claim: Cycle training strategies that simulate paired cross-domain data can compensate for the scarcity of matched
      speech-singing corpora in voice conversion training.
    source: §2.3
    evidence: Cycle training strategies that simulate paired cross-domain data can compensate for the scarcity of
      matched speech-singing corpora in voice conversion training.
    confidence: high
    relevance: low
  - claim_id: zero_shot_singing_voice_conversion_with_speech_prompts_achieves_lower
    role: supports
    claim: Zero-shot singing voice conversion with speech prompts achieves lower timbre similarity scores than same-domain
      (singing-to-singing) conversion, indicating that the cross-modal gap is not fully closed by embedding alignment
      alone.
    source: §3.3, Table 1, Table 2
    evidence: Zero-shot singing voice conversion with speech prompts achieves lower timbre similarity scores than
      same-domain (singing-to-singing) conversion, indicating that the cross-modal gap is not fully closed by embedding
      alignment alone.
    confidence: high
    relevance: low
  - claim_id: automated_speaker_similarity_metrics_capture_relative_improvements_from_cross_domain
    role: supports
    claim: Automated speaker similarity metrics capture relative improvements from cross-domain adaptation that
      are not clearly reflected in small-panel subjective timbre similarity ratings.
    source: §3.2, §3.3, Table 3
    evidence: Automated speaker similarity metrics capture relative improvements from cross-domain adaptation that
      are not clearly reflected in small-panel subjective timbre similarity ratings.
    confidence: high
    relevance: low
  limitations:
  - The subjective evaluation relies on only 10 volunteers, producing confidence intervals that overlap between
    all three systems on both MOS-n and MOS-ts. The claimed superiority of SSANSVC-stage2 over CoMoSVC in naturalness
    and similarity is not statistically robust at this sample size.
  - The model's loss function addresses only mel reconstruction; the authors note that timbre loss and lyrics recognition
    loss (to reduce CER) are absent and represent the primary direction for future improvement. The two-stage training
    procedure also introduces significant complexity and training cost compared to the CoMoSVC baseline. Evaluation
    is restricted to Mandarin speech and singing datasets, so generalisation to other languages and vocal styles
    is untested. The dependency on NUS-48E — one of few available paired speech-singing corpora — limits reproducibility
    in languages where such data does not exist.
  caveats: []
- id: interspeech-2025-0948
  published_date: "2025-08-17"
  entry_date: '2026-07-29'
  year: 2025
  venue: Interspeech
  task:
  - VC
  architecture:
  - VAE
  - diffusion
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  - evaluation_caution
  current_role: active_evidence
  method_family:
  - vae_and_quantized_ssl_representations
  - diffusion_with_ssl_conditioning
  claims:
  - claim_id: natural_language_prompts_enable_more_flexible_and_subjectively_accurate_emotion
    role: supports
    claim: Natural language prompts enable more flexible and subjectively accurate emotion control in voice conversion
      than numeric intensity values or reference audio selection.
    source: §1, §3.4
    evidence: Natural language prompts enable more flexible and subjectively accurate emotion control in voice conversion
      than numeric intensity values or reference audio selection.
    confidence: high
    relevance: low
  - claim_id: a_diffusion_based_mapping_from_text_embeddings_to_speech_emotion
    role: supports
    claim: A diffusion-based mapping from text embeddings to speech emotion embeddings is sufficient to replace
      reference audio at inference time without significant quality loss.
    source: §2.1, §3.2, Table 1
    evidence: A diffusion-based mapping from text embeddings to speech emotion embeddings is sufficient to replace
      reference audio at inference time without significant quality loss.
    confidence: high
    relevance: medium
  - claim_id: joint_training_of_a_text_to_emotion_mapper_with_reference
    role: supports
    claim: Joint training of a text-to-emotion mapper with reference emotion embeddings improves prosody naturalness
      over direct prediction from text alone.
    source: §3.3, Table 1
    evidence: Joint training of a text-to-emotion mapper with reference emotion embeddings improves prosody naturalness
      over direct prediction from text alone.
    confidence: high
    relevance: medium
  - claim_id: preserving_speaker_identity_during_emotional_pitch_manipulation_requires_an_explicit
    role: supports
    claim: Preserving speaker identity during emotional pitch manipulation requires an explicit F0 constraint in
      the speaker encoder; adversarial training alone is insufficient.
    source: §2.3, §3.3, Table 1
    evidence: Preserving speaker identity during emotional pitch manipulation requires an explicit F0 constraint
      in the speaker encoder; adversarial training alone is insufficient.
    confidence: high
    relevance: medium
  - claim_id: mixed_emotion_synthesis_remains_harder_to_control_than_single_category
    role: supports
    claim: Mixed-emotion synthesis remains harder to control than single-category emotion intensity across both
      subjective and objective metrics.
    source: §3.4, Table 2, Table 3
    evidence: Mixed-emotion synthesis remains harder to control than single-category emotion intensity across both
      subjective and objective metrics.
    confidence: high
    relevance: medium
  limitations:
  - Training and evaluation are conducted entirely on TextrolSpeech, a single corpus with a limited speaker set.
    Generalisation to out-of-domain speakers, languages, or acoustic conditions is untested, and all reported numbers
    should be interpreted within that constraint.
  - The evaluation uses only 25 listeners for subjective MOS across 132 utterances — a borderline sample size that
    may limit statistical reliability. The mixed-emotion accuracy (61.3%) is notably lower than single-attribute
    control, and the system's handling of complex emotional blends (e.g., contempt with happiness) is not analysed
    in depth. The discrete HuBERT token approach for linguistic content may introduce quantisation artefacts not
    reported in the paper. Future real-time or streaming deployment, mentioned in the conclusion as a direction,
    is not addressed in the current architecture.
  caveats: []
- id: interspeech-2025-0973
  published_date: "2025-08-17"
  entry_date: '2026-07-29'
  year: 2025
  venue: Interspeech
  task:
  - TTS
  - evaluation
  architecture: []
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  - evaluation_caution
  - infrastructure
  current_role: active_evidence
  method_family: []
  claims:
  - claim_id: automatic_mos_predictors_trained_primarily_on_english_data_underperform_on
    role: supports
    claim: Automatic MOS predictors trained primarily on English data underperform on Spanish TTS, and language-specific
      fine-tuning provides meaningful improvement even with small datasets.
    source: §4.2, Table 2
    evidence: Automatic MOS predictors trained primarily on English data underperform on Spanish TTS, and language-specific
      fine-tuning provides meaningful improvement even with small datasets.
    confidence: high
    relevance: medium
  - claim_id: low_level_local_acoustic_features_from_ssl_encoder_layers_are
    role: supports
    claim: Low-level local acoustic features from SSL encoder layers are more informative for naturalness prediction
      than higher-level contextual representations from deeper transformer blocks.
    source: §4.2, Figure 2
    evidence: Low-level local acoustic features from SSL encoder layers are more informative for naturalness prediction
      than higher-level contextual representations from deeper transformer blocks.
    confidence: high
    relevance: high
  - claim_id: the_scarcity_of_samples_rated_near_mos_4_0_in
    role: supports
    claim: The scarcity of samples rated near MOS 4.0 in evaluation datasets introduces a systematic prediction
      bias toward the dataset mean, suggesting label distribution matters as much as dataset size for MOS predictor
      quality.
    source: §4.2
    evidence: The scarcity of samples rated near MOS 4.0 in evaluation datasets introduces a systematic prediction
      bias toward the dataset mean, suggesting label distribution matters as much as dataset size for MOS predictor
      quality.
    confidence: high
    relevance: low
  - claim_id: lightweight_downstream_models_trained_on_frozen_ssl_representations_can_achieve
    role: supports
    claim: Lightweight downstream models trained on frozen SSL representations can achieve MOS prediction performance
      comparable to fine-tuned specialist models, despite having fewer than half the parameters.
    source: §4.2, Table 2
    evidence: Lightweight downstream models trained on frozen SSL representations can achieve MOS prediction performance
      comparable to fine-tuned specialist models, despite having fewer than half the parameters.
    confidence: high
    relevance: high
  limitations:
  - Most audio samples received only a single rating, making inter-rater reliability estimates noisy at the individual-item
    level. MOS scores near 4.0 are underrepresented, creating a systematic bias in model predictions toward the
    dataset mean. All models exhibit this bias toward predicting mean MOS, suggesting alternative loss functions
    (e.g., focal loss, distribution-matching losses) may help. The geographic distribution is skewed toward Argentine
    Spanish, limiting generalization to other Spanish dialects. The dataset does not include modern large-scale
    zero-shot or LLM-based TTS systems (VALL-E, Voicebox, etc.).
  caveats: []
- id: interspeech-2025-0998
  published_date: "2025-08-17"
  entry_date: '2026-07-29'
  year: 2025
  venue: Interspeech
  task:
  - VC
  architecture:
  - diffusion
  - GAN
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: active_evidence
  method_family:
  - diffusion_with_ssl_conditioning
  - gan_decoders_for_ssl_units
  claims:
  - claim_id: multi_stage_speech_restoration_pipelines_that_separate_noise_suppression_from
    role: supports
    claim: Multi-stage speech restoration pipelines that separate noise suppression from speaker-guided generation
      outperform single-stage generative approaches under severe degradation conditions.
    source: §4.3, Table 3
    evidence: GSR+VC substantially outperforms standalone GSR or standalone VC across all metrics on both VCTK-DEMAND
      and UNIVERSE; the gain is largest on the UNIVERSE set, which simulates more severe distortions including band-limiting,
      reverberation, codec artefacts, and packet drops.
    confidence: high
    relevance: medium
  - claim_id: self_supervised_discrete_speech_representations_provide_more_robust_content_features
    role: supports
    claim: Self-supervised discrete speech representations provide more robust content features for voice conversion
      than raw mel-spectrograms when the input speech is degraded.
    source: §4.3, Table 3
    evidence: VC (SSL) using HuBERT+VQ consistently outperforms VC (Mel) using direct mel-spectrogram input, with
      the gap widening on the more challenging UNIVERSE dataset where VC (Mel) shows significant quality degradation.
    confidence: high
    relevance: high
  - claim_id: diffusion_based_voice_conversion_models_cannot_reliably_handle_degraded_input
    role: complicates
    claim: Diffusion-based voice conversion models cannot reliably handle degraded input without a dedicated pre-processing
      stage, even when conditioned on clean speaker embeddings.
    source: §4.3, Table 3
    evidence: VC (SSL) in standalone mode achieves lower scores than GSR+VC on both evaluation sets; the VC module
      performs markedly worse when the input is noisy without the GSR front-end, confirming that speaker-embedding
      guidance alone does not compensate for noisy content features.
    confidence: high
    relevance: low
  - claim_id: enrollment_dependent_speaker_guidance_for_speech_restoration_limits_applicability_to
    role: complicates
    claim: Enrollment-dependent speaker guidance for speech restoration limits applicability to settings where clean
      reference speech from the same speaker is available in advance.
    source: §3, §4.1
    evidence: The system assumes short, uncorrelated segments of clean speech are obtained beforehand for speaker
      embedding extraction; the paper does not evaluate performance when such enrollment audio is unavailable or
      mismatched.
    confidence: high
    relevance: medium
  limitations:
  - The system requires a clean enrollment utterance from the target speaker, which may not be available in all
    real-world scenarios. No comparison with Miipher could be performed due to unavailability of that model, leaving
    the relationship between this approach and the closest prior work unquantified. Evaluation uses only automated
    perceptual quality metrics (NISQA, UTMOS, WV-MOS, DNSMOS) without any human listening tests, so it is unclear
    whether the metric gains translate to perceived quality improvements. Both the GSR and VC models are trained
    on separate datasets, and the system has not been evaluated on the joint training configuration proposed as
    future work. Performance on languages other than English is not examined.
  caveats: []
- id: interspeech-2025-1101
  published_date: "2025-08-17"
  entry_date: '2026-07-29'
  year: 2025
  venue: Interspeech
  task:
  - VC
  architecture:
  - diffusion
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  - evaluation_caution
  current_role: active_evidence
  method_family:
  - diffusion_with_ssl_conditioning
  claims:
  - claim_id: diffusion_based_voice_conversion_systems_can_achieve_strong_emotion_controllability
    role: supports
    claim: Diffusion-based voice conversion systems can achieve strong emotion controllability in zero-shot settings
      when combined with mutual-information disentanglement and inference-time guidance.
    source: §3.2, §3.3, Table 2
    evidence: Diffusion-based voice conversion systems can achieve strong emotion controllability in zero-shot settings
      when combined with mutual-information disentanglement and inference-time guidance.
    confidence: high
    relevance: low
  - claim_id: disentangling_speaker_identity_and_emotion_via_mutual_information_minimisation_improves
    role: supports
    claim: Disentangling speaker identity and emotion via mutual information minimisation improves emotion controllability
      in voice conversion without requiring parallel or speaker-specific training data.
    source: §2.1.4, §3.3, Table 2
    evidence: Disentangling speaker identity and emotion via mutual information minimisation improves emotion controllability
      in voice conversion without requiring parallel or speaker-specific training data.
    confidence: high
    relevance: low
  - claim_id: in_emotional_voice_conversion_autoencoder_based_methods_tend_to_achieve
    role: complicates
    claim: In emotional voice conversion, autoencoder-based methods tend to achieve higher emotion accuracy than
      GAN-based methods, but at the cost of substantially lower naturalness and higher speech distortion.
    source: §3.2, Table 1
    evidence: In emotional voice conversion, autoencoder-based methods tend to achieve higher emotion accuracy than
      GAN-based methods, but at the cost of substantially lower naturalness and higher speech distortion.
    confidence: high
    relevance: low
  - claim_id: classifier_free_style_guidance_applied_to_emotion_representations_at_inference
    role: supports
    claim: Classifier-free-style guidance applied to emotion representations at inference time provides a direct
      lever for trading naturalness against emotion controllability in diffusion-based EVC.
    source: §2.1.3, §3.3, Table 2
    evidence: Classifier-free-style guidance applied to emotion representations at inference time provides a direct
      lever for trading naturalness against emotion controllability in diffusion-based EVC.
    confidence: high
    relevance: medium
  - claim_id: training_on_large_scale_in_the_wild_emotional_corpora_enables
    role: supports
    claim: Training on large-scale in-the-wild emotional corpora enables zero-shot generalisation to speakers absent
      from training, even when evaluation is conducted on acted-speech datasets with different recording conditions.
    source: §3.4, §4
    evidence: Training on large-scale in-the-wild emotional corpora enables zero-shot generalisation to speakers
      absent from training, even when evaluation is conducted on acted-speech datasets with different recording
      conditions.
    confidence: high
    relevance: low
  limitations:
  - 'The comparison between ZSDEVC and EMOCONV-DIFF in Table 1 is not fully fair: EMOCONV-DIFF is evaluated in a
    seen-speaker scenario while ZSDEVC operates zero-shot. The naturalness gap may reflect this experimental asymmetry
    rather than a fundamental quality deficit.'
  - The model does not address intensity control within a target emotion category — prior work (Emovox) provides
    per-dimension arousal/valence control that ZSDEVC does not directly expose during inference. Evaluation covers
    only five emotion categories (angry, happy, sad, neutral, surprise) and excludes neutral-to-emotional conversion.
    Real-time or streaming use cases are not addressed. Results on out-of-domain acted speech (ESD) and in-the-wild
    speech (MSP-Podcast) show consistent trends, but the system's behaviour on highly expressive or non-English
    speech is untested.
  caveats: []
- id: interspeech-2025-1106
  published_date: "2025-08-17"
  entry_date: '2026-07-29'
  year: 2025
  venue: Interspeech
  task:
  - codec
  architecture:
  - VAE
  - GAN
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: active_evidence
  method_family:
  - vae_and_quantized_ssl_representations
  - gan_decoders_for_ssl_units
  claims:
  - claim_id: explicit_speaker_perturbation_during_codec_training_is_more_effective_for
    role: supports
    claim: Explicit speaker perturbation during codec training is more effective for speaker disentanglement than
      relying on implicit information bottleneck alone.
    source: §3.3, Table 2; §3.4
    evidence: LSCodec achieves higher target-speaker SECS (0.852 at 50Hz) and lower speaker probing accuracy than
      TiCodec 1VQ (SECS 0.714), which uses an implicit VQ bottleneck without any perturbation, even though LSCodec
      operates at lower bitrate (0.45 kbps vs 0.75 kbps).
    confidence: high
    relevance: low
  - claim_id: single_codebook_discrete_speech_codecs_can_match_or_exceed_multi
    role: supports
    claim: Single-codebook discrete speech codecs can match or exceed multi-codebook acoustic codec reconstruction
      quality at ultra-low bitrates when content and speaker information are explicitly decoupled.
    source: §3.2, Table 1
    evidence: LSCodec-50Hz (V=300, 0.45 kbps) achieves WER 3.33% and MOS 4.49 on LibriTTS test-clean, outperforming
      all single-codebook baselines including WavTokenizer-small (WER 7.86, MOS 4.14 at 0.48 kbps) and multi-codebook
      SemantiCodec (WER 4.16 at 0.63 kbps).
    confidence: high
    relevance: low
  - claim_id: reducing_codec_frame_rate_through_temporal_downsampling_degrades_content_intelligibility
    role: complicates
    claim: Reducing codec frame rate through temporal downsampling degrades content intelligibility without proportional
      improvement in speaker disentanglement.
    source: §3.2, Table 1; §3.3, Table 2
    evidence: Halving the frame rate from 50Hz to 25Hz reduces bitrate from 0.45 to 0.25 kbps but increases reconstruction
      WER from 3.33% to 5.46% and VC WER from 4.04% to 6.32%, while reconstruction SECS changes only marginally
      (0.954 to 0.945).
    confidence: high
    relevance: low
  - claim_id: an_auxiliary_ssl_token_prediction_objective_is_necessary_for_maintaining
    role: refines
    claim: An auxiliary SSL token prediction objective is necessary for maintaining content intelligibility when
      an information bottleneck is used to remove speaker timbre.
    source: §3.5, Table 3
    evidence: Ablating the SSL token prediction loss increases VAE-stage WER from 4.96% to 11.22% while SECS remains
      essentially unchanged (0.811 to 0.811), confirming that the SSL prediction task guides content encoding independently
      of the speaker removal objective.
    confidence: high
    relevance: high
  - claim_id: multi_stage_codec_training_establishing_a_continuous_disentangled_space_before
    role: supports
    claim: Multi-stage codec training, establishing a continuous disentangled space before quantization, improves
      both content preservation and speaker disentanglement compared to direct VQ training.
    source: §3.5, Table 3
    evidence: Skipping stage 1 (VAE pre-training) and training VQ-VAE directly degrades WER from 3.39% to 3.84%
      and SECS from 0.817 to 0.800 in the VQ-VAE stage, confirming that continuous-space initialization benefits
      discrete representation quality.
    confidence: high
    relevance: low
  limitations:
  - The model is evaluated exclusively on English LibriTTS data, leaving multilingual and cross-lingual generalization
    untested. The vocoder (CTX-vec2wav alpha) is trained on a fixed 24 kHz corpus, so quality at other sampling
    rates or in noisy conditions is unclear. Speaker probing uses a single X-vector classifier on LibriTTS speakers,
    which may not detect all forms of residual speaker information. The paper notes that stronger perturbation methods,
    better content preservation at 25Hz, and scaling to larger data are open directions. No code or pre-trained
    models are released.
  caveats: []
- id: interspeech-2025-1210
  published_date: "2025-08-17"
  entry_date: '2026-07-29'
  year: 2025
  venue: Interspeech
  task:
  - VC
  architecture:
  - diffusion
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: active_evidence
  method_family:
  - diffusion_with_ssl_conditioning
  claims:
  - claim_id: dual_granularity_emotion_feature_extraction_combining_utterance_level_and_frame
    role: supports
    claim: Dual-granularity emotion feature extraction (combining utterance-level and frame-level representations)
      improves emotion discriminability in voice conversion compared to single-scale approaches.
    source: §2.1.3, §3.3.3, Table 3
    evidence: DiffEmotionVC's dual-granularity emotion encoder achieves 80% ECA and 0.78 Pearson Corr on the ESD
      dataset; ablation confirms removing the emotion encoder is the most damaging intervention, dropping Corr to
      0.38.
    confidence: high
    relevance: low
  - claim_id: orthogonality_constraints_on_emotion_speaker_and_content_feature_spaces_provide
    role: supports
    claim: Orthogonality constraints on emotion, speaker, and content feature spaces provide a stable and effective
      disentanglement mechanism for emotional voice conversion.
    source: §2.2.2, §3.3.3, Table 3
    evidence: Removing orthogonal loss reduces SECS from 0.73 to 0.70 and Corr from 0.78 to 0.72; the paper explicitly
      motivates orthogonal loss as a remedy for the training instability of the mutual information loss used in
      prior work.
    confidence: high
    relevance: low
  - claim_id: diffusion_based_evc_systems_achieve_strong_overall_emotion_accuracy_but
    role: complicates
    claim: Diffusion-based EVC systems achieve strong overall emotion accuracy but struggle to discriminate between
      high-arousal emotions sharing similar arousal-valence profiles.
    source: §3.3.1
    evidence: DiffEmotionVC reaches 80% ECA overall but the paper notes difficulty distinguishing happy, surprised,
      and angry, attributing this to insufficient emotional diversity in the ESD training data rather than a fundamental
      model limitation.
    confidence: high
    relevance: medium
  - claim_id: discretisation_of_continuous_speech_representations_degrades_emotion_voice_conversion_by
    role: complicates
    claim: Discretisation of continuous speech representations degrades emotion voice conversion by introducing
      content-emotion feature leakage.
    source: §3.3.2, Table 2
    evidence: Replacing continuous ContentVec with VQ-ContentVec drops UTMOS from 4.04 to 2.54 and Corr from 0.78
      to 0.50; SpeechTokenizer RVQ1 discrete features produce the worst performance (UTMOS 1.79), demonstrating
      that discrete tokens cause timbre and emotion entanglement.
    confidence: high
    relevance: low
  - claim_id: cross_attention_fusion_outperforms_additive_fusion_for_integrating_heterogeneous_speech
    role: supports
    claim: Cross-attention fusion outperforms additive fusion for integrating heterogeneous speech features in voice
      conversion systems.
    source: §3.3.3, Table 3
    evidence: Ablation replacing gated cross-attention with simple additive fusion reduces UTMOS from 4.04 to 3.26,
      a 19% degradation in predicted audio quality.
    confidence: high
    relevance: low
  limitations:
  - Evaluation is limited to the ESD dataset (five emotions, primarily Mandarin Chinese; the ablation table specifically
    targets the zh-Angry subset), restricting generalisability to other languages and broader emotion categories.
    The model size is unreported, making deployment trade-off analysis impossible. Distinguishing between high-arousal
    emotions (happy, surprised, angry) remains unresolved; the paper identifies more naturalistic emotional data
    as the likely remedy but leaves this to future work. No code is released, and there is no cross-lingual evaluation.
  caveats: []
- id: interspeech-2025-1229
  published_date: "2025-08-17"
  entry_date: '2026-07-29'
  year: 2025
  venue: Interspeech
  task:
  - VC
  architecture:
  - flow-matching
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  - evaluation_caution
  current_role: active_evidence
  method_family:
  - flow_matching_with_ssl_conditioning
  claims:
  - claim_id: in_context_learning_with_flow_matching_can_enable_voice_conversion
    role: supports
    claim: In-context learning with flow matching can enable voice conversion systems to simultaneously transform
      speaker timbre and preserve background sounds without an explicit separation step.
    source: §4.2, Table 2
    evidence: E2E-BPVC achieves BS-MOS 4.60 and SS-MOS 4.02 comparable to Denoise-VC II (BS-MOS 4.65, SS-MOS 4.07)
      using a single model without a denoising module, validated by 12 human raters on LibriTTS test-clean with
      synthetic noise and music backgrounds at 7 and 12 dB SNR.
    confidence: high
    relevance: low
  - claim_id: standard_voice_conversion_evaluation_metrics_speaker_similarity_character_error_rate
    role: complicates
    claim: Standard voice conversion evaluation metrics (speaker similarity, character error rate, speech quality)
      are insufficient for assessing systems that operate on speech with background sounds.
    source: §4.1, Table 1
    evidence: ECAPA-TDNN speaker similarity and ASR-based CER are degraded by background sound in both source and
      converted audio, causing clean-output systems to appear relatively stronger on objective metrics despite failing
      entirely on background preservation (ICL-VC BS-MOS 0.70). The authors explicitly note that objective metrics
      "do not adequately reflect the capabilities" of background-preserving systems.
    confidence: high
    relevance: low
  - claim_id: noise_robust_self_supervised_speech_representations_improve_content_disentanglement_in
    role: supports
    claim: Noise-robust self-supervised speech representations improve content disentanglement in voice conversion
      systems trained on speech with background sounds.
    source: §4.3, Table 3
    evidence: Replacing HuBERT with WavLM as the semantic token backbone reduces CER from 10.27 to 7.99 under noisy
      evaluation conditions, and training k-means on noisy speech (rather than clean speech) further reduces CER
      to 7.99 versus 9.22, demonstrating that noise robustness at the representation level propagates to improved
      content disentanglement.
    confidence: high
    relevance: high
  - claim_id: designing_a_single_voice_conversion_model_to_handle_background_preservation
    role: complicates
    claim: Designing a single voice conversion model to handle background preservation introduces a competing objective
      that slightly degrades clean-speech conversion quality relative to a clean-speech-only system.
    source: §4.1, Table 1
    evidence: E2E-BPVC achieves SIM 0.849 and CER 3.18 on clean-speech conversion, modestly below ICL-VC (SIM 0.873,
      CER 2.37), reflecting a trade-off introduced by the shared background-preservation training objective.
    confidence: high
    relevance: low
  limitations:
  - The evaluation relies exclusively on LibriTTS with synthetically added noise and music at controlled SNRs (7
    dB and 12 dB), without testing on real-world recordings where backgrounds are natural, non-stationary, or below
    7 dB SNR. The comparison set is limited to two baselines sharing the same ICL-VC foundation, so performance
    relative to independent VC architectures is unknown. The paper also acknowledges that the current experiments
    use small-scale training data, and that scaling up is needed to improve robustness. The model offers no controllability
    over whether to preserve or suppress background sounds, though the authors identify this as future work.
  caveats: []
- id: interspeech-2025-1236
  published_date: "2025-08-17"
  entry_date: '2026-07-29'
  year: 2025
  venue: Interspeech
  task:
  - TTS
  architecture:
  - flow-matching
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  - evaluation_caution
  current_role: active_evidence
  method_family:
  - flow_matching_with_ssl_conditioning
  claims:
  - claim_id: auxiliary_alignment_losses_applied_at_intermediate_transformer_layers_can_substantially
    role: supports
    claim: Auxiliary alignment losses applied at intermediate transformer layers can substantially accelerate convergence
      in flow-matching TTS training without modifying the inference pipeline.
    source: §2, §4.1, Table 1
    evidence: A-DMA reduces WER from 2.68% to 1.97% and improves SIM from 0.60 to 0.62 for F5-TTS on LibriSpeech-PC
      test-clean in a low-resource (0.6kh) setting, with the convergence curve showing that A-DMA matches the baseline's
      terminal WER in roughly half the training steps.
    confidence: high
    relevance: medium
  - claim_id: in_transformer_based_generative_models_text_semantic_and_speaker_acoustic
    role: refines
    claim: In transformer-based generative models, text-semantic and speaker-acoustic alignment supervision are
      best applied at different network depths rather than at the same layer.
    source: §4.2, Table 2
    evidence: Layer-wise ablation on F5-TTS Small (18 DiT layers) shows that applying CTC text alignment at layer
      8 and HuBERT speech alignment at layer 12 (WER 2.226%) outperforms both same-layer dual alignment at layer
      8 (WER 3.063%) and same-layer alignment at layer 12 (WER 2.688%), supporting a modality-specific depth hypothesis.
    confidence: high
    relevance: medium
  - claim_id: frozen_self_supervised_speech_encoders_such_as_hubert_are_effective
    role: supports
    claim: Frozen self-supervised speech encoders such as HuBERT are effective perceptual teachers for improving
      speaker similarity in flow-matching TTS through cosine-similarity distillation.
    source: §4.3, Table 3
    evidence: Negative cosine similarity to HuBERT-large last-layer features at the 12th transformer layer improves
      SIM from 0.578 to 0.609 and WER from 7.474% to 3.521%; WavLM achieves similar gains (SIM 0.612, WER 3.992%).
      L1 distance and BCE-sigmoid variants are less effective.
    confidence: high
    relevance: high
  - claim_id: training_time_alignment_acceleration_for_diffusion_tts_primarily_improves_text
    role: complicates
    claim: Training-time alignment acceleration for diffusion TTS primarily improves text intelligibility and speaker
      similarity rather than overall speech naturalness, as measured by automatic MOS prediction.
    source: §4.2, Table 2, §6
    evidence: UTMOS scores across all A-DMA ablation conditions (Table 2) range from 4.01 to 4.12, close to the
      4.04 baseline, while WER and SIM show large improvements. The paper reports no human listening test, leaving
      perceptual naturalness gains unconfirmed.
    confidence: high
    relevance: low
  limitations:
  - The evaluation covers only automated metrics (WER via Whisper-large-V3, SIM via WavLM speaker verification,
    UTMOS as an automatic MOS proxy); no human listening evaluation is reported. All A-DMA experiments use the low-resource
    0.6kh training regime; whether doubled convergence speed holds at the 100kh scale used by high-resource baselines
    is untested. The paper also does not address model efficiency at inference, explicitly leaving inference acceleration
    as future work.
  caveats: []
- id: interspeech-2025-1394
  published_date: "2025-08-17"
  entry_date: '2026-07-29'
  year: 2025
  venue: Interspeech
  task:
  - TTS
  architecture:
  - transformer-enc-dec
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: active_evidence
  method_family:
  - transformer_ssl_encoders_and_adapters
  claims:
  - claim_id: self_supervised_distillation_with_emotion_specific_inductive_biases_can_learn
    role: supports
    claim: Self-supervised distillation with emotion-specific inductive biases can learn speaker-independent emotion
      embeddings more effectively than GRL-based or VQ-based disentanglement.
    source: §3.4, Table 1
    evidence: DiEmo-TTS surpasses Trans-GRL, Trans-VQ, and Trans-Ort on eMOS across all four emotion categories
      in subjective evaluation, without requiring explicit speaker labels during emotion encoder training.
    confidence: high
    relevance: high
  - claim_id: formant_based_speaker_perturbation_is_more_effective_for_disentangling_speaker
    role: supports
    claim: Formant-based speaker perturbation is more effective for disentangling speaker identity from emotion
      than non-targeted noise augmentation in cross-speaker emotion transfer.
    source: §3.6, Table 2
    evidence: Replacing formant perturbation with MUSAN/RIR noise augmentation in the ablation increases WER and
      degrades SECS, whereas formant perturbation exploits the timbre-formant correlation to distort identity while
      preserving emotional expression.
    confidence: high
    relevance: medium
  - claim_id: multi_factor_conditioning_via_shared_attention_mechanisms_produces_better_balance
    role: supports
    claim: Multi-factor conditioning via shared attention mechanisms produces better balance between speaker fidelity
      and emotional expressiveness than concatenation conditioning in transformer-based TTS.
    source: §3.6, Table 2
    evidence: Adding the DCT block improves eMOS from 3.89 to 4.07 while keeping sMOS comparable in the ablation;
      the block applies weight-sharing multi-head attention per style signal and fuses outputs via MLP.
    confidence: high
    relevance: medium
  - claim_id: achieving_speaker_independent_emotion_representations_via_self_supervised_methods_still
    role: complicates
    claim: Achieving speaker-independent emotion representations via self-supervised methods still depends on labeled
      auxiliary data for emotion space definition.
    source: §4, §3.2
    evidence: The emotion clustering step requires a WavLM-based emotional attribute predictor fine-tuned on the
      MSP-Podcast labeled corpus; the authors identify dependence on this predictor as a limitation for generalising
      to unseen speakers.
    confidence: high
    relevance: high
  limitations:
  - Evaluation uses only 2 target speakers (one male, one female from ESD), both trained exclusively on neutral
    utterances. Generalisation to speakers trained on mixed emotional data or to out-of-domain voices is not assessed.
  - WER of 16.16% indicates non-trivial intelligibility degradation compared to clean TTS; the source of this degradation
    (emotion conditioning, model architecture, or dataset characteristics) is not explicitly discussed. The emotional
    attribute predictor is fine-tuned on MSP-Podcast, a separate labeled corpus, introducing a labeled-data dependency
    that the authors acknowledge as a target for future unsupervised replacement. Scalability to languages beyond
    English and to larger multi-emotion multi-speaker datasets remains undemonstrated.
  caveats: []
- id: interspeech-2025-1397
  published_date: "2025-08-17"
  entry_date: '2026-07-29'
  year: 2025
  venue: Interspeech
  task:
  - singing
  - VC
  architecture:
  - diffusion
  relevance: medium
  evidence_role:
  - architecture_variant
  current_role: active_evidence
  method_family:
  - diffusion_with_ssl_conditioning
  claims:
  - claim_id: explicit_frequency_domain_decomposition_of_the_f0_contour_enables_more
    role: supports
    claim: Explicit frequency-domain decomposition of the F0 contour enables more accurate and controllable singing
      style transfer than implicit style-embedding approaches.
    source: §4.1.1, Table 1
    evidence: VibE-SVC achieves 0.700 style accuracy in style-only conversion versus 0.213-0.525 for SoVITS baselines
      using direct style embeddings, on the VocalSet straight/vibrato benchmark.
    confidence: high
    relevance: medium
  - claim_id: adversarial_training_on_the_target_frequency_band_of_the_f0
    role: supports
    claim: Adversarial training on the target frequency band of the F0 contour improves singing style accuracy without
      degrading naturalness.
    source: §4.3, Table 1
    evidence: Removing the multi-period discriminator reduces style accuracy from 0.700 to 0.625 and MOS from 4.124
      to 4.016 in the style-only conversion experiment.
    confidence: high
    relevance: medium
  - claim_id: increasing_style_transfer_accuracy_in_singing_voice_conversion_trades_off
    role: complicates
    claim: Increasing style transfer accuracy in singing voice conversion trades off against naturalness, and explicit
      disentanglement does not fully eliminate this tension.
    source: §4.1.2, Figure 3
    evidence: Figure 3 shows a consistent inverse correlation between MOS and style accuracy across all baselines
      and VibE-SVC; the highest-accuracy model (VibE-SVC) has lower naturalness than the highest-naturalness baseline
      (SoVITS with style embedding, 0.213 style accuracy).
    confidence: high
    relevance: low
  - claim_id: the_effective_granularity_of_f0_based_singing_style_disentanglement_via
    role: refines
    claim: The effective granularity of F0-based singing style disentanglement via DWT is sensitive to decomposition
      level, with an optimal level that captures vibrato without including unrelated high-frequency content.
    source: §4.3, Table 3
    evidence: Style accuracy is 0.163 at DWT level 3 (vibrato information absent), 0.694 at level 4 (optimal), and
      drops slightly at level 5 due to inclusion of irrelevant high-frequency components.
    confidence: high
    relevance: medium
  - claim_id: vibrato_extent_in_singing_voice_conversion_can_be_controlled_continuously
    role: supports
    claim: Vibrato extent in singing voice conversion can be controlled continuously at inference time by scalar
      multiplication of the isolated high-frequency F0 component, without retraining.
    source: §4.2, Table 2, Figure 5
    evidence: Scaling the high-frequency F0 contour from 0.1 to 2.0 produces style accuracy ranging from 0.066 to
      0.928; frame-level control is also demonstrated by applying scaling at specific target frame indices.
    confidence: high
    relevance: low
  limitations:
  - The model is restricted to two singing styles (straight and vibrato) derived from VocalSet, and does not address
    other common techniques such as falsetto, breathy voice, belt, or melisma. Generalization to other datasets
    and languages is untested. The DWT wavelet type (db1) and decomposition level (4) are dataset-specific hyperparameters
    that may need re-tuning for other corpora or style targets. Human evaluation uses at minimum 20 Amazon MTurk
    raters per model, which is on the low end for resolving the small MOS differences reported. The speaker similarity
    scores in the timbre and style conversion setting (SMOS 3.196) are noticeably lower than ground truth (3.549),
    suggesting that joint timbre and style conversion remains a meaningful challenge.
  caveats: []
- id: interspeech-2025-1434
  published_date: "2025-08-17"
  entry_date: '2026-07-29'
  year: 2025
  venue: Interspeech
  task:
  - VC
  architecture:
  - diffusion
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: active_evidence
  method_family:
  - diffusion_with_ssl_conditioning
  claims:
  - claim_id: full_utterance_time_reversal_can_serve_as_an_effective_signal
    role: supports
    claim: Full-utterance time reversal can serve as an effective signal-level data augmentation for speaker representation
      learning in voice conversion, as it suppresses phonemic content while retaining speaker-discriminative tonal
      features.
    source: §3.1, Table 1
    evidence: A perceptual study shows 80.3% speaker identification accuracy from time-reversed speech; Table 1
      confirms complete reversal achieves 100% WER (full linguistic removal) alongside the highest cosine speaker
      similarity score (0.96), higher than any short-time reversal window.
    confidence: high
    relevance: low
  - claim_id: fusing_speaker_embeddings_from_augmented_training_signals_with_conventional_embeddings
    role: supports
    claim: Fusing speaker embeddings from augmented training signals with conventional embeddings improves speaker
      similarity in zero-shot diffusion-based voice conversion.
    source: §4.3, Table 2
    evidence: Adding reversed-speech speaker embeddings via a weighted fusion layer (α = β = 0.5) improves objective
      speaker similarity by 4.16% on average across DiffHierVC and DDDM-VC; DDDM-VC objective SPK-SIM rises from
      0.70 to 0.79 and subjective MUSHRA from 77.46 to 78.61.
    confidence: high
    relevance: low
  - claim_id: the_effectiveness_of_speaker_embedding_augmentation_in_voice_conversion_varies
    role: complicates
    claim: The effectiveness of speaker embedding augmentation in voice conversion varies substantially across backbone
      architectures, complicating claims of generalisability.
    source: §4.3, Table 2
    evidence: For DiffVC, the augmentation improves subjective speaker similarity (50.12 to 53.42) but reduces objective
      similarity (0.75 to 0.71), while DiffHierVC shows objective improvement but negligible subjective change;
      only DDDM-VC shows consistent gains on both metrics.
    confidence: high
    relevance: low
  - claim_id: improving_speaker_disentanglement_through_augmentation_does_not_necessarily_trade_off
    role: supports
    claim: Improving speaker disentanglement through augmentation does not necessarily trade off against generated
      speech quality in diffusion-based voice conversion.
    source: §4.3, Table 2
    evidence: DDDM-VC+Ours improves both WV-MOS (3.84 to 3.91) and UTMOS (3.21 to 3.55) alongside speaker similarity
      gains, indicating that stronger speaker conditioning from the STR augmentation does not degrade synthesis
      quality.
    confidence: high
    relevance: low
  limitations:
  - 'Objective and subjective speaker similarity disagree for DiffVC: the augmentation reduces objective similarity
    (0.75 to 0.71) while improving subjective similarity (50.12 to 53.42). This discrepancy limits confidence in
    the metric-level generalisation claim across all diffusion backbones. *(§4.3, Table 2)*'
  - 'The perceptual study supporting the STR principle is small: 25 participants and 6 speakers, all in English.
    Whether the tonal-pattern preservation property holds equally for tonal languages (Mandarin, Thai) or heavily
    inflected languages is untested. The approach has been evaluated only on diffusion-based VC systems; compatibility
    with flow-matching and codec-based VC architectures remains unexplored. No code is publicly released, limiting
    reproducibility. The weighted fusion coefficients (α and β) are set empirically to 0.5; the paper does not explore
    learned dynamic weighting conditioned on the input utterance.'
  caveats: []
- id: interspeech-2025-1440
  published_date: "2025-08-17"
  entry_date: '2026-07-29'
  year: 2025
  venue: Interspeech
  task:
  - codec
  architecture:
  - VAE
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  - evaluation_caution
  current_role: active_evidence
  method_family:
  - vae_and_quantized_ssl_representations
  claims:
  - claim_id: self_supervised_disentanglement_of_speech_into_content_speaker_and_prosody
    role: supports
    claim: Self-supervised disentanglement of speech into content, speaker, and prosody streams can match or exceed
      supervised codec quality at significantly lower bitrate.
    source: §4.1, Table 1, Table 2
    evidence: Self-supervised disentanglement of speech into content, speaker, and prosody streams can match or
      exceed supervised codec quality at significantly lower bitrate.
    confidence: high
    relevance: high
  - claim_id: codec_coding_efficiency_is_more_sensitive_to_information_factorisation_than
    role: supports
    claim: Codec coding efficiency is more sensitive to information factorisation than to raw model capacity or
      bitrate allocation.
    source: §4.1, Table 1
    evidence: Codec coding efficiency is more sensitive to information factorisation than to raw model capacity
      or bitrate allocation.
    confidence: high
    relevance: low
  - claim_id: routing_wavlm_supervision_to_the_decoder_rather_than_the_encoder
    role: supports
    claim: Routing WavLM supervision to the decoder rather than the encoder during training improves speaker-content
      disentanglement for voice conversion.
    source: §2.5, §4.2
    evidence: Routing WavLM supervision to the decoder rather than the encoder during training improves speaker-content
      disentanglement for voice conversion.
    confidence: high
    relevance: high
  - claim_id: ultra_low_bitrate_codecs_below_0_5_kbps_can_achieve
    role: supports
    claim: Ultra-low-bitrate codecs (below 0.5 kbps) can achieve MUSHRA scores competitive with codecs operating
      at 2–3 kbps when disentanglement is applied to reduce frame-level redundancy.
    source: §4.1, Table 2
    evidence: Ultra-low-bitrate codecs (below 0.5 kbps) can achieve MUSHRA scores competitive with codecs operating
      at 2–3 kbps when disentanglement is applied to reduce frame-level redundancy.
    confidence: high
    relevance: medium
  limitations:
  - The demo and code availability are not confirmed in the paper or metadata. Reproducibility relies on external
    checkpoints for baselines — FACodec and SpeechTokenizer results are inferred from official checkpoints under
    potentially different conditions than the re-trained TiCodec and DAC baselines.
  - Evaluation is restricted to English (LibriSpeech and VCTK). Generalisation to other languages, accents, or spontaneous-speech
    domains is untested. The prosody encoder's low-mel-bin design is validated empirically via t-SNE visualisation
    but without a formal mutual information analysis. It is unclear how much prosody actually remains once the speaker
    and content encoders are also active during decoding — partial speaker clustering in Fig. 2 suggests the separation
    is not complete.
  caveats: []
- id: interspeech-2025-1478
  published_date: "2025-08-17"
  entry_date: '2026-07-29'
  year: 2025
  venue: Interspeech
  task:
  - TTS
  architecture:
  - hybrid
  relevance: medium
  evidence_role:
  - architecture_variant
  current_role: active_evidence
  method_family:
  - hybrid_semantic_acoustic_ssl_systems
  claims:
  - claim_id: replacing_resnet_18_av_hubert_visual_encoders_with_mobile_video
    role: supports
    claim: Replacing ResNet-18/AV-HuBERT visual encoders with mobile video networks reduces lip-to-speech system
      complexity by an order of magnitude at the cost of intelligibility degradation.
    source: §4.1, Table 1
    evidence: LightL2S uses MoViNet-A0 instead of ResNet-18-based AV-HuBERT, reducing inference from 32–34 GMacs
      to 0.8 GMacs on LRS3 while WER increases from 27–30% (AV-HuBERT-based methods) to 64.8%.
    confidence: high
    relevance: high
  - claim_id: multi_resolution_spectrogram_discriminators_substantially_improve_speech_naturalness_in_ddsp
    role: supports
    claim: Multi-resolution spectrogram discriminators substantially improve speech naturalness in DDSP-based lip-to-speech
      synthesis, beyond spectral regression alone.
    source: §4.3, Table 3
    evidence: Removing the adversarial GAN loss from LightL2S collapses UTMOS from 2.93 to 1.58 and SECS from 0.72
      to 0.63 on LRS3, while computational cost remains identical at 0.8 GMacs.
    confidence: high
    relevance: medium
  - claim_id: objective_quality_metrics_can_diverge_from_human_perceptual_judgments_in
    role: complicates
    claim: Objective quality metrics can diverge from human perceptual judgments in lip-to-speech synthesis, making
      WER and UTMOS insufficient as sole quality signals.
    source: §4.2, Table 2
    evidence: NaturalL2S exceeds Ground Truth on UTMOS (3.66 vs. 3.59) but scores lower in subjective naturalness
      MOS (4.10 vs. 4.51); LightL2S achieves the highest speaker similarity MOS (3.72) despite being outranked on
      SECS by several baselines.
    confidence: high
    relevance: medium
  - claim_id: efficient_transformer_variants_zipformer_can_substitute_standard_conformer_backbones_in
    role: refines
    claim: Efficient transformer variants (Zipformer) can substitute standard Conformer backbones in visual speech
      modelling with simultaneous improvements in quality and computational cost.
    source: §4.3, Table 3
    evidence: Replacing Zipformer with a Conformer backbone in LightL2S increases GMacs from 0.80 to 1.09 while
      reducing UTMOS from 2.93 to 2.74 and raising WER from 64.8% to 66.3% on LRS3.
    confidence: high
    relevance: medium
  limitations:
  - 'Intelligibility remains a major limitation: LightL2S achieves WER 64.8% on LRS3, more than double the 27–30%
    of AV-HuBERT-based methods. This gap is attributed to the absence of large-scale pre-trained visual-acoustic
    features and the lack of text supervision, and is explicitly left as an open problem. At this WER level, practical
    deployment for communication assistance — the paper''s stated motivation — requires further work.'
  - Evaluation is confined to English TED/TEDx video from LRS3; generalisation to other languages, accents, spontaneous
    speech, and noisy real-world conditions is untested. The system uses a reference speaker embedding extracted
    at inference time, meaning an enrolment clip must be available, which may be impractical for truly in-the-wild
    edge scenarios. Parameter count is not reported, making direct comparison with published efficient TTS systems
    difficult.
  caveats: []
- id: interspeech-2025-1531
  published_date: "2025-08-17"
  entry_date: '2026-07-29'
  year: 2025
  venue: Interspeech
  task:
  - singing
  - VC
  architecture:
  - VAE
  - GAN
  relevance: medium
  evidence_role:
  - architecture_variant
  current_role: active_evidence
  method_family:
  - vae_and_quantized_ssl_representations
  - gan_decoders_for_ssl_units
  claims:
  - claim_id: in_singing_voice_conversion_reducing_the_dimensionality_of_ssl_embeddings
    role: supports
    claim: In singing voice conversion, reducing the dimensionality of SSL embeddings through random channel selection
      proportionally reduces timbre leakage while preserving sufficient phonetic content for accurate reconstruction.
    source: §4. Results, Table 1
    evidence: SSL-128-Emb (128 of 768 HuBERT dimensions) achieves SMOS 3.920 vs. SSL-Emb's 2.750 on the Chinese
      test set, while maintaining CMOS 4.244 vs. 4.366 for full embeddings. The pattern holds for SSL-Soft and ContentVec
      variants.
    confidence: high
    relevance: high
  - claim_id: discrete_token_based_content_representations_for_voice_conversion_cannot_generalize
    role: complicates
    claim: Discrete token-based content representations for voice conversion cannot generalize to phonetic inventories
      outside the training distribution.
    source: §4. Results, Table 2
    evidence: SSL-Token CMOS drops from 4.086 on Chinese to 3.210 on other languages (English, Korean, Vietnamese,
      Japanese, Cantonese), demonstrating that k-means quantization of HuBERT with up to 10,000 clusters still loses
      phonetic detail that is language-specific. Continuous dimension-reduced embeddings maintain CMOS above 4.17
      across both conditions.
    confidence: high
    relevance: low
  - claim_id: self_supervised_speech_embedding_dimensions_contain_roughly_proportional_timbre_and
    role: supports
    claim: Self-supervised speech embedding dimensions contain roughly proportional timbre and content signal, such
      that uniform random subsampling functions as an effective form of timbre disentanglement.
    source: §2.2 Proposed Content Encoder, §4. Results
    evidence: The uniform-distribution assumption underlying random dimension selection is empirically supported
      by consistent improvements across three distinct SSL embedding types (HuBERT, SSL-Soft, ContentVec), each
      responding to dimension reduction in the same direction.
    confidence: high
    relevance: high
  - claim_id: supervised_disentanglement_methods_for_ssl_content_encoders_in_voice_conversion
    role: refines
    claim: Supervised disentanglement methods for SSL content encoders in voice conversion can be matched or exceeded
      by unsupervised dimensionality reduction at equivalent embedding sizes.
    source: §4. Results, Table 1, Table 2
    evidence: SSL-256-Emb surpasses SSL-Soft-Emb (which uses supervised soft target training) at the same 256-dimensional
      size on both SSIM and SMOS metrics in Chinese. ContentVEC-256-Emb matches token-based singer similarity while
      outperforming token-based CMOS on cross-lingual evaluation.
    confidence: high
    relevance: high
  limitations:
  - All training data is proprietary (200h internal Chinese singing corpus). No public benchmark or standard SVC
    evaluation set is used, making direct comparison to other published SVC systems infeasible.
  - The geometric assumption that timbre and content are uniformly distributed across SSL embedding dimensions is
    intuitive but unverified; the actual structure of HuBERT's representation space may be far from uniform, and
    it is unclear whether the technique generalises beyond informal Chinese singing. The optimal dimension count
    d is determined by grid search over 128 and 256 with no principled selection criterion. As d approaches zero,
    content reconstruction quality must eventually degrade, but the paper does not characterise this failure regime.
    The multilingual test set covers a limited set of languages, and the cross-lingual results do not control for
    dataset recording conditions.
  caveats: []
- id: interspeech-2025-1595
  published_date: "2025-08-17"
  entry_date: '2026-07-29'
  year: 2025
  venue: Interspeech
  task:
  - SCA
  architecture:
  - autoregressive-LM
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: active_evidence
  method_family:
  - autoregressive_ssl_conditioned_models
  claims:
  - claim_id: gradual_modality_transition_during_fine_tuning_rather_than_abrupt_full
    role: supports
    claim: Gradual modality transition during fine-tuning, rather than abrupt full-speech training, improves LLM
      adaptation from text to speech units.
    source: §3.2, Table 2
    evidence: Gradual modality transition during fine-tuning, rather than abrupt full-speech training, improves
      LLM adaptation from text to speech units.
    confidence: high
    relevance: medium
  - claim_id: the_benefit_of_interleaved_speech_text_training_is_amplified_in
    role: supports
    claim: The benefit of interleaved speech-text training is amplified in low-resource language settings where
      speech-domain supervision is scarce.
    source: §4, Table 2
    evidence: The benefit of interleaved speech-text training is amplified in low-resource language settings where
      speech-domain supervision is scarce.
    confidence: high
    relevance: medium
  - claim_id: applying_word_aligned_text_interleaving_to_both_source_and_target
    role: supports
    claim: Applying word-aligned text interleaving to both source and target sequences simultaneously is necessary;
      one-sided interleaving provides substantially weaker adaptation.
    source: §4, Table 3
    evidence: Applying word-aligned text interleaving to both source and target sequences simultaneously is necessary;
      one-sided interleaving provides substantially weaker adaptation.
    confidence: high
    relevance: medium
  - claim_id: the_modality_gap_between_speech_and_text_in_llm_based
    role: supports
    claim: The modality gap between speech and text in LLM-based S2ST manifests as both a length disparity and a
      representation distance, and scheduled training addresses both.
    source: §3.1, Figure 3
    evidence: The modality gap between speech and text in LLM-based S2ST manifests as both a length disparity and
      a representation distance, and scheduled training addresses both.
    confidence: high
    relevance: medium
  limitations:
  - The system uses a unit-based HiFi-GAN vocoder that maps semantic units to single-speaker synthesized English
    speech (from CVSS-C, which uses a canonical TTS voice). Speaker identity is not preserved, so the evaluation
    reflects translation accuracy and audio naturalness but not voice characteristics. Results may not generalise
    to multi-speaker or voice-preserving S2ST settings.
  - The evaluation is limited to seven language pairs from a single corpus (CVSS-C) and uses a 1B-parameter LLM.
    It is unclear whether the scheduling strategy scales to larger models or to more diverse multilingual corpora.
    The method relies on CTC-based word alignments, which requires a separately fine-tuned SSL encoder; its applicability
    when high-quality alignments are unavailable is not explored. The paper also does not compare against systems
    using acoustic (codec) units for speech generation, which would require a different vocoder setup. Future work
    suggested by the authors includes extension to spoken dialogue systems.
  caveats: []
- id: interspeech-2025-1625
  published_date: "2025-08-17"
  entry_date: '2026-07-29'
  year: 2025
  venue: Interspeech
  task:
  - VC
  architecture:
  - GAN
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: active_evidence
  method_family:
  - gan_decoders_for_ssl_units
  claims:
  - claim_id: active_adversarial_defense_mechanisms_can_prevent_voice_conversion_models_from
    role: supports
    claim: Active adversarial defense mechanisms can prevent voice conversion models from extracting speaker characteristics
      without sacrificing perceptual audio quality.
    source: §3.2, Table 1
    evidence: Mimic Blocker achieves PESQ 3.633 and STOI 0.942 on VCTK while maintaining ASR above 0.88 in both
      white-box and black-box conditions, compared to PESQ 1.61-1.99 for prior baselines at equivalent or lower
      ASR levels.
    confidence: high
    relevance: low
  - claim_id: pretrained_self_supervised_speech_representations_can_serve_as_model_agnostic
    role: supports
    claim: Pretrained self-supervised speech representations can serve as model-agnostic attack targets for VC defense,
      enabling consistent protection across different VC architectures without retraining.
    source: §3.1, §3.2, Table 1, Table 2
    evidence: Training with WavLM (white-box) and HuBERT (black-box) as feature extractors, Mimic Blocker achieves
      ASR of 0.88 on FreeVC and 0.99 on TriAAN-VC with the same trained model, demonstrating cross-architecture
      transfer.
    confidence: high
    relevance: high
  - claim_id: waveform_domain_adversarial_perturbation_preserves_speech_naturalness_better_than_spectrogram
    role: supports
    claim: Waveform-domain adversarial perturbation preserves speech naturalness better than spectrogram-domain
      approaches in voice conversion defense.
    source: §2.2, §3.2, §3.3, Table 1
    evidence: Prior defenses applying attacks in the Mel-spectrogram domain with vocoder reconstruction achieve
      PESQ 1.61-1.99; Mimic Blocker's direct waveform noise injection achieves PESQ 3.633, and 99.7% of evaluators
      judge the perturbed waveform as perceptually identical to the original speaker.
    confidence: high
    relevance: low
  - claim_id: active_vc_defense_training_does_not_require_target_speaker_data
    role: refines
    claim: Active VC defense training does not require target speaker data, relaxing a key dependency of earlier
      adversarial protection methods.
    source: §2.3, §1
    evidence: Unlike DYV and RW-VoiceShield, which require style-and-target speech pairs during training, Mimic
      Blocker uses only the original speaker's own speech by maximizing the L2 distance between SSL embeddings of
      original and perturbed waveforms via a self-supervised loss.
    confidence: high
    relevance: medium
  limitations:
  - The subjective evaluation involves only 6 evaluators assessing 50 speech pairs from VCTK. This panel size is
    small by standard listening-test norms and may not capture evaluator diversity or yield statistically robust
    naturalness estimates.
  - Evaluation is limited to two VC models (FreeVC, TriAAN-VC) on a single English dataset; performance against
    other VC architectures and non-English speakers is unknown. Training requires approximately 25 hours on an NVIDIA
    RTX 4090, which could be a practical barrier if per-speaker models are needed at scale. The white-box and black-box
    definitions are modified from prior work to accommodate the model-agnostic design, making direct numerical comparison
    with all earlier methods imprecise. It remains unclear how robust the defense is against adaptive attacks specifically
    designed to counter WavLM/HuBERT-based perturbations.
  caveats: []
- id: interspeech-2025-1639
  published_date: "2025-08-17"
  entry_date: '2026-07-29'
  year: 2025
  venue: Interspeech
  task:
  - codec
  - VC
  architecture:
  - GAN
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: active_evidence
  method_family:
  - gan_decoders_for_ssl_units
  claims:
  - claim_id: targeted_distillation_into_specific_rvq_layers_can_isolate_fine_grained
    role: supports
    claim: Targeted distillation into specific RVQ layers can isolate fine-grained speaking style attributes (such
      as vocal effort) independently of semantic content in neural speech codecs.
    source: §2.3, §3.1, Table 2
    evidence: LombardTokenizer conditions the second RVQ layer via cosine distillation from Lombard speech encoders
      while keeping RVQ layer 1 semantically focused via mHuBERT; the resulting system achieves WER 10.97% and EER
      6.67% on vocal effort conversion on AVID, substantially outperforming FreeVC (WER 20.35%, EER 16.67%) while
      retaining comparable synthesis quality (PESQ 3.07 vs. SpeechTokenizer 3.01).
    confidence: high
    relevance: medium
  - claim_id: disentanglement_constraints_in_neural_codecs_reduce_unconstrained_reconstruction_quality_relative
    role: complicates
    claim: Disentanglement constraints in neural codecs reduce unconstrained reconstruction quality relative to
      codecs without regularisation.
    source: §3.1, Table 1
    evidence: EnCodec (no disentanglement) achieves PESQ 3.32 and STOI 0.94 on LibriSpeech, while LombardTokenizer's
      dual distillation (semantic and Lombard) yields PESQ 3.07 and STOI 0.93, consistent with SpeechTokenizer's
      3.01; the quality gap widens on in-distribution neutral speech but narrows on expressive zero-shot data.
    confidence: high
    relevance: medium
  - claim_id: providing_a_specialized_style_encoder_to_a_voice_conversion_model
    role: complicates
    claim: Providing a specialized style encoder to a voice conversion model architecture does not guarantee that
      the encoder's information will be effectively exploited for style control.
    source: §3.2, Table 2, Figure 2
    evidence: FreeVC modified with the same Lombard encoder (FVClmb) fails to produce statistically distinguishable
      vocal effort levels in perceptual evaluation, and degrades WER on FLombard (32.63%) relative to the standard
      speaker-encoder variant (24.04%), while LT1 using the same encoder via RVQ distillation achieves WER 17.74%
      and accurate perceptual effort transfer.
    confidence: high
    relevance: low
  - claim_id: zero_shot_generalization_of_speaking_style_transfer_across_languages_is
    role: supports
    claim: Zero-shot generalization of speaking style transfer across languages is achievable when codec disentanglement
      is guided by multilingual self-supervised representations.
    source: §2.1, §3.2, Table 2
    evidence: LombardTokenizer trained on English AVID achieves WER 15.77% on the unseen French FLombard dataset
      (zero-shot), compared to FreeVC's 24.04–32.63%; the paper attributes part of this advantage to replacing HuBERT
      with multilingual mHuBERT in the semantic RVQ layer.
    confidence: high
    relevance: high
  - claim_id: codec_level_disentanglement_enables_robust_cross_speaker_style_transfer_with
    role: supports
    claim: Codec-level disentanglement enables robust cross-speaker style transfer with low speaker identity leakage
      under both intra-speaker and inter-speaker conditions.
    source: §3.2, Table 2, Figure 2
    evidence: LombardTokenizer inter-speaker vocal effort conversion produces perceptual distributions not significantly
      different from intra-speaker conversion (Dunn's test), with EER remaining below 8% on AVID for both LT1 (6.67%)
      and LT2 (7.83%), indicating preserved speaker identity.
    confidence: high
    relevance: low
  limitations:
  - The zero-shot evaluation is from English training to French test data, and both languages are Indo-European
    — the claim of cross-lingual generalisation has not been tested on typologically distant languages.
  - The model is evaluated on controlled intensity datasets (AVID, FLombard) and trained on studio-quality recordings;
    performance in real-world noisy conditions or spontaneous Lombard speech is unknown. The AVID dataset uses instructed
    intensity levels rather than naturally elicited Lombard speech, which may reduce ecological validity. Synthesis
    quality is measured with PESQ and STOI, which emphasise intelligibility and signal fidelity but may not capture
    naturalness of expressive speech styles. The study does not address how many simultaneous disentanglement targets
    an RVQ structure can support before quality degrades substantially.
  caveats: []
- id: interspeech-2025-1776
  published_date: "2025-08-17"
  entry_date: '2026-07-29'
  year: 2025
  venue: Interspeech
  task:
  - TTS
  architecture:
  - hybrid
  relevance: medium
  evidence_role:
  - architecture_variant
  current_role: active_evidence
  method_family:
  - hybrid_semantic_acoustic_ssl_systems
  claims:
  - claim_id: multi_task_joint_training_on_synthesis_editing_and_continuation_tasks
    role: supports
    claim: Multi-task joint training on synthesis, editing, and continuation tasks improves speech synthesis quality
      over single-task training in non-autoregressive codec models.
    source: §3.2, §3.3, Tables 1, 3
    evidence: In all three tokenizer configurations (ST, STDAC, HuDAC), multi-task SpeechSEC consistently outperforms
      the corresponding single-task baseline on MOS, voice preservation, WER, and CER, with gains confirmed by ablation.
    confidence: high
    relevance: low
  - claim_id: in_multi_task_speech_generation_training_editing_tasks_primarily_contribute
    role: refines
    claim: In multi-task speech generation training, editing tasks primarily contribute intelligibility improvements
      while continuation tasks primarily contribute acoustic quality and voice preservation.
    source: §3.3, Table 3
    evidence: Ablation removing the editing task increases WER by up to 4.4 points with minimal audio quality change;
      removing continuation degrades MOS by up to 0.18 and voice preservation by up to 0.06 with smaller intelligibility
      effects.
    confidence: high
    relevance: medium
  - claim_id: non_autoregressive_masked_token_prediction_frameworks_can_unify_speech_synthesis
    role: supports
    claim: Non-autoregressive masked token prediction frameworks can unify speech synthesis, editing, and continuation
      tasks through task-specific input conditioning within a single model.
    source: §2, §3.2, Table 2
    evidence: SpeechSEC handles all three tasks with a shared Conformer backbone, differentiating tasks via a Task
      Register embedding and per-task masking strategies, achieving competitive quality on editing (MOS 3.93) and
      continuation (MOS 3.63) alongside synthesis.
    confidence: high
    relevance: medium
  - claim_id: the_choice_of_semantic_and_acoustic_token_extractor_significantly_affects
    role: complicates
    claim: The choice of semantic and acoustic token extractor significantly affects absolute synthesis quality
      in codec-based TTS, even when model architecture and training are held constant.
    source: §3.2, Table 1
    evidence: With the same SpeechSEC architecture and training scheme, MOS ranges from 3.65 (STDAC) to 4.20 (ST)
      across the three tokenizer configurations, and voice preservation from 0.61 to 0.72, indicating that tokenizer
      quality is a dominant factor.
    confidence: high
    relevance: low
  limitations:
  - The evaluation is restricted to LibriTTS-R, a clean studio-quality English corpus, leaving generalization to
    noisy, spontaneous, or multilingual speech untested. Cross-system comparisons with SoundStorm use independently
    reported numbers from separate evaluations, weakening the claim of surpassing prior state of the art. Model
    parameter count is not reported, preventing meaningful comparisons of capacity-normalised performance. Speech
    continuation lacks intelligibility metrics (WER, CER) by design, limiting interpretability of those results.
    The paper does not evaluate the editing task on real-world editing scenarios beyond random masking.
  caveats: []
- id: interspeech-2025-1779
  published_date: "2025-08-17"
  entry_date: '2026-07-29'
  year: 2025
  venue: Interspeech
  task:
  - VC
  architecture:
  - flow-matching
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: active_evidence
  method_family:
  - flow_matching_with_ssl_conditioning
  claims:
  - claim_id: rectified_flow_achieves_comparable_or_better_sample_quality_to_diffusion
    role: supports
    claim: Rectified flow achieves comparable or better sample quality to diffusion for voice conversion at substantially
      fewer inference steps.
    source: §4.2, Table 1
    evidence: Rectified flow achieves comparable or better sample quality to diffusion for voice conversion at substantially
      fewer inference steps.
    confidence: high
    relevance: low
  - claim_id: conditioning_speaker_embeddings_on_concurrent_content_and_pitch_features_improves
    role: supports
    claim: Conditioning speaker embeddings on concurrent content and pitch features improves speaker similarity
      in zero-shot voice conversion.
    source: §3.1, §4.2, Table 1
    evidence: Conditioning speaker embeddings on concurrent content and pitch features improves speaker similarity
      in zero-shot voice conversion.
    confidence: high
    relevance: low
  - claim_id: a_single_step_ode_solver_can_match_the_quality_of
    role: supports
    claim: A single-step ODE solver can match the quality of dozens of diffusion steps when the flow trajectories
      are sufficiently linear.
    source: §4.2, Table 1
    evidence: A single-step ODE solver can match the quality of dozens of diffusion steps when the flow trajectories
      are sufficiently linear.
    confidence: high
    relevance: medium
  - claim_id: recursive_rectification_retraining_on_model_generated_samples_yields_only_marginal
    role: supports
    claim: Recursive rectification (retraining on model-generated samples) yields only marginal improvements when
      the initial model is already well-trained.
    source: §4.2, Table 2
    evidence: Recursive rectification (retraining on model-generated samples) yields only marginal improvements
      when the initial model is already well-trained.
    confidence: high
    relevance: medium
  limitations:
  - The model is trained and evaluated on a 38-hour clean subset of LibriTTS with a single-GPU budget, and evaluated
    only on same-domain test utterances. Generalisation to out-of-domain or expressive speech is neither tested
    nor discussed.
  - Speaker similarity scores (SECS ~0.84) are competitive but still well below the vocoder upper bound (0.987),
    suggesting substantial headroom. WER and CER remain measurably above the ground truth pipeline (CER 2.12 vs.
    0.52 for GT), indicating some content leakage. The pitch extraction step adds over 1 second of preprocessing
    latency for a ~9-second utterance, which limits real-time applicability. The paper does not report real-time
    factor (RTF) or measure latency end-to-end. It is also unclear how performance scales with reference speech
    duration — the evaluation uses a fixed 4.7-second target clip.
  caveats: []
- id: interspeech-2025-2043
  published_date: "2025-08-17"
  entry_date: '2026-07-29'
  year: 2025
  venue: Interspeech
  task:
  - VC
  architecture:
  - transformer-enc-dec
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: active_evidence
  method_family:
  - transformer_ssl_encoders_and_adapters
  claims:
  - claim_id: training_free_voice_conversion_based_on_distribution_matching_in_self
    role: complicates
    claim: Training-free voice conversion based on distribution matching in self-supervised embedding subspaces
      can achieve speaker similarity and content preservation comparable to trained codec-based systems when reference
      audio is limited to a few seconds.
    source: §4.4, Table 2
    evidence: Training-free voice conversion based on distribution matching in self-supervised embedding subspaces
      can achieve speaker similarity and content preservation comparable to trained codec-based systems when reference
      audio is limited to a few seconds.
    confidence: high
    relevance: high
  - claim_id: the_nearest_neighbour_approach_in_knn_style_voice_conversion_degrades
    role: supports
    claim: The nearest-neighbour approach in kNN-style voice conversion degrades significantly in cross-lingual
      settings because phoneme-level proximity in the embedding space conflates content with language-specific pronunciation.
    source: §1, §4.4, Table 3
    evidence: The nearest-neighbour approach in kNN-style voice conversion degrades significantly in cross-lingual
      settings because phoneme-level proximity in the embedding space conflates content with language-specific pronunciation.
    confidence: high
    relevance: low
  - claim_id: non_uniform_variance_structure_in_self_supervised_speech_representations_makes
    role: supports
    claim: Non-uniform variance structure in self-supervised speech representations makes global optimal transport
      maps suboptimal; factorizing the embedding space by variance before applying transport improves both content
      preservation and numerical stability.
    source: §3
    evidence: Non-uniform variance structure in self-supervised speech representations makes global optimal transport
      maps suboptimal; factorizing the embedding space by variance before applying transport improves both content
      preservation and numerical stability.
    confidence: high
    relevance: high
  - claim_id: a_content_speaker_similarity_trade_off_is_inherent_to_wavlm
    role: complicates
    claim: A content-speaker-similarity trade-off is inherent to WavLM-based voice conversion and can be navigated
      via the block-size hyperparameter of a factorized transport scheme.
    source: §4.4, Table 2
    evidence: A content-speaker-similarity trade-off is inherent to WavLM-based voice conversion and can be navigated
      via the block-size hyperparameter of a factorized transport scheme.
    confidence: high
    relevance: high
  limitations:
  - The Gaussian assumption underlying MKL-VC is specific to WavLM-Large embeddings, as demonstrated empirically.
    The authors explicitly note that for any new encoder the assumption must be verified from scratch, making the
    method non-portable without additional analysis effort.
  - The method is evaluated only against objective metrics (WER, CER, X-vector cosine similarity) on LibriSpeech
    and FLEURS, and a small subjective ranking with six experts. No MOS or MUSHRA scores are reported, and the subjective
    evaluation is a ranking rather than an absolute naturalness measure, making it difficult to position MKL-VC
    on the standard VC evaluation scale.
  - Diff-VC was evaluated on only 855 of 7800 samples due to compute constraints; its baseline number may therefore
    be unreliable for fair comparison. SinkVC is a re-implementation rather than the official release, introducing
    uncertainty about whether the reported gap to MKL-VC reflects the method or the reimplementation fidelity.
  - No code or demo is linked in the paper; reproducibility depends on the authors releasing code.
  caveats: []
- id: interspeech-2025-2075
  published_date: "2025-08-17"
  entry_date: '2026-07-29'
  year: 2025
  venue: Interspeech
  task:
  - codec
  architecture: []
  relevance: medium
  evidence_role:
  - architecture_variant
  current_role: active_evidence
  method_family: []
  claims:
  - claim_id: multi_granularity_quantization_codebooks_encode_paralinguistic_and_prosodic_features_more
    role: supports
    claim: Multi-granularity quantization codebooks encode paralinguistic and prosodic features more efficiently
      than single-resolution codebooks at equivalent or higher bitrates.
    source: §6.2, Table 3
    evidence: SVCs (k=500 per codebook, ~544 bits/s on Expresso) outperform frame-level k=2000 baselines (~548 bits/s)
      on all emotion sub-categories and prominence classification, with angry speech F1 rising from 0.298 to 0.614.
    confidence: high
    relevance: medium
  - claim_id: pooling_continuous_speech_representations_before_discretization_retains_more_paralinguistic_and
    role: supports
    claim: Pooling continuous speech representations before discretization retains more paralinguistic and prosodic
      information than pooling after discretization.
    source: §6.1, Table 2
    evidence: Pre-pooling consistently outperforms post-pooling in both utterance-level SER (0.5074 vs 0.2834 accuracy)
      and word-level prominence classification (0.3423 vs 0.1210 F-score) across single-level and multi-level codebook
      conditions.
    confidence: high
    relevance: medium
  - claim_id: dsu_based_expressive_speech_resynthesis_retains_a_large_gap_in
    role: complicates
    claim: DSU-based expressive speech resynthesis retains a large gap in style fidelity relative to continuous-feature
      resynthesis, even with improved codebook designs.
    source: §6.2, §6.3, Table 4
    evidence: SVCs achieve 41.22% style classification accuracy on Expresso versus 74.72% for continuous features
      and 88.42% for ground truth, a 33-point gap despite SVCs outperforming all other DSU baselines tested.
    confidence: high
    relevance: medium
  - claim_id: forced_alignment_requirements_constrain_multi_granularity_dsu_methods_to_text
    role: complicates
    claim: Forced-alignment requirements constrain multi-granularity DSU methods to text-paired speech settings,
      limiting their applicability to spontaneous or unlabelled data.
    source: §4.1, §7
    evidence: Phone, word, and utterance boundaries for SVC segmentation are derived from forced alignments using
      the Montreal Forced Aligner and HTK; no unsupervised segmentation is evaluated, and the authors list automatic
      segmentation as future work.
    confidence: high
    relevance: medium
  limitations:
  - The method depends on forced alignments at three linguistic levels (phone, word, utterance), which requires
    text transcriptions and limits use in low-resource or spontaneous speech domains. The authors acknowledge this
    and identify unsupervised or automatic segmentation as the most promising avenue for removing this constraint.
  - Evaluation uses automated quality proxies only (UTMOS, style classifier accuracy, WER). The authors note that
    human listening tests and qualitative error analysis are needed and explicitly defer them to future work.
  - The merged frame-level DSU representation discards part of the factorized structure SVCs provide. Whether architectures
    that natively consume multi-stream DSUs would yield further gains is unanswered.
  - Finally, all codebook vocabulary sizes are fixed at k=500. The paper suggests that per-level vocabulary tuning
    could improve task adaptation, but this is not explored empirically.
  caveats: []
- id: interspeech-2025-2660
  published_date: "2025-08-17"
  entry_date: '2026-07-29'
  year: 2025
  venue: Interspeech
  task:
  - SCA
  architecture:
  - transformer-enc-dec
  relevance: medium
  evidence_role:
  - architecture_variant
  - acceleration_evidence
  - control_evidence
  current_role: active_evidence
  method_family:
  - transformer_ssl_encoders_and_adapters
  claims:
  - claim_id: acoustic_only_voice_activity_projection_can_predict_next_speaker_identity
    role: supports
    claim: Acoustic-only voice activity projection can predict next-speaker identity in triadic conversation with
      accuracy above a last-speaker baseline, though gains over baseline are modest in spontaneous overlapping speech.
    source: §3.2, Table 4
    evidence: Acoustic-only voice activity projection can predict next-speaker identity in triadic conversation
      with accuracy above a last-speaker baseline, though gains over baseline are modest in spontaneous overlapping
      speech.
    confidence: high
    relevance: medium
  - claim_id: the_complexity_of_multi_party_turn_taking_increases_faster_than
    role: supports
    claim: The complexity of multi-party turn-taking increases faster than linearly with group size, making state-space
      design a fundamental constraint for VAP-style approaches beyond dyadic settings.
    source: §2.1, §4
    evidence: The complexity of multi-party turn-taking increases faster than linearly with group size, making state-space
      design a fundamental constraint for VAP-style approaches beyond dyadic settings.
    confidence: high
    relevance: low
  - claim_id: the_type_of_conversation_spontaneous_discussion_versus_structured_attentive_listening
    role: supports
    claim: The type of conversation — spontaneous discussion versus structured attentive listening — substantially
      affects VAP accuracy, reflecting differences in overlapping speech prevalence.
    source: §3.2, Table 4
    evidence: The type of conversation — spontaneous discussion versus structured attentive listening — substantially
      affects VAP accuracy, reflecting differences in overlapping speech prevalence.
    confidence: high
    relevance: medium
  - claim_id: spoken_dialogue_systems_operating_in_multi_party_scenarios_require_turn
    role: supports
    claim: Spoken dialogue systems operating in multi-party scenarios require turn-taking models trained on matched
      multi-party data, as dyadic training corpora are unlikely to transfer directly.
    source: §2.2, §4
    evidence: Spoken dialogue systems operating in multi-party scenarios require turn-taking models trained on matched
      multi-party data, as dyadic training corpora are unlikely to transfer directly.
    confidence: high
    relevance: low
  limitations:
  - The TEIDAN corpus is small (under 5 hours), monolingual (Japanese), and not yet publicly released. The reported
    accuracy gains are measured on the same corpus used for training and testing, with no held-out language or domain.
    These factors limit the generalisability of the empirical claims.
  - 'The state-space reduction from 4 to 2 sub-state bins per speaker may lose prediction granularity relative to
    dyadic VAP; the paper does not quantify this trade-off directly. The acoustic-only constraint is explicit and
    acknowledged: gaze, gesture, syntax, and prosody are known turn-yielding cues that are not modelled. Scaling
    beyond three participants is expected to be problematic — the paper estimates 65,536 states for four-participant
    conversation under the original 4-bin scheme — and may require architectural alternatives to VAP. Only next-speaker
    prediction during mutual silence is evaluated; prediction during overlap, which is prevalent in spontaneous
    data, is left for future work.'
  caveats: []
- id: interspeech-2025-2684
  published_date: "2025-08-17"
  entry_date: '2026-07-29'
  year: 2025
  venue: Interspeech
  task:
  - VC
  architecture:
  - flow-matching
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: active_evidence
  method_family:
  - flow_matching_with_ssl_conditioning
  claims:
  - claim_id: explicit_disentanglement_of_content_and_prosody_into_separate_discrete_token
    role: supports
    claim: Explicit disentanglement of content and prosody into separate discrete token spaces via different SSL
      representations enables independent control of speaking style in zero-shot voice conversion.
    source: §2.1, §3.2, Table 2
    evidence: Using HuBERT (content) and ContentVec with SimVQ bottleneck (prosody) as separate extractors, with
      a prosody mask transformer for reference-guided style transfer, yields F0 Corr of 0.941 for prosody preservation
      and 0.822 for prosody conversion on ESD/VCTK, while ablations show that using a single mel-based prosody input
      significantly degrades performance.
    confidence: high
    relevance: high
  - claim_id: flow_matching_with_in_context_learning_achieves_competitive_zero_shot
    role: supports
    claim: Flow matching with in-context learning achieves competitive zero-shot voice conversion quality at substantially
      fewer parameters than larger autoregressive systems.
    source: §3.2, Table 1
    evidence: Discl-VC (131M params) achieves MOS 4.31, UTMOS 4.079, and SECS 0.929 on VCTK zero-shot VC, surpassing
      Vevo (922M params) on naturalness and matching speaker similarity, demonstrating that compact flow matching
      transformers are competitive with large-scale AR systems for VC.
    confidence: high
    relevance: low
  - claim_id: prosody_transfer_from_a_reference_speaker_introduces_a_trade_off
    role: complicates
    claim: Prosody transfer from a reference speaker introduces a trade-off with target speaker identity preservation.
    source: §3.2, Table 2
    evidence: In the prosody conversion task, Discl-VC achieves lower SECS (0.847) than Vevo (0.892) despite better
      UTMOS and WER, suggesting that explicitly replacing source prosody with reference tokens disrupts timbre modeling
      and degrades speaker similarity relative to a system where prosody and identity are more entangled.
    confidence: high
    relevance: medium
  - claim_id: simvq_freezing_codebook_vectors_and_learning_a_linear_projection_reduces
    role: supports
    claim: SimVQ (freezing codebook vectors and learning a linear projection) reduces codebook collapse in VQ-based
      speech discretisation.
    source: §3.3, Table 3
    evidence: Ablation replacing SimVQ with standard VQ causes degradation across all metrics (UTMOS 4.042 vs 4.079,
      WER 2.182% vs 1.946%, F0 Corr 0.970 vs 0.973, SECS 0.928 vs 0.929), attributed to codebook collapse during
      training.
    confidence: high
    relevance: medium
  - claim_id: ssl_based_prosody_extraction_for_vc_can_retain_speaker_correlated
    role: complicates
    claim: SSL-based prosody extraction for VC can retain speaker-correlated information when the same-speaker assumption
      holds during training, causing train-inference mismatch.
    source: §3.3, Table 3
    evidence: Ablation without ContentVec (using mel spectrogram first 20 dimensions as prosody input) shows the
      largest performance drop (UTMOS 4.034, WER 3.147%), attributed to prosody tokens containing residual speaker
      information because source and target audio come from the same speaker during training, creating a mismatch
      with cross-speaker inference.
    confidence: high
    relevance: high
  limitations:
  - Evaluation is limited to English and uses only two baselines (Vevo and FAcodec) on a relatively small test set
    (450 pairs for zero-shot VC). The prosody conversion task's lower speaker similarity suggests that the prosody
    mask transformer introduces identity leakage when reference prosody comes from a different speaker. The paper
    does not report results on multilingual speakers or noisy conditions, leaving generalisation untested. The two-stage
    training procedure adds complexity and requires the stage-1 encoder to be frozen before the prosody mask transformer
    can be trained.
  caveats: []
- id: interspeech-2025-cho25c_interspeech
  published_date: "2025-08-17"
  entry_date: '2026-07-29'
  year: 2025
  venue: Interspeech
  task:
  - VC
  architecture:
  - hybrid
  relevance: medium
  evidence_role:
  - architecture_variant
  current_role: minor
  method_family:
  - hybrid_semantic_acoustic_ssl_systems
  claims:
  - claim_id: voice_conversion_models_trained_exclusively_on_human_speech_are_insufficient
    role: complicates
    claim: Voice conversion models trained exclusively on human speech are insufficient for targets with broader-than-human
      frequency range and non-periodic vocalizations.
    source: §3
    evidence: Monster sounds span a wider frequency range than human speech, contain intricate audio-effect details,
      and include non-speech elements (unvoiced breathing, extreme vocal expressions) absent from standard human-speech
      training corpora; standard VC models trained on these corpora fail to reproduce the target timbres.
    confidence: high
    relevance: low
  - claim_id: replacing_f0_estimation_with_frame_level_energy_as_the_primary
    role: supports
    claim: Replacing f0 estimation with frame-level energy as the primary prosodic cue enables voice conversion
      for target voices where pitch periodicity cannot be reliably estimated.
    source: §4
    evidence: The H2NH prior encoder substitutes frame-level energy for f0 because f0 estimation degrades severely
      for non-human monster sounds; the system still produces satisfactory conversions in the real-time demonstration.
    confidence: high
    relevance: low
  - claim_id: extending_audio_bandwidth_to_44_1_khz_and_incorporating_high
    role: supports
    claim: Extending audio bandwidth to 44.1 kHz and incorporating high-temporal-resolution discriminators improves
      voice conversion quality for transient-rich non-human acoustic targets.
    source: §4
    evidence: The H2NH model uses a 5 ms hop-length STFT, Mel spectrograms covering 0-22.05 kHz, FDRL with multiple
      hop lengths, and a DAC discriminator specifically to capture fine temporal details and broad frequency content
      of monster vocalizations that standard 22.05 kHz pipelines cannot reproduce.
    confidence: high
    relevance: low
  limitations:
  - This paper contains only a surface-level architectural description; training data, hyperparameters, speaker
    or monster coverage, and quantitative evaluation metrics are all deferred to the companion Interspeech 2025
    research paper. Claims about conversion fidelity cannot be independently verified from this paper alone.
  - The system targets a predefined set of monster voices for a specific gaming title; generalisation to arbitrary
    non-human voices, varied recording conditions, or diverse languages is not assessed. No ablation studies, failure
    cases, or degradation conditions are described. The demonstration runs locally within a VM environment, and
    no information is given about latency, throughput, or hardware requirements that would inform deployment feasibility.
  caveats: []
- id: '2508.15565'
  published_date: "2025-08-21"
  entry_date: '2026-07-29'
  year: 2025
  venue: arXiv
  task:
  - VC
  architecture:
  - GAN
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: active_evidence
  method_family:
  - gan_decoders_for_ssl_units
  claims:
  - claim_id: anonymizing_adversarial_utterances_toward_a_pseudo_speaker_derived_from_the
    role: supports
    claim: Anonymizing adversarial utterances toward a pseudo-speaker derived from the batch mean of adversarial
      embeddings improves identity unlinkability without requiring a designated real speaker.
    source: §IV-A, Table I
    evidence: Anonymizing adversarial utterances toward a pseudo-speaker derived from the batch mean of adversarial
      embeddings improves identity unlinkability without requiring a designated real speaker.
    confidence: high
    relevance: medium
  - claim_id: in_feedforward_adversarial_anonymization_training_with_a_composite_loss_covering
    role: supports
    claim: In feedforward adversarial anonymization, training with a composite loss covering untargeted attack,
      perceptual quality, and unlinkability objectives achieves better balance between de-identification and identity
      unlinkability than untargeted attack alone.
    source: §IV-A, Table I
    evidence: In feedforward adversarial anonymization, training with a composite loss covering untargeted attack,
      perceptual quality, and unlinkability objectives achieves better balance between de-identification and identity
      unlinkability than untargeted attack alone.
    confidence: high
    relevance: medium
  - claim_id: adversarial_perturbation_based_voice_anonymization_degrades_substantially_under_strong_adaptive
    role: supports
    claim: Adversarial perturbation-based voice anonymization degrades substantially under strong adaptive attacks
      such as quantization, revealing a fundamental vulnerability that is independent of the training strategy.
    source: §V-F, Table III
    evidence: Adversarial perturbation-based voice anonymization degrades substantially under strong adaptive attacks
      such as quantization, revealing a fundamental vulnerability that is independent of the training strategy.
    confidence: high
    relevance: medium
  - claim_id: speaker_adversarial_perturbations_transfer_less_effectively_to_black_box_speaker
    role: supports
    claim: Speaker adversarial perturbations transfer less effectively to black-box speaker extractors with strong
      verification capability, indicating that white-box EER gains overstate real-world privacy protection.
    source: §V-G, Table II
    evidence: Speaker adversarial perturbations transfer less effectively to black-box speaker extractors with strong
      verification capability, indicating that white-box EER gains overstate real-world privacy protection.
    confidence: high
    relevance: medium
  limitations:
  - De-identification and identity unlinkability EERs drop sharply under adaptive attacks (quantization reduces
    de-id EER from 46.79% to 7.63%), and black-box transfer to the ResNet100 extractor similarly reduces EERs to
    single digits. The method's privacy guarantees are therefore limited to white-box threat models, which may not
    reflect realistic deployment conditions.
  - Identity unlinkability generalises less well than de-identification to out-of-domain datasets, particularly
    LibriSpeech and AIShell, suggesting the batch mean loss is sensitive to the speaker distribution at training
    time. The method also degrades speech intelligibility (WER increases across all ASR test sets compared to original
    speech), which limits applicability in voice assistant or captioning contexts. Future directions include improving
    transferability against black-box extractors, enhancing resilience to adaptive attacks, and strengthening out-of-domain
    unlinkability.
  caveats: []
- id: '2508.16188'
  published_date: "2025-08-22"
  entry_date: '2026-07-29'
  year: 2025
  venue: EMNLP
  task:
  - TTS
  - SCA
  architecture:
  - autoregressive-LM
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: active_evidence
  method_family:
  - autoregressive_ssl_conditioned_models
  claims:
  - claim_id: full_face_visual_features_encode_complementary_emotional_information_not_captured
    role: supports
    claim: Full-face visual features encode complementary emotional information not captured by audio alone, and
      combining both modalities yields meaningfully higher emotion recognition accuracy than either modality individually.
    source: §3, Table 1
    evidence: Full-face visual features encode complementary emotional information not captured by audio alone,
      and combining both modalities yields meaningfully higher emotion recognition accuracy than either modality
      individually.
    confidence: high
    relevance: medium
  - claim_id: prefix_based_visual_fusion_using_compressed_q_former_query_latents
    role: supports
    claim: Prefix-based visual fusion using compressed Q-Former query latents integrates more effectively into autoregressive
      speech LMs than direct feature concatenation, which collapses perplexity.
    source: §5.1, Table 3
    evidence: Prefix-based visual fusion using compressed Q-Former query latents integrates more effectively into
      autoregressive speech LMs than direct feature concatenation, which collapses perplexity.
    confidence: high
    relevance: medium
  - claim_id: visual_guidance_during_training_improves_not_only_classification_accuracy_but
    role: supports
    claim: Visual guidance during training improves not only classification accuracy but also the emotional alignment
      of generated speech in conversational settings.
    source: §5.3, Table 7
    evidence: Visual guidance during training improves not only classification accuracy but also the emotional alignment
      of generated speech in conversational settings.
    confidence: high
    relevance: medium
  - claim_id: emotion_controllability_through_prompt_only_label_manipulation_is_insufficient_when
    role: supports
    claim: Emotion controllability through prompt-only label manipulation is insufficient when a model is trained
      on data where input and response emotions are correlated; in-context demonstrations are necessary to override
      this bias.
    source: §5.3, Figure 5
    evidence: Emotion controllability through prompt-only label manipulation is insufficient when a model is trained
      on data where input and response emotions are correlated; in-context demonstrations are necessary to override
      this bias.
    confidence: high
    relevance: medium
  limitations:
  - The expressive generation evaluation relies on a third-party model (Qwen2-Audio) to classify emotion in synthesised
    speech — there is no human perceptual evaluation of naturalness, similarity, or overall quality. The fine-tuning
    dataset is small (4,859 synthetic pairs from IEMOCAP), and the system is evaluated on a closed four-class emotion
    taxonomy, limiting generalisability to naturalistic conversation.
  - The decoding strategy enforces structural constraints on SpiritLM's interleaved token types (style → pitch →
    semantic), which adds inference complexity and may suppress the model's capacity for stylistic diversity beyond
    the four emotion classes. The AVSR results (3.5% WER clean) lag specialist systems by more than 2.5 percentage
    points, a gap the paper attributes to design choices rather than fundamental limitations — though this distinction
    is not tested directly. Cross-language and multi-speaker generalisation beyond IEMOCAP's two-speaker dyadic
    setup are not evaluated.
  caveats: []
- id: '2508.16790'
  published_date: "2025-08-22"
  entry_date: '2026-07-29'
  year: 2025
  venue: arXiv
  task:
  - codec
  - TTS
  architecture:
  - diffusion
  - transformer-enc-dec
  - autoregressive-LM
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: influential
  method_family:
  - diffusion_with_ssl_conditioning
  - transformer_ssl_encoders_and_adapters
  - autoregressive_ssl_conditioned_models
  claims:
  - claim_id: text_conditioning_in_the_codec_decoder_rather_than_in_the
    role: supports
    claim: Text conditioning in the codec decoder, rather than in the language model alone, is a viable lever for
      achieving extreme compression rates in speech tokenization without adversarial training.
    source: §3.1, Table 4
    evidence: Text conditioning in the codec decoder, rather than in the language model alone, is a viable lever
      for achieving extreme compression rates in speech tokenization without adversarial training.
    confidence: high
    relevance: low
  - claim_id: a_single_end_to_end_training_objective_flow_matching_loss
    role: supports
    claim: A single end-to-end training objective (flow-matching loss) is sufficient to jointly optimise quantization
      and reconstruction in a speech codec, eliminating the need for multi-stage pipelines.
    source: §3.1, §4.2.2
    evidence: A single end-to-end training objective (flow-matching loss) is sufficient to jointly optimise quantization
      and reconstruction in a speech codec, eliminating the need for multi-stage pipelines.
    confidence: high
    relevance: low
  - claim_id: the_reconstruction_generation_gap_the_degradation_in_intelligibility_when_tokens
    role: supports
    claim: The reconstruction-generation gap — the degradation in intelligibility when tokens are predicted by a
      language model rather than encoding reference speech — varies substantially across tokenizer architectures
      and is not captured by reconstruction metrics alone.
    source: §4.3, Figure 3
    evidence: The reconstruction-generation gap — the degradation in intelligibility when tokens are predicted by
      a language model rather than encoding reference speech — varies substantially across tokenizer architectures
      and is not captured by reconstruction metrics alone.
    confidence: high
    relevance: medium
  - claim_id: lower_token_rates_in_speech_tokenizers_can_improve_autoregressive_tts
    role: supports
    claim: Lower token rates in speech tokenizers can improve autoregressive TTS intelligibility by shortening prediction
      sequences and reducing error accumulation, particularly on linguistically challenging inputs.
    source: §4.3, Table 5
    evidence: Lower token rates in speech tokenizers can improve autoregressive TTS intelligibility by shortening
      prediction sequences and reducing error accumulation, particularly on linguistically challenging inputs.
    confidence: high
    relevance: medium
  - claim_id: binary_spherical_quantization_without_a_commitment_loss_achieves_stable_end
    role: supports
    claim: Binary Spherical Quantization without a commitment loss achieves stable end-to-end training of a speech
      codec and produces superior representations to standard VQ under equal codebook sizes.
    source: §4.2.2, Table 4
    evidence: Binary Spherical Quantization without a commitment loss achieves stable end-to-end training of a speech
      codec and produces superior representations to standard VQ under equal codebook sizes.
    confidence: high
    relevance: low
  limitations:
  - 'TaDiCodec''s text-aware decoder is not a general-purpose audio codec: it requires a transcript at both training
    and inference time. The strong reconstruction and TTS results are conditional on text availability; performance
    at 6.25 Hz without text conditioning is not competitive (WER exceeds 10% at 12.5 Hz without text, per Table
    4). This limits applicability to codec-transmission, speech enhancement, or any scenario where transcriptions
    are unavailable.'
  - The diffusion decoder introduces multi-step inference latency. At 32 steps, decoding speed is acceptable for
    generation but higher than GAN vocoders; reducing to 5 steps degrades quality noticeably. The authors propose
    distillation as future work but have not yet demonstrated single-step performance.
  - The system has been validated on TTS only; whether TaDiCodec tokens are useful for spoken dialogue modeling,
    speech understanding, or other downstream tasks remains untested. The prompt mechanism, while effective, means
    full reconstruction quality depends on a speech prefix being available — similar to zero-shot TTS conditioning,
    but potentially limiting in some SCA deployment scenarios.
  caveats: []
- id: '2411.19770'
  published_date: "2025-08-28"
  entry_date: '2026-07-29'
  year: 2025
  venue: arXiv
  task:
  - VC
  architecture:
  - diffusion
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: active_evidence
  method_family:
  - diffusion_with_ssl_conditioning
  claims:
  - claim_id: contrastive_training_with_noise_augmented_views_enforces_noise_invariant_speaker
    role: supports
    claim: Contrastive training with noise-augmented views enforces noise-invariant speaker representations and
      substantially improves one-shot VC robustness at low SNR without adding inference cost.
    source: §III.A, Tables I-II
    evidence: Noro's dual-branch reference encoding module and noise-agnostic contrastive speaker loss hold SECS
      at 80.09 and CER at 4.66 in 0-5 dB noise, versus 77.28 and 7.26 for the baseline; weight sharing means inference
      architecture is unchanged.
    confidence: high
    relevance: medium
  - claim_id: standard_one_shot_voice_conversion_systems_degrade_substantially_when_reference
    role: complicates
    claim: Standard one-shot voice conversion systems degrade substantially when reference speech contains background
      noise, even after data augmentation training.
    source: §III.A.2, Table II
    evidence: The diffusion-based baseline trained on LibriLight with no noise-robustness mechanism shows CMOS dropping
      from 3.29 to 2.09 and SMOS from 3.02 to 2.75 under 0-5 dB noisy reference conditions.
    confidence: high
    relevance: low
  - claim_id: voice_conversion_reference_encoders_trained_on_large_scale_speech_data
    role: supports
    claim: Voice conversion reference encoders trained on large-scale speech data develop speaker representations
      competitive with dedicated self-supervised speaker models.
    source: §III.B.2, Table III
    evidence: VC-SPK2VEC (the Noro baseline reference encoder repurposed as a speaker encoder, 72.4M params, trained
      on LibriLight 60k hr) achieves 5.32% EER on VoxCeleb1 under SUPERB, outperforming wav2vec 2.0 Base (6.02%),
      wav2vec 2.0 Large (5.65%), and HuBERT Large (5.98%).
    confidence: high
    relevance: high
  - claim_id: speaker_noise_disentanglement_in_voice_conversion_benefits_from_training_objectives
    role: refines
    claim: Speaker-noise disentanglement in voice conversion benefits from training objectives that explicitly align
      clean and noisy representations of the same speaker, beyond simple noise augmentation.
    source: §III.A.2, Figure 2
    evidence: t-SNE visualisations show that the baseline (trained with augmentation but no contrastive alignment)
      produces clearly separated clean/noisy representation clusters, while Noro's contrastive loss causes them
      to mix, correlating with the performance gap under noise.
    confidence: high
    relevance: low
  limitations:
  - The evaluation uses reference speech from VCTK (studio-recorded English) with synthetically added noise from
    DEMAND; real-world recordings that mix noise with reverberation, codec compression, or far-field capture may
    behave differently. The test set is small (ten source-reference pairs per condition for subjective evaluation,
    150 for objective), and the paper does not report statistical significance for differences between Noro and
    the baseline in clean conditions where they are nearly equal. The contrastive loss relies on speaker labels
    during training, so the approach does not extend directly to fully unsupervised training on unlabelled data.
    The secondary VC-SPK2VEC finding is evaluated only under one SUPERB protocol; performance on other speaker tasks
    (diarization, speaker counting) is not explored.
  caveats: []
- id: '2507.14534'
  published_date: "2025-08-30"
  entry_date: '2026-07-29'
  year: 2025
  venue: arXiv
  task:
  - VC
  architecture:
  - GAN
  - hybrid
  relevance: medium
  evidence_role:
  - architecture_variant
  - acceleration_evidence
  - control_evidence
  - evaluation_caution
  current_role: active_evidence
  method_family:
  - gan_decoders_for_ssl_units
  - hybrid_semantic_acoustic_ssl_systems
  claims:
  - claim_id: streaming_voice_conversion_quality_degrades_significantly_when_offline_hubert_representations
    role: supports
    claim: Streaming voice conversion quality degrades significantly when offline HuBERT representations are naively
      replaced by causal alternatives, but knowledge distillation into an Emformer backbone can recover content
      accuracy with acceptable latency.
    source: §III.B, Table III
    evidence: Streaming voice conversion quality degrades significantly when offline HuBERT representations are
      naively replaced by causal alternatives, but knowledge distillation into an Emformer backbone can recover
      content accuracy with acceptable latency.
    confidence: high
    relevance: high
  - claim_id: causal_temporal_upsampling_via_pixel_shuffle_eliminates_the_checkerboard_artifacts
    role: supports
    claim: Causal temporal upsampling via pixel shuffle eliminates the checkerboard artifacts introduced by zero-padding
      non-causal vocoders, without sacrificing subjective quality.
    source: §III.D, Table III
    evidence: Causal temporal upsampling via pixel shuffle eliminates the checkerboard artifacts introduced by zero-padding
      non-causal vocoders, without sacrificing subjective quality.
    confidence: high
    relevance: medium
  - claim_id: explicit_style_modeling_with_clustering_based_vector_quantization_improves_zero
    role: supports
    claim: Explicit style modeling with clustering-based vector quantization improves zero-shot speaker similarity
      in streaming VC beyond what timbre embeddings alone provide.
    source: §III.C, Table I
    evidence: Explicit style modeling with clustering-based vector quantization improves zero-shot speaker similarity
      in streaming VC beyond what timbre embeddings alone provide.
    confidence: high
    relevance: low
  - claim_id: online_voice_conversion_systems_can_achieve_speaker_similarity_comparable_to
    role: supports
    claim: Online voice conversion systems can achieve speaker similarity comparable to offline systems when style
      transfer is modeled at chunk level rather than at the global utterance level.
    source: §IV.B, Table I
    evidence: Online voice conversion systems can achieve speaker similarity comparable to offline systems when
      style transfer is modeled at chunk level rather than at the global utterance level.
    confidence: high
    relevance: low
  limitations:
  - '- Evaluation is in English only; cross-lingual style transfer is untested. - The reference speaker must be
    fully available before streaming begins, limiting applications where reference is also captured in real-time.
    - CER is slightly higher than StreamVC because StreamVC reuses source pitch; Conan introduces some pitch variation
    that ASR penalizes as CER. - Model size is not reported; the cost of the Emformer + main model + CSV in production
    deployment is unclear. - Perceptual evaluation used 15 listeners per pair on Amazon Mechanical Turk style tasks;
    larger-scale evaluation would be needed to establish statistical robustness.'
  caveats: []
- id: '2509.00503'
  published_date: "2025-08-30"
  entry_date: '2026-07-29'
  year: 2025
  venue: arXiv
  task:
  - VC
  - codec
  architecture:
  - autoregressive-LM
  - transformer-enc-dec
  relevance: medium
  evidence_role:
  - architecture_variant
  current_role: active_evidence
  method_family:
  - autoregressive_ssl_conditioned_models
  - transformer_ssl_encoders_and_adapters
  claims:
  - claim_id: adaptive_entropy_based_segmentation_of_discrete_speech_tokens_preserves_more
    role: supports
    claim: Adaptive entropy-based segmentation of discrete speech tokens preserves more task-relevant linguistic
      information than fixed-length downsampling at equivalent compression ratios.
    source: §6.1, Table 3; §6.2, Table 4
    evidence: Adaptive entropy-based segmentation of discrete speech tokens preserves more task-relevant linguistic
      information than fixed-length downsampling at equivalent compression ratios.
    confidence: high
    relevance: medium
  - claim_id: optimal_token_granularity_differs_systematically_between_understanding_and_generation_tasks
    role: supports
    claim: 'Optimal token granularity differs systematically between understanding and generation tasks: recognition-oriented
      tasks (ASR, ST) benefit from moderate compression near phoneme rate, while voice conversion requires finer
      token density to maintain acoustic fidelity.'
    source: §5.1, Table 1; §6.3, Table 5
    evidence: 'Optimal token granularity differs systematically between understanding and generation tasks: recognition-oriented
      tasks (ASR, ST) benefit from moderate compression near phoneme rate, while voice conversion requires finer
      token density to maintain acoustic fidelity.'
    confidence: high
    relevance: low
  - claim_id: ssl_derived_semantic_tokens_at_standard_rates_50_hz_contain
    role: supports
    claim: SSL-derived semantic tokens at standard rates (50 Hz) contain substantial redundancy that can be removed
      without degrading — and occasionally improving — downstream task performance.
    source: §6.1, Table 3
    evidence: SSL-derived semantic tokens at standard rates (50 Hz) contain substantial redundancy that can be removed
      without degrading — and occasionally improving — downstream task performance.
    confidence: high
    relevance: high
  - claim_id: entropy_boundaries_in_compressed_token_sequences_align_with_linguistically_meaningful
    role: supports
    claim: 'Entropy boundaries in compressed token sequences align with linguistically meaningful units: 15 Hz compression
      achieves 83.2% phoneme boundary alignment, while 7 Hz aligns primarily with word boundaries (89.7%).'
    source: Appendix C.3, Table 11
    evidence: 'Entropy boundaries in compressed token sequences align with linguistically meaningful units: 15 Hz
      compression achieves 83.2% phoneme boundary alignment, while 7 Hz aligns primarily with word boundaries (89.7%).'
    confidence: high
    relevance: medium
  limitations:
  - 'Voice conversion quality degrades noticeably with compression: entropy-guided 15 Hz already falls below HuBERT
    deduplicated (Q-MOS 3.85 vs 4.12), and the gap widens at higher compression. For generation tasks, this framework
    does not improve on simply using deduplicated tokens — only for understanding tasks does compression help.'
  - The framework is evaluated exclusively on HuBERT-derived tokens; whether the entropy-based approach transfers
    to other SSL features (WavLM, w2v-BERT) or supervised tokenizers (S3, FACodec) is untested. The entropy LLM
    requires pre-training on 20k hours of speech, adding a pipeline step beyond vanilla k-means clustering. The
    method is tested on English only, and its behaviour on morphologically complex or tonal languages — where token
    redundancy patterns may differ — remains unknown. All evaluations use the English MLS training corpus, so generalisation
    across domains and recording conditions is unexplored.
  caveats: []
- id: '2509.00675'
  published_date: "2025-08-31"
  entry_date: '2026-07-29'
  year: 2025
  venue: arXiv
  task:
  - TTS
  architecture:
  - transformer-enc-dec
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  - evaluation_caution
  current_role: active_evidence
  method_family:
  - transformer_ssl_encoders_and_adapters
  claims:
  - claim_id: speaker_specific_phrasing_behaviour_is_a_substantive_source_of_variance
    role: complicates
    claim: Speaker-specific phrasing behaviour is a substantive source of variance in RP insertion that generic
      multi-speaker models fail to capture, and modeling it explicitly improves both objective and subjective phrasing
      quality.
    source: §5.1.2, Table 3, Table 5
    evidence: Speaker-specific phrasing behaviour is a substantive source of variance in RP insertion that generic
      multi-speaker models fail to capture, and modeling it explicitly improves both objective and subjective phrasing
      quality.
    confidence: high
    relevance: medium
  - claim_id: phoneme_level_language_models_outperform_subword_level_models_on_phrase
    role: supports
    claim: Phoneme-level language models outperform subword-level models on phrase break prediction, even at smaller
      model sizes, because phoneme representations carry acoustic information more directly relevant to pause insertion
      than subword tokens.
    source: §5.1.3, Table 4
    evidence: Phoneme-level language models outperform subword-level models on phrase break prediction, even at
      smaller model sizes, because phoneme representations carry acoustic information more directly relevant to
      pause insertion than subword tokens.
    confidence: high
    relevance: medium
  - claim_id: scaling_subword_plms_from_base_to_large_yields_diminishing_returns
    role: supports
    claim: Scaling subword PLMs from BASE to LARGE yields diminishing returns for phrasing tasks, suggesting a representational
      ceiling specific to this task modality.
    source: §5.1.3, Table 4
    evidence: Scaling subword PLMs from BASE to LARGE yields diminishing returns for phrasing tasks, suggesting
      a representational ceiling specific to this task modality.
    confidence: high
    relevance: medium
  - claim_id: pre_trained_speaker_verification_embeddings_capture_prosodic_and_fluency_related
    role: supports
    claim: Pre-trained speaker verification embeddings capture prosodic and fluency-related characteristics that
      transfer to phrasing models via few-shot adaptation without fine-tuning, enabling reasonable generalization
      to unseen speakers.
    source: §5.2.2, Table 7
    evidence: Pre-trained speaker verification embeddings capture prosodic and fluency-related characteristics that
      transfer to phrasing models via few-shot adaptation without fine-tuning, enabling reasonable generalization
      to unseen speakers.
    confidence: high
    relevance: medium
  - claim_id: f0_5_score_and_naturalness_mos_can_diverge_for_phrasing
    role: supports
    claim: F0.5 score and naturalness MOS can diverge for phrasing models using different PLMs, indicating that
      objective phrasing accuracy does not fully predict perceived speech naturalness.
    source: §5.2.3, Table 8
    evidence: F0.5 score and naturalness MOS can diverge for phrasing models using different PLMs, indicating that
      objective phrasing accuracy does not fully predict perceived speech naturalness.
    confidence: high
    relevance: medium
  limitations:
  - Training data is exclusively from LibriTTS-R audiobook readings. The resulting phrasing models are likely miscalibrated
    for spontaneous speech, conversational TTS, or out-of-domain styles; generalization is explicitly flagged by
    the authors as an open problem.
  - 'Additional limitations: the paper cannot disentangle the contributions of phoneme vs. subword information within
    phoneme-level PLMs (since their pre-training also includes grapheme-level objectives), leaving the mechanism
    of improvement partially unclear. The embedding adapter for unseen speakers assumes that the mapping from PSVM
    embeddings to trained embeddings is approximately injective and learnable with a small network — this assumption
    may not hold for speakers whose acoustic characteristics fall outside the training distribution.'
  caveats: []
- id: '2509.01391'
  published_date: "2025-09-01"
  entry_date: '2026-07-29'
  year: 2025
  venue: arXiv
  task:
  - TTS
  architecture:
  - transformer-enc-dec
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  - evaluation_caution
  current_role: minor
  method_family:
  - transformer_ssl_encoders_and_adapters
  claims:
  - claim_id: ssl_derived_discrete_tokens_can_substitute_phoneme_sequences_as_input
    role: supports
    claim: SSL-derived discrete tokens can substitute phoneme sequences as input representations in a TTS front-end
      without a substantial naturalness penalty in automatic evaluation.
    source: §V.B, Table II
    evidence: SSL-derived discrete tokens can substitute phoneme sequences as input representations in a TTS front-end
      without a substantial naturalness penalty in automatic evaluation.
    confidence: high
    relevance: high
  - claim_id: g2p_free_tts_pipelines_that_learn_text_to_token_mappings
    role: supports
    claim: G2P-free TTS pipelines that learn text-to-token mappings from paired speech data avoid the language-specific
      resource burden of phoneme dictionaries and morphological analysers.
    source: §I, §III
    evidence: G2P-free TTS pipelines that learn text-to-token mappings from paired speech data avoid the language-specific
      resource burden of phoneme dictionaries and morphological analysers.
    confidence: high
    relevance: medium
  - claim_id: in_an_ssl_token_based_tts_pipeline_the_spectral_predictor
    role: supports
    claim: In an SSL-token-based TTS pipeline, the spectral predictor contributes more to naturalness differences
      than the text-to-token mapping stage.
    source: §V.B, Table II
    evidence: In an SSL-token-based TTS pipeline, the spectral predictor contributes more to naturalness differences
      than the text-to-token mapping stage.
    confidence: high
    relevance: high
  - claim_id: ssl_based_speech_representations_preserve_sufficient_acoustic_quality_through_a
    role: supports
    claim: SSL-based speech representations preserve sufficient acoustic quality through a discretise-then-synthesise
      pipeline to remain competitive with G2P-derived representations on codec-style quality metrics.
    source: §V.C, Table II
    evidence: SSL-based speech representations preserve sufficient acoustic quality through a discretise-then-synthesise
      pipeline to remain competitive with G2P-derived representations on codec-style quality metrics.
    confidence: high
    relevance: high
  limitations:
  - All evaluations use automatic metrics only (UTMOS, CER, WARP-Q, SDR) on 100 utterances from a single speaker
    subset of JVS. No subjective MOS or preference tests are reported; the conclusions about naturalness parity
    are therefore tentative.
  - The system is demonstrated exclusively on Japanese and does not yet extend to multilingual settings. The T5
    tokenizer is language-specific (tohoku-BERTv3), which the authors acknowledge as a barrier to multilingual scalability.
    Future directions involve BPE tokenizers (mT5, ByT5) to reduce this dependency. The oracle's counter-intuitive
    lower UTMOS than the proposed system is left unexplained and may indicate an interaction between SSL token sequence
    statistics and FastSpeech 2's duration predictor. Individual contribution analysis of duration, pitch, and accent
    inputs to FastSpeech 2 is identified as missing.
  caveats: []
- id: '2509.03292'
  published_date: "2025-09-03"
  entry_date: '2026-07-29'
  year: 2025
  venue: arXiv
  task:
  - evaluation
  architecture:
  - hybrid
  relevance: medium
  evidence_role:
  - architecture_variant
  current_role: active_evidence
  method_family:
  - hybrid_semantic_acoustic_ssl_systems
  claims:
  - claim_id: self_supervised_audio_representations_trained_on_natural_audio_can_generalise
    role: supports
    claim: Self-supervised audio representations trained on natural audio can generalise to predicting perceptual
      quality dimensions of synthetic speech and audio without synthetic training examples.
    source: §IV.C, Table I
    evidence: Self-supervised audio representations trained on natural audio can generalise to predicting perceptual
      quality dimensions of synthetic speech and audio without synthetic training examples.
    confidence: high
    relevance: high
  - claim_id: triplet_loss_applied_to_intermediate_embeddings_with_score_based_sampling
    role: supports
    claim: Triplet loss applied to intermediate embeddings with score-based sampling improves rank correlation with
      human perceptual ratings over MSE-only training objectives.
    source: §III, §IV.C, Table II
    evidence: Triplet loss applied to intermediate embeddings with score-based sampling improves rank correlation
      with human perceptual ratings over MSE-only training objectives.
    confidence: high
    relevance: medium
  - claim_id: multi_axis_perceptual_quality_prediction_benefits_from_task_specific_attention
    role: supports
    claim: Multi-axis perceptual quality prediction benefits from task-specific attention heads over a shared backbone,
      reflecting the partial independence of production and content quality dimensions.
    source: §II.C
    evidence: Multi-axis perceptual quality prediction benefits from task-specific attention heads over a shared
      backbone, reflecting the partial independence of production and content quality dimensions.
    confidence: high
    relevance: medium
  - claim_id: domain_shift_between_natural_and_synthetic_audio_is_more_problematic
    role: supports
    claim: Domain shift between natural and synthetic audio is more problematic for absolute score prediction than
      for rank-based ordering of perceptual quality.
    source: §IV.C, Table II
    evidence: Domain shift between natural and synthetic audio is more problematic for absolute score prediction
      than for rank-based ordering of perceptual quality.
    confidence: high
    relevance: medium
  limitations:
  - The model is trained exclusively on 2,700 natural audio samples, yet evaluated entirely on synthetic audio.
    The challenge protocol validates that this generalises in rank-order terms, but the large CE MSE gap (3.991
    vs. 1.142 baseline) indicates that absolute score calibration is unreliable for content enjoyment specifically.
  - Evaluation is limited to the AudioMOS Challenge 2025 protocol; no comparison against other published automatic
    quality predictors (UTMOS, DNSMOS) is included, making it difficult to assess how this approach stands relative
    to the broader field. The model relies on BEATs pretrained on AudioSet rather than speech-specific SSL models,
    and the authors note that integrating WavLM for speech segments is deferred to future work. Batch size of 1
    due to variable-length audio constrains training efficiency. Code and pretrained weights are not released as
    of the preprint date.
  caveats: []
- id: '2509.05359'
  published_date: "2025-09-03"
  entry_date: '2026-07-29'
  year: 2025
  venue: arXiv
  task:
  - SCA
  architecture:
  - autoregressive-LM
  relevance: medium
  evidence_role:
  - architecture_variant
  current_role: active_evidence
  method_family:
  - autoregressive_ssl_conditioned_models
  claims:
  - claim_id: discrete_vocabulary_size_is_the_dominant_design_variable_in_speech
    role: supports
    claim: Discrete vocabulary size is the dominant design variable in speech LM continual pre-training, with compact
      cluster sets substantially outperforming fine-grained discretizations regardless of encoder choice.
    source: §3.1, Table 1
    evidence: WavLM NLL degrades from 2.05 (k=500) to 4.01 (k=2500) and 4.21 (k=5000) for SmolLM-135M at step 300;
      the same degradation pattern holds for HuBERT, XLS-R, and Wav2Vec 2.0 across all model sizes.
    confidence: high
    relevance: medium
  - claim_id: optimal_discrete_vocabulary_granularity_for_speech_lm_pre_training_is
    role: complicates
    claim: 'Optimal discrete vocabulary granularity for speech LM pre-training is not fixed across model capacities:
      larger models tolerate higher cluster counts where smaller models fail.'
    source: §3.2, Table 2
    evidence: SmolLM-135M degrades sharply at k=1,000 (NLL=2.19) versus k=500 (NLL=2.05), while SmolLM-1.7B maintains
      stable NLL of 1.83 to 1.94 across k=125 to k=1,000, suggesting capacity-dependent discretization regimes.
    confidence: high
    relevance: medium
  - claim_id: domain_matching_between_the_clustering_training_corpus_and_the_target
    role: supports
    claim: Domain matching between the clustering training corpus and the target speech domain is a critical factor
      for both discrete unit quality and robustness to acoustic perturbations.
    source: §3.3, Table 3
    evidence: LibriHeavy-trained k-means (read English, in-domain with LibriSpeech) achieves NLL=2.62 clean and
      2.70 under heavy noise; GigaSpeech-trained k-means degrades from 3.07 to 3.09 and CommonVoice from 2.85 to
      3.11, showing that exposure to noisier audio during k-means construction does not improve robustness.
    confidence: high
    relevance: medium
  - claim_id: self_supervised_discrete_speech_units_encode_phoneme_level_distinctions_without
    role: supports
    claim: Self-supervised discrete speech units encode phoneme-level distinctions without explicit phonetic supervision,
      as evidenced by their systematic alignment with forced-alignment phoneme sequences.
    source: §3.5, Figure 2
    evidence: Phoneme confusion matrices on LibriSpeech test-clean (WavLM, k=125) show a strong diagonal structure
      where vowels and consonants each specialize into distinct unit subsets; the pattern persists across all four
      k-means data sources and holds at k=250.
    confidence: high
    relevance: high
  limitations:
  - Evaluation uses only NLL on LibriSpeech, a proxy metric. No downstream task evaluation (spoken question answering,
    spoken language understanding, ASR) is included, so the practical implications of the observed NLL differences
    for real applications remain untested.
  - The study uses the SmolLM family exclusively, trained only on read English speech from LibriSpeech. Whether
    the encoder and cluster-count findings generalize to larger models, multilingual data, conversational speech,
    or production-scale training runs is unknown. The 1.7B model encounters out-of-memory limitations at k=2,500
    and above, so the full picture for large models at high cluster counts is missing. The relationship between
    cluster utilization efficiency and downstream performance is posited but not directly measured.
  caveats: []
- id: '2509.04072'
  published_date: "2025-09-04"
  entry_date: '2026-07-29'
  year: 2025
  venue: arXiv
  task:
  - TTS
  architecture:
  - autoregressive-LM
  - flow-matching
  relevance: medium
  evidence_role:
  - architecture_variant
  - evaluation_caution
  current_role: active_evidence
  method_family:
  - autoregressive_ssl_conditioned_models
  - flow_matching_with_ssl_conditioning
  claims:
  - claim_id: narrative_aware_segmentation_of_audiobook_data_into_character_quotation_and
    role: supports
    claim: Narrative-aware segmentation of audiobook data into character quotation and narration subsets yields
      training material with measurably higher emotional diversity than standard sentence-level audiobook splits.
    source: §3.2, Table 5
    evidence: Narrative-aware segmentation of audiobook data into character quotation and narration subsets yields
      training material with measurably higher emotional diversity than standard sentence-level audiobook splits.
    confidence: high
    relevance: medium
  - claim_id: flow_matching_tts_models_show_larger_expressivity_gains_from_fine
    role: supports
    claim: Flow-matching TTS models show larger expressivity gains from fine-tuning on targeted expressive speech
      data than autoregressive models with equivalent training setups.
    source: §4.3, Table 2
    evidence: Flow-matching TTS models show larger expressivity gains from fine-tuning on targeted expressive speech
      data than autoregressive models with equivalent training setups.
    confidence: high
    relevance: medium
  - claim_id: conditioning_tts_synthesis_on_surrounding_narrative_context_rather_than_only
    role: complicates
    claim: Conditioning TTS synthesis on surrounding narrative context rather than only the target utterance text
      improves contextual appropriateness of synthesized speech at the cost of modest intelligibility degradation.
    source: §4.3, Table 2
    evidence: Conditioning TTS synthesis on surrounding narrative context rather than only the target utterance
      text improves contextual appropriateness of synthesized speech at the cost of modest intelligibility degradation.
    confidence: high
    relevance: medium
  - claim_id: current_open_source_tts_systems_are_substantially_less_expressive_than
    role: supports
    claim: Current open-source TTS systems are substantially less expressive than human audiobook narrators on contextual
      benchmarks, even when naturalness scores (MOS) are comparable.
    source: §5.2, Table 4
    evidence: Current open-source TTS systems are substantially less expressive than human audiobook narrators on
      contextual benchmarks, even when naturalness scores (MOS) are comparable.
    confidence: high
    relevance: medium
  - claim_id: llm_extracted_speech_delivery_pseudo_labels_verbs_and_adverbs_from
    role: supports
    claim: LLM-extracted speech-delivery pseudo-labels (verbs and adverbs) from narrative prose are reliable enough
      to serve as training signals for expressive TTS, achieving high precision when confidence-filtered.
    source: §3.3, Figure 2
    evidence: LLM-extracted speech-delivery pseudo-labels (verbs and adverbs) from narrative prose are reliable
      enough to serve as training signals for expressive TTS, achieving high precision when confidence-filtered.
    confidence: high
    relevance: medium
  limitations:
  - The LibriQuotetest ground-truth is recorded by amateur LibriVox volunteers, who themselves show insufficient
    expressivity for many quotations. ContextMOS scores for ground-truth (3.55 average) are only marginally above
    the best TTS systems, raising questions about whether the benchmark captures professional audiobook narration
    standards or amateur reading behaviour.
  - The training set is not WER-filtered, meaning a small proportion of transcription errors may remain. The gender
    distribution of LibriVox speakers is unknown, introducing potential bias. Experiments use only English fiction;
    cross-lingual and non-fiction applicability is untested. The contextual conditioning results are encouraging
    but use only one-paragraph context windows — longer narrative context might yield larger gains. LibriQuote's
    utility for neural audio codec training (noted as future work) remains to be demonstrated.
  caveats: []
- id: '2509.04667'
  published_date: "2025-09-04"
  entry_date: '2026-07-29'
  year: 2025
  venue: arXiv
  task:
  - VC
  architecture:
  - GAN
  - hybrid
  relevance: medium
  evidence_role:
  - architecture_variant
  - acceleration_evidence
  - control_evidence
  current_role: active_evidence
  method_family:
  - gan_decoders_for_ssl_units
  - hybrid_semantic_acoustic_ssl_systems
  claims:
  - claim_id: limited_lookahead_in_causal_speech_encoders_substantially_improves_linguistic_content
    role: supports
    claim: Limited lookahead in causal speech encoders substantially improves linguistic content preservation with
      minimal latency penalty compared to purely causal encoders.
    source: §V.A, §V.B, Tables I–II
    evidence: Wave+CL accuracy improves from 53.16% at zero lookahead to 78.99% at 140ms, while end-to-end latency
      increases from 84.3ms to 203ms; extending to 280ms adds only 0.7pp accuracy with 120ms additional delay.
    confidence: high
    relevance: low
  - claim_id: better_content_encoding_in_streaming_anonymization_introduces_a_fundamental_tension
    role: complicates
    claim: 'Better content encoding in streaming anonymization introduces a fundamental tension: improved linguistic
      clarity reduces speaker anonymization strength under adversarial threat models.'
    source: §V.E, Table III
    evidence: Adding the contextual layer drops lazy-informed EER from 36.61% to 20.35% at zero lookahead, meaning
      the cleaner representations are more discriminative for speaker recognition attacks, directly trading anonymization
      quality for intelligibility.
    confidence: high
    relevance: low
  - claim_id: token_quantization_via_k_means_clustering_achieves_near_chance_speaker
    role: supports
    claim: Token quantization via k-means clustering achieves near-chance speaker verification performance in streaming
      anonymization by removing fine-grained speaker cues from content representations.
    source: §V.C, §V.D, §V.E, Tables IV–V
    evidence: Applying a 256-centroid k-means bottleneck raises lazy-informed EER from ~12% to ~47% (Wave+CL, 140ms
      lookahead), at a cost of WER rising from 2.09% to 9.52% and MOS falling from 3.79 to 3.22.
    confidence: high
    relevance: low
  - claim_id: streaming_voice_anonymization_systems_can_approach_offline_pipeline_anonymization_performance
    role: supports
    claim: Streaming voice anonymization systems can approach offline-pipeline anonymization performance when evaluated
      under the lazy-informed threat scenario, while retaining real-time latency.
    source: §V.F, Table VI
    evidence: DarkStream achieves 22.68% semi-informed EER in streaming mode (140ms lookahead), matching VoicePrivacy
      2024 baselines B3 (26.28%) and B5a (22.09%) that require full-utterance processing.
    confidence: high
    relevance: low
  - claim_id: direct_waveform_synthesis_in_streaming_voice_conversion_maintains_acceptable_naturalness
    role: complicates
    claim: Direct waveform synthesis in streaming voice conversion maintains acceptable naturalness without mel-spectrogram
      intermediate representations, but k-means quantization introduced for privacy causes perceivable quality degradation
      beyond what objective metrics capture.
    source: §V.D, Table V
    evidence: MOS drops from 3.79 (Wave+CL) to 3.22 (Wave+CL+KMeans) with quantization; WER degrades only modestly,
      indicating that intelligibility metrics underestimate the perceptual impact of quantization artifacts.
    confidence: high
    relevance: low
  limitations:
  - DarkStream does not explicitly disentangle static speaker traits (accent, age, sex) from dynamic attributes
    (emotion, speaking style), leaving indirect identity cues potentially intact. The semi-informed EER of 22.68%
    remains well above chance, indicating meaningful residual linkability for well-resourced adversaries.
  - 'The privacy/quality trade-off exposed by the quantization ablation is steep: each MOS point recovered (by disabling
    k-means) costs roughly 30pp EER under the lazy-informed scenario. Systems requiring both high quality and robust
    anonymization against semi-informed attackers have no current solution in this architecture. Comparison of perceptual
    quality against the offline VoicePrivacy baselines is not reported, so whether DarkStream''s naturalness advantage
    over batch-processing pipelines is real remains an open question. Evaluation is limited to English (LibriTTS),
    and generalization to accented or code-switched speech is untested.'
  caveats: []
- id: '2509.06074'
  published_date: "2025-09-07"
  entry_date: '2026-07-29'
  year: 2025
  venue: EMNLP
  task:
  - TTS
  architecture:
  - transformer-enc-dec
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  - evaluation_caution
  current_role: active_evidence
  method_family:
  - transformer_ssl_encoders_and_adapters
  claims:
  - claim_id: word_level_semantic_and_prosodic_context_modeling_in_dialogue_history
    role: supports
    claim: Word-level semantic and prosodic context modeling in dialogue history improves prosody expressiveness
      in conversational TTS over utterance-level context encoders.
    source: §3.5, Table 1; §3.6, Table 2
    evidence: MFCIG-CSS achieves N-DMOS 3.980 and P-DMOS 3.899 on DailyTalk, outperforming seven baselines that
      operate at the utterance level by 0.122 and 0.104 respectively; ablation removing both graph modules causes
      the largest performance collapse across all metrics.
    confidence: high
    relevance: low
  - claim_id: graph_neural_networks_with_sequential_cross_turn_aggregation_can_encode
    role: supports
    claim: Graph neural networks with sequential cross-turn aggregation can encode complementary semantic and prosodic
      interaction patterns from multimodal dialogue history for prosody-aware TTS.
    source: §3.6, Table 2
    evidence: 'SIG and PIG are independently ablated: removing SIG reduces N-DMOS by 0.147 and P-DMOS by 0.106;
      removing PIG produces comparable drops. Both modules contribute distinctly and their combined use provides
      the strongest performance.'
    confidence: high
    relevance: low
  - claim_id: objective_energy_error_and_subjective_prosody_quality_metrics_can_rank
    role: complicates
    claim: Objective energy error and subjective prosody quality metrics can rank conversational TTS systems differently.
    source: §3.5, Table 1
    evidence: MFCIG-CSS achieves best N-DMOS (3.980) and P-DMOS (3.899) on DailyTalk but ranks second on MAE-E (0.314
      vs. 0.310 for MSRGCN-CSS), indicating that signal-level energy accuracy does not fully predict human prosody
      quality judgments.
    confidence: high
    relevance: medium
  - claim_id: prosody_modeling_gains_demonstrated_with_acoustic_feature_based_tts_backbones
    role: complicates
    claim: Prosody modeling gains demonstrated with acoustic-feature-based TTS backbones may not transfer to codec-token-based
      or flow-based architectures.
    source: §5 Limitations
    evidence: MFCIG-CSS is validated only on a FastSpeech 2 backbone; the authors explicitly identify extension
      to VITS-based architectures and discrete token-based speech encoders as future work, acknowledging that the
      current validation scope limits generalizability claims.
    confidence: high
    relevance: low
  limitations:
  - MFCIG-CSS is evaluated on a single English dialogue dataset (DailyTalk, approximately 20 hours) with a FastSpeech
    2 backbone, leaving generalization to other languages, longer conversations, noisy conditions, and modern autoregressive
    or codec-based TTS systems untested. The interaction graphs operate on frame-averaged acoustic features from
    Wav2Vec 2.0 and do not yet model fine-grained intra-word acoustic cues such as emotion, emphasis, or pauses.
    Extension to VITS-based architectures and discrete token-based speech encoders is the primary open direction
    identified by the authors.
  caveats: []
- id: '2509.09174'
  published_date: "2025-09-11"
  entry_date: '2026-07-29'
  year: 2025
  venue: arXiv
  task:
  - SCA
  architecture:
  - autoregressive-LM
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: active_evidence
  method_family:
  - autoregressive_ssl_conditioned_models
  claims:
  - claim_id: decoupling_semantic_training_objectives_from_acoustic_token_prediction_substantially_reduces
    role: supports
    claim: Decoupling semantic training objectives from acoustic token prediction substantially reduces knowledge
      degradation in speech-to-speech LLMs.
    source: §5.1, Table 4
    evidence: EchoX's Echo training, which generates speech targets from the model's own semantic hidden states,
      raises average QA accuracy from 24.3 (T2C without Echo) to 37.1 (EchoX-3B) on the same data, compared to 12.8
      for direct interleaved training.
    confidence: high
    relevance: medium
  - claim_id: unit_based_speech_token_compression_via_language_model_segmentation_improves
    role: supports
    claim: Unit-based speech token compression via language-model segmentation improves downstream accuracy and
      reduces sequence length without sacrificing audio quality.
    source: §5.3, Table 5, Figure 7
    evidence: Unit language achieves 4.57 length ratio vs. 9.31 for raw units while improving accuracy on all three
      QA benchmarks and maintaining comparable audio quality in spectral comparison.
    confidence: high
    relevance: medium
  - claim_id: streaming_inference_in_speech_llms_can_be_achieved_with_minimal
    role: supports
    claim: Streaming inference in speech LLMs can be achieved with minimal accuracy degradation when the segmentation
      boundary is determined by semantic similarity rather than fixed length.
    source: §5.4, Table 6
    evidence: EchoX's cosine-similarity trigger reduces first-token latency from 138 to 27 tokens at 3B scale with
      less than 1.5 percentage points of accuracy drop on any benchmark.
    confidence: high
    relevance: low
  - claim_id: training_data_efficiency_in_speech_llms_may_depend_more_on
    role: complicates
    claim: Training data efficiency in speech LLMs may depend more on the training paradigm than on data volume.
    source: §4.2, Table 2
    evidence: EchoX achieves competitive performance on spoken QA against models trained on millions of hours using
      only approximately 6,200 hours, but this result holds specifically for factual QA and has not been tested
      on broader spoken dialogue tasks.
    confidence: high
    relevance: medium
  - claim_id: speech_naturalness_and_response_helpfulness_are_not_jointly_optimised_by
    role: complicates
    claim: Speech naturalness and response helpfulness are not jointly optimised by the same training signal in
      speech-to-speech LLMs.
    source: §Appendix C, Figure 8
    evidence: Human evaluation shows EchoX wins clearly on helpfulness but performs only competitively on naturalness,
      reflecting a training objective focused on semantic correctness rather than prosodic quality.
    confidence: high
    relevance: medium
  limitations:
  - Speech quality is assessed only via brief spectral comparison (Figure 7) and a 5-rater human study. No perceptual
    quality metric (MOS, DNSMOS) or automatic speech recognition accuracy on generated audio is reported as a primary
    evaluation result, making it difficult to characterise the system's output quality independently of QA accuracy.
  - The evaluation benchmarks are limited to factual knowledge QA (Llama Questions, Web Questions, TriviaQA). It
    is unclear whether Echo training retains its advantage on open-ended dialogue, instruction following, or longer-form
    conversational tasks. The human evaluation was conducted with only five raters on one dataset (AlpacaEval),
    which limits statistical confidence in the naturalness comparison. The model trains on synthesised assistant
    audio from GPT-SoVITS, which may introduce a fixed timbre bias and limit voice diversity. The streaming threshold
    and window size are fixed hyperparameters with no ablation reported on their sensitivity.
  caveats: []
- id: '2509.09201'
  published_date: "2025-09-11"
  entry_date: '2026-07-29'
  year: 2025
  venue: arXiv
  task:
  - codec
  - VC
  architecture:
  - GAN
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: active_evidence
  method_family:
  - gan_decoders_for_ssl_units
  claims:
  - claim_id: neural_audio_codecs_can_disentangle_speech_and_background_sound_in
    role: supports
    claim: Neural audio codecs can disentangle speech and background sound in the representation domain, enabling
      downstream tasks to selectively access either component without explicit front-end separation.
    source: §4.3, §4.5, §6.1, Table 4, Table 6
    evidence: DeCodec's SOP+RST design achieves orthogonal speech and background subspaces; recombining representations
      enables one-shot VC (SPK-SIM 0.83, WER 50.46) and speech enhancement (DNSMOS OVL 3.39) without denoising pre-processing,
      outperforming cascaded StoRM+SpeechTokenizer on ASR (WER* 26.7 vs. 34.5).
    confidence: high
    relevance: medium
  - claim_id: semantic_guidance_from_self_supervised_features_in_the_first_quantiser
    role: supports
    claim: Semantic guidance from self-supervised features in the first quantiser layer improves the noise-robustness
      of codec representations used for downstream ASR.
    source: §6.1.4, Table 5
    evidence: Ablation-3 (SOP+RST without SG) achieves WER* 41.9 on noisy speech; adding HuBERT-L9 guidance reduces
      it to 25.8 (causal) and 23.6 (non-causal), confirming that semantic anchoring helps the quantiser concentrate
      linguistic content in a background-invariant first layer.
    confidence: high
    relevance: high
  - claim_id: explicit_disentanglement_constraints_in_neural_codecs_introduce_a_reconstruction_quality
    role: complicates
    claim: Explicit disentanglement constraints in neural codecs introduce a reconstruction quality trade-off relative
      to purely reconstruction-optimised designs.
    source: §6.1.1, Table 2
    evidence: DeCodec's mel distance on clean speech (0.89) is worse than DAC (0.65) and HiFi-Codec (0.75), which
      use no orthogonality or disentanglement losses, suggesting that forcing orthogonal subspace separation increases
      spectral distortion.
    confidence: high
    relevance: medium
  - claim_id: hierarchical_codec_disentanglement_enables_controllable_background_sound_handling_in_zero
    role: supports
    claim: Hierarchical codec disentanglement enables controllable background sound handling in zero-shot TTS without
      retraining the downstream generation model.
    source: §6.2.2, Table 7
    evidence: VALL-E trained on DeCodec tokens achieves MOS 4.09 with background preserved (BPMOS 4.19) vs. MOS
      3.96 without, controlled solely by whether the BRVQ-1:8 background tokens are included at inference; the TTS
      model itself was trained only on clean speech.
    confidence: high
    relevance: low
  - claim_id: voice_conversion_on_noisy_speech_via_representation_recombination_produces_high
    role: complicates
    claim: Voice conversion on noisy speech via representation recombination produces high WER even when speaker
      similarity is well-preserved, due to voiced/unvoiced segment boundary mismatches between source and reference
      utterances.
    source: §6.1.3, Table 4
    evidence: DeCodec one-shot VC achieves SPK-SIM 0.83 (matching the reference ceiling of 0.69 being clearly exceeded)
      but WER 50.46, which the authors attribute to structural mismatch in speech segment timing rather than semantic
      content corruption.
    confidence: high
    relevance: low
  limitations:
  - Subjective evaluations use panels of 10-12 volunteers on 300-clip test sets, which is small for MOS-based conclusions.
    The noisy speech test set is constructed by synthetically mixing clean LibriSpeech with DNS-Noise at controlled
    SNRs (-5 to 20 dB), which may not reflect the diversity of real-world background conditions.
  - The model operates at 16 kHz and is trained on a speech-dominant corpus; performance on music or non-speech
    audio types is not evaluated despite the universal codec framing. The high VC WER (50.46) under noisy conditions
    indicates that the current representation recombination approach for voice conversion degrades intelligibility
    significantly and would likely need additional design work before practical deployment. Model size is not reported,
    making direct efficiency comparisons with baselines difficult.
  caveats: []
- id: '2509.09550'
  published_date: "2025-09-11"
  entry_date: '2026-07-29'
  year: 2025
  venue: arXiv
  task:
  - codec
  architecture:
  - hybrid
  - GAN
  relevance: medium
  evidence_role:
  - architecture_variant
  current_role: active_evidence
  method_family:
  - hybrid_semantic_acoustic_ssl_systems
  - gan_decoders_for_ssl_units
  claims:
  - claim_id: fsq_based_neural_audio_codecs_develop_inherent_redundancy_in_their
    role: supports
    claim: FSQ-based neural audio codecs develop inherent redundancy in their code representations, enabling multiple
      independently trained encoders to produce radically different code sequences that nonetheless decode to perceptually
      equivalent audio.
    source: §4.2, Table 2
    evidence: Encoder distillation experiment on NeuCodec showing only 2% element-wise code match between original
      (635M params) and distilled (42M params) encoders, while cosine similarity between pre-quantization projections
      is 0.73 and reconstruction metrics (WER, CER, STOI, PESQ) remain within 0.5% WER of each other.
    confidence: high
    relevance: medium
  - claim_id: fsq_codecs_are_substantially_more_robust_to_bit_level_transmission
    role: supports
    claim: FSQ codecs are substantially more robust to bit-level transmission errors than RVQ codecs of comparable
      codebook size, maintaining intelligibility at bit-flip rates an order of magnitude higher than the RVQ failure
      threshold.
    source: §5, Figure 3
    evidence: Binary symmetric channel simulation on LibriSpeech test-clean showing FSQ codecs (NeuCodec, Distill-NeuCodec,
      StableCodec) maintain stable STOI and PESQ up to 10% bit-flip probability, whereas RVQ codecs (EnCodec, DAC)
      experience sharp quality collapse above 1% bit-flip rate.
    confidence: high
    relevance: medium
  - claim_id: the_perturbation_robustness_advantage_of_fsq_over_rvq_is_structural
    role: refines
    claim: 'The perturbation robustness advantage of FSQ over RVQ is structural: FSQ''s fixed-grid quantization
      maps bit-flip perturbations to bounded, predictable steps in the embedding space, whereas RVQ code perturbations
      produce arbitrarily large embedding-space displacements.'
    source: §4.2, §6
    evidence: Implicit codebook confusion matrices show that 93% of level predictions between original and distilled
      encoders are either correct or off by exactly one neighbouring level, confirming local structure in the FSQ
      embedding space.
    confidence: high
    relevance: medium
  - claim_id: fsq_based_codec_architectures_that_achieve_competitive_intelligibility_metrics_often
    role: complicates
    claim: FSQ-based codec architectures that achieve competitive intelligibility metrics often rely on very large
      pretrained semantic encoders, concentrating most of the parameter budget in a component that contributes indirectly
      to the quantization benefit.
    source: §3, §4.1, Table 2
    evidence: Wav2Vec2-BERT-large accounts for 600M of NeuCodec's 635M total parameters; replacing it with DistilHuBERT
      reduces the model to 42M parameters with only a 0.5% WER increase, suggesting the semantic representation
      is partially substitutable without sacrificing the FSQ robustness property.
    confidence: high
    relevance: low
  limitations:
  - Evaluation uses WER, CER, STOI, and PESQ but no perceptual naturalness metric (MOS or MUSHRA), making it impossible
    to assess how NeuCodec compares to state-of-the-art codecs on speech generation quality. The perturbation experiment
    uses a binary symmetric channel that does not reflect real-world transmission protocols (e.g., structured burst
    errors or packet loss), so robustness claims cannot be extrapolated directly to deployment scenarios.
  - Training includes 1,000 hours of proprietary data, partially limiting reproducibility. The model is evaluated
    only at 16kHz and 24kHz; the 24kHz upsampling decoder is trained on only 2,600 hours of data, and high-fidelity
    (48kHz) audio is not addressed. The paper does not evaluate whether FSQ robustness properties are preserved
    when the codec is used as a tokenizer for downstream TTS or language model tasks, which is the stated application
    motivation.
  caveats: []
- id: '2509.11425'
  published_date: "2025-09-14"
  entry_date: '2026-07-29'
  year: 2025
  venue: arXiv
  task:
  - codec
  - TTS
  architecture:
  - GAN
  - autoregressive-LM
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: active_evidence
  method_family:
  - gan_decoders_for_ssl_units
  - autoregressive_ssl_conditioned_models
  claims:
  - claim_id: injecting_multimodal_guidance_directly_into_the_encoder_latent_space_of
    role: supports
    claim: Injecting multimodal guidance directly into the encoder latent space of a neural codec improves reconstruction
      quality beyond similarity-based supervision of the quantized layer.
    source: §2.3.1, §3.2.1, Table 2
    evidence: FuseCodec-Fusion, which fuses global semantic and contextual vectors into Z via additive fusion, achieves
      WER 3.99, ViSQOL 3.47, and PESQ 3.13 on LibriSpeech test-clean, outperforming FuseCodec-Distill (ViSQOL 3.43,
      PESQ 3.06) and FuseCodec-ContextAlign, both of which only supervise Q(1) without modifying the latent.
    confidence: high
    relevance: low
  - claim_id: neural_codecs_trained_with_joint_semantic_and_contextual_supervision_generalize
    role: supports
    claim: Neural codecs trained with joint semantic and contextual supervision generalize to unseen languages without
      multilingual training data.
    source: §3.3, Table 5
    evidence: FuseCodec, trained exclusively on English LibriSpeech train-clean-100, achieves the best WER and ViSQOL
      in the majority of 7 tested languages from Multilingual LibriSpeech and outperforms all baselines on PESQ
      by at least 0.3 in most languages.
    confidence: high
    relevance: medium
  - claim_id: codec_representations_enriched_with_semantic_and_contextual_signals_support_stronger
    role: supports
    claim: Codec representations enriched with semantic and contextual signals support stronger downstream task
      generalization (emotion recognition, audio event classification) than acoustic-only codecs.
    source: §3.2.2, Table 3
    evidence: On CodecSUPERB at 4 kbps, FuseCodec-Fusion achieves emotion recognition accuracy of 73.96% and audio
      signal quality 0.785, versus 66.18% and 0.697 for EnCodec at 6 kbps. All FuseCodec variants exceed SpeechTokenizer,
      EnCodec, and DAC on audio signal quality.
    confidence: high
    relevance: low
  - claim_id: fine_grained_temporal_alignment_between_text_tokens_and_acoustic_frames
    role: complicates
    claim: Fine-grained temporal alignment between text tokens and acoustic frames improves local interpretability
      but is constrained relative to global supervision strategies.
    source: §2.3.3, §3.2.1, Table 2
    evidence: FuseCodec-ContextAlign, which aligns contextual embeddings to RVQ tokens via a windowed similarity
      matching algorithm, achieves WER 4.15 and ViSQOL 3.18, lagging FuseCodec-Fusion (WER 3.99, ViSQOL 3.47) and
      FuseCodec-Distill (ViSQOL 3.43). The paper attributes this to constrained local alignment limiting global
      contextual guidance.
    confidence: high
    relevance: medium
  - claim_id: distilling_both_semantic_self_supervised_speech_model_and_contextual_language
    role: supports
    claim: Distilling both semantic (self-supervised speech model) and contextual (language model) signals into
      codec token supervision outperforms semantic-only distillation for perceptual naturalness.
    source: §3.2.1, Table 2
    evidence: FuseCodec-Distill achieves UTMOS 3.65 and Similarity 0.996 on LibriSpeech test-clean, while codecs
      using only semantic distillation (SpeechTokenizer, Mimi, X-Codec2) score below 3.55 UTMOS and fail to consistently
      match speaker similarity.
    confidence: high
    relevance: high
  limitations:
  - The codec is trained on LibriSpeech train-clean-100 (100 hours, English read speech), which limits conclusions
    about robustness to spontaneous speech, diverse accents, or noisy conditions. Multilingual generalization results
    are promising but the training data and BERT model are English-only, leaving the mechanism behind cross-lingual
    transfer unclear. The ASR transcription step (wav2vec 2.0) introduces an error-prone intermediate representation
    during training; the effect of ASR errors on contextual embedding quality is not quantified. Model size is not
    reported, making it difficult to assess computational cost relative to baselines. The TTS evaluation compares
    only against other codec-based systems trained on LibriTTS and does not benchmark against the strongest current
    TTS models.
  caveats: []
- id: '2509.13068'
  published_date: "2025-09-16"
  entry_date: '2026-07-29'
  year: 2025
  venue: arXiv
  task:
  - TTS
  - codec
  - VC
  architecture:
  - autoregressive-LM
  - VAE
  relevance: medium
  evidence_role:
  - architecture_variant
  - acceleration_evidence
  - control_evidence
  current_role: active_evidence
  method_family:
  - autoregressive_ssl_conditioned_models
  - vae_and_quantized_ssl_representations
  claims:
  - claim_id: cascaded_residual_codec_architectures_can_enforce_attribute_disentanglement_through_structure
    role: supports
    claim: Cascaded residual codec architectures can enforce attribute disentanglement through structure rather
      than through adversarial training objectives.
    source: §2.1, §3.3.3, Table 3
    evidence: MSR-Codec achieves clean separation of timbre, prosody, and semantic content by having each stream
      operate on residuals from the previous stage, without adversarial disentanglement loss; VC experiments confirm
      independent manipulation of each attribute.
    confidence: high
    relevance: low
  - claim_id: explicit_prosodic_supervision_in_a_dedicated_codec_stream_promotes_measurable
    role: supports
    claim: Explicit prosodic supervision in a dedicated codec stream promotes measurable disentanglement of pitch
      from speaker identity.
    source: §2.1.2, §3.3.3, Table 3
    evidence: VQ1 (prosody stream) is trained with MSE loss against ground-truth F0 and spectral energy; prosody-only
      VC achieves low ΔF0,tar (12.3-14.2 Hz) while maintaining high SIM-src (0.59-0.64), confirming that prosody
      and timbre are independently manipulable.
    confidence: high
    relevance: low
  - claim_id: disentangled_codec_designs_can_achieve_competitive_speaker_similarity_at_lower
    role: supports
    claim: Disentangled codec designs can achieve competitive speaker similarity at lower bitrates than undifferentiated
      RVQ codecs.
    source: §3.3.1, Table 1
    evidence: MSR-Codec-424 achieves SPK-SIM 0.80 at 424 bps, higher than WavTokenizer (0.67 at 900 bps) and X-Codec
      (0.72 at 1000 bps), attributed to the time-invariant timbre stream which preserves speaker identity without
      scaling with utterance length.
    confidence: high
    relevance: low
  - claim_id: data_efficient_tts_systems_built_on_factorized_codec_representations_can
    role: supports
    claim: Data-efficient TTS systems built on factorized codec representations can achieve competitive intelligibility
      relative to larger models trained on more data.
    source: §3.3.2, Table 2
    evidence: The 0.2B MSR-Codec-524 TTS model trained on 45k hours achieves WER 3.07% on Seed-TTS-eval English,
      outperforming Llasa-1B trained on 250k hours (WER 3.22%) and FireRedTTS-0.4B trained on 150k hours (WER 3.82%).
    confidence: high
    relevance: low
  - claim_id: signal_fidelity_codec_metrics_stoi_pesq_and_speaker_similarity_diverge
    role: complicates
    claim: Signal-fidelity codec metrics (STOI, PESQ) and speaker similarity diverge at low bitrates, making holistic
      quality assessment difficult.
    source: §3.3.1, Table 1
    evidence: MSR-Codec-424 achieves the highest SPK-SIM (0.80) among codecs at comparable bitrates but lower STOI
      (0.84) and PESQ-WB (1.82) than some baselines, indicating that speaker preservation and signal-level fidelity
      are optimized differently by the multi-stream design.
    confidence: high
    relevance: low
  limitations:
  - No subjective listening test (MOS/MUSHRA) is reported for any condition; all quality comparisons rely on automatic
    metrics (UTMOS, STOI, PESQ, WER, SPK-SIM). Conclusions about perceived naturalness cannot be confirmed from
    the available data.
  - 'The VC evaluation protocol is small in scope: 8 target speakers from VCTK and 100 source utterances from LibriTTS.
    Generalisation to more diverse speakers, accents, or noisy conditions is not assessed. The FreGAN vocoder operates
    at 16 kHz, and the Mel-spectrogram-based pipeline may impose a quality ceiling relative to waveform-domain codecs.
    The TTS model is evaluated only on English; the codec was trained on Mandarin and English data but multilingual
    TTS capability is not demonstrated. Model size figures for the codec itself are not reported; only the TTS model
    size (0.2B) is provided.'
  caveats: []
- id: '2509.15626'
  published_date: "2025-09-19"
  entry_date: '2026-07-29'
  year: 2025
  venue: arXiv
  task:
  - TTS
  architecture:
  - VAE
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  - evaluation_caution
  current_role: active_evidence
  method_family:
  - vae_and_quantized_ssl_representations
  claims:
  - claim_id: impression_leakage_in_controllable_tts_arises_structurally_when_a_single
    role: supports
    claim: Impression leakage in controllable TTS arises structurally when a single reference utterance is used
      for both speaker identity and style conditioning during training.
    source: §2, §5.2, Table 2
    evidence: VIC-base achieves ∆V = 0.22 (significantly different from zero), confirming that the reference audio
      biases synthesized VI toward the reference's inherent voice impression even when GRL and high-rate dropout
      are applied.
    confidence: high
    relevance: medium
  - claim_id: decoupled_training_using_distinct_utterances_for_speaker_and_style_conditioning
    role: supports
    claim: Decoupled training using distinct utterances for speaker and style conditioning reduces style leakage
      without changing the underlying TTS architecture.
    source: §4.1, §5.2, Table 2
    evidence: VIC-dis uses a separate utterance from the same speaker for reference conditioning during training,
      reducing ∆V from 0.22 to 0.14 with a statistically significant improvement in RVI-MSE (0.61 to 0.51).
    confidence: high
    relevance: medium
  - claim_id: eliminating_the_speaker_reference_and_conditioning_solely_on_a_style
    role: supports
    claim: Eliminating the speaker reference and conditioning solely on a style vector achieves the strongest leakage
      reduction, at a moderate cost in speaker fidelity.
    source: §4.2, §5.2, Table 2
    evidence: VIC-srf reduces ∆V to 0.05 (not significantly different from zero), while SECS drops to 0.72, which
      remains above the cross-speaker bound of 0.63 but below the same-speaker bound of 0.81.
    confidence: high
    relevance: medium
  - claim_id: llm_based_tts_systems_conditioned_on_natural_language_prompts_are
    role: complicates
    claim: LLM-based TTS systems conditioned on natural language prompts are inadequate for fine-grained numerical
      voice impression control.
    source: §5.1, §5.2, Table 2, Table 3
    evidence: Qwen3-TTS (zero-shot) achieves VI-MSE of 0.82 vs. 0.39 for the VITS-based baseline; fine-tuning worsens
      controllability further (VI-MSE 0.87), and text-VI entanglement is confirmed empirically by punctuation biasing
      predicted impressions.
    confidence: high
    relevance: medium
  - claim_id: perceptual_voice_impression_annotation_can_achieve_inter_annotator_agreement_comparable
    role: supports
    claim: Perceptual voice impression annotation can achieve inter-annotator agreement comparable to other subjective
      speech tasks, supporting the validity of VI corpora.
    source: §3, Table 1
    evidence: Krippendorff's alpha averages 0.470 across 10 VI dimensions, exceeding reported agreement for speech
      emotion recognition (0.442) and singing voice preference (0.153).
    confidence: high
    relevance: medium
  limitations:
  - The LibriTTS-VI corpus contains only 130 manually annotated utterances from 130 speakers; the remaining LibriTTS-R
    annotations are estimated by the VIE trained with data augmentation. Several VI dimensions fall below the reliable
    inter-annotator agreement threshold (alpha below 0.667), particularly J) Cold-Warm (0.197) and D) Calm-Restless
    (0.295), limiting annotation reliability for these dimensions.
  - Evaluation is confined to the LibriTTS-R audiobook domain, which is relatively clean and controlled. Generalisation
    of both the VIE and the VIC systems to spontaneous or noisy speech is untested. The subjective MOS evaluation
    uses only two speakers and the authors note potential interface bias due to rating-scale anchoring. The comparison
    with Qwen3-TTS relies on a single external system evaluated in a zero-shot configuration for which the model
    was not explicitly designed.
  caveats: []
- id: '2509.16195'
  published_date: "2025-09-19"
  entry_date: '2026-07-29'
  year: 2025
  venue: arXiv
  task:
  - codec
  - VC
  architecture:
  - hybrid
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  - evaluation_caution
  current_role: active_evidence
  method_family:
  - hybrid_semantic_acoustic_ssl_systems
  claims:
  - claim_id: single_codebook_binary_quantization_at_sub_1_kbps_bitrates_can
    role: supports
    claim: Single-codebook binary quantization at sub-1 kbps bitrates can match or exceed multi-codebook streaming
      codecs on naturalness and intelligibility in speech resynthesis.
    source: §4.1, Table 2
    evidence: FocalCodec-S@50-65k achieves UTMOS 3.85 and dWER 3.68% at 0.80 kbps with a single codebook of 65,536
      entries, outperforming Mimi6 (0.83 kbps, 6 codebooks, UTMOS 3.44, dWER 4.77%) on both metrics.
    confidence: high
    relevance: low
  - claim_id: multi_stage_causal_distillation_of_self_supervised_speech_encoders_preserves
    role: supports
    claim: Multi-stage causal distillation of self-supervised speech encoders preserves hybrid acoustic-semantic
      representations for downstream tasks under streaming constraints.
    source: §4.2, Table 3
    evidence: FocalCodec-Stream variants trained via four-stage WavLM distillation outperform acoustic streaming
      codecs (EnCodec, AudioDec, HILCodec) on ASR, keyword spotting, and intent classification despite operating
      at lower bitrates; the 65k variant matches or surpasses PAST on all discriminative and generative tasks except
      ASR.
    confidence: high
    relevance: high
  - claim_id: supervised_domain_specific_fine_tuning_of_hybrid_codecs_achieves_strong
    role: complicates
    claim: Supervised domain-specific fine-tuning of hybrid codecs achieves strong in-domain intelligibility at
      the cost of multilingual generalization.
    source: §4.1
    evidence: PAST, fine-tuned on English data, achieves the lowest English dWER among streaming codecs (4.04%)
      but degrades severely on multilingual MLS (49.35% dWER), whereas FocalCodec-Stream (trained on English-only
      Libri-Light but without supervised task fine-tuning) retains competitive multilingual performance (19.88%
      dWER).
    confidence: high
    relevance: medium
  - claim_id: voice_conversion_quality_in_streaming_codecs_requires_joint_optimization_of
    role: supports
    claim: Voice conversion quality in streaming codecs requires joint optimization of intelligibility and speaker
      fidelity; gains on one metric alone are insufficient for practical use.
    source: §4.1, Table 2
    evidence: Mimi6 achieves competitive speaker similarity (91.3%) in one-shot VC on VCTK but at 110% dWER, while
      PAST achieves lower dWER (18.28%) at only 68.5% speaker similarity; FocalCodec-S@50-65k is the only streaming
      codec to simultaneously achieve high values on both (dWER 22.71%, Sim 92.5%).
    confidence: high
    relevance: low
  - claim_id: a_lightweight_refiner_module_bridging_causal_and_full_context_feature
    role: supports
    claim: A lightweight refiner module bridging causal and full-context feature distributions substantially improves
      perceptual quality in distilled streaming codecs.
    source: §4.3, Table 4
    evidence: Ablation on FocalCodec-S@50-4k shows that removing the refiner degrades UTMOS from 3.87 to 3.84 and
      dWER from 4.39% to 4.65%; omitting Stage 4 fine-tuning (which jointly trains the refiner) has a larger effect,
      raising dWER to 5.05% and reducing speaker similarity from 96.3% to 95.8%.
    confidence: high
    relevance: low
  limitations:
  - Training data is limited to English (LibriTTS, Libri-Light), so the multilingual robustness observed on MLS
    reflects generalization rather than explicit multilingual training. A performance gap with the non-streaming
    FocalCodec@50 remains across most metrics, particularly in ASR WER (17% vs. 15.33%) and SI error rate (2.18%
    vs. 0.35%), reflecting the inherent cost of the 80 ms latency budget. Downstream generative tasks (TTS, speech
    language modeling) are left for future work, so performance of the codec's discrete representations in autoregressive
    modeling pipelines is not yet demonstrated.
  - The paper does not provide listening tests or crowd-sourced MOS, relying instead on the automatic UTMOS predictor
    for naturalness assessment.
  caveats: []
- id: '2509.20378'
  published_date: "2025-09-20"
  entry_date: '2026-07-29'
  year: 2025
  venue: arXiv
  task:
  - TTS
  architecture:
  - autoregressive-LM
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: active_evidence
  method_family:
  - autoregressive_ssl_conditioned_models
  claims:
  - claim_id: sub_sentence_emotion_conditioning_produces_more_accurate_emotional_dynamics_in
    role: supports
    claim: Sub-sentence emotion conditioning produces more accurate emotional dynamics in TTS than sentence-level
      conditioning.
    source: §4.1, Table 1
    evidence: Emo-FiLM outperforms CosyVoice2 by 9.1% on DTW and 12.7% on ESD DTW, with higher EMOS on both test
      sets, when comparing word-level emotion labels against global prompt conditioning.
    confidence: high
    relevance: medium
  - claim_id: feature_wise_linear_modulation_is_an_effective_mechanism_for_injecting
    role: supports
    claim: Feature-wise Linear Modulation is an effective mechanism for injecting word-level conditioning signals
      into the text representations of autoregressive LLM-TTS systems.
    source: §4.2, Table 2
    evidence: Ablation replacing the FiLM layer with simple addition increases FEDD DTW from 49.62 to 70.5, demonstrating
      that the non-linear affine modulation is essential beyond mere feature fusion.
    confidence: high
    relevance: medium
  - claim_id: self_supervised_speech_emotion_models_provide_useful_word_aligned_supervision
    role: supports
    claim: Self-supervised speech emotion models provide useful word-aligned supervision for fine-grained emotional
      TTS when combined with forced alignment.
    source: §2.1, §4.2, Table 2
    evidence: emotion2vec frame-level features aligned via MFA and pooled to word boundaries form the annotation
      backbone; removing word-level data tuning causes the most severe degradation (FEDD DTW 49.62 to 133.97).
    confidence: high
    relevance: high
  - claim_id: frame_averaged_emotion_similarity_metrics_are_insufficient_for_evaluating_intra
    role: complicates
    claim: Frame-averaged emotion similarity metrics are insufficient for evaluating intra-utterance emotional dynamics
      in TTS output.
    source: §3.4
    evidence: The authors note that Emo SIM averages frame-level emotion vectors, potentially obscuring dynamic
      information, and introduce DTW as a complementary metric that more sensitively distinguishes the models' ability
      to track emotional transitions.
    confidence: high
    relevance: medium
  - claim_id: explicit_emotion_classification_as_an_auxiliary_training_objective_improves_fine
    role: supports
    claim: Explicit emotion classification as an auxiliary training objective improves fine-grained emotion control
      in multi-task TTS training.
    source: §4.2, Table 2
    evidence: Removing the emotion classification loss increases FEDD DTW from 49.62 to 73.96, a degradation comparable
      to removing the FiLM layer entirely.
    confidence: high
    relevance: medium
  limitations:
  - The FEDD dataset is small (1,000 utterances, 5 speakers) and constructed partly via concatenation of emotionally
    distinct segments, which may not capture naturally occurring emotional transitions. The method is not evaluated
    for spontaneous or conversational speech, where emotion boundaries are less well-defined. The backbone (CosyVoice2)
    is frozen, so the approach inherits any limitations of that system in speaker diversity or prosodic range. Model
    size and speaker generalization are not reported.
  caveats: []
- id: '2509.17006'
  published_date: "2025-09-21"
  entry_date: '2026-07-29'
  year: 2025
  venue: arXiv
  task:
  - codec
  architecture:
  - GAN
  relevance: medium
  evidence_role:
  - architecture_variant
  current_role: active_evidence
  method_family:
  - gan_decoders_for_ssl_units
  claims:
  - claim_id: explicit_semantic_acoustic_disentanglement_in_neural_codecs_improves_reconstruction_quality
    role: supports
    claim: Explicit semantic-acoustic disentanglement in neural codecs improves reconstruction quality at comparable
      bitrates relative to conventional residual quantization approaches.
    source: §3.2, Table 1
    evidence: MBCodec (16 codebooks, 50Hz, 8.8kbps) achieves PESQ 3.83 and MUSHRA 85.9, versus EnCodec at the same
      bitrate scoring PESQ 2.78 and MUSHRA 85.3; at 4.4kbps, MBCodec (16 codebooks, 25Hz) scores PESQ 3.64 versus
      SpeechTokenizer at PESQ 1.26 and MUSHRA 79.0.
    confidence: high
    relevance: medium
  - claim_id: pqmf_based_frequency_subband_supervision_substantially_improves_spectral_reconstruction_fidelity
    role: supports
    claim: PQMF-based frequency subband supervision substantially improves spectral reconstruction fidelity in multi-codebook
      codec training.
    source: §3.4, Table 3
    evidence: Ablation shows removing PQMF supervision degrades PESQ from 3.83 to 2.34 (a 39% drop) and increases
      Mel-spectrogram distance from 2.34 to 3.45 in the best MBCodec configuration.
    confidence: high
    relevance: low
  - claim_id: non_uniform_codebook_dropout_distributions_that_concentrate_sampling_on_early
    role: supports
    claim: Non-uniform codebook dropout distributions that concentrate sampling on early RVQ layers outperform uniform
      dropout, better matching the hierarchical information density of residual quantization.
    source: §2.2, §3.4, Table 2, Table 3
    evidence: Half-Gaussian adaptive dropout achieves PESQ 3.28 versus exponential decay at 3.02 and chi-squared
      at 2.75; removing adaptive dropout entirely drops SI-SDR from 7.94 to 7.32 and increases STFT distance from
      0.08 to 0.22.
    confidence: high
    relevance: medium
  - claim_id: ultra_low_bitrate_neural_audio_compression_imposes_a_persistent_quality
    role: complicates
    claim: Ultra-low bitrate neural audio compression imposes a persistent quality gap relative to ground truth,
      even with strong disentanglement strategies.
    source: §3.2, Table 1
    evidence: MBCodec at 2.2kbps achieves MUSHRA 82.8 versus ground truth MUSHRA 90.7 (a gap of 7.9 points) and
      PESQ 2.98 versus 4.64, despite the 170x compression ratio and explicit semantic-acoustic disentanglement.
    confidence: high
    relevance: medium
  limitations:
  - The paper does not evaluate MBCodec as a token representation for downstream TTS or speech LM inference, which
    is presented as a primary motivation. The codec is trained and evaluated on speech-dominant data; generalisation
    to music or general audio is not assessed. No parameter count is reported, making direct efficiency comparison
    with other codecs difficult. The comparison set is limited to DAC, EnCodec, and SpeechTokenizer; more recent
    codecs such as Mimi or BiCodec are not included. As a preprint, no independent reproduction or external listening
    evaluation is available.
  caveats: []
- id: '2509.17143'
  published_date: "2025-09-21"
  entry_date: '2026-07-29'
  year: 2025
  venue: arXiv
  task:
  - VC
  architecture:
  - autoregressive-LM
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: active_evidence
  method_family:
  - autoregressive_ssl_conditioned_models
  claims:
  - claim_id: in_zero_shot_voice_conversion_temporally_coarser_syllabic_representations_reduce
    role: supports
    claim: In zero-shot voice conversion, temporally coarser syllabic representations reduce pitch leakage from
      linguistic features more effectively than standard frame-aligned SSL features, enabling cleaner prosody control
      at the cost of intelligibility.
    source: §2.1, §4, Table 2
    evidence: MaskVCT-Spk using SylBoost syllabic tokens achieves the lowest FPC (0.167) among all tested systems,
      indicating near-complete pitch independence from the source, while MaskVCT-All with continuous features retains
      more pitch correlation (FPC 0.417).
    confidence: high
    relevance: high
  - claim_id: multiple_classifier_free_guidance_weights_applied_to_distinct_conditioning_factors
    role: supports
    claim: Multiple classifier-free guidance weights applied to distinct conditioning factors in a single masked
      generative model enable user-configurable inference-time trade-offs between intelligibility, pitch fidelity,
      and speaker similarity.
    source: §2.5, §3.4, Table 2
    evidence: MaskVCT defines triple CFG weights (w_all, w_spk, w_ling) over speaker, pitch, and linguistic conditions
      within one trained model; sweeping these weights continuously interpolates between MaskVCT-All (WER 4.68%,
      S-SIM 0.865) and MaskVCT-Spk (WER 6.47%, S-SIM 0.895).
    confidence: high
    relevance: low
  - claim_id: syllabic_speech_representations_that_suppress_pitch_leakage_in_voice_conversion
    role: complicates
    claim: Syllabic speech representations that suppress pitch leakage in voice conversion also degrade content
      intelligibility through syllable misreadings caused by coarse temporal quantisation.
    source: §4, §5, Table 2
    evidence: MaskVCT-Spk achieves the highest speaker similarity (S-SIM 0.895) but the highest WER (6.47%) among
      systems tested, substantially above FACodec (3.55%) and FreeVC (3.96%); the conclusion section attributes
      misreadings to K-means syllable mapping errors in SylBoost.
    confidence: high
    relevance: low
  - claim_id: masked_non_autoregressive_codec_models_can_match_or_exceed_autoregressive
    role: supports
    claim: Masked non-autoregressive codec models can match or exceed autoregressive and diffusion-based baselines
      on speaker similarity and quality in zero-shot VC while operating with fewer discrete tokens per utterance.
    source: §3.4, §4, Table 2
    evidence: MaskVCT-Spk (2048 tokens) achieves higher S-SIM (0.895) and SS-MOS (3.69) than MaskGCT-S2A (8192 tokens,
      S-SIM 0.863, SS-MOS 3.02) and competitive UTMOS (3.17 vs. 3.24).
    confidence: high
    relevance: low
  limitations:
  - Syllabic tokens introduce misreadings that WER alone cannot fully diagnose; the authors acknowledge the K-means
    quantiser cannot recover from incorrect syllable boundary assignments, and propose future work to address this
    with a trainable VQ module.
  - The model is English-only. Accent conversion experiments are limited to L2-ARCTIC and test only two conversion
    directions; how well the syllabic pitch-stripping generalises to tonal languages (where pitch is phonemic) is
    unexplored. The 511-pair test set is relatively small for statistical confidence, especially given the reported
    confidence intervals overlap for several key comparisons. No code or trained checkpoint is publicly released,
    limiting reproducibility.
  caveats: []
- id: '2509.19928'
  published_date: "2025-09-24"
  entry_date: '2026-07-29'
  year: 2025
  venue: arXiv
  task:
  - TTS
  - evaluation
  architecture: []
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  - infrastructure
  current_role: active_evidence
  method_family: []
  claims:
  - claim_id: acoustic_proxy_metrics_for_prosodic_variation_correlate_weakly_with_human
    role: supports
    claim: Acoustic proxy metrics for prosodic variation correlate weakly with human perception, while distance
      measures computed over discretized self-supervised speech tokens correlate substantially better.
    source: §4.1, Table 1
    evidence: Averaged Pearson correlation with human PMOS ratings is r̄ = 0.30 for log F0 RMSE and r̄ = 0.66 for
      MCD, versus r̄ = 0.77 for the proposed token-edit-distance metric, aggregated via Fisher's Z transformation
      across 1000 samples and 2000 ratings.
    confidence: high
    relevance: high
  - claim_id: autoregressive_generation_provides_greater_output_diversity_in_prosody_than_non
    role: supports
    claim: Autoregressive generation provides greater output diversity in prosody than non-autoregressive flow-matching
      models with implicit text-speech alignment, but this advantage does not extend to non-autoregressive masked
      generative modeling.
    source: §4.4-4.5, Table 2
    evidence: On the DS-WED benchmark across LibriSpeech test-clean and Seed-TTS test-en, three AR systems (XTTS-v2,
      CosyVoice, CosyVoice 2) outperform three flow-matching NAR systems (E2 TTS, F5-TTS, ZipVoice), while the masked
      generative modeling system (MaskGCT) surpasses all AR systems on LibriSpeech and remains competitive on Seed-TTS
      despite training on the same Emilia corpus as the flow-matching systems.
    confidence: high
    relevance: medium
  - claim_id: explicit_duration_control_during_inference_is_a_significant_but_not
    role: complicates
    claim: Explicit duration control during inference is a significant, but not sufficient, factor in restoring
      prosodic diversity to non-autoregressive TTS systems with implicit alignment.
    source: §4.5, Table 3
    evidence: Applying duration perturbation (0.8-1.2x) to two flow-matching/MGM NAR systems increases DS-WED diversity
      by 13.8-28.5%, but the perturbed flow-matching system (F5-TTS) still lags behind AR and MGM systems evaluated
      without perturbation, suggesting the shortfall is architectural (implicit alignment) rather than purely a
      missing duration-control step.
    confidence: high
    relevance: medium
  - claim_id: preference_optimization_that_targets_one_quality_dimension_intelligibility_can_measurably
    role: complicates
    claim: Preference optimization that targets one quality dimension (intelligibility) can measurably suppress
      output diversity along an unrelated axis (prosody) as a side effect.
    source: §4.5, Table 4
    evidence: Applying DPO for intelligibility to CosyVoice 2 and MaskGCT reduces DS-WED prosody-diversity scores
      by 18.8% and 2.9-3.2% respectively across two test sets, with no diversity-specific reward term in the alignment
      objective.
    confidence: high
    relevance: medium
  - claim_id: a_general_purpose_large_audio_language_model_with_strong_multimodal
    role: contradicts
    claim: A general-purpose large audio language model with strong multimodal reasoning capability can serve as
      a reliable automatic judge of fine-grained prosodic variation.
    source: §4.5, Table 5
    evidence: Gemini 2.5 Pro, prompted to rate relative prosodic difference within groups of five samples, achieves
      only a weak correlation with human PMOS ratings (r̄ = 0.27) and unstable, wide-confidence-interval correlations
      with objective acoustic and token-based metrics.
    confidence: high
    relevance: medium
  limitations:
  - The authors state that DS-WED has only been validated on English speech; cross-lingual applicability is untested.
    The benchmark itself covers only seven open-source systems and two evaluation corpora (LibriSpeech test-clean,
    Seed-TTS test-en), so its paradigm-level conclusions (AR vs. flow-matching vs. MGM) rest on a small number of
    representative systems per category rather than an exhaustive sweep. The DPO and duration-perturbation exploration
    each cover only two systems, limiting how confidently the reported trends generalize across the broader space
    of NAR TTS architectures. The paper also does not report whether code or the ProsodyEval dataset itself will
    be released beyond the audio-sample demo page, which affects reproducibility of the correlation analysis.
  caveats: []
- id: '2509.20485'
  published_date: "2025-09-24"
  entry_date: '2026-07-29'
  year: 2025
  venue: arXiv
  task:
  - evaluation
  - TTS
  architecture: []
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  - evaluation_caution
  current_role: active_evidence
  method_family: []
  claims:
  - claim_id: conditioning_a_discrete_token_likelihood_model_on_the_input_text
    role: supports
    claim: Conditioning a discrete-token likelihood model on the input text produces a substantially more discriminative
      and human-aligned intelligibility signal than unconditional token language modeling.
    source: §V.A, Table 1
    evidence: TTScore-int reaches utterance-level correlation with WER up to 0.78 on VoiceMOS22 (system-level up
      to 0.96), compared to at most 0.44/0.55 for an identically-trained unconditional token-LM (uLM) baseline and
      0.32/0.42 for SpeechLMScore.
    confidence: high
    relevance: medium
  - claim_id: mos_predictors_trained_directly_on_mos_labeled_data_generalize_poorly
    role: supports
    claim: MOS predictors trained directly on MOS-labeled data generalize poorly across benchmarks relative to reference-free
      metrics with a narrower, aspect-specific training objective.
    source: §V.B, Table 2
    evidence: UTMOS, trained on VoiceMOS22 data, achieves high correlation with MOS on VoiceMOS22 but notably lower
      correlation on the out-of-distribution SOMOS dataset, while TTScore-int retains stronger correlation with
      MOS on SOMOS.
    confidence: high
    relevance: medium
  - claim_id: reference_free_text_conditioned_prosody_scoring_can_track_prosodic_appropriateness
    role: supports
    claim: Reference-free, text-conditioned prosody scoring can track prosodic appropriateness where reference-dependent
      F0-based metrics fail, including on benchmarks lacking aligned reference audio.
    source: §VI.C, Tables 3-5
    evidence: TTScore-pro shows positive correlation with MOS on both SOMOS and VoiceMOS22 and correct-sign correlation
      with TTS Arena ELO scores, while F0-RMSE shows near-zero or negative correlation with MOS and an unexpected
      positive (wrong-sign) correlation with ELO.
    confidence: high
    relevance: medium
  - claim_id: objective_prosody_metrics_that_isolate_a_single_acoustic_dimension_pitch
    role: complicates
    claim: Objective prosody metrics that isolate a single acoustic dimension (pitch) correlate only moderately
      with overall perceived naturalness, since naturalness depends on additional prosodic factors beyond F0.
    source: §VI.C, Table 3
    evidence: TTScore-pro's correlations with MOS remain in the 0.2-0.5 range even where it clearly outperforms
      F0-RMSE and F0-correlation baselines, and the paper attributes the residual gap to prosody encompassing rhythm,
      energy, and phrasing beyond the F0-based prosody tokens used here.
    confidence: high
    relevance: medium
  limitations:
  - The method targets only F0-related prosody; rhythm, energy, and phrasing, which the authors acknowledge are
    also core prosodic dimensions, are not modeled or evaluated.
  - The text-to-token generators are trained exclusively on a large-scale, clean, English-only corpus (LibriSpeech-960),
    so the metric's reliability in noisy, spontaneous, or non-English settings is untested. Because the token distributions
    are learned from data, the authors note the approach may inherit whatever biases exist in that training corpus.
    The paper also does not report whether code will be released, and the demo/code availability fields could not
    be confirmed from the paper text.
  caveats: []
- id: '2509.21968'
  published_date: "2025-09-26"
  entry_date: '2026-07-29'
  year: 2025
  venue: arXiv
  task:
  - codec
  architecture:
  - GAN
  - VAE
  relevance: medium
  evidence_role:
  - architecture_variant
  - evaluation_caution
  current_role: active_evidence
  method_family:
  - gan_decoders_for_ssl_units
  - vae_and_quantized_ssl_representations
  claims:
  - claim_id: a_shared_single_codebook_can_be_structured_with_overlapping_nested
    role: supports
    claim: A shared single codebook can be structured with overlapping, nested domain partitions rather than rigid
      disjoint splits, improving both reconstruction and downstream generation quality relative to a rigid-split
      design.
    source: §4.4, Tables 3–4
    evidence: Ablating codebook design at fixed codebook size and data scale, the nested codebook reduces reconstruction
      WER from 4.21 (rigid-split) to 3.99 and downstream TTS generation WER from 6.26 to 4.99 on LibriSpeech-PC
      test-clean.
    confidence: high
    relevance: medium
  - claim_id: distilling_frame_level_representations_from_multiple_domain_specific_self_supervised
    role: supports
    claim: Distilling frame-level representations from multiple domain-specific self-supervised teacher models into
      one acoustic codec can improve reconstruction and generation quality across all the covered domains simultaneously,
      not just the domain of a single teacher.
    source: §3.3, §4.4, Tables 3–4
    evidence: Adding multi-domain distillation (WavLM for speech, MuQ for vocal/music, BEATs for sound) on top of
      the nested codebook further raises the speech-partition code-usage ratio from 37.1% to 38.2% and improves
      downstream TTS generation WER from 4.99 to 4.51.
    confidence: high
    relevance: high
  - claim_id: unifying_multiple_audio_domains_into_a_single_shared_quantization_codebook
    role: complicates
    claim: Unifying multiple audio domains into a single, shared quantization codebook does not close the gap with
      domain-specific single-layer codecs on every reconstruction metric, even when the unified model uses a larger
      codebook and lower token rate.
    source: §4.2, Table 1
    evidence: On LibriSpeech test-clean, AUV's PESQ-WB (2.40) and SPK-SIM (0.81) trail dedicated speech codecs such
      as DAC (4.01, 0.95) and are roughly on par with, not clearly better than, BigCodec and UniCodec.
    confidence: high
    relevance: medium
  - claim_id: larger_unified_codebooks_intended_to_accommodate_more_audio_domains_can
    role: complicates
    claim: Larger unified codebooks intended to accommodate more audio domains can hurt downstream autoregressive
      generation quality by increasing the modeling burden on the generative model consuming the tokens, even when
      reconstruction quality is unaffected.
    source: §4.3, Table 4
    evidence: Scaling the codebook from 16,384 to 20,480 entries (C0 vs. C2) yields slightly worse downstream generation
      WER (4.51 vs. 4.89) despite improved reconstruction metrics; training an autoregressive model on codes from
      a still-larger 131,072-entry codebook (MagiCodec) reportedly failed outright.
    confidence: high
    relevance: medium
  limitations:
  - The paper reports no total parameter count for the AUV encoder-decoder, limiting direct efficiency comparison
    with baselines whose sizes are known. Downstream generative evaluation is restricted to a single autoregressive
    TTS backbone (EmoVoice) trained on a comparatively small 1K-hour subset of LibriSpeech, so it is unclear whether
    the reconstruction and generation gains hold at larger downstream training scales or with non-autoregressive
    generators. Speaker similarity in the downstream TTS setting remains low in absolute terms (SPK-SIM ≈ 0.43–0.44)
    across all codecs tested, including AUV, suggesting the codec-level improvements shown here do not yet translate
    into strong speaker fidelity for generated speech. The domain-label input used during training but withheld
    at inference creates a train/inference mismatch whose effect on codebook index selection is only indirectly
    probed via the reported index-distribution statistics, not directly ablated.
  caveats: []
- id: '2509.22167'
  published_date: "2025-09-26"
  entry_date: '2026-07-29'
  year: 2025
  venue: arXiv
  task:
  - TTS
  - codec
  architecture:
  - VAE
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: active_evidence
  method_family:
  - vae_and_quantized_ssl_representations
  claims:
  - claim_id: regularizing_a_continuous_vae_latent_space_toward_semantic_structure_derived
    role: supports
    claim: Regularizing a continuous VAE latent space toward semantic structure derived from self-supervised speech
      features mitigates the reconstruction-generation trade-off in latent-based non-autoregressive TTS.
    source: §4.1, Table 1
    evidence: At a fixed 64-dimensional latent size, adding cosine-similarity alignment to WavLM features reduces
      WER from 2.65% (vanilla VAE) to 2.10% while raising speaker similarity from 0.59 to 0.64 on LibriSpeech-PC
      test-clean.
    confidence: high
    relevance: high
  - claim_id: semantic_alignment_regularization_of_the_generation_target_accelerates_training_convergence
    role: supports
    claim: Semantic alignment regularization of the generation target accelerates training convergence of latent
      diffusion/flow-matching TTS models.
    source: §4.1, Figure 3
    evidence: Training-step comparisons show Semantic-VAE-based F5-TTS reaching lower WER and higher SIM than both
      the vanilla-VAE and mel-spectrogram baselines at the same number of training steps.
    confidence: high
    relevance: medium
  - claim_id: the_intelligibility_versus_speaker_similarity_trade_off_in_vae_latent
    role: refines
    claim: The intelligibility-versus-speaker-similarity trade-off in VAE latent representations is governed by
      which layer of a self-supervised model is used for semantic supervision, not eliminated outright by adding
      semantic alignment.
    source: §4.3, Table 3
    evidence: Ablating SSL layer choice shows the final layer of WavLM/HuBERT substantially degrades SIM (0.50-0.58)
      despite comparable WER, while an intermediate layer (WavLM layer 23) gives the best joint WER/SIM balance.
    confidence: high
    relevance: high
  - claim_id: cosine_similarity_alignment_losses_to_self_supervised_speech_features_preserve
    role: supports
    claim: Cosine-similarity alignment losses to self-supervised speech features preserve useful semantic structure
      in a latent space more effectively than L1 or L2 distance losses to the same features.
    source: §4.3, Table 3
    evidence: Negative cosine alignment achieves 2.10% WER / 0.64 SIM, compared to 3.12% WER / 0.47 SIM for L1 alignment
      and 4.37% WER / 0.48 SIM for L2 alignment under otherwise identical settings.
    confidence: high
    relevance: high
  - claim_id: adding_a_semantic_regularization_objective_to_a_vae_s_training
    role: complicates
    claim: Adding a semantic regularization objective to a VAE's training loss does not necessarily degrade the
      representation's reconstruction fidelity, contrary to the general expectation that additional regularization
      terms cost reconstruction quality.
    source: §4.2, Table 2
    evidence: Semantic-VAE reconstruction metrics on LibriTTS test-other (PESQ 3.74, STOI 0.96, UTMOS 3.56) are
      nearly identical to an unregularized vanilla VAE of the same architecture (PESQ 3.75, STOI 0.97, UTMOS 3.57).
    confidence: high
    relevance: medium
  limitations:
  - The high-resource baseline numbers in Table 1 (CosyVoice, FireRedTTS, E2 TTS, F5-TTS at 100k-580k training hours)
    are quoted directly from their original papers rather than re-run under matched data and compute, so any comparison
    against Semantic-VAE's low-resource (0.6k-hour) results is illustrative context only, not a controlled comparison.
    The method is validated on English speech only, at a single latent configuration (64 dimensions, 40Hz), and
    on two NAR backbones (F5-TTS, E2 TTS); generalization to AR codec-token TTS systems, other languages, or other
    latent dimensionalities/frame rates is untested. The paper also does not report inference latency or discuss
    how the added SSL model changes deployment cost of the VAE training pipeline (inference itself is unaffected,
    since the SSL branch is training-only).
  caveats: []
- id: '2509.24773'
  published_date: "2025-09-29"
  entry_date: '2026-07-29'
  year: 2025
  venue: arXiv
  task:
  - TTS
  architecture:
  - flow-matching
  relevance: medium
  evidence_role:
  - architecture_variant
  current_role: active_evidence
  method_family:
  - flow_matching_with_ssl_conditioning
  claims:
  - claim_id: joint_training_of_two_video_conditioned_audio_generation_tasks_with
    role: supports
    claim: Joint training of two video-conditioned audio generation tasks with different output characteristics
      does not need to degrade either task's performance, provided the conditioning pathways for each task's signals
      are architecturally separated.
    source: §4.4, Figure 4
    evidence: Learning curves comparing joint (V2S+VisualTTS), sound-only, and speech-only training at matched epochs
      show the joint-trained model reaches performance parity with each single-task model on FAD/DeSync (sound)
      and WER/UTMOS (speech), directly contrasting with a prior report that joint training of these tasks causes
      mutual degradation.
    confidence: high
    relevance: medium
  - claim_id: the_choice_between_cross_attention_and_channel_concatenation_conditioning_in
    role: refines
    claim: The choice between cross-attention and channel-concatenation conditioning in a DiT-based generative model
      should be determined by whether the conditioning signal is temporally dense/aligned or sparse/global, not
      fixed uniformly across all condition types.
    source: §4.3, Figure 3
    evidence: A 5-variant ablation shows concatenation-based conditioning improves WER and lip-sync error for temporally-dense
      signals (transcript, sound/speech sync cues), while cross-attention conditioning improves FAD and video-audio
      semantic alignment for the sparse, global video-semantic feature; self-attention weight visualization shows
      a diagonal locality bias explaining the concatenation result.
    confidence: high
    relevance: medium
  - claim_id: feature_level_synthesis_of_paired_training_examples_constructed_by_mixing
    role: supports
    claim: Feature-level synthesis of paired training examples, constructed by mixing independently-sourced single-modality
      data in representation space, can substitute for scarce real multimodal joint-generation training data.
    source: §4.4, Table 4
    evidence: Fine-tuning on synthetic sound-speech mixtures alone yields the best WER (19.4) and UTMOS (2.07) on
      the V2C-Animation joint-generation benchmark among all training-data configurations tested, exceeding fine-tuning
      on the small (8k-sample) real V2C corpus alone (36.5 WER).
    confidence: high
    relevance: medium
  - claim_id: unified_video_conditioned_generation_systems_that_depend_on_frozen_pretrained
    role: complicates
    claim: Unified video-conditioned generation systems that depend on frozen pretrained feature extractors for
      their conditioning signals inherit those extractors' representational limits, and systems adapted with synthetic
      joint data remain bounded by the distributional gap between synthetic and real joint data.
    source: §Limitations
    evidence: The authors state that reliance on frozen pretrained extractors (CLIP, Synchformer, AV-HuBERT) may
      limit capture of fine-grained audio-visual nuances, and that generation quality on real joint scenarios is
      inherently constrained by the distributional gap of the synthetic training data used for joint-generation
      fine-tuning.
    confidence: high
    relevance: medium
  limitations:
  - Beyond the frozen-extractor and synthetic-data-gap limitations the authors state directly, the joint sound-speech
    generation results are evaluated on a single benchmark (V2C-Animation) drawn entirely from animated content,
    and the joint-generation comparison is against pipeline baselines composed by the authors rather than against
    other end-to-end unified systems, since none are publicly available for this exact three-way task. Code was
    not released at the time of publication ("Demos and code will be soon released" per the paper), though a project
    page with generated examples is available.
  caveats: []
- id: '2509.26276'
  published_date: "2025-09-30"
  entry_date: '2026-07-29'
  year: 2025
  venue: arXiv
  task:
  - TTS
  - SCA
  architecture:
  - autoregressive-LM
  relevance: medium
  evidence_role:
  - architecture_variant
  - control_evidence
  current_role: active_evidence
  method_family:
  - autoregressive_ssl_conditioned_models
  claims:
  - claim_id: acoustic_consistency_in_autoregressive_speech_language_models_can_be_improved
    role: supports
    claim: Acoustic consistency in autoregressive speech language models can be improved through LM-side embedding
      initialization and auxiliary training objectives, independent of parameter count.
    source: §4, Table 1
    evidence: The 0.7B speech-only CAST model scores 90.8 speaker and 90.0 gender consistency on SALMON, exceeding
      the 7B SpiritLM (81.0/85.0) and 7B Twist (71.0/70.0), despite using an order of magnitude fewer parameters
      and the same evaluation protocol.
    confidence: high
    relevance: medium
  - claim_id: interleaving_text_and_speech_tokens_in_a_shared_vocabulary_language
    role: complicates
    claim: Interleaving text and speech tokens in a shared-vocabulary language model improves semantic and lexical
      grounding at the cost of acoustic consistency.
    source: §4, Tables 1-2
    evidence: The 1.0B interleaved model drops SALMON consistency by 5-9 points across every measured factor relative
      to the 1.0B speech-only model (e.g., speaker 83.5 vs. 90.0) while sWUGGY rises from 67.0 to 73.7 and semantic-acoustic
      alignment scores improve by 6-7.5 points.
    confidence: high
    relevance: medium
  - claim_id: auxiliary_objectives_that_require_the_model_to_plan_coarse_content
    role: supports
    claim: Auxiliary objectives that require the model to plan coarse content structure before predicting fine acoustic
      detail materially contribute to speaker- and gender-identity stability beyond what embedding initialization
      alone provides.
    source: §4, Table 4
    evidence: Removing the delayed coarse-label and next-code auxiliary losses drops speaker consistency from 90.8
      to 83.5 in the 0.7B model and from 90.0 to 83.0 in the 1.0B model, with background and room consistency declining
      more moderately.
    confidence: high
    relevance: medium
  - claim_id: initializing_speech_token_embeddings_from_self_supervised_acoustic_features_biases
    role: refines
    claim: Initializing speech-token embeddings from self-supervised acoustic features biases the resulting representation
      toward content structure at the expense of fine-grained prosodic detail.
    source: §4, Table 3
    evidence: Linear probes on frozen LM features show the SSL-initialized variant gains on average +16% relative
      accuracy on content-leaning tasks (ESC-50, US8K, RAVDESS) while losing about 7% relative accuracy on prosody-leaning
      tasks (VIVAE, EMOVO), compared to a randomly initialized baseline.
    confidence: high
    relevance: high
  limitations:
  - The paper evaluates consistency exclusively through likelihood-based pairwise preference scoring (SALMON) rather
    than through generated-sample quality or human listening tests; no MOS or subjective naturalness evaluation
    is reported, so it is untested whether the consistency gains translate into perceptibly better generations.
    The comparison against Twist, SpiritLM, LAST, and Flow-SLM baselines is confounded by differences in tokenizer,
    training data, and architecture beyond scale, which the paper acknowledges only implicitly by noting these systems
    "differ in tokenization, architecture, and training objectives." The 1.0B ablation shows a single anomalous
    result for room consistency that the authors attribute to run instability rather than investigate further. All
    experiments use English-only, largely read-speech and audiobook data (LibriLight, People's Speech), leaving
    open whether the training recipe generalizes to conversational, multilingual, or noisier real-world speech.
  caveats: []
claim_clusters:
- id: masked_and_contrastive_pretraining_learn_reusable_speech_features
  claim: Masked and contrastive self-supervised objectives learn reusable contextual speech representations without
    task labels.
  status: strongly_supported
  confidence: high
  supporting_papers:
  - '2209.03143'
  - interspeech-2025-1776
  - '2411.19770'
  contradicting_papers: []
  refining_papers: []
  caveats:
  - Objective comparisons often change architecture, data, and augmentation together.
  last_reviewed: '2026-07-29'
- id: ssl_features_encode_linguistic_content
  claim: Self-supervised speech features encode phonetic, lexical, and semantic content useful across downstream
    tasks.
  status: strongly_supported
  confidence: high
  supporting_papers:
  - '2104.00355'
  - '2209.03143'
  - '2502.06490'
  - '2502.07243'
  - 2025.iwsds-1.11
  - '2505.13000'
  - '2506.10274'
  - 2025.acl-long.87
  - '2508.04141'
  - '2508.04996'
  - '2508.08399'
  - '2508.11224'
  - interspeech-2025-0166
  - interspeech-2025-0468
  - interspeech-2025-0998
  - interspeech-2025-1229
  - interspeech-2025-1440
  - interspeech-2025-1531
  - '2507.14534'
  - '2509.00503'
  - '2509.01391'
  - '2509.05359'
  - '2509.11425'
  - '2509.16195'
  - '2509.22167'
  contradicting_papers: []
  refining_papers:
  - '2402.13236'
  - 2025.findings-naacl.471
  - '2507.02176'
  - '2507.08012'
  - 2025.acl-long.87
  - interspeech-2025-0506
  - interspeech-2025-1106
  - interspeech-2025-1531
  - interspeech-2025-2043
  - '2509.22167'
  - '2509.26276'
  caveats:
  - Linguistic information varies substantially by model layer and pretraining corpus.
  last_reviewed: '2026-07-29'
- id: ssl_layer_choice_controls_information_content
  claim: Different SSL layers specialize in acoustic, phonetic, semantic, speaker, and task-specific information.
  status: strongly_supported
  confidence: high
  supporting_papers:
  - 2025.findings-naacl.130
  - '2505.13000'
  - '2506.10274'
  - '2508.06890'
  - interspeech-2025-0815
  - interspeech-2025-1639
  - '2509.09201'
  - '2509.17006'
  contradicting_papers: []
  refining_papers:
  - '2509.22167'
  caveats:
  - Layer probes can reflect the probe model and do not guarantee causal usefulness downstream.
  last_reviewed: '2026-07-29'
- id: speaker_and_prosody_leakage_persists_in_ssl_features
  claim: SSL representations retain varying amounts of speaker, pitch, prosody, and paralinguistic information despite
    content-oriented training.
  status: strongly_supported
  confidence: high
  supporting_papers:
  - '2104.00355'
  - '2402.08093'
  - 2025.coling-main.518
  - '2502.06490'
  - '2502.07243'
  - 2025.findings-naacl.130
  - 2025.naacl-short.65
  - 2025.iwsds-1.11
  - '2505.13000'
  - '2507.03912'
  - 2025.acl-long.87
  - '2503.11026'
  - '2508.04141'
  - '2508.04996'
  - '2508.08399'
  - '2508.11224'
  - '2508.11273'
  - interspeech-2025-0166
  - interspeech-2025-0468
  - interspeech-2025-0723
  - interspeech-2025-0815
  - interspeech-2025-1101
  - interspeech-2025-1236
  - interspeech-2025-1394
  - interspeech-2025-1440
  - interspeech-2025-1531
  - interspeech-2025-2075
  - '2411.19770'
  - '2509.20378'
  - '2509.17143'
  contradicting_papers: []
  refining_papers:
  - '2402.13236'
  - '2411.13577'
  - 2025.findings-naacl.471
  - '2507.08012'
  - 2025.acl-long.681
  - 2025.acl-long.87
  - '2508.08399'
  - '2508.11224'
  - interspeech-2025-0305
  - interspeech-2025-1106
  - interspeech-2025-1394
  - interspeech-2025-2043
  - interspeech-2025-2684
  - '2509.26276'
  caveats:
  - Whether leakage is harmful depends on whether the downstream task needs invariance or expressive reconstruction.
  last_reviewed: '2026-07-29'
- id: explicit_disentanglement_improves_controllable_generation
  claim: Explicitly separating content from speaker and style information improves controllable speech generation
    and conversion.
  status: strongly_supported
  confidence: high
  supporting_papers:
  - '2104.00355'
  - '2209.03143'
  - '2308.16692'
  - '2402.08093'
  - '2412.04724'
  - '2502.07243'
  - 2025.findings-naacl.130
  - 2025.acl-demo.37
  - 2025.acl-long.790
  - 2025.acl-long.87
  - '2508.04141'
  - '2508.04996'
  - '2508.06890'
  - '2508.08399'
  - interspeech-2025-0246
  - interspeech-2025-0438
  - interspeech-2025-1229
  - interspeech-2025-1394
  - interspeech-2025-1440
  - interspeech-2025-1531
  - interspeech-2025-1639
  - interspeech-2025-2684
  - '2509.09201'
  - '2509.17006'
  contradicting_papers: []
  refining_papers:
  - '2507.08012'
  - 2025.acl-long.681
  - 2025.acl-long.87
  - interspeech-2025-0115
  - interspeech-2025-0305
  - interspeech-2025-1106
  - interspeech-2025-1397
  - interspeech-2025-1531
  - interspeech-2025-1639
  - interspeech-2025-2043
  - '2509.09201'
  - '2509.17006'
  caveats:
  - Hard separation may discard accent, timing, and expressive cues that carry both content and identity.
  last_reviewed: '2026-07-29'
- id: discrete_ssl_units_enable_language_modeling
  claim: Discrete SSL-derived speech units make speech compatible with token language modeling and textless generation.
  status: strongly_supported
  confidence: high
  supporting_papers:
  - '2104.00355'
  - '2209.03143'
  - '2301.11325'
  - '2305.11000'
  - '2402.05755'
  - '2402.08093'
  - '2408.02622'
  - '2502.06490'
  - '2506.10274'
  - 2025.acl-long.682
  - 2025.acl-long.790
  - 2025.acl-long.817
  - 2025.findings-acl.75
  - '2508.11273'
  - interspeech-2025-0246
  - interspeech-2025-2684
  - '2509.00503'
  - '2509.01391'
  - '2509.05359'
  - '2509.20485'
  contradicting_papers: []
  refining_papers:
  - '2402.13236'
  - '2410.03751'
  - '2502.06490'
  - '2507.03887'
  - 2025.acl-long.681
  - '2509.05359'
  caveats:
  - Quantization introduces codebook, rate, and reconstruction trade-offs and may weaken speaker fidelity.
  last_reviewed: '2026-07-29'
- id: continuous_features_outperform_discrete_units_on_discrimination
  claim: Continuous SSL features often retain more information than discrete units for discriminative and speaker-sensitive
    tasks.
  status: emerging
  confidence: medium
  supporting_papers:
  - '2506.10274'
  contradicting_papers: []
  refining_papers:
  - '2508.08399'
  caveats:
  - The comparison is emerging and depends on rate, layer, quantizer, and downstream model capacity.
  last_reviewed: '2026-07-29'
- id: semantic_acoustic_hierarchies_balance_coherence_and_fidelity
  claim: Combining SSL semantic representations with acoustic codec detail balances linguistic coherence against
    waveform fidelity.
  status: strongly_supported
  confidence: high
  supporting_papers:
  - '2209.03143'
  - '2301.11325'
  - '2305.09636'
  - '2402.13236'
  - '2410.03751'
  - '2502.06490'
  - '2502.17239'
  - '2505.13000'
  - '2506.10274'
  - '2412.18603'
  - 2025.acl-long.682
  - '2508.14049'
  - '2508.04141'
  - interspeech-2025-0468
  - interspeech-2025-1639
  - '2509.09174'
  - '2509.09201'
  - '2509.11425'
  - '2509.16195'
  - '2509.17006'
  - '2509.22167'
  contradicting_papers: []
  refining_papers:
  - '2410.03751'
  - '2411.13577'
  - '2502.06490'
  - iclr-2025-dGSOn7sdWg
  - interspeech-2025-1236
  - '2509.22167'
  - '2509.26276'
  caveats:
  - Semantic and acoustic information are not cleanly separable for prosody and non-verbal vocal events.
  last_reviewed: '2026-07-29'
- id: ssl_pretraining_improves_low_resource_transfer
  claim: SSL pretraining improves data efficiency and transfer in low-resource or few-shot speech tasks.
  status: emerging
  confidence: medium
  supporting_papers:
  - '2402.05755'
  - '2509.00675'
  contradicting_papers: []
  refining_papers: []
  caveats:
  - Reported gains vary with language coverage and the amount of supervised adaptation.
  last_reviewed: '2026-07-29'
- id: multilingual_ssl_supports_cross_lingual_transfer
  claim: Multilingual SSL representations support cross-lingual transfer, but performance remains sensitive to pretraining-language
    coverage.
  status: strongly_supported
  confidence: high
  supporting_papers:
  - '2212.04356'
  - '2306.12925'
  - 2025.findings-acl.75
  - '2503.11026'
  - 2025.iwslt-1.5
  - '2508.14049'
  - interspeech-2025-0143
  - interspeech-2025-0166
  - interspeech-2025-1639
  contradicting_papers: []
  refining_papers:
  - 2025.findings-acl.75
  - '2503.11026'
  caveats:
  - High-resource languages dominate most pretraining mixtures and evaluation suites.
  last_reviewed: '2026-07-29'
- id: ssl_features_improve_noise_and_domain_robustness
  claim: Pretrained speech representations improve robustness to noise, channel, speaker, and domain mismatch.
  status: strongly_supported
  confidence: high
  supporting_papers:
  - '2204.02152'
  - '2506.10274'
  - '2508.04996'
  - '2508.11273'
  - interspeech-2025-0506
  - interspeech-2025-0998
  - interspeech-2025-1531
  - '2509.09201'
  - '2509.21968'
  contradicting_papers: []
  refining_papers:
  - interspeech-2025-0506
  caveats:
  - Robustness on simulated corruption may not transfer to spontaneous or atypical speech.
  last_reviewed: '2026-07-29'
- id: pretraining_scale_benefits_are_not_universal
  claim: Larger SSL pretraining scale generally improves transfer, but simple scaling laws do not reliably predict
    all speech-model regimes.
  status: contested
  confidence: medium
  supporting_papers: []
  contradicting_papers:
  - 2025.findings-acl.631
  refining_papers: []
  caveats:
  - Architecture, data quality, tokenization, and downstream adaptation can dominate parameter or compute scale.
  last_reviewed: '2026-07-29'
- id: task_adaptation_remains_necessary
  claim: Task-specific fine-tuning, adapters, prompting, or layer weighting remain necessary to realize SSL feature
    quality downstream.
  status: strongly_supported
  confidence: high
  supporting_papers:
  - '2408.02622'
  - '2507.04349'
  - 2025.findings-acl.75
  contradicting_papers: []
  refining_papers: []
  caveats:
  - Adaptation results are difficult to compare because parameter budgets and supervision differ.
  last_reviewed: '2026-07-29'
- id: ssl_representations_enable_unified_speech_models
  claim: SSL representations help unify speech understanding, generation, translation, and dialogue within shared
    models.
  status: strongly_supported
  confidence: high
  supporting_papers:
  - '2410.03751'
  - '2503.11026'
  - '2508.09600'
  contradicting_papers: []
  refining_papers:
  - '2509.24773'
  caveats:
  - Unified coverage can trade against specialized task performance and frequently relies on proprietary data.
  last_reviewed: '2026-07-29'
- id: ssl_conditioning_improves_speech_generation
  claim: SSL content and semantic features provide effective conditioning for TTS, voice conversion, and speech-to-speech
    generation.
  status: strongly_supported
  confidence: high
  supporting_papers:
  - '2104.00355'
  - '2402.08093'
  - 2025.coling-main.518
  - 2025.naacl-demo.12
  - 2025.acl-long.87
  - '2508.04996'
  - '2508.11273'
  - interspeech-2025-0438
  - interspeech-2025-0998
  - interspeech-2025-1229
  - interspeech-2025-1236
  - interspeech-2025-1440
  - interspeech-2025-1625
  - interspeech-2025-2684
  - '2509.01391'
  - '2509.20378'
  - '2509.21968'
  - '2509.22167'
  contradicting_papers: []
  refining_papers:
  - '2507.08012'
  - '2507.04349'
  - interspeech-2025-0305
  - interspeech-2025-1531
  - interspeech-2025-2043
  - interspeech-2025-2684
  caveats:
  - Generation gains depend on the acoustic decoder and may not isolate the representation contribution.
  last_reviewed: '2026-07-29'
- id: ssl_units_support_low_bitrate_neural_codecs
  claim: SSL-derived discrete units support low-bitrate speech coding while retaining task-relevant linguistic information.
  status: strongly_supported
  confidence: high
  supporting_papers:
  - '2104.00355'
  - '2209.03143'
  - '2402.08093'
  - '2402.13236'
  - '2411.19842'
  - '2505.13000'
  - '2508.11224'
  - interspeech-2025-0246
  - interspeech-2025-0468
  - interspeech-2025-1440
  - interspeech-2025-1531
  - '2509.09201'
  - '2509.11425'
  - '2509.21968'
  - '2509.22167'
  contradicting_papers: []
  refining_papers:
  - '2402.13236'
  - 2025.acl-long.87
  - interspeech-2025-2043
  caveats:
  - Content-oriented units can underrepresent speaker identity, prosody, and fine acoustic detail.
  last_reviewed: '2026-07-29'
- id: automatic_judges_are_unreliable_for_fine_grained_prosody
  claim: General-purpose audio-language-model judges are not yet reliable substitutes for fine-grained human evaluation
    of prosody and paralinguistics.
  status: contested
  confidence: medium
  supporting_papers:
  - '2507.16632'
  contradicting_papers:
  - '2509.19928'
  refining_papers:
  - '2507.16632'
  caveats:
  - Capabilities may improve rapidly and reliability can differ between rubric-based and raw-acoustic judgments.
  last_reviewed: '2026-07-29'
method_families:
- id: autoregressive_ssl_conditioned_models
  name: Autoregressive SSL-conditioned speech models
  summary: Autoregressive language models consume continuous or discrete self-supervised speech representations
    for generation, dialogue, and multimodal reasoning.
  papers:
  - '2209.03143'
  - '2301.11325'
  - '2305.09636'
  - '2305.11000'
  - '2306.12925'
  - '2402.05755'
  - '2402.08093'
  - '2408.02622'
  - '2409.00750'
  - '2409.03283'
  - '2410.00037'
  - '2411.13577'
  - '2502.04128'
  - '2502.07243'
  - '2502.17239'
  - '2503.01710'
  - '2504.10344'
  - iclr-2025-dGSOn7sdWg
  - 2025.naacl-demo.12
  - 2025.naacl-long.484
  - '2505.13000'
  - '2412.18603'
  - '2507.16632'
  - 2025.acl-long.681
  - 2025.acl-long.682
  - 2025.acl-long.817
  - 2025.acl-long.997
  - 2025.findings-acl.631
  - 2025.findings-acl.71
  - 2025.iwslt-1.5
  - '2508.14049'
  - '2508.04141'
  - '2508.07375'
  - '2508.09600'
  - interspeech-2025-0310
  - interspeech-2025-1595
  - '2508.16188'
  - '2508.16790'
  - '2509.00503'
  - '2509.05359'
  - '2509.04072'
  - '2509.09174'
  - '2509.11425'
  - '2509.13068'
  - '2509.20378'
  - '2509.17143'
  - '2509.26276'
  open_questions:
  - Which SSL representation rate and layer best balance semantic reasoning, acoustic detail, and sequence cost?
- id: transformer_ssl_encoders_and_adapters
  name: Transformer SSL encoders and adapters
  summary: Transformer encoders learn contextual speech features through masked or contrastive pretraining and feed
    task-specific decoders or adapters.
  papers:
  - '2212.04356'
  - '2411.13577'
  - '2411.19842'
  - '2409.09098'
  - 2025.findings-naacl.130
  - '2507.00808'
  - '2507.08012'
  - 2025.findings-acl.75
  - '2508.05385'
  - '2508.06890'
  - '2508.07273'
  - '2508.11273'
  - interspeech-2025-0115
  - interspeech-2025-0166
  - interspeech-2025-0203
  - interspeech-2025-0246
  - interspeech-2025-0305
  - interspeech-2025-0383
  - interspeech-2025-0506
  - interspeech-2025-0723
  - interspeech-2025-1394
  - interspeech-2025-2043
  - interspeech-2025-2660
  - '2508.16790'
  - '2509.00503'
  - '2509.00675'
  - '2509.01391'
  - '2509.06074'
  open_questions:
  - How much task adaptation is needed before pretrained features outperform purpose-built supervised encoders?
- id: hybrid_semantic_acoustic_ssl_systems
  name: Hybrid semantic–acoustic SSL systems
  summary: Hybrid systems combine SSL semantic features with acoustic, codec, speaker, or prosodic pathways to retain
    complementary information.
  papers:
  - '2409.03283'
  - '2409.06666'
  - '2410.00037'
  - '2411.13577'
  - 2025.coling-main.518
  - '2502.07243'
  - '2502.17239'
  - 2025.naacl-short.65
  - '2507.16632'
  - 2025.acl-demo.37
  - 2025.acl-long.682
  - 2025.sigdial-1.21
  - '2508.04141'
  - interspeech-2025-0438
  - interspeech-2025-0656
  - interspeech-2025-0756
  - interspeech-2025-1478
  - interspeech-2025-1776
  - interspeech-2025-cho25c_interspeech
  - '2507.14534'
  - '2509.03292'
  - '2509.04667'
  - '2509.09550'
  - '2509.16195'
  open_questions:
  - Can semantic and acoustic pathways be separated without losing timing, style, or speaker cues needed downstream?
- id: gan_decoders_for_ssl_units
  name: GAN decoders for SSL units
  summary: Adversarial waveform decoders reconstruct or convert speech from discrete or continuous SSL-derived content
    representations.
  papers:
  - '2104.00355'
  - '2305.02765'
  - '2308.16692'
  - 2025.findings-naacl.130
  - 2025.naacl-short.65
  - '2507.18897'
  - 2025.acl-long.682
  - '2508.06890'
  - interspeech-2025-0468
  - interspeech-2025-0815
  - interspeech-2025-0998
  - interspeech-2025-1106
  - interspeech-2025-1531
  - interspeech-2025-1625
  - interspeech-2025-1639
  - '2508.15565'
  - '2507.14534'
  - '2509.04667'
  - '2509.09201'
  - '2509.09550'
  - '2509.11425'
  - '2509.17006'
  - '2509.21968'
  open_questions:
  - Which adversarial objectives preserve perceptual detail without reintroducing source-speaker leakage?
- id: flow_matching_with_ssl_conditioning
  name: Flow matching with SSL conditioning
  summary: Flow-matching generators synthesize acoustic representations or waveforms from compact SSL content features
    with parallel decoding.
  papers:
  - '2409.03283'
  - '2412.04724'
  - 2025.coling-main.518
  - '2502.07243'
  - '2502.17239'
  - '2507.03887'
  - '2507.04349'
  - '2507.16632'
  - 2025.acl-long.790
  - 2025.acl-long.87
  - '2503.11026'
  - '2508.14049'
  - '2508.04996'
  - '2508.09600'
  - interspeech-2025-0203
  - interspeech-2025-0305
  - interspeech-2025-1229
  - interspeech-2025-1236
  - interspeech-2025-1779
  - interspeech-2025-2684
  - '2509.04072'
  - '2509.24773'
  open_questions:
  - Do flow-matching gains persist when SSL layer, decoder capacity, and inference budget are matched?
- id: vae_and_quantized_ssl_representations
  name: VAE and quantized SSL representations
  summary: VAE-derived systems compress, discretize, or factor contextual speech representations for generation
    and conversion.
  papers:
  - '2104.00355'
  - '2305.02765'
  - '2411.19842'
  - '2504.10344'
  - '2508.08399'
  - interspeech-2025-0433
  - interspeech-2025-0468
  - interspeech-2025-0723
  - interspeech-2025-0815
  - interspeech-2025-0948
  - interspeech-2025-1106
  - interspeech-2025-1440
  - interspeech-2025-1531
  - '2509.13068'
  - '2509.15626'
  - '2509.21968'
  - '2509.22167'
  open_questions:
  - When does quantization improve modelability enough to offset its loss of speaker and acoustic detail?
- id: diffusion_with_ssl_conditioning
  name: Diffusion with SSL conditioning
  summary: Diffusion systems denoise acoustic or latent speech representations conditioned on pretrained semantic,
    speaker, or content features.
  papers:
  - interspeech-2025-0816
  - interspeech-2025-0948
  - interspeech-2025-0998
  - interspeech-2025-1101
  - interspeech-2025-1210
  - interspeech-2025-1397
  - interspeech-2025-1434
  - '2508.16790'
  - '2411.19770'
  open_questions:
  - How few denoising steps can preserve the benefits of SSL conditioning in real-time settings?
reassessment_queue:
- id: continuous_features_outperform_discrete_units_on_discrimination
  type: claim_status
  reason: Only a small set of matched comparisons separates representation form from rate and model capacity.
  trigger: Matched continuous-versus-discrete studies reproduce the result across discriminative and generative
    tasks.
  due: 2026-10
  current_assessment: emerging
  watch_for:
  - Rate-controlled representation studies
  - Speaker-sensitive task comparisons
- id: pretraining_scale_benefits_are_not_universal
  type: claim_status
  reason: Speech scaling behavior varies across architectures, tokenizers, and academic compute regimes.
  trigger: Independent scaling studies converge on predictive compute/data boundaries across speech architectures.
  due: 2026-10
  current_assessment: contested
  watch_for:
  - Reproduced speech scaling laws
  - Data-quality-adjusted scaling studies
- id: automatic_judges_are_unreliable_for_fine_grained_prosody
  type: benchmark_validity
  reason: The evidence directly disputes general-purpose audio-language-model reliability for fine prosodic judgments.
  trigger: Listener-calibrated benchmarks demonstrate reliable raw-acoustic and paralinguistic scoring.
  due: 2026-10
  current_assessment: contested
  watch_for:
  - Listener-calibrated audio judges
  - Fine-grained prosody benchmarks
- id: hybrid_semantic_acoustic_ssl_systems
  type: method_family
  reason: The family combines several placements of semantic and acoustic pathways.
  trigger: Enough matched systems exist to split parallel streams, hierarchical tokens, and layer-conditioned decoders.
  due: 2026-10
  current_assessment: active_evidence
  watch_for:
  - Matched semantic-integration ablations
  - Layer-wise information probes
- id: speaker_and_prosody_leakage_persists_in_ssl_features
  type: claim_status
  reason: The desired level of invariance differs across recognition, TTS, VC, and dialogue.
  trigger: Task-conditioned studies quantify beneficial versus harmful leakage across applications.
  due: 2026-10
  current_assessment: strongly_supported
  watch_for:
  - Task-specific invariance studies
  - Causal representation interventions
open_questions:
- Which SSL objective and model layer best balances linguistic abstraction with speaker, prosodic, and acoustic
  information?
- When should speech representations remain continuous, and when does quantization improve downstream modelability
  enough to justify information loss?
- How should multilingual pretraining data be allocated to avoid systematic degradation for low-resource languages
  and accents?
- Can one pretrained speech representation serve recognition, synthesis, conversion, dialogue, and evaluation without
  task-specific compromises?
- How much downstream performance comes from representation quality versus decoder capacity, prompting, adapters,
  or supervised fine-tuning?
- What evaluation protocol can measure semantic content, acoustic fidelity, speaker information, robustness, and
  generative usefulness jointly?
trend_notes:
- SSL speech representations have shifted from recognition-oriented encoders toward core interfaces for generative
  and conversational models.
- Discrete semantic units became increasingly common after 2022 as speech language modeling expanded.
- Recent systems increasingly combine SSL semantic features with separate acoustic codec or generative-decoder pathways.
- Multilingual and multimodal pretraining expanded rapidly in 2024–2025, while balanced low-resource coverage remains
  unresolved.
- Evaluation is moving beyond linear probes toward generation quality, robustness, information leakage, and task-specific
  adaptation behavior.
