Append-only chronological log of changes to the wiki. Entry types: ingest (new paper page), review (quality review of existing paper page), integrate (concept YAML updated from paper pages), render (Overview and In Depth concept pages regenerated from YAML), query (research question answered and filed back), misc (structural fixes, metadata corrections, and other changes that don’t fit the above). Most recent entries are at the top.
2026-07-24
- misc | removed obsolete
evidence/directory after Overview and In Depth production migration; deleted the legacy speech-to-speech and flow-matching evidence dossiers | runtime: codex | provider: openai | model: gpt-5 - render | 1 concept | formats: both | mode: full | runtime: codex | provider: openai | model: gpt-5
- render | 4 concepts | formats: both | mode: full | runtime: codex | provider: openai | model: gpt-5
2026-07-22
- integrate | phase 2 only | disentanglement (100 papers) | 17 claim_clusters | 10 method_families | 7 reassessment_queue items | first Phase 2 synthesis run for this concept, following the completed 100/100 Phase 1 pass (batches 1-5); used a metadata-first + relevance-filtered-claims condensation strategy (per-paper metadata dump plus high/medium-relevance claim text extracted to a scratch file) rather than reading the full ~4900-line YAML at once, per the evaluation-metrics Phase 2 precedent. An initial pass during this session left 23/100 papers (23%) with an empty method_family, a noticeably higher outlier rate than the other closed concepts (5/97, 4/29, ~14/285); after a mid-task course correction, 7 of those 23 were reclassified for consistency with papers already included on similar evidentiary grounds (e.g. FreeCodec’s self-reported incomplete separation was already included in the RVQ-distillation family, so DiffStyleTTS’s self-reported prosody/timbre non-separation, VibeVoice’s and PerformSinger’s inherited-not-novel codec splits, and ChiReSSD’s/Conan’s inherited-mechanism applications were brought in on the same basis), landing at a final 16/100 (16%) empty rate. 10 method_families: explicit_multiencoder_factorization (30, largest/most heterogeneous, flagged as a reassessment_queue mega-family watch), rvq_semantic_acoustic_distillation (19), data_training_recipe_disentanglement (11), adversarial_grl_disentanglement (7), asr_text_alignment_disentanglement (7), discrete_bottleneck_disentanglement (7), signal_level_prosody_decomposition (6), orthogonality_constraint_disentanglement (5), ssl_statistical_normalization_disentanglement (4), mutual_information_minimization (2, thin, flagged in reassessment_queue); 14 papers hold dual family membership (e.g. NaturalSpeech 3/FACodec spans explicit factorization and GRL; XY-Tokenizer spans RVQ distillation and ASR/text-alignment). 17 claim_clusters: 12 strongly_supported, 3 emerging, 2 contested. Headline strongly_supported clusters: disentanglement_reconstruction_fidelity_tradeoff (10 supporting, incorporating DeCodec’s explicit reconstruction-quality trade-off), explicit_disentanglement_enables_independent_multiattribute_control (12 supporting), emotion_speaker_disentanglement_enables_crossspeaker_transfer (7 supporting). The 2 contested clusters directly surface the complicating evidence flagged during Phase 1: grl_disentanglement_leakage_vs_quality_tradeoff (GRL reduces leakage per 5 papers, but PeriodCodec reports a GRL-induced MOS regression and DiEmoTTS directly critiques GRL/VQ disentanglement) and disentanglement_oriented_codecs_fail_on_pitch (interspeech-2025-0115/AnCoGen’s finding that even SpeechTokenizer- and Mimi-style disentanglement-oriented codecs fail to cleanly separate pitch, contradicted by three papers using dedicated signal-level F0 mechanisms rather than generic representation-level disentanglement, with FreeCodec’s self-reported incomplete t-SNE separation noted as a refining data point in both this cluster and the ssl_content_representations_incomplete_disentanglement cluster). 7 reassessment_queue items: 2 method_family threshold-watches (explicit_multiencoder_factorization’s mega-family risk; mutual_information_minimization’s 2-paper thinness), 3 contested/thin claim_status watches (the two contested clusters above, plus objective_metrics_dont_predict_subjective_disentanglement_quality’s 3-paper emerging base), 1 fast-emerging claim_status watch (prompt_based_conditioning_causes_attribute_leakage, 3 papers), 1 paper_role watch (Vevo2’s un-ablated bottleneck tokenizer). health_check —module integrate —concept disentanglement —phase 2 passed 0 errors, 16 warnings (all method_family_coverage, matching the 16 legitimate outliers). | runtime: claude-code | provider: anthropic | model: claude-sonnet-5
- integrate | 20 papers | disentanglement | Phase 1 batch 1 | first-ever integration pass for this concept, new
wiki/_claims/disentanglement.yamlcreated from scratch; Q3-scoped (published_date < 2025-10-01) oldest-first, 100 in-scope candidates found (110 total minus 10 Q4 2025+ papers deferred to a later pass), 0 Tier 2 skipped, batch 1 processed the 20 oldest: 2010.05646, 2104.00355, 2209.03143, 2304.09116, 2308.16692, 2312.01479, 2402.05755, 2403.03100, 2406.02430, 2408.16532, 2411.09943, 2412.04724, 2409.09098, 2025.coling-main.352, 2502.06490, 2502.07243, 2503.01710, 2025.findings-naacl.130, 2025.naacl-long.242, 2505.07916; 18 papers use the legacy bare-claims format (role inferred from wording, evidence synthesized from Method/Key Results/Limitations sections per docs/schemas/claims.md compatibility rules) and 2 use the structured bold-prefix/blockquote format (2409.09098 AccentBox, 2025.findings-naacl.130 DiVISe); 95 claims extracted across 20 papers, all with extractable source citations (no “not specified” sources needed); notable negative/complicating evidence surfaced: WavTokenizer (2408.16532) shows semantic richness without explicit disentanglement, DiffStyleTTS (2025.coling-main.352) explicitly reports failure to disentangle timbre from prosody, Spark-TTS (2503.01710) documents a speaker-similarity cost from single-stream disentanglement; Phase 1 only, no claim_clusters/method_families synthesis this batch; 80 in-scope papers remain for follow-up batches (next oldest: 2506.10274, 2507.01348, 2506.23325, 2507.09070, 2507.12197, …); health_check —module integrate —concept disentanglement —phase 1 passed 0 errors, 0 warnings | runtime: claude-code | provider: anthropic | model: claude-sonnet-5 - integrate | 20 papers | disentanglement | Phase 1 batch 2 | continued oldest-first from batch 1 (which ended at 2505.07916); 100 in-scope candidates total (published_date < 2025-10-01), 20 already integrated in batch 1, 0 Tier 2 skipped, batch 2 processed the next 20 oldest: 2506.10274, 2507.01348, 2506.23325, 2507.09070, 2507.12197, 2025.acl-demo.37, 2025.acl-long.1043, 2025.acl-long.346, 2025.acl-long.87, 2025.ccl-1.77, 2508.02038, 2508.02849, 2508.04141, 2507.20091, 2508.06890, 2508.07426, 2508.08095, 2508.08399, 2508.11224, interspeech-2025-0115; 6 papers use the structured bold-prefix/blockquote claims format (2507.01348 SpeechAccentLLM, 2506.23325 XY-Tokenizer, 2507.09070 SemAlignVC, 2507.12197 QTTS, 2025.acl-long.87 Takin-VC, 2025.ccl-1.77 HFSD-V2C) and 14 use the legacy bare-claims format (role inferred from wording, evidence synthesized from Method/Key Results/Limitations sections); 92 claims extracted across 20 papers, all with extractable source citations (no “not specified” sources needed); notable negative/complicating evidence surfaced: Bringing Interpretability to Neural Audio Codecs (interspeech-2025-0115) directly contradicts the premise that disentanglement-oriented codecs (SpeechTokenizer, Mimi) cleanly separate pitch from other attributes, and Exploring Disentangled Neural Speech Codecs (2508.08399) documents a speaker-identity-fidelity cost from quantizing an otherwise unsupervised disentangled codec; Phase 1 only, no claim_clusters/method_families synthesis this batch; 60 in-scope papers remain for follow-up batches (next oldest: interspeech-2025-0196, interspeech-2025-0203, interspeech-2025-0246, interspeech-2025-0347, interspeech-2025-0383, …); paper_count updated 20 -> 40 | runtime: claude-code | provider: anthropic | model: claude-sonnet-5
- integrate | 20 papers | disentanglement | Phase 1 batch 3 | continued oldest-first from batch 2 (which ended at interspeech-2025-0115); 100 in-scope candidates total (published_date < 2025-10-01), 40 already integrated in batches 1-2, 0 Tier 2 skipped, batch 3 processed the next 20 oldest: interspeech-2025-0196, interspeech-2025-0203, interspeech-2025-0246, interspeech-2025-0347, interspeech-2025-0383, interspeech-2025-0433, interspeech-2025-0438, interspeech-2025-0455, interspeech-2025-0464, interspeech-2025-0468, interspeech-2025-0575, interspeech-2025-0596, interspeech-2025-0656, interspeech-2025-0723, interspeech-2025-0815, interspeech-2025-0816, interspeech-2025-0948, interspeech-2025-1081, interspeech-2025-1101, interspeech-2025-1106; 8 papers use the structured bold-prefix/blockquote claims format (interspeech-2025-0347 PeriodCodec, interspeech-2025-0383 Likability VC, interspeech-2025-0433 H2NH-VC, interspeech-2025-0438 LinearVC, interspeech-2025-0464 PACE/VALL-E-X-VC, interspeech-2025-0656 EEG-VC, interspeech-2025-1081 SNCR-VC, interspeech-2025-1106 LSCodec) and 12 use the legacy bare-claims format (role inferred from wording, evidence synthesized from Method/Key Results/Limitations sections per docs/schemas/claims.md compatibility rules); 90 claims extracted across 20 papers, all with extractable source citations (no “not specified” sources needed); notable negative/complicating evidence surfaced: PeriodCodec (interspeech-2025-0347) reports GRL-based pitch disentanglement lowering MOS rather than improving it, Non-AR Expressive VC (interspeech-2025-0815) documents an acknowledged and unresolved intelligibility regression from discrete-unit speaker disentanglement, and LinearVC (interspeech-2025-0438) provides clean SVD/constrained-transformation evidence for the orthogonal content-speaker subspace hypothesis; several papers (SPCODEC, VoiceMark, DualCodec, PromptEVC, ClapFM-EVC) use disentanglement only as a peripheral mechanism within a broader VC/codec/TTS system rather than as the primary contribution, and were scored with correspondingly lower paper-level relevance; Phase 1 only, no claim_clusters/method_families synthesis this batch; 40 in-scope papers remain for follow-up batches (next oldest: interspeech-2025-1115, interspeech-2025-1192, interspeech-2025-1210, interspeech-2025-1394, interspeech-2025-1397, …); paper_count updated 40 -> 60; health_check —module integrate —concept disentanglement —phase 1 passed 0 errors, 0 warnings | runtime: claude-code | provider: anthropic | model: claude-sonnet-5
- integrate | 20 papers | disentanglement | Phase 1 batch 4 | continued oldest-first from batch 3 (which ended at interspeech-2025-1106); re-derived the candidate list fresh from frontmatter at batch start rather than trusting the carried-over hint, confirming 40 in-scope candidates remained (published_date < 2025-10-01, non-Tier-2) after the 60 papers integrated in batches 1-3; note: an earlier invocation of this exact batch was cut off by a session limit before any write occurred (paper_count/len(papers) both still 60, no batch 4 log entry present — a clean pre-write state with nothing to resume), so this run is a fresh restart rather than a continuation; 0 Tier 2 skipped, batch 4 processed the next 20 oldest: interspeech-2025-1115, interspeech-2025-1192, interspeech-2025-1210, interspeech-2025-1394, interspeech-2025-1397, interspeech-2025-1434, interspeech-2025-1440, interspeech-2025-1531, interspeech-2025-1538, interspeech-2025-1638, interspeech-2025-1639, interspeech-2025-1684, interspeech-2025-1779, interspeech-2025-2586, interspeech-2025-2684, interspeech-2025-cho25c_interspeech, 2508.15565, 2508.15931, 2508.16188, 2508.16332; 6 papers use the legacy bare-claims format (interspeech-2025-1440 FreeCodec, interspeech-2025-1779 ReFlow-VC, 2508.15565 Any-to-any Speaker Attribute Perturbation, 2508.15931 QvTAD, 2508.16188 Seeing is Believing/AVLM, 2508.16332 Vevo2) and 14 use the structured bold-prefix/blockquote format; 91 claims extracted across 20 papers, all with extractable source citations (no “not specified” sources needed); several papers (interspeech-2025-1779 ReFlow-VC, interspeech-2025-cho25c_interspeech H2NH demo, 2508.15565, 2508.15931, 2508.16188, 2508.16332) use disentanglement only as a peripheral or tangential mechanism rather than the primary contribution and were scored with correspondingly lower paper-level relevance (low/medium) and current_role: minor; notable negative/complicating evidence surfaced: FreeCodec (interspeech-2025-1440) self-reports via t-SNE visualisation that its own prosody/speaker separation is incomplete, LombardTokenizer (interspeech-2025-1639) documents that injecting a specialised style encoder into an existing VC architecture (FreeVC) fails to produce controllable style transfer even though the same encoder succeeds via dedicated RVQ-layer distillation, and Discl-VC (interspeech-2025-2684) shows a same-speaker training assumption causes residual speaker leakage into prosody tokens (the largest ablation degradation of any component tested); Phase 1 only, no claim_clusters/method_families synthesis this batch; 20 in-scope papers remain for a follow-up batch 5 (next oldest: 2508.17031, 2508.19205, 2411.19770, 2507.14534, 2509.00503, …); paper_count updated 60 -> 80; health_check —module integrate —concept disentanglement —phase 1 passed 0 errors, 0 warnings | runtime: claude-code | provider: anthropic | model: claude-sonnet-5
- integrate | 20 papers | disentanglement | Phase 1 batch 5 (final) | continued oldest-first from batch 4 (which ended at 2508.16332); re-derived the candidate list fresh from frontmatter at batch start, confirming exactly 20 in-scope candidates remained (published_date < 2025-10-01, non-Tier-2) after the 80 papers integrated in batches 1-4, matching the expected count exactly (no scope miscount); batch 5 processed all 20 remaining oldest: 2508.17031, 2508.19205, 2411.19770, 2507.14534, 2509.00503, 2506.21619, 2509.09201, 2509.13068, 2509.14684, 2509.15626, 2509.16010, 2509.17006, 2509.17516, 2509.18060, 2509.19231, 2509.19883, 2509.22718, 2505.10599, 2509.22062, 2509.24650; 13 papers use the structured bold-prefix/blockquote claims format (2411.19770 Noro, 2509.09201 DeCodec, 2509.13068 MSR-Codec, 2509.14684 DAIEN-TTS, 2509.15626 LibriTTS-VI, 2509.16010 Fed-PISA, 2509.17006 MBCodec, 2509.17516 Audiobook-CC, 2509.18060 TMD-TTS, 2509.19231 ChiReSSD, 2509.19883 CoMelSinger, 2509.22718 PerformSinger, 2505.10599 UDDETTS, 2509.22062 CaT-TTS, 2509.24650 VoxCPM) and 5 use the legacy bare-claims format (2508.17031 RephraseTTS, 2508.19205 VibeVoice, 2507.14534 Conan, 2509.00503 Entropy-based Coarse Compression, 2506.21619 IndexTTS2), role inferred from wording and evidence synthesized from Method/Key Results/Limitations sections per docs/schemas/claims.md compatibility rules for the legacy set; 91 claims extracted across 20 papers, all with extractable source citations (no not-specified sources needed); several papers in this batch contribute genuinely central disentanglement mechanisms rather than peripheral mentions: DeCodec (SOP+RST speech/background subspace orthogonality with swap training), MSR-Codec (cascaded residual structural disentanglement as an explicit adversarial-training-free alternative to NaturalSpeech 3), Noro (noise-agnostic contrastive speaker loss, t-SNE-validated), LibriTTS-VI (impression leakage reduced via training-data decoupling alone, no architecture change), Fed-PISA (identity/style split via non-overlapping-gradient LoRA adapters), MBCodec (PQMF frequency-subband codebook disentanglement), and CoMelSinger (direct diagnostic evidence, generalizing beyond its own singing-synthesis framing, that prompt-based masked generative TTS models leak prosody from acoustic prompts); several other papers use disentanglement only peripherally (VibeVoice, Conan, Entropy-based Compression, PerformSinger, UDDETTS) and were scored with correspondingly lower paper-level relevance and current_role: minor; this closes Phase 1 for disentanglement entirely: 100/100 in-scope Q3-and-earlier candidates now integrated, 0 remain in scope; paper_count updated 80 -> 100; Phase 1 only, no claim_clusters/method_families synthesis this batch (Phase 2 planned as a separate follow-up); health_check —module integrate —concept disentanglement —phase 1 passed 0 errors, 0 warnings | runtime: claude-code | provider: anthropic | model: claude-sonnet-5
2026-07-21
-
integrate | phase 2 only | evaluation-metrics (285 papers) | 31 claim_clusters (31 new) | 17 method_families (17 new) | 7 reassessment_queue items added | first Phase 2 synthesis run for this concept, largest Phase 2 pass to date (vs. 97/60/29 for flow-matching/speech-to-speech/rlhf-speech); processed via a metadata-first + relevance-filtered-claims condensation strategy (the full 285-paper YAML exceeds single-context Read limits) rather than reading the whole file at once. First pass produced only 10 method_families covering 59/285 papers (21%), which a course-correction from the orchestrator correctly flagged as an incomplete synthesis (the other three closed concepts left only single-digit outliers); the taxonomy was reworked around evaluation-methodology TYPE (which metric or evaluation dimension a paper’s claim concerns: WER/CER, speaker-similarity, MOS/PESQ-style divergence, prosody, codec multi-metric behavior, LLM-judge reliability, bias/fairness, low-resource/community, clinical/accessibility, data curation, watermarking) rather than TTS architecture, since this concept is cross-cutting and most of its 285 papers are TTS/VC/SCA/codec system papers whose relevance is a single evaluation_caution finding rather than a dedicated methodology contribution. Final coverage: 271/285 papers (95%), 14 legitimate outliers (pure architecture/theory/efficiency papers with no evaluation-methodology content: e.g. Tacotron 2, HiFi-GAN, AudioLDM, GOAT, model-quantization and inference-acceleration papers) — matching the coverage ratio of the three previously-closed concepts. 31 claim_clusters: 23 strongly_supported, 7 emerging, 1 contested (llm_audio_judge_reliability_is_contested: LLM/audio-LM judges approximate human evaluation well for instruction-following, style adherence, and checklist-rubric dialogue scoring, but fail for fine-grained prosody and non-verbal/raw-waveform cues). Headline strongly_supported clusters recur across nearly every metric family in the corpus: wer_cer_unreliable_intelligibility_proxy (8 supporting), embedding_speaker_similarity_diverges_from_human_perception (11 supporting after cross-checking the fuller paper set), automatic_mos_predictors_diverge_from_subjective_mos (8 supporting) — together the single most cross-cutting finding in the concept is that no automatic objective metric reliably substitutes for targeted human perceptual judgment. Two families are very large umbrella categories by design (objective_perceptual_quality_metric_divergence, 86 papers; asr_wer_intelligibility_evaluation, 79 papers) spanning many unrelated architectures that happen to share the same evaluation-methodology observation, flagged in reassessment_queue as candidates for future sub-splitting rather than force-fragmenting now. 7 reassessment_queue items: 5 claim_status entries thin on independent replication (fullduplex automatic metrics without human raters, automated-annotation-consistency possible shared lineage, deception-rate evaluation, reference-free task-specific metrics, watermark robustness under realistic transforms) plus 2 method_family-threshold/structural watches (embedding-distributional-distance metrics at 4 papers, and the two mega-families’ internal heterogeneity). health_check —module integrate —concept evaluation-metrics passed 0 errors, 14 warnings (all method_family_coverage, the 14 legitimate outliers). | runtime: claude-code | provider: anthropic | model: claude-sonnet-5
-
integrate | evaluation-metrics | 25 papers integrated (phase 1 only) | Q3-scoped (published_date < 2025-10-01) oldest-first continuation, 260 -> 285 of 286 in-scope candidates, CLOSING Phase 1 batch for this concept; processed 2509.17006 through 2510.00264 (2509.17006, 2509.17021, 2509.17988, 2509.18060, 2509.18470, 2509.18806, 2509.18823, 2509.19231, 2509.19668, 2509.19812, 2509.19928, 2509.20086, 2509.20321, 2509.20410, 2509.20485, 2506.21875, 2509.21968, 2509.22243, 2509.23147, 2510.02352, 2509.23938, 2509.24457, 2509.26276, 2509.26542, 2510.00264); the standing exclusion 2207.12598 (Classifier-Free Diffusion Guidance) was re-confirmed one final time (re-derived independently from frontmatter, off-topic ImageNet diffusion paper with
task: []and no speech content) and excluded again for the same reason as batches 1-13, permanently out of scope for this concept; no other exclusions this batch; 1 paper uses the legacy bare-claims format (2509.19668, Selective CFG for Zero-shot TTS) handled per docs/schemas/claims.md compatibility rules with evidence synthesized from Key Results/Novelty Assessment and role inferred from wording rather than defaulted; 24 use the structured bold-prefix/blockquote format; 0 in-scope Q3 candidates remain besides the permanent exclusion, Phase 1 for evaluation-metrics is now fully closed (285/286 in-scope papers integrated, 1 permanently excluded) | phase 2 synthesis not yet run for this concept, planned as a separate follow-up | runtime: claude-code | provider: anthropic | model: claude-sonnet-5 -
integrate | evaluation-metrics | 20 papers integrated (phase 1 only) | Q3-scoped (published_date < 2025-10-01) oldest-first continuation, 240 -> 260 of 286 in-scope candidates; processed 2509.11425 through 2509.20378 (2509.11425, 2508.18240, 2509.12171, 2509.14270, 2504.20581, 2509.13667, 2509.13670, 2509.13989, 2509.14684, 2509.14882, 2509.15085, 2509.15253, 2509.15462, 2505.17093, 2509.15626, 2509.15629, 2509.15969, 2509.16195, 2509.16589, 2509.20378); the standing exclusion 2207.12598 (Classifier-Free Diffusion Guidance) was re-confirmed as still the oldest unintegrated in-scope candidate (re-derived independently from frontmatter, off-topic ImageNet diffusion paper with
task: []and no speech content) and excluded again for the same reason as batches 1-12, not replaced since the 20-entry cap was reached from the next oldest genuine candidate onward; no other exclusions this batch; 1 paper uses the legacy bare-claims format (2509.15969, VoXtream) handled per docs/schemas/claims.md compatibility rules with evidence synthesized from Method/Key Results/Limitations and role inferred from wording (supports/complicates) rather than defaulted; 19 use the structured bold-prefix/blockquote format; 25 in-scope papers remain for a smaller closing Phase 1 batch 14 (next: the oldest unintegrated candidate after 2509.20378, to be re-derived from frontmatter rather than assumed) | phase 2 synthesis not yet run for this concept | runtime: claude-code | provider: anthropic | model: claude-sonnet-5 -
integrate | evaluation-metrics | 20 papers integrated (phase 1 only) | Q3-scoped (published_date < 2025-10-01) oldest-first continuation, 220 -> 240 of 286 in-scope candidates; processed 2508.18006 through 2509.11084 (2508.18006, 2508.20660, 2509.00685, 2509.01391, 2509.02244, 2506.23367, 2509.03292, 2509.05359, 2509.03940, 2509.04072, 2509.04093, 2509.06502, 2509.07376, 2509.09716, 2509.08696, 2506.04077, 2509.09550, 2509.09631, 2509.09748, 2509.11084); the standing exclusion 2207.12598 (Classifier-Free Diffusion Guidance) was re-confirmed as still the oldest unintegrated in-scope candidate (re-derived independently from frontmatter, off-topic ImageNet diffusion paper with
task: []and no speech content) and excluded again for the same reason as batches 1-11, not replaced since the 20-entry cap was reached from the next oldest genuine candidate onward; no other exclusions this batch; 9 papers use the legacy bare-claims format (2508.18006, 2508.20660, 2509.00685, 2509.01391, 2509.02244, 2509.03292, 2509.03940, 2509.04072, 2509.09631) handled per docs/schemas/claims.md compatibility rules with evidence synthesized from Method/Key Results and role inferred from wording (supports/complicates) rather than defaulted; 11 use the structured bold-prefix/blockquote format (2506.23367, 2509.05359, 2509.04093, 2509.06502, 2509.07376, 2509.09716, 2509.08696, 2506.04077, 2509.09550, 2509.09748, 2509.11084); 45 in-scope papers remain for follow-up Phase 1 batches (next: 2509.11425, then 2508.18240, 2509.12171, 2509.14270, …) | phase 2 synthesis not yet run for this concept | runtime: claude-code | provider: anthropic | model: claude-sonnet-5 -
integrate | evaluation-metrics | 20 papers integrated (phase 1 only) | Q3-scoped (published_date < 2025-10-01) oldest-first continuation, 200 -> 220 of 286 in-scope candidates; processed interspeech-2025-2449 through interspeech-2025-gourav25_interspeech, then 2508.13028, 2507.16835, 2508.15442, 2508.15565, 2508.15931, 2508.16188, 2508.17494, 2508.17623 (interspeech-2025-2449, 2536, 2573, 2595, 2660, 2679, 2726, 2739, 2765, 2787, bokkahallisatish25_interspeech, gourav25_interspeech, 2508.13028, 2507.16835, 2508.15442, 2508.15565, 2508.15931, 2508.16188, 2508.17494, 2508.17623); the standing exclusion 2207.12598 (Classifier-Free Diffusion Guidance) was re-confirmed as still the oldest unintegrated in-scope candidate (re-derived independently from frontmatter, off-topic ImageNet diffusion paper with
task: []and no speech content) and excluded again for the same reason as batches 1-10, not replaced since the 20-entry cap was reached from the next oldest genuine candidate onward; no other exclusions this batch; 10 papers use the legacy bare-claims format (interspeech-2025-2449, 2660, 2765, 2508.13028, 2508.15442, 2508.15565, 2508.15931, 2508.16188, 2508.17494, 2508.17623) handled per docs/schemas/claims.md compatibility rules with evidence synthesized from Method/Key Results and role inferred from wording (supports/complicates) rather than defaulted; 10 use the structured bold-prefix/blockquote format; 65 in-scope papers remain for follow-up Phase 1 batches (next: 2508.18006, then 2508.20660, 2509.00685, 2509.01391, …) | phase 2 synthesis not yet run for this concept | runtime: claude-code | provider: anthropic | model: claude-sonnet-5 -
integrate | evaluation-metrics | 20 papers integrated (phase 1 only) | Q3-scoped (published_date < 2025-10-01) oldest-first continuation, 180 -> 200 of 286 in-scope candidates; processed interspeech-2025-1531 through interspeech-2025-2447 (1531, 1536, 1550, 1726, 1747, 1763, 1776, 1819, 1873, 1940, 1993, 2031, 2032, 2043, 2151, 2159, 2189, 2283, 2328, 2447); the standing exclusion 2207.12598 (Classifier-Free Diffusion Guidance) was re-confirmed as still the oldest unintegrated in-scope candidate (re-derived independently from frontmatter, off-topic ImageNet diffusion paper with
task: []and no speech content) and excluded again for the same reason as batches 1-9, not replaced since the 20-entry cap was reached from the next oldest genuine candidate onward; no other exclusions this batch; 2 papers use the legacy bare-claims format (interspeech-2025-1993, interspeech-2025-2043) plus 2 more discovered mid-batch (interspeech-2025-2447, and interspeech-2025-2449 which was read but excluded from this batch’s write since it falls after the 20-cap), handled per docs/schemas/claims.md compatibility rules with evidence synthesized from Method/Key Results and role defaulted to supports unless wording warranted complicates/contradicts; 85 in-scope papers remain for follow-up Phase 1 batches (next: interspeech-2025-2449, then interspeech-2025-2536, interspeech-2025-2573, interspeech-2025-2595, interspeech-2025-2660, …) | phase 2 synthesis not yet run for this concept | runtime: claude-code | provider: anthropic | model: claude-sonnet-5 -
integrate | evaluation-metrics | 20 papers integrated (phase 1 only) | Q3-scoped (published_date < 2025-10-01) oldest-first continuation, 160 -> 180 of 286 in-scope candidates; processed interspeech-2025-0816 through interspeech-2025-1494 (0816, 0854, 0902, 0973, 0984, 0989, 0998, 1020, 1066, 1081, 1084, 1098, 1115, 1122, 1229, 1334, 1364, 1394, 1478, 1494); the standing exclusion 2207.12598 (Classifier-Free Diffusion Guidance) was re-confirmed as still the oldest unintegrated in-scope candidate (re-derived independently from frontmatter, off-topic ImageNet diffusion paper with
task: []and no speech content) and excluded again for the same reason as batches 1-8, not replaced since the 20-entry cap was reached from the next oldest genuine candidate onward; no other exclusions this batch; 105 in-scope papers remain for follow-up Phase 1 batches (next: interspeech-2025-1531, then interspeech-2025-1536, interspeech-2025-1550, interspeech-2025-1726, …) | phase 2 synthesis not yet run for this concept | runtime: claude-code | provider: anthropic | model: claude-sonnet-5 -
integrate | evaluation-metrics | 20 papers integrated (phase 1 only) | Q3-scoped (published_date < 2025-10-01) oldest-first continuation, 140 -> 160 of 286 in-scope candidates; processed interspeech-2025-0310 through interspeech-2025-0779 (0310, 0347, 0355, 0383, 0401, 0406, 0433, 0438, 0464, 0469, 0506, 0551, 0554, 0648, 0656, 0704, 0706, 0739, 0756, 0779); the standing exclusion 2207.12598 (Classifier-Free Diffusion Guidance) was re-confirmed as still the oldest unintegrated in-scope candidate (re-derived independently from frontmatter rather than trusted from prior batch preview text) and excluded again for the same reason as batches 1-7, not replaced since the 20-entry cap was reached from the next oldest genuine candidate onward; no other exclusions this batch; 125 in-scope papers remain for follow-up Phase 1 batches (next: interspeech-2025-0816, then interspeech-2025-0854, interspeech-2025-0902, interspeech-2025-0973, …) | phase 2 synthesis not yet run for this concept | runtime: claude-code | provider: anthropic | model: claude-sonnet-5
-
integrate | evaluation-metrics | 20 papers integrated (phase 1 only) | Q3-scoped (published_date < 2025-10-01) oldest-first continuation, 120 -> 140 of 286 in-scope candidates; processed 2508.06890 through interspeech-2025-0305; the standing exclusion 2207.12598 (Classifier-Free Diffusion Guidance) was re-confirmed as still the oldest unintegrated in-scope candidate and excluded again for the same reason as batches 1-6, not replaced since the 20-entry cap was reached from the next oldest genuine candidate onward; no other exclusions this batch; 145 in-scope papers remain for follow-up Phase 1 batches (next: interspeech-2025-0310, then interspeech-2025-0347, interspeech-2025-0355, interspeech-2025-0383, …) | phase 2 synthesis not yet run for this concept | runtime: claude-code | provider: anthropic | model: claude-sonnet-5
2026-07-20
-
integrate | evaluation-metrics | 20 papers integrated (phase 1 only) | Q3-scoped (published_date < 2025-10-01) oldest-first continuation, 100 -> 120 of 286 in-scope candidates; processed 2025.findings-acl.470 through 2508.06870; the standing exclusion 2207.12598 (Classifier-Free Diffusion Guidance) was re-confirmed as still the oldest unintegrated in-scope candidate and excluded again for the same reason as batches 1-5, not replaced since the 20-entry cap was reached from the next oldest genuine candidate onward; no other exclusions this batch; 166 in-scope papers remain for follow-up Phase 1 batches (next: 2207.12598 - standing exclusion to re-confirm and exclude again, then 2508.06890, 2508.07426, 2508.07273, 2508.07375, …) | phase 2 synthesis not yet run for this concept | runtime: claude-code | provider: anthropic | model: claude-sonnet-5
-
integrate | evaluation-metrics | 20 papers integrated (phase 1 only) | Q3-scoped (published_date < 2025-10-01) oldest-first continuation, 80 -> 100 of 286 in-scope candidates; processed 2507.09282 through 2025.findings-acl.1226; the standing exclusion 2207.12598 (Classifier-Free Diffusion Guidance, off-topic ImageNet diffusion paper, excluded in batches 1-4) was not re-flagged by this batch’s own candidate list, but independent verification confirms it remains absent from the YAML and is still the true oldest unintegrated in-scope candidate overall, unaffected by this batch’s real progress; 186 in-scope papers remain for follow-up Phase 1 batches (next: 2207.12598 - standing exclusion to re-confirm and exclude again, then 2025.findings-acl.470, 2025.findings-acl.534, 2025.findings-acl.71, 2507.17527, …) | phase 2 synthesis not yet run for this concept | runtime: claude-code | provider: anthropic | model: claude-sonnet-5
-
integrate | evaluation-metrics | 20 papers integrated (phase 1 only) | Q3-scoped (published_date < 2025-10-01) oldest-first continuation, 60 -> 80 of 286 in-scope candidates; processed 2025.americasnlp-1.1 through 2507.08319; 1 candidate skipped as scope-mismatch on re-encounter (2207.12598, Classifier-Free Diffusion Guidance: same off-topic ImageNet diffusion paper excluded in batches 1-3, resurfaced again as the oldest unintegrated candidate since it was never written to the YAML, excluded again for the same reason), not replaced in this batch since the 20-entry cap was reached from the next oldest candidate onward; 205 in-scope papers remain for follow-up Phase 1 batches (next: 2507.09282, 2507.09310, 2506.18296, 2507.10985, …) | phase 2 synthesis not yet run for this concept | runtime: claude-code | provider: anthropic | model: claude-sonnet-5
-
integrate | evaluation-metrics | 20 papers integrated (phase 1 only) | first-ever integration pass for this concept, new YAML created; Q3-scoped (published_date < 2025-10-01) oldest-first batch, 20 of 286 in-scope candidates processed (1703.10135 through 2403.16973); 1 candidate skipped as scope-mismatch (2207.12598, Classifier-Free Diffusion Guidance: class-conditional ImageNet paper with no speech-generation content, does not genuinely fit the concept despite being wikilinked from evaluation-metrics), replaced in the batch by the next oldest-first candidate (2403.16973); 265 in-scope papers remain for follow-up Phase 1 batches (next: 2404.03204, 2406.00654, 2406.04904, …) | phase 2 synthesis not yet run for this concept | runtime: claude-code | provider: anthropic | model: claude-sonnet-5
-
misc | related_concepts normalization | 354 paper pages (348 bracket-quoted, 6 YAML block-list) rewritten to canonical bracket-unquoted
related_concepts: [slug-one, slug-two]form viascripts/normalize_related_concepts.py; only therelated_conceptsline(s) touched per file, verified viagit diffandhealth_check.py(0 errors before and after) | runtime: claude-code | provider: anthropic | model: claude-sonnet-5 -
misc | concept cleanup | 21 stale concept pages replaced with pending-integration placeholders; concept index split into rendered and pending sections | runtime: codex | provider: openai | model: gpt-5
-
render | 2 concepts | mode: full | runtime: codex | provider: openai | model: gpt-5
-
integrate | rlhf-speech | 20 papers integrated (phase 1 only) | first-ever integration pass for this concept, new YAML created; Q3-scoped (published_date < 2025-10-01) oldest-first batch, 20 of 29 in-scope candidates processed (2406.00654 through 2508.15442), 2 Tier 2 candidates skipped (2501.12948, 2307.09288); 9 in-scope papers remain for a follow-up Phase 1 batch (2508.16332, 2509.00685, 2509.05863, 2509.14946, 2501.04561, 2509.18531, 2509.18928, 2509.19928, 2509.25416) | phase 2 synthesis not yet run for this concept | runtime: claude-code | provider: anthropic | model: claude-sonnet-5
-
integrate | rlhf-speech | 9 papers integrated (phase 1 only) | second and final Phase 1 batch for this concept, Q3-scoped oldest-first continuation, 20 -> 29 of 29 in-scope candidates; processed 2508.16332 through 2509.25416; no candidates skipped (all 9 confirmed Tier 1, ingested, not previously integrated); Q3-and-earlier Phase 1 backlog for rlhf-speech is now fully closed (29/29) | phase 2 synthesis not yet run for this concept | runtime: claude-code | provider: anthropic | model: claude-sonnet-5
-
integrate | phase 2 only | rlhf-speech (29 papers) | 20 claim_clusters (20 new) | 5 method_families (5 new) | 3 reassessment_queue items added | first Phase 2 synthesis run for this concept; headline strongly_supported cluster rlhf_posttraining_beats_sft_baseline (11 supporting papers) confirms preference/reward-based post-training consistently beats SFT-only baselines; rlhf_reward_hacking_risk strongly_supported across 7 supporting + 2 refining papers with no disputing paper; 0 contested clusters (unlike flow-matching and speech-to-speech), read as too few independent replications yet rather than genuine consensus; 4 papers left as method_family outliers (2025.naacl-demo.12 toolkit/demo, 2025.acl-long.682 survey, 2508.15442 GOAT single-paper GFlowNet paradigm below the 2-paper family threshold, 2509.19928 evaluation-methodology paper with DPO incidental) | runtime: claude-code | provider: anthropic | model: claude-sonnet-5
-
integrate | evaluation-metrics | 20 papers integrated (phase 1 only) | Q3-scoped (published_date < 2025-10-01) oldest-first continuation, 20 -> 40 of 286 in-scope candidates; processed 2404.03204 through 2502.04128; 1 candidate skipped as scope-mismatch on re-encounter (2207.12598, Classifier-Free Diffusion Guidance: same off-topic ImageNet diffusion paper excluded in batch 1, resurfaced as the oldest unintegrated candidate since it was never written to the YAML, excluded again for the same reason), not replaced in this batch since the 20-entry cap was reached from the next oldest candidate onward; 245 in-scope papers remain for follow-up Phase 1 batches (next: 2502.06490, 2025.computel-main.6, 2025.nodalida-1.32, 2504.08528, …) | phase 2 synthesis not yet run for this concept | runtime: claude-code | provider: anthropic | model: claude-sonnet-5
-
integrate | evaluation-metrics | 20 papers integrated (phase 1 only) | Q3-scoped (published_date < 2025-10-01) oldest-first continuation, 40 -> 60 of 286 in-scope candidates; processed 2502.06490 through 2025.iwsds-1.27; 1 candidate skipped as scope-mismatch on re-encounter (2207.12598, Classifier-Free Diffusion Guidance: same off-topic ImageNet diffusion paper excluded in batches 1 and 2, resurfaced again as the oldest unintegrated candidate since it was never written to the YAML, excluded again for the same reason), not replaced in this batch since the 20-entry cap was reached from the next oldest candidate onward; 225 in-scope papers remain for follow-up Phase 1 batches (next: 2025.americasnlp-1.1, 2505.09558, 2505.14648, 2507.06235, …) | phase 2 synthesis not yet run for this concept | runtime: claude-code | provider: anthropic | model: claude-sonnet-5
2026-07-19
- integrate | speech-to-speech | 6 papers integrated (phase 1 only) | Q3-scoped oldest-first continuation, 54 -> 60; processed the final 7 in-scope candidates (2509.17765 through 2509.26542), 1 skipped as scope-mismatch (2509.23938, Easy Turn: turn-taking detection module with no speech-generation stage of its own, does not fit any of the 3 S2S sub-paradigms, same reasoning as the VAP/DiVA precedents); Q3-and-earlier Phase 1 backlog for speech-to-speech is now fully closed | phase 2 synthesis not yet run for this concept
- integrate | speech-to-speech | 19 papers integrated (phase 1 only) | Q3-scoped oldest-first continuation, 35 -> 54; processed the next 20 oldest-first candidates (2025.findings-acl.75 through 2508.18240), 1 skipped as scope-mismatch (interspeech-2025-2660, Triadic VAP: acoustic-only turn-taking predictor with no speech-generation stage, does not fit any of the 3 S2S sub-paradigms, same reasoning as the DiVA precedent); 7 in-scope papers remain, continuing oldest-first from 2509.17765 (2025-09-22) through 2509.26542 (2025-09-30) | phase 2 synthesis deferred
- integrate | speech-to-speech | 19 papers integrated (phase 1 only) | Q3-scoped oldest-first continuation, 16 -> 35 of 47 in-scope candidates; 1 paper (2025.acl-long.388, DiVA) skipped as scope-mismatch (speech-in/text-out, no speech output, does not fit any of the 3 S2S sub-paradigms); 27 in-scope papers remain, continuing oldest-first from 2025.findings-acl.75 (2025-07-27) through 2509.26542 (2025-09-30) | phase 2 synthesis deferred
- integrate | flow-matching | 20 papers integrated (phase 1 only) | oldest-first continuation of round-1 production run (34 -> 54 of 112 candidates) | phase 2 synthesis deferred
- integrate | flow-matching | 20 papers integrated (phase 1 only) | oldest-first continuation of round-1 production run (54 -> 74 of 109 candidates) | phase 2 synthesis deferred
- integrate | flow-matching | 20 papers integrated (phase 1 only) | Q3-scoped oldest-first continuation, 74 -> 94; 3 in-scope papers remain (2509.24650, 2509.24773, 2509.26514); 11 Q4-2025-or-later candidates explicitly excluded per user-supplied scope cutoff | phase 2 synthesis deferred
- integrate | 3 papers | 1 concepts updated | 0 clusters updated | 0 reassessments checked (phase 1 only; flow-matching Q3-and-earlier backlog closed, 97/97)
- integrate | phase 2 only | flow-matching (97 papers) | 34 claim_clusters (14 new, 12 updated) | 7 method_families (2 new) | 3 reassessment_queue items added
- integrate | speech-to-speech | 16 papers integrated (phase 1 only, first pass, new YAML created) | Q3-scoped oldest-first batch, 0 -> 16 of 71 in-scope candidates; 4 papers skipped as Tier 2 despite appearing on the supplied ID list (2308.11596, 2408.05211, 2410.21276, 2501.01957); 55 in-scope papers remain, continuing oldest-first from 2501.06282 | phase 2 synthesis deferred
- integrate | phase 2 only | speech-to-speech (60 papers) | 16 claim_clusters (16 new) | 5 method_families (5 new) | 5 reassessment_queue items added | first Phase 2 synthesis run for this concept; surfaced staged_pretraining_effect_on_instruction_following as contested (SLAM-Omni vs Baichuan-Audio) and cascade_outperforms_e2e_on_benchmarks as strongly_supported (10 supporting + 5 refining papers); flagged a second duplicate paper pair (2412.15649 / 2025.findings-acl.115, SLAM-Omni) alongside the already-known LLaMA-Omni 2 pair
2026-07-18
- misc | 2407.05361 | Emilia (re-ingested, merged with arXiv:2501.15907 full version) | canonical ID stays 2407.05361 (SLT 2024 published paper); content updated to reflect the 216k-hour Emilia-Large scale and expanded scaling-law/multilingual experiments from the full arXiv version, which is not separately tracked (marked duplicate, canonical_id: 2407.05361)
- ingest | 2506.15556 | PredGen: Accelerated Inference of Large Language Models through Input-Time Speculation for Real-Time Speech Interaction | arXiv 2025
- ingest | 2510.07096 | Modeling Sarcastic Speech: Semantic and Prosodic Cues in a Speech Synthesis Framework | arXiv 2025
- ingest | 2510.06917 | SHANKS: Simultaneous Hearing and Thinking for Spoken Language Models | arXiv 2025
- ingest | 2510.06927 | Position: Towards Responsible Evaluation for Text-to-Speech | arXiv 2025
- ingest | 2510.07881 | CS3-Bench: Evaluating and Enhancing Speech-to-Speech LLMs for Mandarin-English Code-Switching | arXiv 2025
- ingest | 2510.08373 | DialoSpeech: Dual-Speaker Dialogue Generation with LLM and Flow Matching | arXiv 2025
- ingest | 2510.08392 | MeanVC: Lightweight and Streaming Zero-Shot Voice Conversion via Mean Flows | arXiv 2025
- ingest | 2510.07978 | VoiceAgentBench: Are Voice Assistants ready for agentic tasks? | arXiv 2025
- ingest | 2510.09061 | O_O-VC: Synthetic Data-Driven One-to-One Alignment for Any-to-Any Voice Conversion | EMNLP 2025
- ingest | 2510.09016 | DiTSinger: Scaling Singing Voice Synthesis with Diffusion Transformer and Implicit Alignment | arXiv 2025
- ingest | 2506.12311 | Phonikud: Hebrew Grapheme-to-Phoneme Conversion for Real-Time Text-to-Speech | arXiv 2025
- ingest | 2510.09424 | The Speech-LLM Takes It All: A Truly Fully End-to-End Spoken Dialogue State Tracking Approach | arXiv 2025
- ingest | 2510.09592 | Mind-Paced Speaking: A Dual-Brain Approach to Real-Time Reasoning in Spoken Language Models | arXiv 2025
- ingest | 2510.09245 | SynthVC: Leveraging Synthetic Data for End-to-End Low Latency Streaming Voice Conversion | arXiv 2025
- ingest | 2510.10003 | MTP-S2UT: Enhancing Speech-to-Speech Translation Quality with Multi-token Prediction | arXiv 2025
- ingest | 2510.10774 | ParsVoice: A Large-Scale Multi-Speaker Persian Speech Corpus for Text-to-Speech Synthesis | arXiv 2025
- ingest | 2510.11646 | BridgeCode: A Dual Speech Representation Paradigm for Autoregressive Zero-Shot Text-to-Speech Synthesis | arXiv 2025
- ingest | 2510.11124 | Perturbation Self-Supervised Representations for Cross-Lingual Emotion TTS: Stage-Wise Modeling of Emotion and Speaker | arXiv 2025
- ingest | 2510.12964 | VCTR: A Transformer-Based Model for Non-parallel Voice Conversion | arXiv 2025
- ingest | 2510.12995 | Continuous-Token Diffusion for Speaker-Referenced TTS in Multimodal LLMs | arXiv 2025
- ingest | 2510.13221 | Acoustic Teleportation via Disentangled Neural Audio Codec Representations | arXiv 2025
- ingest | 2510.13293 | Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS Models | arXiv 2025
- ingest | 2510.13194 | StressTransfer: Stress-Aware Speech-to-Speech Translation with Emphasis Preservation | arXiv 2025
- ingest | 2510.15364 | LDCodec: A high quality neural audio codec with low-complexity decoder | arXiv 2025
- ingest | 2510.15227 | LongCat-Audio-Codec: An Audio Tokenizer and Detokenizer Solution Designed for Speech Large Language Models | arXiv 2025
- ingest | 2510.16841 | SAC: Neural Speech Codec with Semantic-Acoustic Dual-Stream Quantization | arXiv 2025
- ingest | 2510.16718 | U-Codec: Ultra Low Frame-rate Neural Speech Codec for Fast High-fidelity Speech Generation | arXiv 2025
- ingest | 2503.06211 | Late Fusion and Multi-Level Fission Amplify Cross-Modal Transfer in Text-Speech LMs | arXiv 2025
- ingest | 2510.18308 | ParaStyleTTS: Toward Efficient and Robust Paralinguistic Style Control for Expressive Text-to-Speech Generation | arXiv 2025
- ingest | 2506.23670 | Efficient Interleaved Speech Modeling through Knowledge Distillation | arXiv 2025
- ingest | 2510.19509 | Which Evaluation for Which Model? A Taxonomy for Speech Model Assessment | arXiv 2025
- ingest | 2510.10785 | FAC-FACodec: Controllable Zero-Shot Foreign Accent Conversion with Factorized Speech Codec | arXiv 2025
2026-07-17
- ingest | 2509.22062 | Comprehend and Talk: Text to Speech Synthesis via Dual Language Modeling | arXiv 2025
- ingest | 2509.22167 | Semantic-VAE: Semantic-Alignment Latent Representation for Better Speech Synthesis | arXiv 2025
- ingest | 2509.22243 | FLEXI: Benchmarking Full-duplex Human-LLM Speech Interaction | arXiv 2025
- ingest | 2509.23147 | BFA: Real-time Multilingual Text-to-speech Forced Alignment | arXiv 2025
- ingest | 2510.02352 | Evaluating Bias in Spoken Dialogue LLMs for Real-World Decisions and Recommendations | arXiv 2025
- ingest | 2509.23938 | Easy Turn: Integrating Acoustic and Linguistic Modalities for Robust Turn-Taking in Full-Duplex Spoken Dialogue Systems | arXiv 2025
- ingest | 2509.24457 | Assessing speech quality metrics for evaluation of neural audio codecs under clean speech conditions | arXiv 2025
- ingest | 2509.24570 | ISSE: An Instruction-Guided Speech Style Editing Dataset And Benchmark | arXiv 2025
- ingest | 2509.24650 | VoxCPM: Tokenizer-Free TTS for Context-Aware Speech Generation and True-to-Life Voice Cloning | arXiv 2025
- ingest | 2509.24773 | VSSFlow: Unifying Video-conditioned Sound and Speech Generation via Joint Learning | arXiv 2025
- ingest | 2509.25131 | MGM-Omni: Scaling Omni LLMs to Personalized Long-Horizon Speech | arXiv 2025
- ingest | 2509.25416 | Emotion-Aligned Generation in Diffusion Text to Speech Models via Preference-Guided Optimization | arXiv 2025
- ingest | 2509.26276 | Optimizing Speech Language Models for Acoustic Consistency | arXiv 2025
- ingest | 2509.26514 | BatonVoice: An Operationalist Framework for Enhancing Controllable Speech Synthesis with Linguistic Intelligence from LLMs | arXiv 2025
- ingest | 2509.26542 | Voice Evaluation of Reasoning Ability: Diagnosing the Modality-Induced Performance Gap | arXiv 2025
- ingest | 2510.00264 | Baseline Systems For The 2025 Low-Resource Audio Codec Challenge | arXiv 2025
- ingest | 2510.00499 | MOSS-Speech: Towards True Speech-to-Speech Models Without Text Guidance | arXiv 2025
- ingest | 2510.00743 | From Scores to Preferences: Redefining MOS Benchmarking for Speech Quality Reward Modeling | arXiv 2025
- ingest | 2025.vlsp-1.15 | Twinkle-VC: A Robust and High-Quality Zero-Shot Voice Conversion System for the VLSP 2025 Shared Task | VLSP 2025
- ingest | 2025.vlsp-1.14 | ViettelRoar: Voice conversion approach for VLSP 2025 | VLSP 2025
- ingest | 2025.vlsp-1.13 | The 2025 VLSP Task on Vietnamese Voice Conversion: Overview and Preliminary Results | VLSP 2025
- ingest | 2510.05150 | Chronological Thinking in Full-Duplex Spoken Dialogue Language Models | arXiv 2025
- ingest | 2510.02066 | Chain-of-Thought Reasoning in Streaming Full-Duplex End-to-End Spoken Dialogue Systems | arXiv 2025
- ingest | 2510.01722 | Emotional Text-To-Speech Based on Mutual-Information-Guided Emotion-Timbre Disentanglement | arXiv 2025
- ingest | 2510.01903 | MelTok: 2D Tokenization for Single-Codebook Audio Compression | arXiv 2025
- ingest | 2510.02044 | Stream RAG: Instant and Accurate Spoken Dialogue Systems with Streaming Tool Usage | arXiv 2025
- ingest | 2510.03111 | Evaluation of preprocessing pipelines in the creation of in-the-wild TTS datasets | arXiv 2025
- ingest | 2510.03735 | Soft Disentanglement in Frequency Bands for Neural Audio Codecs | arXiv 2025
- ingest | 2510.04738 | Speak, Edit, Repeat: High-Fidelity Voice Editing and Zero-Shot TTS with Cross-Attentive Mamba | arXiv 2025
- ingest | 2510.05619 | Teaching Machines to Speak Using Articulatory Control | arXiv 2025
- ingest | 2510.05984 | ECTSpeech: Enhancing Efficient Speech Synthesis via Easy Consistency Tuning | arXiv 2025
- ingest | 2510.05799 | Data-efficient Targeted Token-level Preference Optimization for LLM-based Text-to-Speech | arXiv 2025
2026-07-16
- ingest | 2506.21875 | WildSpeech-Bench: Benchmarking End-to-End SpeechLLMs in the Wild | arXiv 2025
- ingest | 2509.22727 | DiaMoE-TTS: A Unified IPA-Based Dialect TTS Framework with Mixture-of-Experts and Parameter-Efficient Zero-Shot Adaptation | arXiv 2025
- ingest | 2509.19928 | Measuring Prosody Diversity in Zero-Shot TTS: A New Metric, Benchmark, and Exploration | arXiv 2025
- ingest | 2509.20086 | OLaPh: Optimal Language Phonemizer | arXiv 2025
- ingest | 2509.20321 | Conversational Speech Reveals Structural Robustness Failures in SpeechLLM Backbones | arXiv 2025
- ingest | 2509.20410 | Phoenix-VAD: Streaming Semantic Endpoint Detection for Full-Duplex Speech Interaction | arXiv 2025
- ingest | 2509.20485 | Objective Evaluation of Prosody and Intelligibility in Speech Synthesis via Conditional Prediction of Discrete Tokens | arXiv 2025
- ingest | 2509.22718 | PerformSinger: Multimodal Singing Voice Synthesis Leveraging Synchronized Lip Cues from Singing Performance Videos | arXiv 2025
- ingest | 2505.10599 | UDDETTS: Unifying Discrete and Dimensional Emotions for Controllable Emotional Text-to-Speech | arXiv 2025
- ingest | 2509.20802 | SPADE: Structured Pruning and Adaptive Distillation for Efficient LLM-TTS | arXiv 2025
- ingest | 2509.21968 | AUV: Teaching Audio Universal Vector Quantization with Single Nested Codebook | arXiv 2025
2026-07-15
- ingest | 2509.19231 | Finding My Voice: Generative Reconstruction of Disordered Speech for Automated Clinical Evaluation | arXiv 2025
- ingest | 2509.19592 | Frame-Stacked Local Transformers For Efficient Multi-Codebook Speech Generation | arXiv 2025
- ingest | 2509.19812 | Efficient Speech Watermarking for Speech Synthesis via Progressive Knowledge Distillation | arXiv 2025
- ingest | 2509.19883 | CoMelSinger: Discrete Token-Based Zero-Shot Singing Synthesis With Structured Melody Control and Guidance | arXiv 2025
- ingest | 2509.18823 | Towards Evaluating Generative Audio: Insights from Neural Audio Codec Embedding Distances | arXiv 2025
- ingest | 2509.18928 | Direct Preference Optimization for Speech Autoregressive Diffusion Models | arXiv 2025
- ingest | 2509.19025 | Enhancing Noise Robustness for Neural Speech Codecs through Resource-Efficient Progressive Quantization Perturbation Simulation | arXiv 2025
- ingest | 2509.19186 | Improving Test-Time Performance of RVQ-based Neural Codecs | arXiv 2025
- ingest | 2509.18470 | Discrete-Time Diffusion-Like Models for Speech Synthesis | arXiv 2025
- ingest | 2501.04561 | OpenOmni: Advancing Open-Source Omnimodal Large Language Models with Progressive Multimodal Alignment and Real-Time Self-Aware Emotional Speech Synthesis | arXiv 2025
- ingest | 2509.18531 | No Verifiable Reward for Prosody: Toward Preference-Guided Prosody Learning in TTS | arXiv 2025
- ingest | 2509.18806 | Rethinking the joint estimation of magnitude and phase for time-frequency domain neural vocoders | arXiv 2025
- ingest | 2509.17516 | Audiobook-CC: Controllable Long-context Speech Generation for Multicast Audiobook | arXiv 2025
- ingest | 2509.17765 | Qwen3-Omni Technical Report | arXiv 2025
- ingest | 2509.17988 | Nord-Parl-TTS: Finnish and Swedish TTS Dataset from Parliament Speech | arXiv 2025
- ingest | 2509.18060 | TMD-TTS: A Unified Tibetan Multi-Dialect Text-to-Speech Framework for Ü-Tsang, Amdo and Kham Speech Dataset Generation | arXiv 2025
- misc | removed
wiki/venues/(27 per-venue-year pages + index.md) and the ingest-time instructions that auto-updated them — most were thin paper listings that accumulated one row per ingest with no synthesis; venue trend reports will instead be generated on demand for venues with enough papers to support real synthesis (e.g. a full conference), not auto-maintained per paper - integrate | 20 papers | 1 concepts updated | 20 clusters updated | 0 reassessments checked
2026-07-14
- ingest | 2509.13989 | Do You Hear What I Mean? Quantifying the Instruction-Perception Gap in Instruction-Guided Expressive Text-To-Speech Systems | arXiv 2025
- ingest | 2509.14579 | Cross-Lingual F5-TTS: Towards Language-Agnostic Voice Cloning and Speech Synthesis | arXiv 2025
- ingest | 2509.14684 | DAIEN-TTS: Disentangled Audio Infilling for Environment-Aware Text-to-Speech Synthesis | arXiv 2025
- ingest | 2509.14784 | MELA-TTS: Joint transformer-diffusion model with representation alignment for speech synthesis | arXiv 2025
- ingest | 2509.14946 | SynParaSpeech: Automated Synthesis of Paralinguistic Datasets for Speech Generation and Understanding | arXiv 2025
- ingest | 2509.15085 | Real-Time Streaming Mel Vocoding with Generative Flow Matching | arXiv 2025
- ingest | 2509.15253 | Emotion-Aware Speech Generation with Character-Specific Voices for Comics | arXiv 2025
- ingest | 2509.15462 | A Novel Semantic Compression Approach for Ultra-low Bandwidth Voice Communication | arXiv 2025
- ingest | 2505.17093 | P2VA: Converting Persona Descriptions into Voice Attributes for Fair and Controllable Text-to-Speech | arXiv 2025
- ingest | 2509.15626 | LibriTTS-VI: A Public Corpus and Novel Methods for Efficient Voice Impression Control | arXiv 2025
- ingest | 2509.15629 | The Singing Voice Conversion Challenge 2025: From Singer Identity Conversion To Singing Style Conversion | arXiv 2025
- ingest | 2509.15845 | Deep Dubbing: End-to-End Auto-Audiobook System with Text-to-Timbre and Context-Aware Instruct-TTS | arXiv 2025
- ingest | 2509.16010 | Fed-PISA: Federated Voice Cloning via Personalized Identity-Style Adaptation | arXiv 2025
- ingest | 2509.16195 | FocalCodec-Stream: Streaming Low-Bitrate Speech Coding via Causal Distillation | arXiv 2025
- ingest | 2509.16589 | Benchmarking Contextual and Paralinguistic Reasoning in Speech-LLMs: A Case Study with In-the-Wild Data | EMNLP 2025
- ingest | 2509.20378 | Beyond Global Emotion: Fine-Grained Emotional Speech Synthesis with Dynamic Word-Level Modulation | arXiv 2025
- ingest | 2509.17006 | MBCodec:Thorough disentangle for high-fidelity audio compression | arXiv 2025
- ingest | 2509.17021 | Bridging the gap between training and inference in LM-based TTS models | arXiv 2025
- ingest | 2509.17143 | MaskVCT: Masked Voice Codec Transformer for Zero-Shot Voice Conversion With Increased Controllability via Multiple Guidances | arXiv 2025
- ingest | 2509.14882 | Llama-Mimi: Exploring the Limits of Flattened Speech Language Modeling | arXiv 2025
2026-07-13
- ingest | 2509.09201 | DeCodec: Rethinking Audio Codecs as Universal Disentangled Representation Learners | arXiv 2025
- ingest | 2509.09550 | Finite Scalar Quantization Enables Redundant and Transmission-Robust Neural Audio Compression at Low Bit-rates | arXiv 2025
- ingest | 2509.09748 | DiTReducio: A Training-Free Acceleration for DiT-Based TTS via Progressive Calibration | arXiv 2025
- ingest | 2509.11084 | Length-Aware Rotary Position Embedding for Text-Speech Alignment | arXiv 2025
- ingest | 2509.11425 | FuseCodec: Semantic-Contextual Fusion and Supervision for Neural Codecs | arXiv 2025
- ingest | 2508.18240 | MTalk-Bench: Evaluating Speech-to-Speech Models in Multi-Turn Dialogues via Arena-style and Rubrics Protocols | arXiv 2025
- ingest | 2509.12171 | Preservation of Language Understanding Capabilities in Speech-aware Large Language Models | arXiv 2025
- ingest | 2509.14270 | SpeechWeave: Diverse Multilingual Synthetic Text & Audio Data Generation Pipeline for Training Text to Speech Models | ACL 2025
- ingest | 2509.12831 | A Lightweight Pipeline for Noisy Speech Voice Cloning and Accurate Lip Sync Synthesis | arXiv 2025
- ingest | 2509.13068 | MSR-Codec: A Low-Bitrate Multi-Stream Residual Codec for High-Fidelity Speech Generation with Information Disentanglement | arXiv 2025
- ingest | 2412.16846 | KALL-E:Autoregressive Speech Synthesis with Next-Distribution Prediction | arXiv 2025
- ingest | 2504.20581 | ClonEval: An Open Voice Cloning Benchmark | arXiv 2025
- ingest | interspeech-2025-cho25c_interspeech | Unleashing the Inner Monster: Demonstrating High-Fidelity Human to Non-Human Voice Conversion | Interspeech 2025
- ingest | interspeech-2025-gourav25_interspeech | Code Mix TTS: An Approach to Infer Human Like Speech for Multi-Lingual Input Texts | Interspeech 2025
- ingest | interspeech-2025-raju25_interspeech | End-to-End Indian Language Dubbing with Zero-Shot Speaker Preservation | Interspeech 2025
- ingest | 2509.13667 | A Distilled Low-Latency Neural Vocoder with Explicit Amplitude and Phase Prediction | arXiv 2025
- ingest | 2509.13670 | A High-Quality and Low-Complexity Streamable Neural Speech Codec with Knowledge Distillation | arXiv 2025
- misc | 2509.13785 | Ingest reverted: pure ASR + speaker-diarization challenge summary (MER/tcpMER only), no TTS/VC/SCA generative component; filter-stage false accept, same pattern as FAMA; user-confirmed reject after reading the PDF | arXiv 2025
2026-07-12
- ingest | 2506.23367 | You Sound a Little Tense: L2 Tailored Clear TTS Using Durational Vowel Properties | arXiv 2025
- ingest | 2509.05359 | An Empirical Analysis of Discrete Unit Representations in Speech Language Modeling Pre-training | arXiv 2025
- ingest | 2509.04093 | Open-Source Full-Duplex Conversational Datasets for Natural and Interactive Speech Synthesis | arXiv 2025
- ingest | 2509.04667 | DarkStream: real-time speech anonymization with low latency | arXiv 2025
- ingest | 2509.04685 | Say More with Less: Variable-Frame-Rate Speech Tokenization via Adaptive Clustering and Implicit Duration Coding | arXiv 2025
- ingest | 2509.04702 | OleSpeech-IV: A Large-Scale Multispeaker and Multilingual Conversational Speech Dataset with Diverse Topics | arXiv 2025
- ingest | 2509.05863 | LatinX: Aligning a Multilingual TTS Model with Direct Preference Optimization | arXiv 2025
- ingest | 2509.06074 | Multimodal Fine-grained Context Interaction Graph Modeling for Conversational Speech Synthesis | EMNLP 2025
- ingest | 2509.06502 | FireRedChat: A Pluggable, Full-Duplex Voice Interaction System with Cascaded and Semi-Cascaded Implementations | arXiv 2025
- ingest | 2509.07038 | Controllable Singing Voice Synthesis using Phoneme-Level Energy Sequence | arXiv 2025
- ingest | 2509.07376 | Progressive Facial Granularity Aggregation with Bilateral Attribute-based Enhancement for Face-to-Speech Synthesis | EMNLP 2025
- ingest | 2509.09716 | VStyle: A Benchmark for Voice Style Adaptation with Spoken Instructions | arXiv 2025
- ingest | 2509.08379 | LatentVoiceGrad: Nonparallel Voice Conversion with Latent Diffusion/Flow-Matching Models | arXiv 2025
- ingest | 2509.08696 | Accelerating Diffusion Transformer-Based Text-to-Speech with Transformer Layer Caching | arXiv 2025
- ingest | 2506.04077 | A Novel Data Augmentation Approach for Automatic Speaking Assessment on Opinion Expressions | arXiv 2025
- ingest | 2509.09174 | EchoX: Towards Mitigating Acoustic-Semantic Gap via Echo Training for Speech-to-Speech LLMs | arXiv 2025
2026-07-05
- ingest | interspeech-2025-2328 | A Watermark for Auto-Regressive Speech Generation Models | Interspeech 2025
- ingest | interspeech-2025-2536 | The Text-to-speech in the Wild (TITW) Database | Interspeech 2025
- ingest | interspeech-2025-2564 | Towards a Japanese Full-duplex Spoken Dialogue System | Interspeech 2025
- ingest | interspeech-2025-2573 | SawtArabi: A Benchmark Corpus for Arabic TTS. Standard, Dialectal and Code-Switching | Interspeech 2025
- ingest | interspeech-2025-2586 | Spotlight-TTS: Spotlighting the Style via Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech | Interspeech 2025
- ingest | interspeech-2025-2595 | Harnessing Text-to-Speech Voice Cloning Models for Improved Audiological Speech Assessment | Interspeech 2025
- ingest | interspeech-2025-2679 | Can We Reconstruct a Dysarthric Voice with the Large Speech Model Parler TTS? | Interspeech 2025
- ingest | interspeech-2025-2684 | Discl-VC: Disentangled Discrete Tokens and In-Context Learning for Controllable Zero-Shot Voice Conversion | Interspeech 2025
- ingest | interspeech-2025-2726 | DS-Codec: Dual-Stage Training with Mirror-to-NonMirror Architecture Switching for Speech Codec | Interspeech 2025
- ingest | interspeech-2025-2739 | AF-Vocoder: Artifact-Free Neural Vocoder with Global Artifact Filter | Interspeech 2025
- ingest | interspeech-2025-2787 | Towards Adaptable and Intelligible Speech Synthesis in Noisy Environments | Interspeech 2025
- ingest | interspeech-2025-2815 | From Static to Dynamic: Enhancing AAC with Generative Imagery and Zero-Shot TTS | Interspeech 2025
- ingest | interspeech-2025-bokkahallisatish25_interspeech | Hear Me Out: Interactive evaluation and bias discovery platform for speech-to-speech conversational AI | Interspeech 2025
- ingest | 2507.16835 | Evaluating Speech-to-Text x LLM x Text-to-Speech Combinations for AI Interview Systems | arXiv 2025
- ingest | 2411.19770 | Noro: Noise-Robust One-shot Voice Conversion with Hidden Speaker Representation Learning | arXiv 2025
- ingest | 2025.clicit-1.27 | Veras Audire Et Reddere Voces: A Corpus of Prosodically-Correct Latin Poetic Audio from Large-Language-Model TTS | workshop 2025
- misc | 2025.clicit-1.81 | FAMA ingest reverted: paper is a pure ASR/speech-translation model with no TTS/VC/SCA component, outside wiki scope; moved to review_queue.md for a final corpus-inclusion decision | workshop 2025
2026-07-04
- ingest | interspeech-2025-1776 | SpeechSEC: A Unified Multi-Task Framework for Speech Synthesis, Editing, and Continuation | Interspeech 2025
- ingest | interspeech-2025-1819 | Comparative Analysis of Fast and High-Fidelity Neural Vocoders for Low-Latency Streaming Synthesis in Resource-Constrained Environments | Interspeech 2025
- ingest | interspeech-2025-1873 | Can AI Understand Mandarin Speech Prosody? A Framework and Benchmark Showcase | Interspeech 2025
- ingest | interspeech-2025-1940 | Investigating Stochastic Methods for Prosody Modeling in Speech Synthesis | Interspeech 2025
- ingest | interspeech-2025-2031 | Kinship in Speech: Leveraging Linguistic Relatedness for Zero-Shot TTS in Indian Languages | Interspeech 2025
- ingest | interspeech-2025-2032 | ExagTTS: An Approach Towards Controllable Word Stress Incorporated TTS for Exaggerated Synthesized Speech Aiding Second Language Learners | Interspeech 2025
- ingest | interspeech-2025-2075 | Segmentation-Variant Codebooks for Preservation of Paralinguistic and Prosodic Information | Interspeech 2025
- ingest | interspeech-2025-2151 | FaVC: A Validated, Transcribed, Parallel Farsi Speech Dataset for Voice Conversion | Interspeech 2025
- ingest | interspeech-2025-2159 | Generating Consistent Prosodic Patterns from Open-Source TTS Systems | Interspeech 2025
- ingest | interspeech-2025-2189 | ProMode: A Speech Prosody Model Conditioned on Acoustic and Textual Inputs | Interspeech 2025
- ingest | interspeech-2025-2283 | Pairwise Evaluation of Accent Similarity in Speech Synthesis | Interspeech 2025
2026-07-03
- ingest | interspeech-2025-1020 | Learning Optimal Prosody Embedding Codebook based on F0 and Energy | Interspeech 2025
- ingest | interspeech-2025-1081 | Speaker Normalization and Content Restoration for Zero-Shot Voice Conversion with Attention-Enhanced Discriminator | Interspeech 2025
- ingest | interspeech-2025-1084 | Efficient Streaming TTS Acoustic Model with Depthwise RVQ Decoding Strategies in a Mamba Framework | Interspeech 2025
- ingest | interspeech-2025-1098 | GST-BERT-TTS: Prosody Prediction Without Accentual Labels For Multi-Speaker TTS Using BERT With Global Style Tokens | Interspeech 2025
- ingest | interspeech-2025-1106 | LSCodec: Low-Bitrate and Speaker-Decoupled Discrete Speech Codec | Interspeech 2025
- ingest | interspeech-2025-1115 | MPE-TTS: Customized Emotion Zero-Shot Text-To-Speech Using Multi-Modal Prompt | Interspeech 2025
- ingest | interspeech-2025-1192 | Voice Impression Control in Zero-Shot TTS | Interspeech 2025
- ingest | interspeech-2025-1210 | DiffEmotionVC: A Dual-Granularity Disentangled Diffusion Framework for Any-to-Any Emotional Voice Conversion | Interspeech 2025
- ingest | interspeech-2025-1229 | E2E-BPVC: End-to-End Background-Preserving Voice Conversion via In-Context Learning | Interspeech 2025
- ingest | interspeech-2025-1236 | Accelerating Diffusion-based Text-to-Speech Model Trainingwith Dual Modality Alignment | Interspeech 2025
- ingest | interspeech-2025-1334 | MiSTR: Multi-Modal iEEG-to-Speech Synthesis with Transformer-Based Prosody Prediction and Neural Phase Reconstruction | Interspeech 2025
- ingest | interspeech-2025-1364 | VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schrödinger Bridge | Interspeech 2025
- ingest | interspeech-2025-1394 | DiEmo-TTS: Disentangled Emotion Representations via Self-Supervised Distillation for Cross-Speaker Emotion Transfer in Text-to-Speech | Interspeech 2025
- ingest | interspeech-2025-1397 | VibE-SVC: Vibrato Extraction with High-frequency F0 Contour for Singing Voice Conversion | Interspeech 2025
- ingest | interspeech-2025-1434 | REWIND: Speech Time Reversal for Enhancing Speaker Representations in Diffusion-based Voice Conversion | Interspeech 2025
- ingest | interspeech-2025-1478 | LightL2S: Ultra-Low Complexity Lip-to-Speech Synthesis for Multi-Speaker Scenarios | Interspeech 2025
- ingest | interspeech-2025-1494 | VisualSpeech: Enhancing Prosody Modeling in TTS Using Video | Interspeech 2025
- ingest | interspeech-2025-1531 | Simple and Effective Content Encoder for Singing Voice Conversion via SSL-Embedding Dimension Reduction | Interspeech 2025
- ingest | interspeech-2025-1536 | Fairness in Dysarthric Speech Synthesis: Understanding Intrinsic Bias in Dysarthric Speech Cloning using F5-TTS | Interspeech 2025
- ingest | interspeech-2025-1538 | StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion | Interspeech 2025
- ingest | interspeech-2025-1550 | ArVoice: A Multi-Speaker Dataset for Arabic Speech Synthesis | Interspeech 2025
- ingest | interspeech-2025-1625 | Mimic Blocker: Self-Supervised Adversarial Training for Voice Conversion Defense with Pretrained Feature Extractors | Interspeech 2025
- ingest | interspeech-2025-1638 | EATS-Speech: Emotion-Adaptive Transformation and Priority Synthesis for Zero-Shot Text-to-Speech | Interspeech 2025
- ingest | interspeech-2025-1639 | LombardTokenizer: Disentanglement and Control of Vocal Effort in a Neural Speech Codec | Interspeech 2025
- ingest | interspeech-2025-1684 | SA-RAS: Speaker-Aware Style Retrieval Augmented Generation for Expressive Zero-Shot Text-to-Speech Synthesis | Interspeech 2025
- ingest | interspeech-2025-1726 | Voice Reconstruction through Large-Scale TTS Models: Comparing Zero-Shot and Fine-tuning Approaches to Personalise TTS in Assistive Communication | Interspeech 2025
- ingest | interspeech-2025-1747 | FasterVoiceGrad: Faster One-step Diffusion-Based Voice Conversion with Adversarial Diffusion Conversion Distillation | Interspeech 2025
- ingest | interspeech-2025-1763 | Vocoder-Projected Feature Discriminator | Interspeech 2025
2026-07-02
- ingest | 2507.09282 | ClaritySpeech: Dementia Obfuscation in Speech | arXiv 2025
- ingest | 2507.09310 | Voice Conversion for Lombard Speaking Style with Implicit and Explicit Acoustic Feature Conditioning | arXiv 2025
- ingest | 2507.10985 | Pronunciation Deviation Analysis Through Voice Cloning and Acoustic Comparison | arXiv 2025
- ingest | 2507.12197 | Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations | arXiv 2025
- ingest | 2507.14988 | DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis | arXiv 2025
- ingest | 2507.15272 | A2TTS: TTS for Low Resource Indian Languages | arXiv 2025
- ingest | 2507.16875 | Technical report: Impact of Duration Prediction on Speaker-specific TTS for Indian Languages | arXiv 2025
- ingest | 2507.21138 | TTS-1 Technical Report | arXiv 2025
- ingest | 2507.18119 | GOAT-SLM: A Spoken Language Model with Paralinguistic and Speaker Characteristic Awareness | arXiv 2025
- ingest | 2507.18897 | HH-Codec: High Compression High-fidelity Discrete Neural Codec for Spoken Language Modeling | arXiv 2025
- ingest | 2507.17527 | Seed LiveInterpret 2.0: End-to-end Simultaneous Speech-to-speech Translation with Your Voice | arXiv 2025
- ingest | 2507.20140 | Do Not Mimic My Voice: Speaker Identity Unlearning for Zero-Shot Text-to-Speech | arXiv 2025
- ingest | 2507.20731 | Learning Neural Vocoder from Range-Null Space Decomposition | arXiv 2025
- ingest | 2025.ccl-1.77 | HFSD-V2C: Zero-Shot Visual Voice Cloning Via Hierarchical Face-Styled Diffusion Model | workshop 2025
- ingest | 2025.icnlsp-1.34 | Beyond Labeled Datasets: Advancing TTS with Direct Preference Optimization on Unlabeled Speech Dataset | workshop 2025
- ingest | 2025.sigdial-1.21 | Transition Relevance Point Detection for Spoken Dialogue Systems with Self-Attention Transformer | workshop 2025
- ingest | 2025.sigdial-1.27 | EmoNews: A Spoken Dialogue System for Expressive News Conversations | workshop 2025
- ingest | 2025.sigdial-1.51 | rrSDS 2.0: Incremental, Modular, Distributed, Multimodal Spoken Dialogue with Robotic Platforms | workshop 2025
- ingest | interspeech-2025-0166 | Frozen Large Language Models Can Perceive Paralinguistic Aspects of Speech | Interspeech 2025
- ingest | interspeech-2025-0305 | DAFMSVC: One-Shot Singing Voice Conversion with Dual Attention Mechanism and Flow Matching | Interspeech 2025
- ingest | interspeech-2025-0347 | PeriodCodec: A Pitch-Controllable Neural Audio Codec Using Periodic Signals for Singing Voice Synthesis | Interspeech 2025
- ingest | interspeech-2025-0355 | Probing the Robustness Properties of Neural Speech Codecs | Interspeech 2025
- ingest | interspeech-2025-0383 | Voice Conversion for Likability Control via Automated Rating of Speech Synthesis Corpora | Interspeech 2025
- ingest | interspeech-2025-0433 | When Humans Growl and Birds Speak: High-Fidelity Voice Conversion from Human to Animal and Designed Sounds | Interspeech 2025
- ingest | interspeech-2025-0438 | LinearVC: Linear Transformations of Self-Supervised Features Through the Lens of Voice Conversion | Interspeech 2025
- ingest | interspeech-2025-0464 | Prosody-Adaptable Audio Codecs for Zero-Shot Voice Conversion via In-Context Learning | Interspeech 2025
- ingest | interspeech-2025-0506 | EnCodecMAE: leveraging neural codecs for universal audio representation learning | Interspeech 2025
- ingest | interspeech-2025-0656 | EEG-based Voice Conversion : Hearing the Voice of Your Brain | Interspeech 2025
- ingest | interspeech-2025-0706 | Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation | Interspeech 2025
- ingest | interspeech-2025-0756 | A-SMiLE: Affective Sparse Mixture-of-Experts Adapter with Multi-Task Learning for Spoken Dialogue Models | Interspeech 2025
- ingest | interspeech-2025-0998 | Voice-ENHANCE: Speech Restoration using a Diffusion-based Voice Conversion Framework | Interspeech 2025
2026-07-01
- ingest | 2025.iwsds-1.27 | A Survey of Recent Advances on Turn-taking Modeling in Spoken Dialogue Systems | IWSDS 2025
- ingest | 2505.15772 | MIKU-PAL: An Automated and Standardized Multi-Modal Method for Speech Paralinguistic and Affect Labeling | arXiv 2025
- ingest | 2507.06235 | Super Kawaii Vocalics: Amplifying the “Cute” Factor in Computer Voice | arXiv 2025
- ingest | 2506.23049 | AURA: Agent for Understanding, Reasoning, and Automated Tool Use in Voice-Driven Tasks | arXiv 2025
- ingest | 2025.acl-long.681 | SIFT-50M: A Large-Scale Multilingual Dataset for Speech Instruction Fine-Tuning | ACL 2025
- ingest | 2025.acl-long.790 | Rhythm Controllable and Efficient Zero-Shot Voice Conversion via Shortcut Flow Matching | ACL 2025
- ingest | 2025.acl-long.817 | SimulS2S-LLM: Unlocking Simultaneous Inference of Speech LLMs for Speech-to-Speech Translation | ACL 2025
- ingest | 2025.acl-long.87 | Takin-VC: Expressive Zero-Shot Voice Conversion via Adaptive Hybrid Content Encoding and Enhanced Timbre Modeling | ACL 2025
- ingest | 2025.acl-long.937 | UniCodec: Unified Audio Codec with Single Domain-Adaptive Codebook | ACL 2025
- ingest | 2025.acl-long.997 | Align-SLM: Textless Spoken Language Models with Reinforcement Learning from AI Feedback | ACL 2025
- ingest | 2025.conll-1.9 | A Linguistically Motivated Analysis of Intonational Phrasing in Text-to-Speech Systems: Revealing Gaps in Syntactic Sensitivity | CoNLL 2025
- ingest | 2025.findings-acl.101 | Chain-Talker: Chain Understanding and Rendering for Empathetic Conversational Speech Synthesis | ACL 2025
- ingest | 2025.findings-acl.115 | SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training | ACL 2025
- ingest | 2025.findings-acl.1226 | PodAgent: A Comprehensive Framework for Podcast Generation | ACL 2025
- ingest | 2025.findings-acl.470 | Does Your Voice Assistant Remember? Analyzing Conversational Context Recall and Utilization in Voice Interaction Models | ACL 2025
- ingest | 2025.findings-acl.534 | Unlocking Speech Instruction Data Potential with Query Rewriting | ACL 2025
- ingest | 2025.findings-acl.631 | Slamming: Training a Speech Language Model on One GPU in a Day | ACL 2025
- ingest | 2025.findings-acl.687 | TCSinger 2: Customizable Multilingual Zero-shot Singing Voice Synthesis | ACL 2025
- ingest | 2025.findings-acl.71 | Data-Centric Improvements for Enhancing Multi-Modal Understanding in Spoken Conversation Modeling | ACL 2025
- ingest | 2025.findings-acl.75 | Leveraging Unit Language Guidance to Advance Speech Modeling in Textless Speech-to-Speech Translation | ACL 2025
- ingest | 2025.findings-ijcnlp.49 | Incorporating Dialogue State Tracking into Japanese Full-duplex Task-oriented Spoken Dialogue Model | ACL 2025
- ingest | 2025.iwslt-1.5 | SSR: Alignment-Aware Modality Connector for Speech Language Models | IWSLT 2025
- ingest | 2025.unlp-1.11 | Context-Aware Lexical Stress Prediction and Phonemization for Ukrainian TTS Systems | workshop 2025
- ingest | 2412.18603 | Long-Form Speech Generation with Spoken Language Models | ICML 2025
- ingest | 2503.11026 | MAVFlow: Preserving Paralinguistic Elements with Conditional Flow Matching for Zero-Shot AV2AV Multilingual Translation | arXiv 2025
- ingest | 2505.15670 | SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model | arXiv 2025
- ingest | 2506.09874 | UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching | arXiv 2025
- ingest | 2506.18296 | JIS: A Speech Corpus of Japanese Idol Speakers with Various Speaking Styles | Interspeech 2025
- ingest | 2507.02176 | Analyzing and Improving Speaker Similarity Assessment for Speech Synthesis | arXiv 2025
- ingest | 2507.00808 | Multi-interaction TTS toward professional recording reproduction | arXiv 2025
- ingest | 2507.01611 | QHARMA-GAN: Quasi-Harmonic Neural Vocoder based on Autoregressive Moving Average Model | arXiv 2025
- ingest | 2507.02380 | JoyTTS: LLM-based Spoken Chatbot With Voice Cloning | arXiv 2025
- ingest | 2507.03887 | Traceable TTS: Toward Watermark-Free TTS with Strong Traceability | arXiv 2025
- ingest | 2507.03912 | Prosody Labeling with Phoneme-BERT and Speech Foundation Models | arXiv 2025
- ingest | 2507.08012 | RepeaTTS: Towards Feature Discovery through Repeated Fine-Tuning | arXiv 2025
- ingest | 2507.04349 | TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet | arXiv 2025
- ingest | 2507.04598 | Multi-Step Prediction and Control of Hierarchical Emotion Distribution in Text-to-Speech Synthesis | arXiv 2025
- ingest | 2507.04817 | Fast-VGAN: Lightweight Voice Conversion with Explicit Control of F0 and Duration Parameters | arXiv 2025
- ingest | 2507.01348 | SpeechAccentLLM: A Unified Framework for Foreign Accent Conversion and Text to Speech | arXiv 2025
- ingest | 2507.06116 | Speech Quality Assessment Model Based on Mixture of Experts: System-Level Performance Enhancement and Utterance-Level Challenge Analysis | arXiv 2025
- ingest | 2506.23325 | XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs | arXiv 2025
- ingest | 2507.07799 | SecureSpeech: Prompt-based Speaker and Content Protection | arXiv 2025
- ingest | 2507.08319 | Active Learning for Text-to-Speech Synthesis with Informative Sample Collection | arXiv 2025
- ingest | 2507.09070 | SemAlignVC: Enhancing zero-shot timbre conversion using semantic alignment | arXiv 2025
2026-06-30
- ingest | iclr-2025-hQvX9MBowC | DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors | ICLR 2025
- ingest | iclr-2025-uxDFlPGRLX | FlowDec: A flow-based full-band general audio codec with high perceptual quality | ICLR 2025
- ingest | 2025.findings-naacl.130 | DiVISe: Direct Visual-Input Speech Synthesis Preserving Speaker Characteristics And Intelligibility | NAACL 2025
- ingest | 2025.findings-naacl.279 | BnTTS: Few-Shot Speaker Adaptation in Low-Resource Setting | NAACL 2025
- ingest | 2025.findings-naacl.38 | Prompt-Guided Selective Masking Loss for Context-Aware Emotive Text-to-Speech | NAACL 2025
- ingest | 2025.naacl-demo.12 | ESPnet-SpeechLM: An Open Speech Language Model Toolkit | NAACL 2025
- ingest | 2025.naacl-demo.21 | ESPnet-SDS: Unified Toolkit and Demo for Spoken Dialogue Systems | NAACL 2025
- ingest | 2025.naacl-long.484 | Behavior-SD: Behaviorally Aware Spoken Dialogue Generation with Large Language Models | NAACL 2025
- ingest | 2025.naacl-long.591 | Robust and Unbounded Length Generalization in Autoregressive Transformer-Based Text-to-Speech | NAACL 2025
- ingest | 2025.naacl-short.65 | kNN Retrieval for Simple and Effective Zero-Shot Multi-speaker Text-to-Speech | NAACL 2025
- ingest | 2025.naacl-short.69 | Developing multilingual speech synthesis system for Ojibwe, Mi’kmaq, and Maliseet | NAACL 2025
- ingest | 2025.iwsds-1.11 | Paralinguistic Attitude Recognition for Spoken Dialogue Systems | IWSDS 2025
2026-06-29
- ingest | 2409.09098 | AccentBox: Towards High-Fidelity Zero-Shot Accent Generation | arXiv 2025
- ingest | 2025.coling-industry.29 | CarMem: Enhancing Long-Term Memory in LLM Voice Assistants through Category-Bounding | COLING Industry Track 2025
- ingest | 2025.chipsal-1.18 | Impacts of Vocoder Selection on Tacotron-based Nepali Text-To-Speech Synthesis | CHiPSAL 2025
- ingest | 2025.coling-main.685 | VoxpopuliTTS: a large-scale multilingual TTS corpus for zero-shot speech generation | COLING 2025
- ingest | 2409.20007 | DeSTA2: Developing Instruction-Following Speech Language Model Without Speech Instruction-Tuning Data | arXiv 2025
- ingest | 2025.computel-main.6 | Evaluating Indigenous language speech synthesis for education: A participatory design workshop on Ojibwe TTS | ComputEL 2025
- ingest | 2025.nodalida-1.32 | Estonian isolated-word text-to-speech synthesiser | NoDaLiDa/Baltic-HLT 2025
- ingest | 2025.naacl-srw.6 | Towards Codec-LM Co-design for Neural Codec Language Models | NAACL 2025
- ingest | 2025.findings-naacl.298 | Gender Bias in Instruction-Guided Speech Synthesis Models | NAACL 2025
- ingest | 2025.findings-naacl.471 | The Role of Prosody in Spoken Question Answering | NAACL 2025
- ingest | 2025.naacl-long.464 | ManaTTS Persian: a recipe for creating TTS datasets for lower resource languages | NAACL 2025
- ingest | 2025.naacl-long.619 | ProSE: Diffusion Priors for Speech Enhancement | NAACL 2025
- ingest | iclr-2025-tQ1PmLfPBL | PeriodWave: Multi-Period Flow Matching for High-Fidelity Waveform Generation | ICLR 2025
- ingest | iclr-2025-cuFzE8Jlvb | Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis | ICLR 2025
- ingest | iclr-2025-dGSOn7sdWg | SyllableLM: Learning Coarse Semantic Units for Speech Language Models | ICLR 2025
- ingest | iclr-2025-868masI331 | HALL-E: Hierarchical Neural Codec Language Model for Minute-Long Zero-Shot Text-to-Speech Synthesis | ICLR 2025
2026-06-28
- review | 2510.05758 | EMORL-TTS: Reinforcement Learning for Fine-Grained Emotion Control in LLM-based TTS | ICASSP 2026
- review | 2510.07979 | IntMeanFlow: Few-step Speech Generation with Integral Velocity Distillation | arXiv 2025
- review | 2510.12210 | DiSTAR: Diffusion over a Scalable Token Autoregressive Representation for Speech Generation | arXiv 2025
- review | 2511.12347 | VoiceCraft-X: Unifying Multilingual, Voice-Cloning Speech Synthesis and Speech Editing | EMNLP 2025
- review | 2512.04720 | M3-TTS: Multi-modal DiT Alignment & Mel-latent for Zero-shot High-fidelity Speech Synthesis | arXiv 2025
- review | 2512.13251 | DisCo-Speech: Controllable Zero-Shot Speech Generation with A Disentangled Speech Codec | arXiv 2025
- review | 2512.14291 | GLM-TTS Technical Report | arXiv 2025
- review | 2601.03888 | IndexTTS 2.5 Technical Report | arXiv 2026
- review | 2601.15621 | Qwen3-TTS Technical Report | arXiv 2026
- review | 2603.08823 | Fish Audio S2 Technical Report | arXiv 2026
- review | 2603.18090 | MOSS-TTS Technical Report | arXiv 2026
- review | 2603.26364 | LLaDA-TTS: Unifying Speech Synthesis and Zero-Shot Editing via Masked Diffusion Modeling | arXiv 2026
- review | 2603.29339 | LongCat-AudioDiT: High-Fidelity Diffusion Text-to-Speech in the Waveform Latent Space | arXiv 2026
- review | 2604.00688 | OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models | arXiv 2026
- review | 2604.01760 | T5Gemma-TTS Technical Report | arXiv 2026
- review | 2604.12438 | An Ultra-Low Latency, End-to-End Streaming Speech Synthesis Architecture via Block-Wise Generation and Depth-Wise Codec Decoding | arXiv 2026
2026-06-25
- review | 2508.11273 | EmoSSLSphere: Multilingual Emotional Speech Synthesis with Spherical Vectors and Discrete Speech Tokens | arXiv 2025
- review | 2508.11326 | MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts | arXiv 2025
- review | 2508.12001 | FNH-TTS: A Fast, Natural, and Human-Like Speech Synthesis System with advanced prosodic modeling based on Mixture of Experts | arXiv 2025
- review | 2508.14049 | MahaTTS: A Unified Framework for Multilingual Text-to-Speech Synthesis | arXiv 2025
- review | 2508.15442 | Mitigating Hallucinations in LM-Based TTS Models via Distribution Alignment Using GFlowNets | EMNLP 2025
- review | 2508.15827 | Mini-Omni-Reasoner: Token-Level Thinking-in-Speaking in Large Speech Models | arXiv 2025
- review | 2508.16332 | Vevo2: A Unified and Controllable Framework for Speech and Singing Voice Generation | arXiv 2025
- review | 2508.19098 | CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis | arXiv 2025
- review | 2508.20660 | CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation | arXiv 2025
- review | 2509.00685 | MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech | arXiv 2025
- review | 2509.02020 | FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot | arXiv 2025
- review | 2509.09631 | DiFlow-TTS: Compact and Low-Latency Zero-Shot Text-to-Speech with Factorized Discrete Flow Matching | arXiv 2025
- review | 2509.15969 | VoXtream: Full-Stream Text-to-Speech with Extremely Low Latency | arXiv 2025
- review | 2509.19668 | Selective Classifier-free Guidance for Zero-shot Text-to-speech | arXiv 2025
- review | 2510.00981 | FlexiCodec: A Dynamic Neural Audio Codec for Low Frame Rates | arXiv 2025
- review | 2510.02848 | Flamed-TTS: Flow Matching Attention-Free Models for Efficient Generating and Dynamic Pacing Zero-shot Text-to-Speech | arXiv 2025
2026-06-24
- render | 1 concepts | mode: full | model: claude-sonnet-4-6
- review | 2025.findings-naacl.184 | Continuous Speech Tokenizer in Text To Speech | NAACL 2025
- review | 2025.naacl-long.110 | WaveFM: A High-Fidelity and Efficient Vocoder Based on Flow Matching | NAACL 2025
- review | 2025.naacl-long.242 | StyleTTS-ZS: Efficient High-Quality Zero-Shot Text-to-Speech Synthesis with Distilled Time-Varying Style Diffusion | NAACL 2025
- review | 2508.06890 | Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody | arXiv 2025
- review | 2508.07302 | XEmoRAG: Cross-Lingual Emotion Transfer with Controllable Intensity Using Retrieval-Augmented Generation | arXiv 2025
- review | 2508.07426 | Scalable Controllable Accented TTS | arXiv 2025
- review | 2508.07711 | Is GAN Necessary for Mel-Spectrogram-based Neural Vocoder? | arXiv 2025
- review | 2508.08095 | Dual Information Speech Language Models for Emotional Conversations | arXiv 2025
- review | 2508.08399 | Exploring Disentangled Neural Speech Codecs from Self-Supervised Representations | arXiv 2025
- review | 2508.08715 | MultiGen: Child-Friendly Multilingual Speech Generator with LLMs | arXiv 2025
- review | 2508.08961 | DualSpeechLM: Towards Unified Speech Understanding and Generation via Dual Speech Token Modeling with Large Language Models | arXiv 2025
- review | 2508.09767 | UtterTune: LoRA-Based Target-Language Pronunciation Edit and Control in Multilingual Text-to-Speech | arXiv 2025
2026-06-23
- review | 2025.acl-long.654 | Language-Codec: Bridging Discrete Codec Representations and Speech Language Models | ACL 2025
- review | 2507.20091 | ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models | arXiv 2025
- review | 2025.acl-long.682 | Recent Advances in Speech Language Models: A Survey | ACL 2025
- review | 2025.acl-long.911 | DNASpeech: A Contextualized and Situated Text-to-Speech Dataset with Dialogues, Narratives and Actions | ACL 2025
- review | 2507.22746 | Next Tokens Denoising for Speech Synthesis | arXiv 2025
- review | 2508.00317 | Advancing Speech Quality Assessment Through Scientific Challenges and Open-source Activities | arXiv 2025
- review | 2025.acl-long.912 | LLaMA-Omni 2: LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech Synthesis | ACL 2025
- review | 2025.acl-short.81 | Zero-Shot Text-to-Speech for Vietnamese | ACL 2025
- review | 2508.01796 | Enhancing Spectrogram Realism in Singing Voice Synthesis via Explicit Bandwidth Extension Prior to Vocoder | arXiv 2025
- review | 2508.02013 | SpeechRole: A Large-Scale Dataset and Benchmark for Evaluating Speech Role-Playing Agents | arXiv 2025
- review | 2025.americasnlp-1.1 | Text-to-speech system for low-resource languages: A case study in Shipibo-Konibo (a Panoan language from Peru) | workshop 2025
- review | 2025.ccl-1.80 | Lao-English Code-Switched Speech Synthesis Via Neural Codec Language Modeling | workshop 2025
- review | 2508.02038 | Marco-Voice Technical Report | arXiv 2025
- review | 2508.02849 | SecoustiCodec: Cross-Modal Aligned Streaming Single-Codecbook Speech Codec | arXiv 2025
- review | 2025.coling-main.352 | DiffStyleTTS: Diffusion-based Hierarchical Prosody Modeling for Text-to-Speech with Diverse and Controllable Styles | workshop 2025
- review | 2025.coling-main.518 | ProsodyFlow: High-fidelity Text-to-Speech through Conditional Flow Matching and Prosody Modeling with Large Speech Language Models | workshop 2025
- review | 2508.03543 | EmoSteer-TTS: Fine-Grained and Training-Free Emotion-Controllable Text-to-Speech via Activation Steering | arXiv 2025
- review | 2508.04141 | Parallel GPT: Harmonizing the Independence and Interdependence of Acoustic and Semantic Information for Zero-Shot Text-to-Speech | arXiv 2025
- review | 2025.emnlp-demos.70 | OpenS2S: Advancing Fully Open-Source End-to-End Empathetic Large Speech Language Model | EMNLP 2025
- review | 2025.emnlp-main.1730 | FillerSpeech: Towards Human-Like Text-to-Speech Synthesis with Filler Insertion and Filler Style Control | EMNLP 2025
- review | 2508.04585 | UniTalker: Conversational Speech-Visual Synthesis | ACM MM 2025
- review | 2508.04996 | REF-VC: Robust, Expressive and Fast Zero-Shot Voice Conversion with Diffusion Transformers | arXiv 2025
- review | 2025.emnlp-main.180 | Scaling Rich Style-Prompted Text-to-Speech Datasets | EMNLP 2025
- review | 2025.emnlp-main.989 | VocalNet: Speech LLMs with Multi-Token Prediction for Faster and High-Quality Generation | EMNLP 2025
- review | 2508.05207 | SpectroStream: A Versatile Neural Codec for General Audio | arXiv 2025
- review | 2508.05385 | A Scalable Pipeline for Enabling Non-Verbal Speech Generation and Understanding | arXiv 2025
- review | 2025.findings-acl.1051 | LLMVoX: Autoregressive Streaming Text-to-Speech Model for Any LLM | ACL 2025
- review | 2025.findings-emnlp.424 | InteractSpeech: A Speech Dialogue Interaction Corpus for Spoken Dialogue Model | EMNLP 2025
- review | 2508.06262 | Llasa+: Free Lunch for Accelerated and Streaming Llama-Based Speech Synthesis | arXiv 2025
- review | 2508.06870 | Text to Speech System for Meitei Mayek Script | arXiv 2025
2026-06-22
- review | interspeech-2025-0469 | Developing High-Quality TTS for Punjabi and Urdu: Benchmarking against MMS Models | Interspeech 2025
- review | interspeech-2025-0973 | A Dataset for Automatic Assessment of TTS Quality in Spanish | Interspeech 2025
- review | 2025.acl-demo.37 | RT-VC: Real-Time Zero-Shot Voice Conversion with Speech Articulatory Coding | ACL 2025
- review | 2412.17048 | Why Do Speech Language Models Fail to Generate Semantically Coherent Outputs? A Modality Evolving Perspective | arXiv 2026
- review | interspeech-2025-0989 | HiFiTTS-2: A Large-Scale High Bandwidth Speech Dataset | Interspeech 2025
- review | interspeech-2025-1034 | Non-Standard Accent TTS Support via Large Multi-Accent Frontend Pronunciation Knowledge Transfer | Interspeech 2025
- review | 2025.acl-industry.42 | Scaling Under-Resourced TTS: A Data-Optimized Framework with Advanced Acoustic Modeling for Thai | ACL 2025
- review | 2502.11128 | FELLE: Autoregressive Speech Synthesis with Token-Wise Coarse-to-Fine Flow Matching | arXiv 2025
- review | interspeech-2025-1066 | Score-Based Training for Energy-Based TTS Models | Interspeech 2025
- review | interspeech-2025-1101 | ZSDEVC: Zero-Shot Diffusion-based Emotional Voice Conversion with Disentangled Mechanism | Interspeech 2025
- review | 2025.acl-long.1252 | Finding A Voice: Exploring the Potential of African American Dialect and Voice Generation for Chatbots | ACL 2025
- review | 2503.04721 | Full-Duplex-Bench: A Benchmark to Evaluate Full-duplex Spoken Dialogue Models on Turn-taking Capabilities | arXiv 2025
- review | interspeech-2025-1122 | BitTTS: Highly Compact Text-to-Speech Using 1.58-bit Quantization and Weight Indexing | Interspeech 2025
- review | interspeech-2025-1344 | Parameter-Efficient Fine-Tuning for Low-Resource Text-to-Speech via Cross-Lingual Continual Learning | Interspeech 2025
- review | 2025.acl-long.1471 | The time scale of redundancy between prosody and linguistic context | ACL 2025
- review | 2504.10352 | Pseudo-Autoregressive Neural Codec Language Models for Efficient Zero-Shot Text-to-Speech Synthesis | arXiv 2025
- review | interspeech-2025-1440 | FreeCodec: A Disentangled Neural Speech Codec with Fewer Tokens | Interspeech 2025
- review | interspeech-2025-1595 | Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs | Interspeech 2025
- review | 2025.acl-long.1498 | Analyzing and Mitigating Inconsistency in Discrete Speech Tokens for Neural Codec Language Models | ACL 2025
- review | 2506.21619 | IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech | arXiv 2025
- review | interspeech-2025-1779 | ReFlow-VC: Zero-shot Voice Conversion Based on Rectified Flow and Speaker Feature Optimization | Interspeech 2025
- review | interspeech-2025-1993 | Defending Unauthorized Voice Cloning with Watermark-Aware Codecs | Interspeech 2025
- review | 2025.acl-long.313 | F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching | ACL 2025
- review | 2507.09318 | ZipVoice-Dialog: Non-Autoregressive Spoken Dialogue Generation with Flow Matching | arXiv 2026
- review | interspeech-2025-2043 | Training-Free Voice Conversion with Factorized Optimal Transport | Interspeech 2025
- review | interspeech-2025-2447 | Accelerating Autoregressive Speech Synthesis Inference With Speech Speculative Decoding | Interspeech 2025
- review | 2025.acl-long.388 | Distilling an End-to-End Voice Assistant Without Instruction Training Data | ACL 2025
- review | 2507.14534 | Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion | arXiv 2025
- review | interspeech-2025-2449 | Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling | Interspeech 2025
- review | 2025.acl-long.598 | Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment | ACL 2025
2026-06-21
- review | interspeech-2025-0047 | Revival with Voice: Multi-modal Controllable Text-to-Speech Synthesis | Interspeech 2025
- review | interspeech-2025-0143 | Multimodal Prosody Modeling: A Use Case for Multilingual Sentence Mode Prediction | Interspeech 2025
- review | interspeech-2025-0196 | SPCODEC: Split and Prediction for Neural Speech Codec | Interspeech 2025
- review | interspeech-2025-0203 | ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech | Interspeech 2025
- review | interspeech-2025-0246 | DC-Spin: A Speaker-invariant Speech Tokenizer for Spoken Language Models | Interspeech 2025
- review | interspeech-2025-0253 | Long-Context Speech Synthesis with Context-Aware Memory | Interspeech 2025
- review | interspeech-2025-0310 | Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models | Interspeech 2025
- review | interspeech-2025-0401 | Enabling the replicability of speech synthesis perceptual evaluations | Interspeech 2025
- review | interspeech-2025-0406 | Zero-Shot Mono-to-Binaural Speech Synthesis | Interspeech 2025
- review | interspeech-2025-0408 | Improving User Impression of Spoken Dialogue Systems by Controlling Para-linguistic Expression Based on Intimacy | Interspeech 2025
- review | interspeech-2025-0551 | Monotonic Attention for Robust Text-to-Speech Synthesis in Large Language Model Frameworks | Interspeech 2025
- review | interspeech-2025-0554 | RapFlow-TTS: Rapid and High-Fidelity Text-to-Speech with Improved Consistency Flow Matching | Interspeech 2025
- review | interspeech-2025-0575 | VoiceMark: Zero-Shot Voice Cloning-Resistant Watermarking Approach Leveraging Speaker-Specific Latents | Interspeech 2025
- review | interspeech-2025-0596 | Facilitating Personalized TTS for Dysarthric Speakers Using Knowledge Anchoring and Curriculum Learning | Interspeech 2025
- review | interspeech-2025-0648 | MIKU-PAL: An Automated and Standardized Multimodal Method for Speech Paralinguistic and Affect Labeling | Interspeech 2025
- review | interspeech-2025-0704 | Differentiable Reward Optimization for LLM based TTS system | Interspeech 2025
- review | interspeech-2025-0723 | Counterfactual Activation Editing for Post-hoc Prosody and Mispronunciation Correction in TTS Models | Interspeech 2025
- review | interspeech-2025-0754 | EME-TTS: Unlocking the Emphasis and Emotion Link in Speech Synthesis | Interspeech 2025
- review | interspeech-2025-0762 | Intrasentential English in Swedish TTS: perceived English-accentedness | Interspeech 2025
- review | interspeech-2025-0779 | Intelligibility of Text-to-Speech Systems for Mathematical Expressions | Interspeech 2025
- review | interspeech-2025-0787 | Gradual modeling of the Lombard effect by modifying speaker embeddings from a Text-To-Speech model | Interspeech 2025
- review | interspeech-2025-0854 | Bridging the TrainingâInference Gap in TTS: Training Strategies for Robust Generative Postprocessing for Low-Resource Speakers | Interspeech 2025
- review | interspeech-2025-0902 | VoiceQualityVC: A Voice Conversion System for Studying the Perceptual Effects of Voice Quality in Speech | Interspeech 2025
- review | interspeech-2025-0948 | PromptEVC: Controllable Emotional Voice Conversion with Natural Language Prompts | Interspeech 2025
2026-06-20
- review | interspeech-2025-0063 | Fine-Tuning Text-to-Speech Diffusion Models Using Reinforcement Learning with Human Feedback | Interspeech 2025
- review | interspeech-2025-0115 | Bringing Interpretability to Neural Audio Codecs | Interspeech 2025
- review | interspeech-2025-0319 | Improving Noise Robustness of LLM-based Zero-shot TTS via Discrete Acoustic Token Denoising | Interspeech 2025
- review | interspeech-2025-0455 | APTTS: Adversarial Post-training in Latent Flow Matching for Fast and High-fidelity Text-to-Speech | Interspeech 2025
- review | interspeech-2025-0468 | DualCodec: A Low-Frame-Rate, Semantically-Enhanced Neural Audio Codec for Speech Generation | Interspeech 2025
- review | interspeech-2025-0669 | PAST: Phonetic-Acoustic Speech Tokenizer | Interspeech 2025
- review | interspeech-2025-0874 | Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model | Interspeech 2025
- review | interspeech-2025-1641 | Robust Neural Codec Language Modeling with Phoneme Position Prediction for Zero-Shot TTS | Interspeech 2025
- review | interspeech-2025-2660 | Triadic Multi-party Voice Activity Projection for Turn-taking in Spoken Dialogue Systems | Interspeech 2025
- review | interspeech-2025-2765 | The State Of TTS: A Case Study with Human Fooling Rates | Interspeech 2025
2026-06-19
- misc | quoted id: fields in frontmatter (327 pages; unquoted arXiv IDs were parsed as floats by YAML)
- misc | added [!info] Citation Stub callout to all 65 Tier 2 stub pages
- misc | papers/index.md: all ID cells →
[[wikilinks]], all titles → markdown links, blank row removed - review | 2025.acl-long.65 | Autoregressive Speech Synthesis without Vector Quantization | ACL 2025
- review | 2025.acl-long.346 | ControlSpeech: Towards Simultaneous and Independent Zero-shot Speaker Cloning and Zero-shot Language Style Control | ACL 2025
- review | interspeech-2025-0739 | FD-Bench: A Full-Duplex Benchmarking Pipeline Designed for Full Duplex Spoken Dialogue Systems | Interspeech 2025
- review | interspeech-2025-0815 | Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion | Interspeech 2025
- review | 2025.emnlp-main.40 | Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey | EMNLP 2025
- review | 2301.02111 | Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers | arXiv 2023
- review | 2308.16692 | SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models | arXiv 2023
- review | 2403.03100 | NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models | arXiv 2024
- review | 2406.02430 | Seed-TTS: A Family of High-Quality Versatile Speech Generation Models | arXiv 2024
- review | 2407.05407 | CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens | arXiv 2024
- review | 2412.10117 | CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models | arXiv 2024
- review | 2502.03930 | DiTAR: Diffusion Transformer Autoregressive Modeling for Speech Generation | arXiv 2025
- review | 2504.12867 | EmoVoice: LLM-based Emotional Text-To-Speech Model with Freestyle Text Prompting | arXiv 2025
- review | 2025.acl-long.1043 | OZSpeech: One-step Zero-shot Speech Synthesis with Learned-Prior-Conditioned Flow Matching | ACL 2025
- review | interspeech-2025-0816 | Bridging Speech and Singing: Multi-stage Speech-Prompted Singing Voice Conversion with Speaker Embedding Adaptation | Interspeech 2025
2026-06-17
- ingest | 1510.08484 | MUSAN: A Music, Speech, and Noise Corpus | arXiv 2015
- ingest | 1607.06450 | Layer Normalization | arXiv 2016
- ingest | 1908.06248 | JVS corpus: free Japanese multi-speaker voice corpus | arXiv 2019
- ingest | 1808.10583 | AISHELL-2: Transforming Mandarin ASR Research Into Industrial Scale | arXiv 2018
- ingest | 2002.05202 | GLU Variants Improve Transformer | arXiv 2020
- ingest | 2001.08361 | Scaling Laws for Neural Language Models | arXiv 2020
- ingest | 2007.10310 | CoVoST 2 and Massively Multilingual Speech-to-Text Translation | arXiv 2020
- ingest | 2302.00482 | Improving and generalizing flow-based generative models with minibatch optimal transport | arXiv 2023
- ingest | 2308.10248 | Steering Language Models With Activation Engineering | arXiv 2023
- ingest | 2308.05725 | EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis | arXiv 2023
- ingest | 2309.16609 | Qwen Technical Report | arXiv 2023
- ingest | 2308.11596 | SeamlessM4T: Massively Multilingual & Multimodal Machine Translation | arXiv 2023
- ingest | 2408.05211 | VITA: Towards Open-Source Interactive Omni Multimodal LLM | arXiv 2024
- ingest | 2408.01800 | MiniCPM-V: A GPT-4V Level MLLM on Your Phone | arXiv 2024
- ingest | 2312.10997 | Retrieval-Augmented Generation for Large Language Models: A Survey | arXiv 2023
- ingest | 2501.07246 | Audio-CoT: Exploring Chain-of-Thought Reasoning in Large Audio Language Model | arXiv 2025
- ingest | 2412.08635 | Multimodal Latent Language Modeling with Next-Token Diffusion | arXiv 2024
- ingest | 2410.19168 | MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark | arXiv 2024
- ingest | 2501.01957 | VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction | arXiv 2025
- ingest | 2503.01743 | Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs | arXiv 2025
- ingest | 2503.19786 | Gemma 3 Technical Report | arXiv 2025
- ingest | 2505.03739 | VITA-Audio: Fast Interleaved Cross-Modal Token Generation for Efficient Large Speech-Language Model | arXiv 2025
- ingest | 2501.15368 | Baichuan-Omni-1.5 Technical Report | arXiv 2025
- ingest | 2506.02863 | CapSpeech: Enabling Downstream Applications in Style-Captioned Text-to-Speech | arXiv 2025
- ingest | 2507.12705 | AudioJudge: Understanding What Works in Large Audio Model Based Speech Evaluation | arXiv 2025
- ingest | 2506.07900 | MiniCPM4: Ultra-Efficient LLMs on End Devices | arXiv 2025
- ingest | 2507.08128 | Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models | arXiv 2025
- ingest | 2510.14664 | SpeechLLM-as-Judges: Towards General and Interpretable Speech Quality Evaluation | arXiv 2025
- ingest | 2511.09690 | Omnilingual ASR: Open-Source Multilingual Speech Recognition for 1600+ Languages | arXiv 2025
- ingest | 2508.13992 | MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence | arXiv 2025
- ingest | 2509.08753 | Streaming Sequence-to-Sequence Learning with Delayed Streams Modeling | arXiv 2025
2026-06-16
- ingest | 2411.09943 | Zero-shot Voice Conversion with Diffusion Transformers | arXiv 2024
- ingest | 2411.17607 | Scaling Speech-Text Pre-training with Synthetic Interleaved Data | arXiv 2024
- ingest | 2411.18803 | TS3-Codec: Transformer-Based Simple Streaming Single Codec | arXiv 2024
- ingest | 2412.04724 | StableVC: Style Controllable Zero-Shot Voice Conversion with Conditional Flow Matching | arXiv 2024
- ingest | 2506.13053 | ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching | arXiv 2025
- ingest | 2505.02625 | LLaMA-Omni2: LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech Synthesis | arXiv 2025
- ingest | 2502.18924 | MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis | arXiv 2025
- ingest | 2507.23159 | Full-Duplex-Bench v1.5: Evaluating Overlap Handling for Full-Duplex Speech Models | arXiv 2025
- ingest | 2506.16381 | InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems | arXiv 2025
- ingest | 2506.10274 | Discrete Audio Tokens: More Than a Survey! | arXiv 2025
- ingest | 2505.13000 | DualCodec: A Low-Frame-Rate, Semantically-Enhanced Neural Audio Codec for Speech Generation | arXiv 2025
- ingest | 2503.14345 | MoonCast: High-Quality Zero-Shot Podcast Generation | arXiv 2025
- ingest | 2508.04195 | NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations | arXiv 2025
- ingest | 2505.09558 | WavReward: Spoken Dialogue Models With Generalist Reward Evaluators | arXiv 2025
- ingest | 2504.10344 | ALMTokenizer: A Low-bitrate and Semantic-rich Audio Codec Tokenizer for Audio Language Modeling | arXiv 2025
- ingest | 2504.02407 | F5R-TTS: Improving Flow-Matching based Text-to-Speech with Group Relative Policy Optimization | arXiv 2025
- ingest | 2511.15848 | Step-Audio-R1 Technical Report | arXiv 2025
- ingest | 2510.07838 | Full-Duplex-Bench-v2: A Multi-Turn Evaluation Framework for Duplex Dialogue Systems with an Automated Examiner | arXiv 2025
- ingest | 2505.14648 | Vox-Profile: A Speech Foundation Model Benchmark for Characterizing Diverse Speaker and Speech Traits | arXiv 2025
2026-06-15
- integrate | 25 papers | 20 concepts updated | 23 digests updated | 8 cross-links added
- ingest | 2301.12503 | AudioLDM: Text-to-Audio Generation with Latent Diffusion Models | ICML 2023
- ingest | 2301.11325 | MusicLM: Generating Music From Text | arXiv 2023
- ingest | 2305.15255 | Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM | arXiv 2023
- ingest | 2312.01479 | OpenVoice: Versatile Instant Voice Cloning | arXiv 2023
- ingest | 2312.15821 | Audiobox: Unified Audio Generation with Natural Language Prompts | arXiv 2023
- ingest | 2401.07333 | ELLA-V: Stable Neural Codec Language Modeling with Alignment-guided Sequence Reordering | arXiv 2024
- ingest | 2402.13236 | Towards audio language modeling — an overview | arXiv 2024
- ingest | 2404.03204 | RALL-E: Robust Codec Language Modeling with Chain-of-Thought Prompting for Text-to-Speech Synthesis | arXiv 2024
- ingest | 2406.00654 | Enhancing Zero-shot Text-to-Speech Synthesis with Human Feedback | arXiv 2024
- ingest | 2406.05551 | Autoregressive Diffusion Transformer for Text-to-Speech Synthesis | arXiv 2024
- ingest | 2408.02622 | Language Model Can Listen While Speaking | arXiv 2024
2026-06-14
- ingest | 2411.01156 | Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis | arXiv 2024
- ingest | 2505.07916 | MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder | arXiv 2025
- ingest | 2410.11190 | Mini-Omni2: Towards Open-source GPT-4o with Vision, Speech and Duplex Capabilities | arXiv 2024
- ingest | 2410.03751 | Recent Advances in Speech Language Models: A Survey | arXiv 2024
- ingest | 2104.00355 | Speech Resynthesis from Discrete Disentangled Self-Supervised Representations | arXiv 2021
- ingest | 2502.06490 | Recent Advances in Discrete Speech Tokens: A Review | arXiv 2025
- ingest | 2105.06337 | Grad-TTS: A Diffusion Probabilistic Model for Text-to-Speech | arXiv 2021
- ingest | 2412.15649 | SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training | arXiv 2024
- ingest | 2502.05512 | IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System | arXiv 2025
- ingest | 2502.07243 | Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement | ICLR 2025
- ingest | 2406.07855 | VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment | arXiv 2024
- ingest | 2410.17799 | OmniFlatten: An End-to-end GPT Model for Seamless Voice Conversation | arXiv 2024
- ingest | 2504.08528 | On The Landscape of Spoken Language Models: A Comprehensive Survey | arXiv 2025
- ingest | 2206.08317 | Paraformer: Fast and Accurate Parallel Transformer for Non-autoregressive End-to-End Speech Recognition | arXiv 2022
- ingest | 2412.19437 | DeepSeek-V3 Technical Report | arXiv 2024
- ingest | 2402.03300 | DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models | arXiv 2024
- ingest | 2310.13289 | SALMONN: Towards Generic Hearing Abilities for Large Language Models | arXiv 2023
- ingest | 1810.04805 | BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding | arXiv 2018
- ingest | 2502.05139 | Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound | arXiv 2025
- ingest | 2307.09288 | Llama 2: Open Foundation and Fine-Tuned Chat Models | arXiv 2023
- ingest | 2312.11805 | Gemini: A Family of Highly Capable Multimodal Models | arXiv 2023
- ingest | 2005.14165 | Language Models are Few-Shot Learners | arXiv 2020
- ingest | 2407.10671 | Qwen2 Technical Report | arXiv 2024
- ingest | 2106.04624 | SpeechBrain: A General-Purpose Speech Toolkit | arXiv 2021
- ingest | 2406.14294 | DASB - Discrete Audio and Speech Benchmark | arXiv 2024
2026-06-13
-
misc | Factor A/B/C label cleanup | 12 files | paper-internal experiment labels from 2412.17048 replaced with descriptive language; citations added where missing
-
integrate | 4 papers (orphan fix) | autoregressive-codec-tts, prosody-control, spoken-language-model, evaluation-metrics, voice-conversion, multilingual-tts updated | interspeech-2025-0253, interspeech-2025-0408, interspeech-2025-0902, 2025.americasnlp-1.1 linked to concepts
-
integrate | 25 papers | 21 concepts updated | 21 digests updated | 133 cross-links added
-
ingest | 2305.09636 | SoundStorm: Efficient Parallel Audio Generation | arXiv 2023
-
ingest | 1712.05884 | Natural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram Predictions | arXiv 2017
-
ingest | 2402.01912 | Natural language guidance of high-fidelity text-to-speech with synthetic annotations | arXiv 2024
-
ingest | 2306.12925 | AudioPaLM: A Large Language Model That Can Speak and Listen | arXiv 2023
-
ingest | 2305.07243 | Better speech synthesis through scaling | arXiv 2023
-
ingest | 1609.03499 | WaveNet: A Generative Model for Raw Audio | arXiv 2016
-
ingest | 2411.19842 | Scaling Transformers for Low-Bitrate High-Quality Speech Coding | arXiv 2024
-
ingest | 2407.08551 | Autoregressive Speech Synthesis without Vector Quantization | arXiv 2024
-
ingest | 1703.10135 | Tacotron: Towards End-to-End Speech Synthesis | arXiv 2017
-
ingest | 2502.17239 | Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction | arXiv 2025
-
ingest | 2402.05755 | Spirit LM: Interleaved Spoken and Written Language Model | arXiv 2024
-
ingest | 2402.08093 | BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data | arXiv 2024
-
ingest | 2507.16632 | Step-Audio 2 Technical Report | arXiv 2025
-
ingest | 2310.00704 | UniAudio: An Audio Foundation Model Toward Universal Audio Generation | arXiv 2023
-
ingest | 2106.15561 | A Survey on Neural Speech Synthesis | arXiv 2021
2026-06-12
- ingest | 2311.07919 | Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models | arXiv 2023
- ingest | 2505.09388 | Qwen3 Technical Report | arXiv 2025
- ingest | 2005.07143 | ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification | arXiv 2020
- ingest | 2302.13971 | LLaMA: Open and Efficient Foundation Language Models | arXiv 2023
- ingest | 2507.06261 | Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities | arXiv 2025
- ingest | 2012.03411 | MLS: A Large-Scale Multilingual Dataset for Speech Research | arXiv 2020
- ingest | 2501.12948 | DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning | arXiv 2025
- ingest | 2312.05187 | Seamless: Multilingual Expressive and Streaming Speech Translation | arXiv 2023
- ingest | 2106.06909 | GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio | arXiv 2021
- ingest | 2309.15505 | Finite Scalar Quantization: VQ-VAE Made Simple | arXiv 2023
- integrate | 25 papers | 19 concepts updated | 4 digests updated | 16 cross-links added
- ingest | 2306.00814 | Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis | arXiv 2023
- ingest | 2407.05361 | Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation | arXiv 2024
- ingest | 2406.18009 | E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS | arXiv 2024
- ingest | 2406.04904 | XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model | arXiv 2024
- ingest | 2409.05377 | BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec | arXiv 2024
- ingest | 2305.02765 | HiFi-Codec: Group-residual Vector quantization for High Fidelity Audio Codec | arXiv 2023
- ingest | 2403.16973 | VoiceCraft: Zero-Shot Speech Editing and Text-to-Speech in the Wild | arXiv 2024
- ingest | 2502.11946 | Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction | arXiv 2025
- ingest | 2501.06282 | MinMo: A Multimodal Large Language Model for Seamless Voice Interaction | arXiv 2025
- ingest | 2303.03926 | Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling | arXiv 2023
2026-06-11
- ingest | 2308.16692 | SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models | arXiv 2023
- ingest | 2503.01710 | Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens | arXiv 2025
- ingest | 2409.00750 | MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer | arXiv 2024
- ingest | 2505.17589 | CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training | arXiv 2025
- ingest | 2408.16725 | Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming | arXiv 2024
- ingest | 2502.04128 | Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis | arXiv 2025
- ingest | 2304.09116 | NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers | arXiv 2023
- ingest | 2406.05370 | VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers | arXiv 2024
- ingest | 2409.06666 | LLaMA-Omni: Seamless Speech Interaction with Large Language Models | arXiv 2024
- ingest | 2411.00774 | Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM | arXiv 2024
- ingest | 2408.16532 | WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling | arXiv 2024
- ingest | 2407.04051 | FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs | arXiv 2024
- ingest | 2305.11000 | SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities | arXiv 2023
- ingest | 2410.17196 | VoiceBench: Benchmarking LLM-Based Voice Assistants | arXiv 2024
- ingest | 2409.03283 | FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications | arXiv 2024
2026-06-10
- integrate | 25 papers | 12 concepts updated | 7 digests updated | 11 cross-links added
- ingest | 2204.02152 | UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022 | arXiv 2022
- ingest | 2504.18425 | Kimi-Audio Technical Report | arXiv 2025
- ingest | 1904.02882 | LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech | arXiv 2019
- ingest | 2407.10759 | Qwen2-Audio Technical Report | arXiv 2024
- ingest | 2403.03100 | NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models | arXiv 2024
- ingest | 2303.08774 | GPT-4 Technical Report | arXiv 2023
- ingest | 2412.02612 | GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot | arXiv 2024
- ingest | 1711.05101 | Decoupled Weight Decay Regularization | arXiv 2017
- ingest | 2410.21276 | GPT-4o System Card | arXiv 2024
- ingest | 2412.15115 | Qwen2.5 Technical Report | arXiv 2024
- ingest | 2207.12598 | Classifier-Free Diffusion Guidance | arXiv 2022
- ingest | 2210.02747 | Flow Matching for Generative Modeling | arXiv 2022
- ingest | 2209.03143 | AudioLM: a Language Modeling Approach to Audio Generation | arXiv 2022
- ingest | 2206.04658 | BigVGAN: A Universal Neural Vocoder with Large-Scale Training | arXiv 2022
- ingest | 1412.6980 | Adam: A Method for Stochastic Optimization | arXiv 2014
2026-06-09
- ingest | 2410.00037 | Moshi: a speech-text foundation model for real-time dialogue | arXiv 2024
- ingest | 2210.13438 | High Fidelity Neural Audio Compression | arXiv 2022
- ingest | 2006.04558 | FastSpeech 2: Fast and High-Quality End-to-End Text to Speech | arXiv 2020
- ingest | 2411.13577 | WavChat: A Survey of Spoken Dialogue Models | arXiv 2024
- ingest | 2503.20215 | Qwen2.5-Omni Technical Report | arXiv 2025
- ingest | 2407.21783 | The Llama 3 Herd of Models | arXiv 2024
- ingest | 1912.06670 | Common Voice: A Massively-Multilingual Speech Corpus | arXiv 2019
- ingest | 2312.15185 | emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation | arXiv 2023
- ingest | 2212.04356 | Robust Speech Recognition via Large-Scale Weak Supervision | arXiv 2022
- ingest | 2010.05646 | HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis | arXiv 2020
2026-06-08
- misc | landing page redesign | index.md editorial rewrite + field snapshot; overview.md renamed to “Field Overview”; start.md navigation hub added
- integrate | seeded 2 evidence digests | transformer-enc-dec-tts (11 papers, 6 claim clusters) | rlhf-speech (15 papers, 7 claim clusters)
2026-06-05
- integrate | 25 papers | 22 concepts updated | 14 digests updated | 10 cross-links added
2026-06-04
- integrate | seed concept stubs | singing (3 papers) | fine-tuning (2 papers)
2026-06-03
- integrate | 26 papers | pass 6 | 18 concepts updated | 19 digests updated | 6 cross-links added
- ingest | interspeech-2025-1993 | Defending Unauthorized Voice Cloning with Watermark-Aware Codecs | Interspeech 2025
- ingest | interspeech-2025-2765 | The State Of TTS: A Case Study with Human Fooling Rates | Interspeech 2025
- ingest | interspeech-2025-0948 | PromptEVC: Controllable Emotional Voice Conversion with Natural Language Prompts | Interspeech 2025
- ingest | 2508.08095 | Dual Information Speech Language Models for Emotional Conversations | arXiv 2025
2026-06-02
- ingest | interspeech-2025-0203 | ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control | Interspeech 2025
- ingest | interspeech-2025-0196 | SPCODEC: Split and Prediction for Neural Speech Codec | Interspeech 2025
- ingest | 2503.04721 | Full-Duplex-Bench: A Benchmark to Evaluate Full-duplex Spoken Dialogue Models | arXiv 2025
- ingest | 2508.08715 | MultiGen: Child-Friendly Multilingual Speech Generator with LLMs | arXiv 2025
- ingest | 2508.09767 | UtterTune: LoRA-Based Target-Language Pronunciation Edit and Control in Multilingual TTS | arXiv 2025
- ingest | 2508.11326 | MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts | arXiv 2025
- ingest | 2504.12867 | EmoVoice: LLM-based Emotional Text-To-Speech Model with Freestyle Text Prompting | arXiv 2025
- ingest | 2508.08961 | DualSpeechLM: Towards Unified Speech Understanding and Generation via Dual Speech Token Modeling | arXiv 2025
- ingest | 2508.08399 | Exploring Disentangled Neural Speech Codecs from Self-Supervised Representations | arXiv 2025
- ingest | 2508.07711 | Is GAN Necessary for Mel-Spectrogram-based Neural Vocoder? | arXiv 2025
- ingest | 2508.07426 | Scalable Controllable Accented TTS | ASRU 2025
- ingest | 2508.07302 | XEmoRAG: Cross-Lingual Emotion Transfer with Controllable Intensity Using Retrieval-Augmented Generation | arXiv 2025
- ingest | 2508.06890 | Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody | arXiv 2025
- ingest | 2508.06870 | Text to Speech System for Meitei Mayek Script | arXiv 2025
- ingest | 2508.05385 | A Scalable Pipeline for Enabling Non-Verbal Speech Generation and Understanding | arXiv 2025
- ingest | 2508.14049 | MahaTTS: A Unified Framework for Multilingual Text-to-Speech Synthesis | arXiv 2025
- ingest | 2508.04585 | UniTalker: Conversational Speech-Visual Synthesis | arXiv 2025
- ingest | 2508.04996 | REF-VC: Robust, Expressive and Fast Zero-Shot Voice Conversion with Diffusion Transformers | arXiv 2025
- ingest | 2508.05207 | SpectroStream: A Versatile Neural Codec for General Audio | arXiv 2025
- ingest | 2507.20091 | ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models | arXiv 2025
- ingest | 2508.00317 | Advancing Speech Quality Assessment Through Scientific Challenges and Open-source Activities | arXiv 2025
- ingest | 2507.22746 | Next Tokens Denoising for Speech Synthesis | arXiv 2025
- ingest | 2508.01796 | Enhancing Spectrogram Realism in Singing Voice Synthesis via Explicit Bandwidth Extension Prior to Vocoder | arXiv 2025
- ingest | 2508.02013 | SpeechRole: A Large-Scale Dataset and Benchmark for Evaluating Speech Role-Playing Agents | arXiv 2025
- ingest | 2508.02849 | SecoustiCodec: Cross-Modal Aligned Streaming Single-Codebook Speech Codec | arXiv 2025
- integrate | 25 papers | 20 concepts updated | 18 digests created | 8 cross-links added
2026-06-01
- integrate | concept page migration: all 21 concept pages rewritten to new research-briefing template (Executive Summary, Current Status, Major Claims, Representative Papers, Relationship to Other Concepts, status vocab)
- integrate | wiki redesign: landing page, folder navigation, venue naming, unified paper card format
- ingest | 2604.12438 | An Ultra-Low Latency, End-to-End Streaming Speech Synthesis Architecture via Block-Wise Generation and Depth-Wise Codec Decoding | arXiv 2026
- ingest | 2025.acl-long.682 | Recent Advances in Speech Language Models: A Survey | ACL 2025
- ingest | 2301.02111 | Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers | arXiv 2023
- ingest | 2025.acl-long.313 | F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching | ACL 2025 (re-ingest)
2026-05-30
- ingest | interspeech-2025-0047 | Revival with Voice: Multi-modal Controllable Text-to-Speech Synthesis | Interspeech 2025
- ingest | interspeech-2025-0063 | Fine-Tuning Text-to-Speech Diffusion Models Using Reinforcement Learning with Human Feedback | Interspeech 2025
- ingest | interspeech-2025-0143 | Multimodal Prosody Modeling: A Use Case for Multilingual Sentence Mode Prediction | Interspeech 2025
- ingest | interspeech-2025-0310 | Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models | Interspeech 2025
- ingest | interspeech-2025-0319 | Improving Noise Robustness of LLM-based Zero-shot TTS via Discrete Acoustic Token Denoising | Interspeech 2025
- ingest | interspeech-2025-0406 | Zero-Shot Mono-to-Binaural Speech Synthesis | Interspeech 2025
- ingest | interspeech-2025-0408 | Improving User Impression of Spoken Dialogue Systems by Controlling Para-linguistic Expression Based on Intimacy | Interspeech 2025
- ingest | interspeech-2025-0455 | APTTS: Adversarial Post-training in Latent Flow Matching for Fast and High-fidelity Text-to-Speech | Interspeech 2025
- ingest | interspeech-2025-0554 | RapFlow-TTS: Rapid and High-Fidelity Text-to-Speech with Improved Consistency Flow Matching | Interspeech 2025
- ingest | interspeech-2025-0551 | Monotonic Attention for Robust Text-to-Speech Synthesis in Large Language Model Frameworks | Interspeech 2025
- ingest | interspeech-2025-0575 | VoiceMark: Zero-Shot Voice Cloning-Resistant Watermarking Approach Leveraging Speaker-Specific Latents | Interspeech 2025
- ingest | interspeech-2025-0596 | Facilitating Personalized TTS for Dysarthric Speakers Using Knowledge Anchoring and Curriculum Learning | Interspeech 2025
- ingest | interspeech-2025-0648 | MIKU-PAL: An Automated and Standardized Multimodal Method for Speech Paralinguistic and Affect Labeling | Interspeech 2025
- ingest | interspeech-2025-0669 | PAST: Phonetic-Acoustic Speech Tokenizer | Interspeech 2025
- ingest | interspeech-2025-0704 | Differentiable Reward Optimization for LLM based TTS system | Interspeech 2025
- ingest | interspeech-2025-0723 | Counterfactual Activation Editing for Post-hoc Prosody and Mispronunciation Correction in TTS Models | Interspeech 2025
- ingest | interspeech-2025-0754 | EME-TTS: Unlocking the Emphasis and Emotion Link in Speech Synthesis | Interspeech 2025
- ingest | interspeech-2025-0762 | Intrasentential English in Swedish TTS: perceived English-accentedness | Interspeech 2025
- ingest | interspeech-2025-0779 | Intelligibility of Text-to-Speech Systems for Mathematical Expressions | Interspeech 2025
- ingest | interspeech-2025-0787 | Gradual modeling of the Lombard effect by modifying speaker embeddings from a Text-To-Speech model | Interspeech 2025
- ingest | interspeech-2025-0469 | Developing High-Quality TTS for Punjabi and Urdu: Benchmarking against MMS Models | Interspeech 2025
- ingest | interspeech-2025-0854 | Bridging the Training–Inference Gap in TTS: Training Strategies for Robust Generative Postprocessing for Low-Resource Speakers | Interspeech 2025
- ingest | interspeech-2025-0973 | A Dataset for Automatic Assessment of TTS Quality in Spanish | Interspeech 2025
- ingest | interspeech-2025-0989 | HiFiTTS-2: A Large-Scale High Bandwidth Speech Dataset | Interspeech 2025
- ingest | interspeech-2025-1034 | Non-Standard Accent TTS Support via Large Multi-Accent Frontend Pronunciation Knowledge Transfer | Interspeech 2025
- ingest | 2025.naacl-long.110 | WaveFM: A High-Fidelity and Efficient Vocoder Based on Flow Matching | NAACL 2025
- ingest | 2025.findings-acl.1051 | LLMVoX: Autoregressive Streaming Text-to-Speech Model for Any LLM | ACL 2025
- ingest | 2025.emnlp-main.180 | Scaling Rich Style-Prompted Text-to-Speech Datasets | EMNLP 2025
- ingest | 2507.09318 | ZipVoice-Dialog: Non-Autoregressive Spoken Dialogue Generation with Flow Matching | arXiv 2026
- ingest | 2025.coling-main.518 | ProsodyFlow: High-fidelity Text-to-Speech through Conditional Flow Matching and Prosody Modeling | workshop 2025
- integrate | 30 papers | 18 concepts updated | 12 cross-links added
2026-05-29
- ingest | 2512.04720 | M3-TTS: Multi-modal DiT Alignment & Mel-latent for Zero-shot High-fidelity Speech Synthesis | arXiv 2025
- ingest | 2511.12347 | VoiceCraft-X: Unifying Multilingual, Voice-Cloning Speech Synthesis and Speech Editing | EMNLP 2025
- ingest | 2509.00685 | MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech | arXiv 2025
- ingest | 2512.13251 | DisCo-Speech: Controllable Zero-Shot Speech Generation with A Disentangled Speech Codec | arXiv 2025
- ingest | 2509.09631 | DiFlow-TTS: Compact and Low-Latency Zero-Shot TTS with Factorized Discrete Flow Matching | arXiv 2025
- ingest | 2603.29339 | LongCat-AudioDiT: High-Fidelity Diffusion Text-to-Speech in the Waveform Latent Space | arXiv 2026
- ingest | 2508.11273 | EmoSSLSphere: Multilingual Emotional Speech Synthesis with Spherical Vectors and Discrete Speech Tokens | arXiv 2025
- ingest | 2604.12438 | An Ultra-Low Latency, End-to-End Streaming Speech Synthesis Architecture via Block-Wise Generation and Depth-Wise Codec Decoding | arXiv 2026
- ingest | 2604.01760 | T5Gemma-TTS Technical Report | arXiv 2026
- ingest | 2508.15442 | Mitigating Hallucinations in LM-Based TTS Models via Distribution Alignment Using GFlowNets | EMNLP 2025
- ingest | 2025.acl-long.654 | Language-Codec: Bridging Discrete Codec Representations and Speech Language Models | ACL 2025
- ingest | 2603.18090 | MOSS-TTS Technical Report | arXiv 2026
- ingest | 2508.04141 | Parallel GPT: Harmonizing the Independence and Interdependence of Acoustic and Semantic Information for Zero-Shot TTS | arXiv 2025
- ingest | 2502.11128 | FELLE: Autoregressive Speech Synthesis with Token-Wise Coarse-to-Fine Flow Matching | arXiv 2025
- ingest | 2603.26364 | LLaDA-TTS: Unifying Speech Synthesis and Zero-Shot Editing via Masked Diffusion Modeling | arXiv 2026
- ingest | 2508.19098 | CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis | arXiv 2025
- ingest | 2508.12001 | FNH-TTS: A Fast, Natural, and Human-Like Speech Synthesis System with advanced prosodic modeling | arXiv 2025
- ingest | 2510.05758 | EMORL-TTS: Reinforcement Learning for Fine-Grained Emotion Control in LLM-based TTS | ICASSP 2026
- ingest | 2601.03888 | IndexTTS 2.5 Technical Report | arXiv 2026
- ingest | 2509.15969 | VoXtream: Full-Stream Text-to-Speech with Extremely Low Latency | arXiv 2025
- ingest | 2510.07979 | IntMeanFlow: Few-step Speech Generation with Integral Velocity Distillation | arXiv 2025
- ingest | 2025.ccl-1.80 | Lao-English Code-Switched Speech Synthesis Via Neural Codec Language Modeling | workshop 2025
- ingest | 2025.coling-main.352 | DiffStyleTTS: Diffusion-based Hierarchical Prosody Modeling for Text-to-Speech with Diverse and Controllable Styles | workshop 2025
- ingest | 2025.acl-long.911 | DNASpeech: A Contextualized and Situated Text-to-Speech Dataset with Dialogues, Narratives and Actions | ACL 2025
- ingest | 2025.acl-short.81 | Zero-Shot Text-to-Speech for Vietnamese | ACL 2025
- ingest | 2025.acl-long.912 | LLaMA-Omni 2: LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech Synthesis | ACL 2025
- integrate | 45 papers | 17 concepts updated | 23 cross-links added
2026-05-28
- ingest | 2406.02430 | Seed-TTS: A Family of High-Quality Versatile Speech Generation Models | arXiv 2024
- ingest | 2407.05407 | CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens | arXiv 2024
- ingest | 2025.acl-long.65 | Autoregressive Speech Synthesis without Vector Quantization | ACL 2025
- ingest | 2412.10117 | CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models | arXiv 2024
- ingest | 2410.06885 | F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching | ACL 2025
- ingest | 2601.15621 | Qwen3-TTS Technical Report | arXiv 2026
- ingest | 2512.14291 | GLM-TTS Technical Report | arXiv 2025
- ingest | 2508.06262 | Llasa+: Free Lunch for Accelerated and Streaming Llama-Based Speech Synthesis | arXiv 2025
- ingest | 2502.03930 | DiTAR: Diffusion Transformer Autoregressive Modeling for Speech Generation | arXiv 2025
- ingest | 2504.10352 | Pseudo-Autoregressive Neural Codec Language Models for Efficient Zero-Shot Text-to-Speech Synthesis | arXiv 2025
- ingest | 2508.16332 | Vevo2: A Unified and Controllable Framework for Speech and Singing Voice Generation | arXiv 2025
- ingest | 2508.02038 | Marco-Voice Technical Report | arXiv 2025
- ingest | 2604.00688 | OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models | arXiv 2026
- ingest | 2508.03543 | EmoSteer-TTS: Fine-Grained and Training-Free Emotion-Controllable Text-to-Speech via Activation Steering | arXiv 2025
- ingest | 2510.02848 | Flamed-TTS: Flow Matching Attention-Free Models for Efficient Generating and Dynamic Pacing Zero-shot Text-to-Speech | arXiv 2025
- ingest | 2506.21619 | IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech | arXiv 2025
- ingest | 2025.naacl-long.242 | StyleTTS-ZS: Efficient High-Quality Zero-Shot Text-to-Speech Synthesis with Distilled Time-Varying Style Diffusion | NAACL 2025
- ingest | 2510.12210 | DiSTAR: Diffusion over a Scalable Token Autoregressive Representation for Speech Generation | arXiv 2025
- ingest | 2025.emnlp-main.40 | Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey | EMNLP 2025
- ingest | 2603.08823 | Fish Audio S2 Technical Report | arXiv 2026
2026-05-27
- ingest | interspeech-2025-0253 | Long-Context Speech Synthesis with Context-Aware Memory | Interspeech 2025
- ingest | 2301.02111 | Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers | arXiv 2023
- ingest | interspeech-2025-0902 | VoiceQualityVC: A Voice Conversion System for Studying the Perceptual Effects of Voice Quality in Speech | Interspeech 2025
- ingest | 2025.emnlp-main.989 | VocalNet: Speech LLMs with Multi-Token Prediction for Faster and High-Quality Generation | EMNLP 2025
- ingest | 2025.acl-long.682 | Recent Advances in Speech Language Models: A Survey | ACL 2025
- ingest | 2025.americasnlp-1.1 | Text-to-speech system for low-resource languages: A case study in Shipibo-Konibo | workshop 2025
- ingest | 2025.emnlp-main.1730 | FillerSpeech: Towards Human-Like Text-to-Speech Synthesis with Filler Insertion and Filler Style Control | EMNLP 2025
- ingest | 2025.findings-naacl.184 | Continuous Speech Tokenizer in Text To Speech | NAACL 2025
- ingest | 2025.emnlp-demos.70 | OpenS2S: Advancing Fully Open-Source End-to-End Empathetic Large Speech Language Model | EMNLP 2025
- integrate | 10 papers | 14 concepts updated | 12 cross-links added
- integrate | 15 papers | 16 concepts updated | 3 cross-links added
2026-05-26
- ingest | 2509.02020 | FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot | arXiv 2025
- ingest | 2507.14534 | Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion | arXiv 2025
- ingest | 2509.19668 | Selective Classifier-free Guidance for Zero-shot Text-to-speech | arXiv 2025
- ingest | 2510.00981 | FlexiCodec: A Dynamic Neural Audio Codec for Low Frame Rates | arXiv 2025
- ingest | 2412.17048 | Why Do Speech Language Models Fail to Generate Semantically Coherent Outputs? | arXiv 2026
- ingest | 2025.findings-emnlp.424 | InteractSpeech: A Speech Dialogue Interaction Corpus for Spoken Dialogue Model | EMNLP 2025
- query | keyword filter expansion review — 11 new terms mapped to concept gaps
- ingest | 2025.acl-demo.37 | RT-VC: Real-Time Zero-Shot Voice Conversion with Speech Articulatory Coding | ACL 2025
- ingest | 2025.acl-industry.42 | Scaling Under-Resourced TTS: A Data-Optimized Framework with Advanced Acoustic Modeling for Thai | ACL 2025
- ingest | 2025.acl-long.1043 | OZSpeech: One-step Zero-shot Speech Synthesis with Learned-Prior-Conditioned Flow Matching | ACL 2025
- ingest | 2025.acl-long.1252 | Finding A Voice: Exploring the Potential of African American Dialect and Voice Generation for Chatbots | ACL 2025
- ingest | 2025.acl-long.1471 | The time scale of redundancy between prosody and linguistic context | ACL 2025
- ingest | 2025.acl-long.1498 | Analyzing and Mitigating Inconsistency in Discrete Speech Tokens for Neural Codec Language Models | ACL 2025
- ingest | 2025.acl-long.313 | F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching | ACL 2025
- ingest | 2025.acl-long.346 | ControlSpeech: Towards Simultaneous and Independent Zero-shot Speaker Cloning and Zero-shot Language Style Control | ACL 2025
- ingest | 2025.acl-long.388 | Distilling an End-to-End Voice Assistant Without Instruction Training Data | ACL 2025
- ingest | 2025.acl-long.598 | Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment | ACL 2025
- ingest | 2025.emnlp-main.180 | Scaling Rich Style-Prompted Text-to-Speech Datasets | EMNLP 2025
- ingest | interspeech-2025-0401 | Enabling the replicability of speech synthesis perceptual evaluations | Interspeech 2025
- ingest | interspeech-2025-0115 | Bringing Interpretability to Neural Audio Codecs | Interspeech 2025
- ingest | interspeech-2025-0468 | DualCodec: A Low-Frame-Rate, Semantically-Enhanced Neural Audio Codec for Speech Generation | Interspeech 2025
- ingest | interspeech-2025-1641 | Robust Neural Codec Language Modeling with Phoneme Position Prediction for Zero-Shot TTS | Interspeech 2025
- ingest | interspeech-2025-2447 | Accelerating Autoregressive Speech Synthesis Inference With Speech Speculative Decoding | Interspeech 2025
- ingest | interspeech-2025-1779 | ReFlow-VC: Zero-shot Voice Conversion Based on Rectified Flow and Speaker Feature Optimization | Interspeech 2025
- ingest | interspeech-2025-0874 | Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model | Interspeech 2025
- ingest | interspeech-2025-0246 | DC-Spin: A Speaker-invariant Speech Tokenizer for Spoken Language Models | Interspeech 2025
- ingest | interspeech-2025-1440 | FreeCodec: A Disentangled Neural Speech Codec with Fewer Tokens | Interspeech 2025
- ingest | interspeech-2025-2043 | Training-Free Voice Conversion with Factorized Optimal Transport | Interspeech 2025
- ingest | interspeech-2025-0816 | Bridging Speech and Singing: Multi-stage Speech-Prompted Singing Voice Conversion with Speaker Embedding Adaptation | Interspeech 2025
- ingest | interspeech-2025-0739 | FD-Bench: A Full-Duplex Benchmarking Pipeline Designed for Full Duplex Spoken Dialogue Systems | Interspeech 2025
- ingest | 2508.15827 | Mini-Omni-Reasoner: Token-Level Thinking-in-Speaking in Large Speech Models | arXiv 2025
- ingest | interspeech-2025-1066 | Score-Based Training for Energy-Based TTS Models | Interspeech 2025
- ingest | interspeech-2025-1122 | BitTTS: Highly Compact Text-to-Speech Using 1.58-bit Quantization and Weight Indexing | Interspeech 2025
- ingest | 2508.20660 | CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation | arXiv 2025
- ingest | interspeech-2025-1344 | Parameter-Efficient Fine-Tuning for Low-Resource Text-to-Speech via Cross-Lingual Continual Learning | Interspeech 2025
- ingest | interspeech-2025-2449 | Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling | Interspeech 2025
- ingest | interspeech-2025-1595 | Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs | Interspeech 2025
- ingest | interspeech-2025-0815 | Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion | Interspeech 2025
- ingest | interspeech-2025-1101 | ZSDEVC: Zero-Shot Diffusion-based Emotional Voice Conversion with Disentangled Mechanism | Interspeech 2025
- ingest | interspeech-2025-2660 | Triadic Multi-party Voice Activity Projection for Turn-taking in Spoken Dialogue Systems | Interspeech 2025
- ingest | 2508.07375 | TurnGuide: Enhancing Meaningful Full Duplex Spoken Interactions via Dynamic Turn-Level Text-Speech Interleaving | arXiv 2025
- ingest | 2508.16790 | TaDiCodec: Text-aware Diffusion Speech Tokenizer for Speech Language Modeling | arXiv 2025
- ingest | interspeech-2025-1289 | Unlocking Temporal Flexibility: Neural Speech Codec with Variable Frame Rate | Interspeech 2025
- ingest | interspeech-2025-0984 | Benchmarking Neural Speech Codec Intelligibility with SITool | Interspeech 2025
- ingest | 2508.07273 | Incorporating Contextual Paralinguistic Understanding in Large Speech-Language Models | arXiv 2025
- ingest | 2508.08957 | QAMRO: Quality-aware Adaptive Margin Ranking Optimization for Human-aligned Assessment of Audio Generation Systems | arXiv 2025
- ingest | 2508.09600 | OSUM-EChat: Enhancing End-to-End Empathetic Spoken Chatbot via Understanding-Driven Spoken Dialogue | arXiv 2025
- ingest | 2508.09702 | : A Multimodal, Multi-Label, Multilingual Prompt Database for Speech Generation | arXiv 2025
- ingest | 2508.11224 | Benchmarking Prosody Encoding in Discrete Speech Tokens | ASRU 2025
- ingest | 2508.13028 | Integrating Feedback Loss from Bi-modal Sarcasm Detector for Sarcastic Speech Synthesis | arXiv 2025
- ingest | 2508.15565 | Any-to-any Speaker Attribute Perturbation for Asynchronous Voice Anonymization | arXiv 2025
- ingest | 2508.15931 | QvTAD: Differential Relative Attribute Learning for Voice Timbre Attribute Detection | arXiv 2025
- ingest | 2508.16188 | Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation | EMNLP 2025
- ingest | 2508.17031 | RephraseTTS: Dynamic Length Text based Speech Insertion with Speaker Style Transfer | arXiv 2025
- ingest | 2508.17494 | Improving French Synthetic Speech Quality via SSML Prosody Control | workshop 2025
- ingest | 2508.17623 | EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Spoken Dialogue Systems | arXiv 2025
- ingest | 2508.18006 | Unseen Speaker and Language Adaptation for Lightweight Text-To-Speech with Adapters | arXiv 2025
- ingest | 2508.19205 | VibeVoice Technical Report | arXiv 2025
- ingest | 2509.00503 | Entropy-based Coarse and Compressed Semantic Speech Representation Learning | arXiv 2025
- ingest | 2509.00675 | Speaker-Conditioned Phrase Break Prediction for Text-to-Speech with Phoneme-Level Pre-trained Language Model | arXiv 2025
- ingest | 2509.01391 | MixedG2P-T5: G2P-free Speech Synthesis for Mixed-script texts using Speech Self-Supervised Learning and Language Model | arXiv 2025
- ingest | 2509.02244 | Spectrogram Patch Codec: A 2D Block-Quantized VQ-VAE and HiFi-GAN for Neural Speech Coding | arXiv 2025
- ingest | 2509.03292 | Improving Perceptual Audio Aesthetic Assessment via Triplet Loss and Self-Supervised Embeddings | arXiv 2025
- ingest | 2509.03940 | VoxRole: A Comprehensive Benchmark for Evaluating Speech-Based Role-Playing Agents | arXiv 2025
- ingest | 2509.04072 | Computational Narrative Understanding for Expressive Text-to-Speech | arXiv 2025
- ingest-batch | 2 ingested, 0 failed
- ingest-batch | 2 ingested, 0 failed
- ingest-batch | 2 ingested, 0 failed
- ingest-batch | 2 ingested, 0 failed
- ingest-batch | 1 ingested, 0 failed
- render | field overview and reader navigation | formats: field-overview, index, start | mode: full | runtime: codex | provider: openai | model: gpt-5
- render | speech-to-speech | formats: in-depth | mode: light | runtime: codex | provider: openai | model: gpt-5
- render | 5 concepts | formats: overview, in-depth, field-overview, navigation | mode: light | runtime: codex | provider: openai | model: gpt-5