Full catalog of all ingested papers. Navigate by topic via Concepts. Use search to find a specific paper by title or author.

IDTitleOrgVenueYearTaskArchitectureIngested
1904.02882LibriTTS: A Corpus Derived from LibriSpeech for Text-to-SGoogle AIarXiv2019TTS2026-06-10
2403.03100NaturalSpeech 3: Zero-Shot Speech Synthesis with FactoriMicrosoftarXiv2024TTS, VCdiffusion, hybrid2026-06-10
2504.18425Kimi-Audio Technical ReportMoonshot AIarXiv2025TTS, VC, SCAautoregressive-LM, flow-matching2026-06-10
2204.02152UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 202University of TokyoarXiv2022evaluation2026-06-10
2509.04072Computational Narrative Understanding for Expressive TearXiv2025TTSautoregressive-LM, flow-matching2026-06-04
2508.15827Mini-Omni-Reasoner: Token-Level Thinking-in-Speaking inarXiv2025SCAautoregressive-LM2026-06-03
interspeech-2025-0739FD-Bench: A Full-Duplex Benchmarking Pipeline Designed for Full Duplex Spoken Dialogue SystemsInterspeech2025SCA, evaluation2026-06-03
interspeech-2025-1993Defending Unauthorized Voice Cloning with Watermark-Aware CodecsThe Chinese University of Hong KongInterspeech2025TTS, VCautoregressive-LM2026-06-03
2508.08095Dual Information Speech Language Models for Emotional ConversationsMashang Consumer Finance Co., Ltd.arXiv2025SCAtransformer-enc-dec2026-06-03
interspeech-2025-0948PromptEVC: Controllable Emotional Voice Conversion with Natural Language PromptsSoutheast UniversityInterspeech2025VCVAE, diffusion2026-06-03
interspeech-2025-0203ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and SpeechKyushu University / University of Tokyo / EverestAI XimalayaInterspeech2025VCflow-matching, transformer-enc-dec2026-06-02
interspeech-2025-0196SPCODEC: Split and Prediction for Neural Speech CodecSamsungInterspeech2025codecGAN2026-06-02
2503.04721Full-Duplex-Bench: A Benchmark to Evaluate Full-duplex Spoken Dialogue Models on Turn-taking CapabilitiesarXiv2025SCA, evaluation2026-06-02
2508.08715MultiGen: Child-Friendly Multilingual Speech Generator with LLMsA*STAR Institute for Infocomm ResearcharXiv2025TTSautoregressive-LM, flow-matching, GAN2026-06-02
2508.09767UtterTune: LoRA-Based Target-Language Pronunciation Edit and Control in Multilingual Text-to-SpeecharXiv2025TTSautoregressive-LM2026-06-02
2508.11326MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-ExpertsKunlun Inc.arXiv2025TTSautoregressive-LM, diffusion2026-06-02
2504.12867EmoVoice: LLM-based Emotional Text-To-Speech Model with Freestyle Text PromptingShanghai Jiao Tong University / Tongyi Speech LabarXiv2025TTSautoregressive-LM, flow-matching2026-06-02
2508.08961DualSpeechLM: Towards Unified Speech Understanding and Generation via Dual Speech Token ModelingarXiv2025TTS, SCA, VCautoregressive-LM2026-06-02
2508.08399Exploring Disentangled Neural Speech Codecs from Self-Supervised RepresentationsMERL / Mitsubishi ElectricarXiv2025codec, VCVAE2026-06-02
2508.07711Is GAN Necessary for Mel-Spectrogram-based Neural Vocoder?arXiv2025TTSGAN2026-06-02
2508.07426Scalable Controllable Accented TTSJohns Hopkins UniversityASRU2025TTStransformer-enc-dec, GAN, VAE2026-06-02
2508.07302XEmoRAG: Cross-Lingual Emotion Transfer with Controllable Intensity Using Retrieval-Augmented GenerationNorthwestern Polytechnical UniversityarXiv2025TTS, VCautoregressive-LM, flow-matching2026-06-02
2508.06890Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit ProsodyPOSTECHarXiv2025VCGAN, transformer-enc-dec2026-06-02
2508.06870Text to Speech System for Meitei Mayek ScriptarXiv2025TTStransformer-enc-dec, GAN2026-06-02
2508.05385A Scalable Pipeline for Enabling Non-Verbal Speech Generation and UnderstandingTsinghua UniversityarXiv2025TTS, SCAtransformer-enc-dec2026-06-02
2508.14049MahaTTS: A Unified Framework for Multilingual Text-to-Speech SynthesisDubverse AIarXiv2025TTSautoregressive-LM, flow-matching2026-06-02
2508.04585UniTalker: Conversational Speech-Visual SynthesisInner Mongolia UniversityarXiv2025TTS, SCAautoregressive-LM, flow-matching2026-06-02
2508.04996REF-VC: Robust, Expressive and Fast Zero-Shot Voice Conversion with Diffusion TransformersNorthwestern Polytechnical UniversityarXiv2025VCflow-matching2026-06-02
2508.05207SpectroStream: A Versatile Neural Codec for General AudioGoogle DeepMindarXiv2025codecGAN2026-06-02
2507.20091ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language ModelsarXiv2025SCA, TTSautoregressive-LM2026-06-02
2508.00317Advancing Speech Quality Assessment Through Scientific Challenges and Open-source ActivitiesNagoya UniversityarXiv2025evaluation2026-06-02
2507.22746Next Tokens Denoising for Speech SynthesisMicrosoftarXiv2025TTShybrid2026-06-02
2508.01796Enhancing Spectrogram Realism in Singing Voice Synthesis via Explicit Bandwidth Extension Prior to VocoderTsinghua UniversityarXiv2025singing, TTSdiffusion, GAN2026-06-02
2508.02013SpeechRole: A Large-Scale Dataset and Benchmark for Evaluating Speech Role-Playing AgentsFudan UniversityarXiv2025SCA, evaluation2026-06-02
2508.02849SecoustiCodec: Cross-Modal Aligned Streaming Single-Codebook Speech CodecarXiv2025codecVAE, transformer-enc-dec2026-06-02
2025.naacl-long.110WaveFM: A High-Fidelity and Efficient Vocoder Based on Flow MatchingTsinghua UniversityNAACL2025TTSflow-matching2026-05-30
2025.findings-acl.1051LLMVoX: Autoregressive Streaming Text-to-Speech Model for Any LLMMBZUAIACL2025TTS, SCAautoregressive-LM2026-05-30
2025.emnlp-main.180Scaling Rich Style-Prompted Text-to-Speech DatasetsUT Austin / NYUEMNLP2025TTS, evaluationautoregressive-LM2026-06-01
2507.09318ZipVoice-Dialog: Non-Autoregressive Spoken Dialogue Generation with Flow MatchingXiaomi Corp.arXiv2026TTS, SCAflow-matching2026-05-30
2025.coling-main.518ProsodyFlow: High-fidelity Text-to-Speech through Conditional Flow Matching and Prosody Modeling with Large Speech Language ModelsZhejiang Universityworkshop2025TTSflow-matching, hybrid2026-05-30
interspeech-2025-0469Developing High-Quality TTS for Punjabi and Urdu: Benchmarking against MMS ModelsUniversity of Engineering and Technology, LahoreInterspeech2025TTS, evaluationtransformer-enc-dec2026-05-30
interspeech-2025-0854Bridging the Training–Inference Gap in TTS: Training Strategies for Robust Generative Postprocessing for Low-Resource SpeakersFraunhofer IISInterspeech2025TTSGAN, flow-matching, transformer-enc-dec2026-05-30
interspeech-2025-0973A Dataset for Automatic Assessment of TTS Quality in SpanishInterspeech2025TTS, evaluation2026-05-30
interspeech-2025-0989HiFiTTS-2: A Large-Scale High Bandwidth Speech DatasetNVIDIAInterspeech2025TTS, evaluationautoregressive-LM2026-05-30
interspeech-2025-1034Non-Standard Accent TTS Support via Large Multi-Accent Frontend Pronunciation Knowledge TransferUniversity of EdinburghInterspeech2025TTStransformer-enc-dec2026-05-30
interspeech-2025-0723Counterfactual Activation Editing for Post-hoc Prosody and Mispronunciation Correction in TTS ModelsKAIST / Samsung ElectronicsInterspeech2025TTStransformer-enc-dec, VAE2026-05-30
interspeech-2025-0754EME-TTS: Unlocking the Emphasis and Emotion Link in Speech SynthesisUCAS HangzhouInterspeech2025TTStransformer-enc-dec2026-05-30
interspeech-2025-0762Intrasentential English in Swedish TTS: perceived English-accentednessKTH / MTMInterspeech2025TTSflow-matching2026-05-30
interspeech-2025-0779Intelligibility of Text-to-Speech Systems for Mathematical ExpressionsEricsson R&DInterspeech2025TTS, evaluation2026-05-30
interspeech-2025-0787Gradual modeling of the Lombard effect by modifying speaker embeddings from a Text-To-Speech modelHEAD acousticsInterspeech2025TTSautoregressive-LM2026-05-30
interspeech-2025-0575VoiceMark: Zero-Shot Voice Cloning-Resistant Watermarking Approach Leveraging Speaker-Specific LatentsTsinghua UniversityInterspeech2025VC, TTSVAE2026-05-30
interspeech-2025-0596Facilitating Personalized TTS for Dysarthric Speakers Using Knowledge Anchoring and Curriculum LearningPOSTECHInterspeech2025TTStransformer-enc-dec2026-05-30
interspeech-2025-0648MIKU-PAL: An Automated and Standardized Multimodal Method for Speech Paralinguistic and Affect LabelingFish AudioInterspeech2025TTS, evaluation2026-05-30
interspeech-2025-0669PAST: Phonetic-Acoustic Speech TokenizerHebrew University of JerusalemInterspeech2025codec, TTShybrid2026-05-30
interspeech-2025-0704Differentiable Reward Optimization for LLM based TTS systemAlibaba GroupInterspeech2025TTSautoregressive-LM, flow-matching2026-05-30
interspeech-2025-0406Zero-Shot Mono-to-Binaural Speech SynthesisGoogleInterspeech2025TTSGAN2026-05-30
interspeech-2025-0408Improving User Impression of Spoken Dialogue Systems by Controlling Para-linguistic Expression Based on IntimacyTohoku UniversityInterspeech2025SCA, TTStransformer-enc-dec, GAN2026-05-30
interspeech-2025-0455APTTS: Adversarial Post-training in Latent Flow Matching for Fast and High-fidelity Text-to-SpeechLG AI ResearchInterspeech2025TTSflow-matching, VAE, hybrid2026-05-30
interspeech-2025-0554RapFlow-TTS: Rapid and High-Fidelity Text-to-Speech with Improved Consistency Flow MatchingNAVER CloudInterspeech2025TTSflow-matching, GAN2026-05-30
interspeech-2025-0551Monotonic Attention for Robust Text-to-Speech Synthesis in Large Language Model FrameworksTencentInterspeech2025TTSautoregressive-LM2026-05-30
interspeech-2025-0047Revival with Voice: Multi-modal Controllable Text-to-Speech SynthesisMeta AIInterspeech2025TTSautoregressive-LM2026-05-30
interspeech-2025-0063Fine-Tuning Text-to-Speech Diffusion Models Using Reinforcement Learning with Human FeedbackInterspeech2025TTSdiffusion2026-05-30
interspeech-2025-0143Multimodal Prosody Modeling: A Use Case for Multilingual Sentence Mode PredictionIdiap Research InstituteInterspeech2025TTS, evaluation2026-05-30
interspeech-2025-0310Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language ModelsInterspeech2025TTS, codecautoregressive-LM2026-05-30
interspeech-2025-0319Improving Noise Robustness of LLM-based Zero-shot TTS via Discrete Acoustic Token DenoisingUSTCInterspeech2025TTSautoregressive-LM2026-05-30
2509.02020FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and ChatbotarXiv2025TTS, SCAautoregressive-LM2026-05-26
2507.14534Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice ConversionarXiv2025VCGAN, hybrid2026-05-26
2509.19668Selective Classifier-free Guidance for Zero-shot Text-to-speecharXiv2025TTSflow-matching2026-05-26
2510.00981FlexiCodec: A Dynamic Neural Audio Codec for Low Frame RatesarXiv2025codec, TTShybrid2026-05-26
2412.17048Why Do Speech Language Models Fail to Generate Semantically Coherent Outputs?arXiv2026SCAautoregressive-LM2026-05-26
2025.findings-emnlp.424InteractSpeech: A Speech Dialogue Interaction Corpus for Spoken Dialogue ModelEMNLP2025SCA, evaluationautoregressive-LM2026-05-26
2025.acl-demo.37RT-VC: Real-Time Zero-Shot Voice Conversion with Speech Articulatory CodingUC BerkeleyACL2025VChybrid2026-05-26
2025.acl-industry.42Scaling Under-Resourced TTS: A Data-Optimized Framework with Advanced Acoustic Modeling for ThaiBeijing Logic Intelligence TechnologyACL2025TTSGAN, transformer-enc-dec2026-05-26
2025.acl-long.1043OZSpeech: One-step Zero-shot Speech Synthesis with Learned-Prior-Conditioned Flow MatchingFPT Software AI CenterACL2025TTSflow-matching, hybrid2026-05-26
2025.acl-long.1252Finding A Voice: Exploring the Potential of African American Dialect and Voice Generation for ChatbotsEmory UniversityACL2025TTS, SCAhybrid2026-05-26
2025.acl-long.1471The time scale of redundancy between prosody and linguistic contextMITACL2025evaluationtransformer-enc-dec2026-05-26
2025.acl-long.1498Analyzing and Mitigating Inconsistency in Discrete Speech Tokens for Neural Codec Language ModelsAlibaba Group / Zhejiang UniversityACL2025TTS, codecautoregressive-LM, GAN2026-05-26
2025.acl-long.313F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow MatchingShanghai Jiao Tong UniversityACL2025TTSflow-matching2026-06-01
2025.acl-long.346ControlSpeech: Towards Simultaneous and Independent ZerZhejiang University / Alibaba Tongyi Speech LabACL2025TTStransformer-enc-dec, hybrid2026-05-26
2025.acl-long.388Distilling an End-to-End Voice Assistant Without InstruACL2025SCAtransformer-enc-dec2026-05-26
2025.acl-long.598Advancing Zero-shot Text-to-Speech Intelligibility acroACL2025TTSautoregressive-LM, flow-matching, hybrid2026-05-26
interspeech-2025-0253Long-Context Speech Synthesis with Context-Aware MemorySouth China University of Technology / Alibaba GroupInterspeech2025TTSautoregressive-LM, hybrid2026-05-27
2301.02111Neural Codec Language Models are Zero-Shot Text to Speech SynthesizersMicrosoftarXiv2023TTS, codecautoregressive-LM2026-06-01
interspeech-2025-0902VoiceQualityVC: A Voice Conversion System for Studying the Perceptual Effects of Voice Quality in SpeechKTH Royal Institute of TechnologyInterspeech2025VCVAE, GAN2026-05-27
2025.emnlp-main.989VocalNet: Speech LLMs with Multi-Token Prediction for Faster and High-Quality GenerationSJTU / Ant Group / Wuhan UniversityEMNLP2025TTS, SCAautoregressive-LM, hybrid2026-05-27
2025.acl-long.682Recent Advances in Speech Language Models: A SurveyCUHK / Tencent / NUSACL2025TTS, SCA, evaluationautoregressive-LM, hybrid2026-06-01
2025.americasnlp-1.1Text-to-speech system for low-resource languages: A case study in Shipibo-KoniboPontificia Universidad Católica del Perúworkshop2025TTStransformer-enc-dec, GAN2026-05-27
2025.emnlp-main.1730FillerSpeech: Towards Human-Like Text-to-Speech Synthesis with Filler Insertion and Filler Style ControlKorea University / Samsung ResearchEMNLP2025TTSflow-matching, transformer-enc-dec2026-05-27
2025.findings-naacl.184Continuous Speech Tokenizer in Text To SpeechCUHK / TencentNAACL2025TTS, codecautoregressive-LM, VAE, flow-matching2026-05-27
2025.emnlp-demos.70OpenS2S: Advancing Fully Open-Source End-to-End Empathetic Large Speech Language ModelInstitute of Automation, Chinese Academy of SciencesEMNLP2025SCAautoregressive-LM, hybrid2026-05-27
2406.02430Seed-TTS: A Family of High-Quality Versatile Speech Generation ModelsByteDancearXiv2024TTS, VCautoregressive-LM, diffusion, hybrid2026-05-28
2407.05407CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic TokensAlibaba GrouparXiv2024TTSautoregressive-LM, flow-matching, hybrid2026-05-28
2025.acl-long.65Autoregressive Speech Synthesis without Vector QuantizationMicrosoft / CUHKACL2025TTSautoregressive-LM, VAE, hybrid2026-05-28
2412.10117CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language ModelsAlibaba GrouparXiv2024TTSautoregressive-LM, flow-matching, hybrid2026-05-28
2601.15621Qwen3-TTS Technical ReportAlibaba / Qwen TeamarXiv2026TTSautoregressive-LM, hybrid2026-05-28
2512.14291GLM-TTS Technical ReportZhipu AI / Tsinghua UniversityarXiv2025TTSautoregressive-LM, diffusion, GAN, hybrid2026-05-28
2508.06262Llasa+: Free Lunch for Accelerated and Streaming Llama-Based Speech SynthesisNorthwestern Polytechnical University / HKUSTarXiv2025TTSautoregressive-LM, hybrid2026-05-28
2502.03930DiTAR: Diffusion Transformer Autoregressive Modeling for Speech GenerationByteDance SeedarXiv2025TTSautoregressive-LM, diffusion, hybrid2026-05-28
2504.10352Pseudo-Autoregressive Neural Codec Language Models for Efficient Zero-Shot Text-to-Speech SynthesisMicrosoft / SJTUarXiv2025TTSautoregressive-LM, hybrid2026-05-28
2508.16332Vevo2: A Unified and Controllable Framework for Speech and Singing Voice GenerationCUHK Shenzhen / ByteDance SeedarXiv2025TTS, VC, singingautoregressive-LM, flow-matching, hybrid2026-05-28
2508.02038Marco-Voice Technical ReportAlibaba International Digital CommercearXiv2025TTS, VCautoregressive-LM, flow-matching, hybrid2026-05-28
2604.00688OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language ModelsXiaomi Corp.arXiv2026TTSdiffusion, hybrid2026-05-28
2508.03543EmoSteer-TTS: Fine-Grained and Training-Free Emotion-Controllable Text-to-Speech via Activation SteeringHKUST (Guangzhou) / Tencent AI LabarXiv2025TTSflow-matching2026-05-28
2510.02848Flamed-TTS: Flow Matching Attention-Free Models for Efficient Generating and Dynamic Pacing Zero-shot Text-to-SpeechFPT Software AI CenterarXiv2025TTSflow-matching, hybrid2026-05-28
2506.21619IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-SpeechbilibiliarXiv2025TTSautoregressive-LM, flow-matching, GAN, hybrid2026-05-28
2025.naacl-long.242StyleTTS-ZS: Efficient High-Quality Zero-Shot Text-to-Speech Synthesis with Distilled Time-Varying Style DiffusionColumbia UniversityNAACL2025TTSdiffusion, GAN, VAE, hybrid2026-05-28
2510.12210DiSTAR: Diffusion over a Scalable Token Autoregressive Representation for Speech GenerationSJTU / ByteDancearXiv2025TTSautoregressive-LM, diffusion, hybrid2026-05-28
2025.emnlp-main.40Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic SurveyHKUST-GZ / University of SurreyEMNLP2025TTS, evaluationautoregressive-LM, flow-matching, diffusion, GAN, VAE, transformer-enc-dec, hybrid2026-05-28
2603.08823Fish Audio S2 Technical ReportFish AudioarXiv2026TTSautoregressive-LM, GAN, hybrid2026-05-28
2509.00685MPO: Multidimensional Preference Optimization for Language Model-based Text-to-SpeechNorthwestern Polytechnical UniversityarXiv2025TTSautoregressive-LM2026-05-29
2511.12347VoiceCraft-X: Unifying Multilingual, Voice-Cloning Speech Synthesis and Speech EditingUniversity of Texas at Austin / AmazonEMNLP2025TTSautoregressive-LM2026-05-29
2512.13251DisCo-Speech: Controllable Zero-Shot Speech GenerationChina Mobile Nineverse AI / Peking UniversityarXiv2025TTS, VC, codecautoregressive-LM, GAN2026-05-29
2509.09631DiFlow-TTS: Compact and Low-Latency Zero-Shot Text-to-Speech with Factorized Discrete Flow MatchingFPT Software AI CenterarXiv2025TTSflow-matching, transformer-enc-dec2026-05-29
2512.04720M3-TTS: Multi-modal DiT Alignment & Mel-latent for Zero-shot High-fidelity SpeecharXiv2025TTSdiffusion, VAE2026-05-29
2603.29339LongCat-AudioDiT: High-Fidelity Diffusion Text-to-Speech in the Waveform Latent SpaceMeituanarXiv2026TTSflow-matching, VAE2026-05-29
2508.11273EmoSSLSphere: Multilingual Emotional Speech Synthesis with Spherical Vectors and Discrete Speech TokensarXiv2025TTStransformer-enc-dec2026-05-29
2604.12438An Ultra-Low Latency, End-to-End Streaming Speech Synthesis Architecture via Block-Wise Generation and Depth-Wise Codec DecodingarXiv2026TTStransformer-enc-dec2026-06-01
2604.01760T5Gemma-TTS Technical ReportarXiv2026TTSautoregressive-LM2026-05-29
2508.15442Mitigating Hallucinations in LM-Based TTS Models via Distribution Alignment Using GFlowNetsEMNLP2025TTSautoregressive-LM2026-05-29
2025.acl-long.654Language-Codec: Bridging Discrete Codec Representations and Speech Language ModelsZhejiang UniversityACL2025TTS, codecGAN, VAE2026-05-29
2603.18090MOSS-TTS Technical ReportShanghai Innovation Institute / Fudan UniversityarXiv2026TTSautoregressive-LM, hybrid2026-05-29
2508.04141Parallel GPT: Harmonizing the Independence and Interdependence of Acoustic and Semantic Information for Zero-Shot Text-to-SpeechSouth China University of TechnologyarXiv2025TTSautoregressive-LM, hybrid2026-05-29
2502.11128FELLE: Autoregressive Speech Synthesis with Token-Wise Coarse-to-Fine Flow MatchingarXiv2025TTSautoregressive-LM, flow-matching2026-05-29
2603.26364LLaDA-TTS: Unifying Speech Synthesis and Zero-Shot Editing via Masked Diffusion ModelingBairong, Inc.arXiv2026TTSdiffusion2026-05-29
2508.19098CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech SynthesisarXiv2025TTSautoregressive-LM, flow-matching, VAE2026-05-29
2508.12001FNH-TTS: A Fast, Natural, and Human-Like Speech Synthesis System with advanced prosodic modeling based on Mixture of ExpertsMegatronixarXiv2025TTSVAE, GAN, hybrid2026-05-29
2510.05758EMORL-TTS: Reinforcement Learning for Fine-Grained Emotion Control in LLM-based TTSHangzhou Institute for Advanced Study, UCASICASSP2026TTSautoregressive-LM2026-05-29
2601.03888IndexTTS 2.5 Technical ReportBilibiliarXiv2026TTSautoregressive-LM, flow-matching, hybrid2026-05-29
2509.15969VoXtream: Full-Stream Text-to-Speech with Extremely Low LatencyKTH Royal Institute of TechnologyarXiv2025TTSautoregressive-LM, hybrid2026-05-29
2510.07979IntMeanFlow: Few-step Speech Generation with Integral Velocity DistillationByteDancearXiv2025TTSflow-matching2026-05-29
2025.ccl-1.80Lao-English Code-Switched Speech Synthesis Via Neural Codec Language ModelingKunming University of Science and Technologyworkshop2025TTSautoregressive-LM, hybrid2026-05-29
2025.coling-main.352DiffStyleTTS: Diffusion-based Hierarchical Prosody Modeling for Text-to-Speech with Diverse and Controllable StylesUniversity of Science and Technology of Chinaworkshop2025TTSdiffusion, transformer-enc-dec2026-05-29
2025.acl-long.911DNASpeech: A Contextualized and Situated Text-to-Speech Dataset with Dialogues, Narratives and ActionsRenmin University of ChinaACL2025TTS, evaluationtransformer-enc-dec2026-05-29
2025.acl-short.81Zero-Shot Text-to-Speech for VietnameseMovian AIACL2025TTS, evaluationautoregressive-LM, transformer-enc-dec2026-05-29
2025.acl-long.912LLaMA-Omni 2: LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech SynthesisChinese Academy of SciencesACL2025SCA, TTSautoregressive-LM, flow-matching, hybrid2026-05-29
interspeech-2025-2765The State Of TTS: A Case Study with Human Fooling RatesIIT MadrasInterspeech2025TTS, evaluation2026-06-03
interspeech-2025-0401Enabling the replicability of speech synthesis perceptuInterspeech2025evaluation2026-06-03
interspeech-2025-0115Bringing Interpretability to Neural Audio CodecsInterspeech2025codectransformer-enc-dec2026-06-03
interspeech-2025-0468DualCodec: A Low-Frame-Rate, Semantically-Enhanced NeurCUHK-SZ / BaiduInterspeech2025codecVAE, GAN2026-06-03
interspeech-2025-1641Robust Neural Codec Language Modeling with Phoneme PosiSamsungInterspeech2025TTSautoregressive-LM2026-06-03
interspeech-2025-2447Accelerating Autoregressive Speech Synthesis InferenceTsinghua / TencentInterspeech2025TTSautoregressive-LM2026-06-03
interspeech-2025-1779ReFlow-VC: Zero-shot Voice Conversion Based on RectifieInterspeech2025VCflow-matching2026-06-03
interspeech-2025-0874Efficient and Direct Duplex Modeling for Speech-to-SpeeInterspeech2025SCAautoregressive-LM, hybrid2026-06-03
interspeech-2025-0246DC-Spin: A Speaker-invariant Speech Tokenizer for SpokeInterspeech2025codectransformer-enc-dec2026-06-03
interspeech-2025-1440FreeCodec: A Disentangled Neural Speech Codec with FeweInterspeech2025codecVAE2026-06-03
interspeech-2025-2043Training-Free Voice Conversion with Factorized Optimal Interspeech2025VCtransformer-enc-dec2026-06-03
interspeech-2025-0816Bridging Speech and Singing: Multi-stage Speech-PrompteInterspeech2025singing, VCdiffusion2026-06-03
interspeech-2025-1066Score-Based Training for Energy-Based TTS ModelsInterspeech2025TTSdiffusion2026-06-03
interspeech-2025-1122BitTTS: Highly Compact Text-to-Speech Using 1.58-bit QuInterspeech2025TTSGAN, transformer-enc-dec2026-06-03
2508.20660CodecBench: A Comprehensive Benchmark for Acoustic andFudan UniversityarXiv2025codec, evaluation2026-06-03
interspeech-2025-1344Parameter-Efficient Fine-Tuning for Low-Resource Text-tAjou UniversityInterspeech2025TTSflow-matching2026-06-03
interspeech-2025-2449Accelerating Flow-Matching-Based Text-to-Speech via EmpInterspeech2025TTSflow-matching2026-06-03
interspeech-2025-1595Scheduled Interleaved Speech-Text Training for Speech-tInterspeech2025TTS, SCAautoregressive-LM2026-06-03
interspeech-2025-0815Towards Better Disentanglement in Non-Autoregressive ZeInterspeech2025VCVAE, GAN2026-06-03
interspeech-2025-1101ZSDEVC: Zero-Shot Diffusion-based Emotional Voice ConveInterspeech2025VCdiffusion2026-06-03
interspeech-2025-2660Triadic Multi-party Voice Activity Projection for Turn-Kyoto UniversityInterspeech2025SCAtransformer-enc-dec2026-06-03
2508.07375TurnGuide: Enhancing Meaningful Full Duplex Spoken IntearXiv2025SCAautoregressive-LM2026-06-04
2508.16790TaDiCodec: Text-aware Diffusion Speech Tokenizer for SpCUHK-SZarXiv2025codecdiffusion, transformer-enc-dec, autoregressive-LM2026-06-04
interspeech-2025-1289Unlocking Temporal Flexibility: Neural Speech Codec witInterspeech2025codechybrid2026-06-04
interspeech-2025-0984Benchmarking Neural Speech Codec Intelligibility with SInterspeech2025codec, evaluation2026-06-04
2508.07273Incorporating Contextual Paralinguistic Understanding iarXiv2025SCAtransformer-enc-dec2026-06-04
2508.08957QAMRO: Quality-aware Adaptive Margin Ranking OptimizatiarXiv2025evaluation2026-06-04
2508.09600OSUM-EChat: Enhancing End-to-End Empathetic Spoken ChatNorthwestern Polytechnical UniversityarXiv2025SCAautoregressive-LM, flow-matching2026-06-04
2508.09702M3PDB: A Multimodal, Multi-Label, Multilingual Prompt Database for Speech GenerationarXiv2025TTS, evaluation2026-06-04
2508.11224Benchmarking Prosody Encoding in Discrete Speech TokensThe University of Tokyo / AISTASRU2025evaluation, TTS2026-06-04
2508.13028Integrating Feedback Loss from Bi-modal Sarcasm DetectoarXiv2025TTStransformer-enc-dec2026-06-04
2508.15565Any-to-any Speaker Attribute Perturbation for AsynchronarXiv2025VCGAN2026-06-04
2508.15931QvTAD: Differential Relative Attribute Learning for VoiQifu TechnologyarXiv2025evaluationtransformer-enc-dec2026-06-04
2508.16188Seeing is Believing: Emotion-Aware Audio-Visual LanguagEMNLP2025TTS, SCAautoregressive-LM2026-06-04
2508.17031RephraseTTS: Dynamic Length Text based Speech InsertionIIT KanpurarXiv2025TTS, VCtransformer-enc-dec, GAN2026-06-04
2508.17494Improving French Synthetic Speech Quality via SSML Prosworkshop2025TTShybrid2026-06-04
2508.17623EMO-Reasoning: Benchmarking Emotional Reasoning CapabilarXiv2025SCA, evaluation2026-06-04
2508.18006Unseen Speaker and Language Adaptation for LightweightAmazonarXiv2025TTSGAN2026-06-04
2508.19205VibeVoice Technical ReportMicrosoft ResearcharXiv2025TTShybrid2026-06-04
2509.00503Entropy-based Coarse and Compressed Semantic Speech ReparXiv2025codecautoregressive-LM, transformer-enc-dec2026-06-04
2509.00675Speaker-Conditioned Phrase Break Prediction for Text-toarXiv2025TTStransformer-enc-dec2026-06-04
2509.01391MixedG2P-T5: G2P-free Speech Synthesis for Mixed-scriptarXiv2025TTStransformer-enc-dec2026-06-04
2509.02244Spectrogram Patch Codec: A 2D Block-Quantized VQ-VAE anarXiv2025codecVAE, GAN2026-06-04
2509.03292Improving Perceptual Audio Aesthetic Assessment via TriarXiv2025evaluationhybrid2026-06-04
2509.03940VoxRole: A Comprehensive Benchmark for Evaluating SpeecarXiv2025SCA, evaluation2026-06-04
2410.00037Moshi: a speech-text foundation model for real-time diaKyutaiarXiv2024SCA, TTSautoregressive-LM, hybrid2026-06-09
2411.13577WavChat: A Survey of Spoken Dialogue ModelsarXiv2024SCAautoregressive-LM, transformer-enc-dec, hybrid2026-06-09
2503.20215Qwen2.5-Omni Technical ReportarXiv2025autoregressive-LM, flow-matching, hybrid2026-06-09
2407.21783The Llama 3 Herd of ModelsMetaarXiv2024autoregressive-LM2026-06-09
1912.06670Common Voice: A Massively-Multilingual Speech CorpusarXiv20192026-06-09
2312.15185emotion2vec: Self-Supervised Pre-Training for Speech EmarXiv20232026-06-09
2212.04356Robust Speech Recognition via Large-Scale Weak SupervisarXiv2022transformer-enc-dec2026-06-09
2010.05646HiFi-GAN: Generative Adversarial Networks for EfficientKakao EnterprisearXiv2020TTSGAN2026-06-09
2210.13438High Fidelity Neural Audio CompressionMeta AIarXiv2022codecGAN, VAE2026-06-09
2006.04558FastSpeech 2: Fast and High-Quality End-to-End Text toMicrosoft Research AsiaarXiv2020TTStransformer-enc-dec2026-06-09
2407.10759Qwen2-Audio Technical ReportAlibaba GrouparXiv20242026-06-10
2303.08774GPT-4 Technical ReportOpenAIarXiv2023autoregressive-LM2026-06-10
2412.02612GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken ChatbotTsinghua University / Zhipu.AIarXiv2024TTS, SCAautoregressive-LM, flow-matching2026-06-10
1711.05101Decoupled Weight Decay RegularizationarXiv20172026-06-10
2410.21276GPT-4o System CardarXiv2024autoregressive-LM2026-06-10
2412.15115Qwen2.5 Technical ReportAlibabaarXiv2024autoregressive-LM2026-06-10
2210.02747Flow Matching for Generative ModelingMeta AI (FAIR)arXiv2022flow-matching2026-06-10
2209.03143AudioLM: a Language Modeling Approach to Audio GeneratiarXiv2022TTS, SCAautoregressive-LM2026-06-10
2206.04658BigVGAN: A Universal Neural Vocoder with Large-Scale TrNVIDIAarXiv2022TTSGAN2026-06-10
2207.12598Classifier-Free Diffusion GuidancearXiv2022diffusion2026-06-10
1412.6980Adam: A Method for Stochastic OptimizationarXiv20142026-06-10
2308.16692SpeechTokenizer: Unified Speech Tokenizer for Speech LaFudan UniversityarXiv2023TTS, codecGAN, VAE2026-06-11
2503.01710Spark-TTS: An Efficient LLM-Based Text-to-Speech ModelHKUSTarXiv2025TTSautoregressive-LM2026-06-11
2409.00750MaskGCT: Zero-Shot Text-to-Speech with Masked GenerativCUHK-SZarXiv2024TTS, VCautoregressive-LM2026-06-11
2505.17589CosyVoice 3: Towards In-the-wild Speech Generation viaAlibabaarXiv2025TTSautoregressive-LM, flow-matching2026-06-11
2408.16725Mini-Omni: Language Models Can Hear, Talk While ThinkiInspiraiarXiv2024SCA, TTSautoregressive-LM2026-06-11
2502.04128Llasa: Scaling Train-Time and Inference-Time Compute foarXiv2025TTSautoregressive-LM2026-06-11
2304.09116NaturalSpeech 2: Latent Diffusion Models are Natural anMicrosoft Research AsiaarXiv2023TTS, VC, singingdiffusion, VAE2026-06-11
2406.05370VALL-E 2: Neural Codec Language Models are Human ParityMicrosoftarXiv2024TTSautoregressive-LM2026-06-11
2409.06666LLaMA-Omni: Seamless Speech Interaction with Large LangICT/CASarXiv2024SCAhybrid2026-06-11
2411.00774Freeze-Omni: A Smart and Low Latency Speech-to-speech DTencent Youtu LabarXiv2024SCAautoregressive-LM, hybrid2026-06-11
2408.16532WavTokenizer: an Efficient Acoustic Discrete Codec TokearXiv2024codecGAN, VAE2026-06-11
2407.04051FunAudioLLM: Voice Understanding and Generation FoundatAlibaba GrouparXiv2024TTShybrid2026-06-11
2305.11000SpeechGPT: Empowering Large Language Models with IntrinFudan UniversityarXiv2023SCA, TTSautoregressive-LM2026-06-11
2410.17196VoiceBench: Benchmarking LLM-Based Voice AssistantsNational University of SingaporearXiv2024evaluation, SCA2026-06-11
2409.03283FireRedTTS: A Foundation Text-To-Speech Framework for IXiaohongshuarXiv2024TTSautoregressive-LM, flow-matching, hybrid2026-06-11
2311.07919Qwen-Audio: Advancing Universal Audio Understanding viaarXiv2023transformer-enc-dec2026-06-12
2505.09388Qwen3 Technical ReportarXiv2025autoregressive-LM2026-06-12
2005.07143ECAPA-TDNN: Emphasized Channel Attention, Propagation aarXiv20202026-06-12
2302.13971LLaMA: Open and Efficient Foundation Language ModelsarXiv2023autoregressive-LM2026-06-12
2507.06261Gemini 2.5: Pushing the Frontier with Advanced ReasoninGooglearXiv2025TTS, SCA2026-06-12
2012.03411MLS: A Large-Scale Multilingual Dataset for Speech ResearXiv20202026-06-12
2501.12948DeepSeek-R1: Incentivizing Reasoning Capability in LLMsDeepSeekarXiv2025autoregressive-LM2026-06-12
2312.05187Seamless: Multilingual Expressive and Streaming SpeecharXiv20232026-06-12
2106.06909GigaSpeech: An Evolving, Multi-domain ASR Corpus with 1arXiv20212026-06-12
2309.15505Finite Scalar Quantization: VQ-VAE Made SimplearXiv2023VAE2026-06-12
2306.00814Vocos: Closing the gap between time-domain and Fourier-arXiv2023TTSGAN2026-06-12
2407.05361Emilia: An Extensive, Multilingual, and Diverse SpeecharXiv20242026-06-12
2406.18009E2 TTS: Embarrassingly Easy Fully Non-Autoregressive ZeMicrosoftarXiv2024TTSflow-matching2026-06-12
2406.04904XTTS: a Massively Multilingual Zero-Shot Text-to-SpeechCoqui.ai / NVIDIA / Cantina.aiarXiv2024autoregressive-LM, GAN2026-06-12
2409.05377BigCodec: Pushing the Limits of Low-Bitrate Neural SpeeUniversity of Tokyo, Microsoft, Keio UniversityarXiv2024codecGAN, VAE2026-06-12
2305.02765HiFi-Codec: Group-residual Vector quantization for HighPeking University / Tencent AI LabarXiv2023codecGAN, VAE2026-06-12
2403.16973VoiceCraft: Zero-Shot Speech Editing and Text-to-SpeecharXiv2024TTSautoregressive-LM2026-06-12
2502.11946Step-Audio: Unified Understanding and Generation in IntStepFunarXiv2025TTS, SCAautoregressive-LM, flow-matching, hybrid2026-06-12
2501.06282MinMo: A Multimodal Large Language Model for Seamless VAlibaba GrouparXiv2025TTS, SCAautoregressive-LM, flow-matching, hybrid2026-06-12
2303.03926Speak Foreign Languages with Your Own Voice: Cross-LingMicrosoftarXiv2023TTS, multilingual-ttsautoregressive-LM2026-06-12
2305.09636SoundStorm: Efficient Parallel Audio GenerationGooglearXiv2023TTS, SCAautoregressive-LM2026-06-13
1712.05884Natural TTS Synthesis by Conditioning WaveNet on Mel SpGooglearXiv2017TTStransformer-enc-dec2026-06-13
2402.01912Natural language guidance of high-fidelity text-to-speeStability AIarXiv2024TTSautoregressive-LM2026-06-13
2306.12925AudioPaLM: A Large Language Model That Can Speak and LiGooglearXiv2023TTS, SCAautoregressive-LM2026-06-13
2305.07243Better speech synthesis through scalingarXiv2023TTSautoregressive-LM, diffusion, VAE2026-06-13
1609.03499WaveNet: A Generative Model for Raw AudioGoogle DeepMindarXiv2016TTSautoregressive-LM2026-06-13
2411.19842Scaling Transformers for Low-Bitrate High-Quality SpeecStability AIarXiv2024codectransformer-enc-dec, VAE2026-06-13
2407.08551Autoregressive Speech Synthesis without Vector QuantizaMicrosoftarXiv2024TTSautoregressive-LM2026-06-13
1703.10135Tacotron: Towards End-to-End Speech SynthesisGooglearXiv2017transformer-enc-dec2026-06-13
2502.17239Baichuan-Audio: A Unified Framework for End-to-End SpeeBaichuan Inc.arXiv2025SCA, TTSautoregressive-LM, flow-matching, hybrid2026-06-13
2402.05755Spirit LM: Interleaved Spoken and Written Language ModeMeta AIarXiv2024SCA, TTSautoregressive-LM2026-06-13
2402.08093BASE TTS: Lessons from building a billion-parameter TexAmazon AGIarXiv2024TTSautoregressive-LM2026-06-13
2507.16632Step-Audio 2 Technical ReportStepFunarXiv2025SCA, TTSautoregressive-LM, flow-matching, hybrid2026-06-13
2310.00704UniAudio: An Audio Foundation Model Toward Universal AuarXiv2023TTS, VC, singingautoregressive-LM, hybrid2026-06-13
2106.15561A Survey on Neural Speech SynthesisMicrosoft Research AsiaarXiv2021TTSautoregressive-LM, flow-matching, diffusion, GAN, VAE, transformer-enc-dec2026-06-13
2411.01156Fish-Speech: Leveraging Large Language Models for AdvanFish AudioarXiv2024TTSautoregressive-LM, GAN2026-06-14
2505.07916MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech withMiniMaxarXiv2025TTSautoregressive-LM, flow-matching, VAE2026-06-14
2410.11190Mini-Omni2: Towards Open-source GPT-4o with Vision, SpeInspirai / Tsinghua UniversityarXiv2024SCAautoregressive-LM2026-06-14
2410.03751Recent Advances in Speech Language Models: A SurveyChinese University of Hong KongarXiv2024SCA, TTSautoregressive-LM2026-06-14
2104.00355Speech Resynthesis from Discrete Disentangled Self-SupeFacebook AI ResearcharXiv2021TTS, VCGAN, VAE2026-06-14
2502.06490Recent Advances in Discrete Speech Tokens: A ReviewSJTU / MSRAarXiv2025TTS, VC, SCA, codecautoregressive-LM, transformer-enc-dec, GAN, VAE2026-06-14
2105.06337Grad-TTS: A Diffusion Probabilistic Model for Text-to-SarXiv2021TTSdiffusion, transformer-enc-dec2026-06-14
2412.15649SLAM-Omni: Timbre-Controllable Voice Interaction SystemSJTU / MicrosoftarXiv2024SCAautoregressive-LM, flow-matching2026-06-14
2502.05512IndexTTS: An Industrial-Level Controllable and EfficienbilibiliarXiv2025TTSautoregressive-LM, GAN2026-06-14
2502.07243Vevo: Controllable Zero-Shot Voice Imitation with Self-Meta AIICLR2025TTS, VCautoregressive-LM, flow-matching, hybrid2026-06-14
2406.07855VALL-E R: Robust and Efficient Zero-Shot Text-to-SpeechMicrosoftarXiv2024TTSautoregressive-LM2026-06-14
2410.17799OmniFlatten: An End-to-end GPT Model for Seamless VoiceAlibaba (Tongyi Lab)arXiv2024SCAautoregressive-LM2026-06-14
2504.08528On The Landscape of Spoken Language Models: A ComprehenarXiv2025SCAautoregressive-LM, transformer-enc-dec2026-06-14
2206.08317Paraformer: Fast and Accurate Parallel Transformer forAlibaba GrouparXiv2022transformer-enc-dec2026-06-14
2412.19437DeepSeek-V3 Technical ReportDeepSeek-AIarXiv20242026-06-14
2402.03300DeepSeekMath: Pushing the Limits of Mathematical ReasonarXiv2024autoregressive-LM2026-06-14
2310.13289SALMONN: Towards Generic Hearing Abilities for Large LaTsinghua University / ByteDancearXiv20232026-06-14
1810.04805BERT: Pre-training of Deep Bidirectional Transformers farXiv20182026-06-14
2502.05139Meta Audiobox Aesthetics: Unified Automatic Quality AssMeta (FAIR)arXiv20252026-06-14
2307.09288Llama 2: Open Foundation and Fine-Tuned Chat ModelsMetaarXiv2023autoregressive-LM2026-06-14
2312.11805Gemini: A Family of Highly Capable Multimodal ModelsarXiv20232026-06-14
2005.14165Language Models are Few-Shot LearnersarXiv2020autoregressive-LM2026-06-14
2407.10671Qwen2 Technical ReportAlibaba GrouparXiv20242026-06-14
2106.04624SpeechBrain: A General-Purpose Speech ToolkitarXiv20212026-06-14
2406.14294DASB - Discrete Audio and Speech BenchmarkarXiv20242026-06-14
2301.12503AudioLDM: Text-to-Audio Generation with Latent DiffusioICML2023diffusion, VAE2026-06-15
2301.11325MusicLM: Generating Music From TextGooglearXiv2023SCAautoregressive-LM2026-06-15
2305.15255Spoken Question Answering and Speech Continuation UsingGoogle ResearcharXiv2023SCAautoregressive-LM2026-06-15
2312.01479OpenVoice: Versatile Instant Voice CloningMIT & MyShell.aiarXiv2023TTS, VCGAN, VAE2026-06-15
2312.15821Audiobox: Unified Audio Generation with Natural LanguagMeta FAIRarXiv2023TTSflow-matching2026-06-15
2401.07333ELLA-V: Stable Neural Codec Language Modeling with AligarXiv2024TTSautoregressive-LM2026-06-15
2402.13236Towards audio language modeling — an overviewarXiv2024TTS, SCA, codec2026-06-15
2404.03204RALL-E: Robust Codec Language Modeling with Chain-of-ThMicrosoftarXiv2024TTSautoregressive-LM2026-06-15
2406.00654Enhancing Zero-shot Text-to-Speech Synthesis with HumanNanyang Technological UniversityarXiv2024TTSautoregressive-LM2026-06-15
2406.05551Autoregressive Diffusion Transformer for Text-to-SpeechCUHK ShenzhenarXiv2024TTShybrid2026-06-15
2408.02622Language Model Can Listen While SpeakingShanghai Jiao Tong University / ByteDancearXiv2024SCA, TTSautoregressive-LM2026-06-15
2411.09943Zero-shot Voice Conversion with Diffusion TransformersNanyang Technological UniversityarXiv2024VCdiffusion, transformer-enc-dec2026-06-16
2411.17607Scaling Speech-Text Pre-training with Synthetic InterleTsinghua University / Zhipu.AIarXiv2024SCAautoregressive-LM, flow-matching2026-06-16
2411.18803TS3-Codec: Transformer-Based Simple Streaming Single CoarXiv2024GAN2026-06-16
2412.04724StableVC: Style Controllable Zero-Shot Voice ConversionNorthwestern Polytechnical University / XimalayaarXiv2024VCflow-matching2026-06-16
2506.13053ZipVoice: Fast and High-Quality Zero-Shot Text-to-SpeecXiaomiarXiv2025TTSflow-matching2026-06-16
2505.02625LLaMA-Omni2: LLM-based Real-time Spoken Chatbot with AuICT/CASarXiv2025SCAautoregressive-LM, flow-matching, hybrid2026-06-16
2502.18924MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion TZhejiang University, ByteDancearXiv2025TTSdiffusion, VAE2026-06-16
2507.23159Full-Duplex-Bench v1.5: Evaluating Overlap Handling forarXiv2025SCA, evaluation2026-06-16
2506.16381InstructTTSEval: Benchmarking Complex Natural-LanguageFudan UniversityarXiv2025TTS, evaluation2026-06-16
2506.10274Discrete Audio Tokens: More Than a Survey!arXiv2025codec, TTS, evaluation2026-06-16
2505.13000DualCodec: A Low-Frame-Rate, Semantically-Enhanced NeurCUHK-SZ / BaiduarXiv2025codec, TTSautoregressive-LM2026-06-16
2503.14345MoonCast: High-Quality Zero-Shot Podcast GenerationarXiv2025TTSautoregressive-LM, flow-matching2026-06-16
2508.04195NVSpeech: An Integrated and Scalable Pipeline for HumanarXiv2025autoregressive-LM, flow-matching2026-06-16
2505.09558WavReward: Spoken Dialogue Models With Generalist RewarZhejiang University / Alibaba GrouparXiv2025SCA, evaluationautoregressive-LM2026-06-16
2504.10344ALMTokenizer: A Low-bitrate and Semantic-rich Audio CodarXiv2025codec, TTS, SCAautoregressive-LM, VAE2026-06-16
2504.02407F5R-TTS: Improving Flow-Matching based Text-to-Speech wTencentarXiv2025TTSflow-matching2026-06-16
2511.15848Step-Audio-R1 Technical ReportStepFunarXiv2025SCAautoregressive-LM2026-06-16
2510.07838Full-Duplex-Bench-v2: A Multi-Turn Evaluation FrameworkarXiv2025SCA, evaluation2026-06-16
2505.14648Vox-Profile: A Speech Foundation Model Benchmark for CharXiv2025evaluation2026-06-16
1510.08484MUSAN: A Music, Speech, and Noise CorpusarXiv20152026-06-17
1908.06248JVS corpus: free Japanese multi-speaker voice corpusarXiv2019TTS, VC2026-06-17
1607.06450Layer NormalizationarXiv20162026-06-17
1808.10583AISHELL-2: Transforming Mandarin ASR Research Into InduarXiv20182026-06-17
2002.05202GLU Variants Improve TransformerarXiv20202026-06-17
2302.00482Improving and generalizing flow-based generative modelsarXiv20232026-06-17
2007.10310CoVoST 2 and Massively Multilingual Speech-to-Text TranarXiv20202026-06-17
2001.08361Scaling Laws for Neural Language ModelsarXiv20202026-06-17
2308.10248Steering Language Models With Activation EngineeringarXiv20232026-06-17
2309.16609Qwen Technical ReportarXiv2023autoregressive-LM2026-06-17
2308.05725EXPRESSO: A Benchmark and Analysis of Discrete ExpressiarXiv20232026-06-17
2308.11596SeamlessM4T: Massively Multilingual & Multimodal MachinarXiv2023transformer-enc-dec2026-06-17
2408.05211VITA: Towards Open-Source Interactive Omni Multimodal LarXiv20242026-06-17
2408.01800MiniCPM-V: A GPT-4V Level MLLM on Your PhonearXiv20242026-06-17
2312.10997Retrieval-Augmented Generation for Large Language ModelarXiv20232026-06-17
2402.07729AIR-Bench: Benchmarking Large Audio-Language Models viaarXiv20242026-06-17
2501.07246Audio-CoT: Exploring Chain-of-Thought Reasoning in LargarXiv20252026-06-17
2412.08635Multimodal Latent Language Modeling with Next-Token DifarXiv2024autoregressive-LM, diffusion, VAE, hybrid2026-06-17
2410.19168MMAU: A Massive Multi-Task Audio Understanding and ReasarXiv20242026-06-17
2501.01957VITA-1.5: Towards GPT-4o Level Real-Time Vision and SpearXiv2025autoregressive-LM, transformer-enc-dec2026-06-17
2503.01743Phi-4-Mini Technical Report: Compact yet Powerful MultiarXiv20252026-06-17
2503.19786Gemma 3 Technical ReportGoogle DeepMindarXiv2025autoregressive-LM2026-06-17
2505.03739VITA-Audio: Fast Interleaved Cross-Modal Token GeneratiarXiv2025autoregressive-LM, hybrid2026-06-17
2501.15368Baichuan-Omni-1.5 Technical ReportBaichuan Inc.arXiv2025autoregressive-LM, flow-matching2026-06-17
2506.02863CapSpeech: Enabling Downstream Applications in Style-CaarXiv2025TTS, evaluationautoregressive-LM, flow-matching2026-06-17
2507.12705AudioJudge: Understanding What Works in Large Audio ModarXiv20252026-06-17
2506.07900MiniCPM4: Ultra-Efficient LLMs on End DevicesarXiv2025autoregressive-LM2026-06-17
2507.08128Audio Flamingo 3: Advancing Audio Intelligence with FularXiv2025autoregressive-LM2026-06-17
2510.14664SpeechLLM-as-Judges: Towards General and InterpretablearXiv20252026-06-17
2508.13992MMAU-Pro: A Challenging and Comprehensive Benchmark forarXiv20252026-06-17
2511.09690Omnilingual ASR: Open-Source Multilingual Speech RecognMeta (FAIR)arXiv2025transformer-enc-dec2026-06-17
2509.08753Streaming Sequence-to-Sequence Learning with Delayed StarXiv2025autoregressive-LM2026-06-17
2409.09098AccentBox: Towards High-Fidelity Zero-Shot Accent GenerUniversity of EdinburgharXiv2025TTStransformer-enc-dec2026-06-29
2025.coling-industry.29CarMem: Enhancing Long-Term Memory in LLM Voice AssistaBMW Group / Univ. Augsburg / TUMCOLING2025SCA2026-06-29
2025.chipsal-1.18Impacts of Vocoder Selection on Tacotron-based Nepali TCHiPSAL2025TTS, evaluationGAN, transformer-enc-dec2026-06-29
2025.coling-main.685VoxpopuliTTS: a large-scale multilingual TTS corpus forZhejiang UniversityCOLING2025TTS2026-06-29
2409.20007DeSTA2: Developing Instruction-Following Speech LanguagearXiv2025SCAtransformer-enc-dec2026-06-29
2025.computel-main.6Evaluating Indigenous language speech synthesis for educComputEL2025TTS, evaluation2026-06-29
2025.nodalida-1.32Estonian isolated-word text-to-speech synthesiserInstitute of the Estonian LanguageNoDaLiDa2025TTS2026-06-29
2025.naacl-srw.6Towards Codec-LM Co-design for Neural Codec Language MoCartesia AI / MIT / CMUNAACL2025TTS, codecautoregressive-LM, hybrid2026-06-29
2025.findings-naacl.298Gender Bias in Instruction-Guided Speech Synthesis ModeNAACL2025TTS, evaluation2026-06-29
2025.findings-naacl.471The Role of Prosody in Spoken Question AnsweringNAACL2025evaluation2026-06-29
2025.naacl-long.464ManaTTS Persian: a recipe for creating TTS datasets forSharif University of TechnologyNAACL2025TTS, evaluation2026-06-29
2025.naacl-long.619ProSE: Diffusion Priors for Speech EnhancementUniversity of MarylandNAACL2025TTSdiffusion, transformer-enc-dec2026-06-29
iclr-2025-tQ1PmLfPBLPeriodWave: Multi-Period Flow Matching for High-Fidelity Waveform GenICLR2025TTSflow-matching2026-06-29
iclr-2025-cuFzE8JlvbContinuous Autoregressive Modeling with Stochastic Monotonic AlignmenThe Hong Kong Polytechnic UniversityICLR2025TTSautoregressive-LM, VAE2026-06-29
iclr-2025-dGSOn7sdWgSyllableLM: Learning Coarse Semantic Units for Speech Language ModelsUniversity of Texas at AustinICLR2025SCAautoregressive-LM2026-06-29
iclr-2025-868masI331HALL-E: Hierarchical Neural Codec Language Model for Minute-Long ZeroICLR2025TTSautoregressive-LM2026-06-29
iclr-2025-hQvX9MBowCDiTTo-TTS: Diffusion Transformers for Scalable Text-to-KRAFTONICLR2025TTSdiffusion, transformer-enc-dec2026-06-30
iclr-2025-uxDFlPGRLXFlowDec: A flow-based full-band general audio codec witMetaICLR2025codecflow-matching, GAN2026-06-30
2025.findings-naacl.130DiVISe: Direct Visual-Input Speech Synthesis PreservingShanghai Jiao Tong UniversityNAACL2025TTStransformer-enc-dec, GAN2026-06-30
2025.findings-naacl.279BnTTS: Few-Shot Speaker Adaptation in Low-Resource SettHishab SingaporeNAACL2025TTSautoregressive-LM, GAN, hybrid2026-06-30
2025.findings-naacl.38Prompt-Guided Selective Masking Loss for Context-Aware POSTECHNAACL2025TTStransformer-enc-dec2026-06-30
2025.naacl-demo.12ESPnet-SpeechLM: An Open Speech Language Model ToolkitCarnegie Mellon UniversityNAACL2025TTS, SCAautoregressive-LM2026-06-30
2025.naacl-demo.21ESPnet-SDS: Unified Toolkit and Demo for Spoken DialoguCarnegie Mellon UniversityNAACL2025SCA2026-06-30
2025.naacl-long.484Behavior-SD: Behaviorally Aware Spoken Dialogue GeneratSeoul National UniversityNAACL2025SCAautoregressive-LM2026-06-30
2025.naacl-long.591Robust and Unbounded Length Generalization in AutoregreGoogle DeepMindNAACL2025TTStransformer-enc-dec2026-06-30
2025.naacl-short.65kNN Retrieval for Simple and Effective Zero-Shot Multi-NAACL2025TTShybrid, GAN2026-06-30
2025.naacl-short.69Developing multilingual speech synthesis system for OjiNAACL2025TTSflow-matching2026-06-30
2025.iwsds-1.11Paralinguistic Attitude Recognition for Spoken DialogueFairy Devices Inc.IWSDS2025SCA2026-06-30
2025.iwsds-1.27A Survey of Recent Advances on Turn-taking Modeling in Université Paris-Saclay, CEA, ListIWSDS2025SCA2026-07-01
2505.15772MIKU-PAL: An Automated and Standardized Multi-Modal MetFish Audio; Carnegie Mellon UniversityarXiv2025TTS, evaluation2026-07-01
2507.06235Super Kawaii Vocalics: Amplifying the “Cute” Factor in arXiv2025TTS2026-07-01
2506.23049AURA: Agent for Understanding, Reasoning, and AutomatedarXiv2025SCAautoregressive-LM2026-07-01
2025.acl-long.681SIFT-50M: A Large-Scale Multilingual Dataset for SpeechAmazon AGIACL2025SCAautoregressive-LM2026-07-01
2025.acl-long.790Rhythm Controllable and Efficient Zero-Shot Voice ConveZhejiang UniversityACL2025VCflow-matching2026-07-01
2025.acl-long.817SimulS2S-LLM: Unlocking Simultaneous Inference of SpeecACL2025SCA, TTSautoregressive-LM2026-07-01
2025.acl-long.87Takin-VC: Expressive Zero-Shot Voice Conversion via AdaACL2025VCflow-matching2026-07-01
2025.acl-long.937UniCodec: Unified Audio Codec with Single Domain-AdaptiACL2025codecVAE2026-07-01
2025.acl-long.997Align-SLM: Textless Spoken Language Models with ReinforACL2025SCAautoregressive-LM2026-07-01
2025.conll-1.9A Linguistically Motivated Analysis of Intonational PhrCoNLL2025TTS, evaluation2026-07-01
2025.findings-acl.101Chain-Talker: Chain Understanding and Rendering for EmpACL2025TTSautoregressive-LM, flow-matching2026-07-01
2025.findings-acl.115SLAM-Omni: Timbre-Controllable Voice Interaction SystemACL2025SCA, TTSautoregressive-LM2026-07-01
2025.findings-acl.1226PodAgent: A Comprehensive Framework for Podcast GeneratACL2025TTS, VChybrid2026-07-01
2025.findings-acl.470Does Your Voice Assistant Remember? Analyzing ConversatSeoul National UniversityACL2025SCA, evaluation2026-07-01
2025.findings-acl.534Unlocking Speech Instruction Data Potential with Query ACL2025SCA2026-07-01
2025.findings-acl.631Slamming: Training a Speech Language Model on One GPU iThe Hebrew University of JerusalemACL2025SCAautoregressive-LM2026-07-01
2025.findings-acl.687TCSinger 2: Customizable Multilingual Zero-shot SingingZhejiang UniversityACL2025singing, TTSflow-matching, VAE2026-07-01
2025.findings-acl.71Data-Centric Improvements for Enhancing Multi-Modal UndACL2025SCAautoregressive-LM2026-07-01
2025.findings-acl.75Leveraging Unit Language Guidance to Advance Speech ModACL2025TTS, SCAtransformer-enc-dec2026-07-01
2025.findings-ijcnlp.49Incorporating Dialogue State Tracking into Japanese FulNTT / Nagoya UniversityACL2025SCAautoregressive-LM2026-07-01
2025.iwslt-1.5SSR: Alignment-Aware Modality Connector for Speech LangIWSLT2025SCAautoregressive-LM2026-07-01
2025.unlp-1.11Context-Aware Lexical Stress Prediction and Phonemizatiworkshop2025TTStransformer-enc-dec2026-07-01
2412.18603Long-Form Speech Generation with Spoken Language ModelsICML2025SCAautoregressive-LM2026-07-01
2503.11026MAVFlow: Preserving Paralinguistic Elements with ConditKAISTarXiv2025VC, TTSflow-matching2026-07-01
2505.15670SALM-Duplex: Efficient and Direct Duplex Modeling for SNVIDIAarXiv2025SCAautoregressive-LM2026-07-01
2506.09874UmbraTTS: Adapting Text-to-Speech to Environmental ContarXiv2025TTSflow-matching2026-07-01
2506.18296JIS: A Speech Corpus of Japanese Idol Speakers with VarNTT CorporationInterspeech2025TTS, VC, evaluation2026-07-01
2507.02176Analyzing and Improving Speaker Similarity Assessment farXiv2025evaluation2026-07-01
2507.00808Multi-interaction TTS toward professional recording repNTTarXiv2025TTStransformer-enc-dec2026-07-01
2507.01611QHARMA-GAN: Quasi-Harmonic Neural Vocoder based on AutoarXiv2025TTSGAN2026-07-01
2507.02380JoyTTS: LLM-based Spoken Chatbot With Voice CloningJD Health International Inc.arXiv2025SCA, TTSautoregressive-LM, flow-matching2026-07-01
2507.03887Traceable TTS: Toward Watermark-Free TTS with Strong TrarXiv2025TTSflow-matching2026-07-01
2507.03912Prosody Labeling with Phoneme-BERT and Speech FoundatioCyberAgentarXiv2025TTS2026-07-01
2507.08012RepeaTTS: Towards Feature Discovery through Repeated FiarXiv2025TTStransformer-enc-dec2026-07-01
2507.04349TTS-CtrlNet: Time varying emotion aligned text-to-speecarXiv2025TTSflow-matching2026-07-01
2507.04598Multi-Step Prediction and Control of Hierarchical EmotiarXiv2025TTStransformer-enc-dec2026-07-01
2507.04817Fast-VGAN: Lightweight Voice Conversion with Explicit CarXiv2025VCGAN2026-07-01
2507.01348SpeechAccentLLM: A Unified Framework for Foreign AccentarXiv2025VC, TTSautoregressive-LM, VAE, GAN2026-07-01
2507.06116Speech Quality Assessment Model Based on Mixture of ExpZhejiang UniversityarXiv2025evaluation2026-07-01
2506.23325XY-Tokenizer: Mitigating the Semantic-Acoustic ConflictFudan UniversityarXiv2025codecGAN, hybrid2026-07-01
2507.07799SecureSpeech: Prompt-based Speaker and Content ProtectiarXiv2025TTSautoregressive-LM2026-07-01
2507.08319Active Learning for Text-to-Speech Synthesis with InforarXiv2025TTStransformer-enc-dec2026-07-01
2507.09070SemAlignVC: Enhancing zero-shot timbre conversion usingMeta / KTH Royal Institute of TechnologyarXiv2025VCautoregressive-LM, flow-matching2026-07-01
2507.09282ClaritySpeech: Dementia Obfuscation in SpeechImperial College LondonarXiv2025TTSautoregressive-LM, diffusion, VAE2026-07-02
2507.09310Voice Conversion for Lombard Speaking Style with ImplicAmazon Alexa / Imperial College LondonarXiv2025VC, TTSVAE2026-07-02
2507.10985Pronunciation Deviation Analysis Through Voice Cloning California State University Long BeacharXiv2025TTS2026-07-02
2507.12197Quantize More, Lose Less: Autoregressive Generation froarXiv2025TTS, singingautoregressive-LM, GAN2026-07-02
2507.14988DMOSpeech 2: Reinforcement Learning for Duration PredicColumbia UniversityarXiv2025TTSflow-matching2026-07-02
2507.15272A2TTS: TTS for Low Resource Indian LanguagesarXiv2025TTSdiffusion2026-07-02
2507.16875Technical report: Impact of Duration Prediction on SpeaarXiv2025TTSflow-matching2026-07-02
2507.21138TTS-1 Technical ReportInworld AIarXiv2025TTSautoregressive-LM2026-07-02
2507.18119GOAT-SLM: A Spoken Language Model with Paralinguistic aTeleAI, China TelecomarXiv2025SCAautoregressive-LM, flow-matching2026-07-02
2507.18897HH-Codec: High Compression High-fidelity Discrete NeuraarXiv2025codecGAN2026-07-02
2507.17527Seed LiveInterpret 2.0: End-to-end Simultaneous Speech-ByteDancearXiv2025TTS, VCautoregressive-LM2026-07-02
2507.20140Do Not Mimic My Voice: Speaker Identity Unlearning for arXiv2025TTSflow-matching2026-07-02
2507.20731Learning Neural Vocoder from Range-Null Space DecomposiarXiv2025TTSGAN2026-07-02
2025.ccl-1.77HFSD-V2C: Zero-Shot Visual Voice Cloning Via Hierarchicworkshop2025TTSdiffusion2026-07-02
2025.icnlsp-1.34Beyond Labeled Datasets: Advancing TTS with Direct Prefworkshop2025TTSautoregressive-LM2026-07-02
2025.sigdial-1.21Transition Relevance Point Detection for Spoken Dialoguworkshop2025SCAhybrid2026-07-02
2025.sigdial-1.27EmoNews: A Spoken Dialogue System for Expressive News Cworkshop2025SCA, TTStransformer-enc-dec2026-07-02
2025.sigdial-1.51rrSDS 2.0: Incremental, Modular, Distributed, Multimodaworkshop2025SCA2026-07-02
interspeech-2025-0166Frozen Large Language Models Can Perceive ParalinguistiInterspeech2025SCAtransformer-enc-dec2026-07-02
interspeech-2025-0305DAFMSVC: One-Shot Singing Voice Conversion with Dual AtInterspeech2025singing, VCflow-matching, transformer-enc-dec2026-07-02
interspeech-2025-0347PeriodCodec: A Pitch-Controllable Neural Audio Codec UsInterspeech2025codec, singingGAN, VAE2026-07-02
interspeech-2025-0355Probing the Robustness Properties of Neural Speech CodeInterspeech2025codec, evaluation2026-07-02
interspeech-2025-0383Voice Conversion for Likability Control via Automated RInterspeech2025VCtransformer-enc-dec2026-07-02
interspeech-2025-0433When Humans Growl and Birds Speak: High-Fidelity Voice Interspeech2025VCVAE2026-07-02
interspeech-2025-0438LinearVC: Linear Transformations of Self-Supervised FeaInterspeech2025VChybrid2026-07-02
interspeech-2025-0464Prosody-Adaptable Audio Codecs for Zero-Shot Voice ConvInterspeech2025VC, codecautoregressive-LM2026-07-02
interspeech-2025-0506EnCodecMAE: leveraging neural codecs for universal audiInterspeech2025codectransformer-enc-dec2026-07-02
interspeech-2025-0656EEG-based Voice Conversion : Hearing the Voice of Your Beijing University of Posts and TelecommunicationsInterspeech2025VChybrid2026-07-02
interspeech-2025-0706Contextual Paralinguistic Data Creation for Multi-ModalInterspeech2025SCA2026-07-02
interspeech-2025-0756A-SMiLE: Affective Sparse Mixture-of-Experts Adapter wiInterspeech2025SCAhybrid2026-07-02
interspeech-2025-0998Voice-ENHANCE: Speech Restoration using a Diffusion-basInterspeech2025VCdiffusion, GAN2026-07-02
interspeech-2025-1020Learning Optimal Prosody Embedding Codebook based on F0Interspeech2025TTS, evaluationVAE2026-07-03
interspeech-2025-1081Speaker Normalization and Content Restoration for Zero-Interspeech2025VCGAN2026-07-03
interspeech-2025-1084Efficient Streaming TTS Acoustic Model with Depthwise RInterspeech2025TTSautoregressive-LM2026-07-03
interspeech-2025-1098GST-BERT-TTS: Prosody Prediction Without Accentual LabeInterspeech2025TTStransformer-enc-dec2026-07-03
interspeech-2025-1106LSCodec: Low-Bitrate and Speaker-Decoupled Discrete SpeInterspeech2025codecVAE, GAN2026-07-03
interspeech-2025-1115MPE-TTS: Customized Emotion Zero-Shot Text-To-Speech UsInterspeech2025TTSdiffusion, autoregressive-LM2026-07-03
interspeech-2025-1192Voice Impression Control in Zero-Shot TTSInterspeech2025TTStransformer-enc-dec2026-07-03
interspeech-2025-1210DiffEmotionVC: A Dual-Granularity Disentangled DiffusioInterspeech2025VCdiffusion2026-07-03
interspeech-2025-1229E2E-BPVC: End-to-End Background-Preserving Voice ConverInterspeech2025VCflow-matching2026-07-03
interspeech-2025-1236Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality AlignmentInterspeech2025TTSflow-matching2026-07-03
interspeech-2025-1334MiSTR: Multi-Modal iEEG-to-Speech Synthesis with TransfInterspeech2025TTStransformer-enc-dec2026-07-03
interspeech-2025-1364VS-Singer: Vision-Guided Stereo Singing Voice SynthesisInterspeech2025singing, TTSdiffusion2026-07-03
interspeech-2025-1394DiEmo-TTS: Disentangled Emotion Representations via SelInterspeech2025TTStransformer-enc-dec2026-07-03
interspeech-2025-1397VibE-SVC: Vibrato Extraction with High-frequency F0 ConKorea UniversityInterspeech2025singing, VCdiffusion2026-07-03
interspeech-2025-1434REWIND: Speech Time Reversal for Enhancing Speaker ReprInterspeech2025VCdiffusion2026-07-03
interspeech-2025-1478LightL2S: Ultra-Low Complexity Lip-to-Speech Synthesis Interspeech2025TTShybrid2026-07-03
interspeech-2025-1494VisualSpeech: Enhancing Prosody Modeling in TTS Using VInterspeech2025TTStransformer-enc-dec2026-07-03
interspeech-2025-1531Simple and Effective Content Encoder for Singing Voice Interspeech2025singing, VCVAE, GAN2026-07-03
interspeech-2025-1536Fairness in Dysarthric Speech Synthesis: Understanding Interspeech2025TTS, evaluationflow-matching2026-07-03
interspeech-2025-1538StarVC: A Unified Auto-Regressive Framework for Joint TInterspeech2025VCautoregressive-LM2026-07-03
interspeech-2025-1550ArVoice: A Multi-Speaker Dataset for Arabic Speech SyntMohamed Bin Zayed University of Artificial IntelligenceInterspeech2025TTS, VC, evaluationtransformer-enc-dec, VAE, GAN2026-07-03
interspeech-2025-1625Mimic Blocker: Self-Supervised Adversarial Training forInterspeech2025VCGAN2026-07-03
interspeech-2025-1638EATS-Speech: Emotion-Adaptive Transformation and PrioriInterspeech2025TTShybrid2026-07-03
interspeech-2025-1639LombardTokenizer: Disentanglement and Control of Vocal GIPSA-lab, Univ. Grenoble AlpesInterspeech2025codec, VCGAN2026-07-03
interspeech-2025-1684SA-RAS: Speaker-Aware Style Retrieval Augmented GeneratInterspeech2025TTShybrid2026-07-03
interspeech-2025-1726Voice Reconstruction through Large-Scale TTS Models: CoInterspeech2025TTS, evaluationhybrid2026-07-03
interspeech-2025-1747FasterVoiceGrad: Faster One-step Diffusion-Based Voice NTT, Inc.Interspeech2025VCdiffusion, GAN2026-07-03
interspeech-2025-1763Vocoder-Projected Feature DiscriminatorNTTInterspeech2025VCGAN, diffusion2026-07-03
interspeech-2025-1776SpeechSEC: A Unified Multi-Task Framework for Speech SyInterspeech2025TTShybrid2026-07-04
interspeech-2025-1819Comparative Analysis of Fast and High-Fidelity Neural VInterspeech2025TTSGAN2026-07-04
interspeech-2025-1873Can AI Understand Mandarin Speech Prosody? A FrameworkInterspeech2025SCA, evaluation2026-07-04
interspeech-2025-1940Investigating Stochastic Methods for Prosody Modeling iInterspeech2025TTStransformer-enc-dec, flow-matching2026-07-04
interspeech-2025-2031Kinship in Speech: Leveraging Linguistic Relatedness foInterspeech2025TTStransformer-enc-dec2026-07-04
interspeech-2025-2032ExagTTS: An Approach Towards Controllable Word Stress IIIIT HyderabadInterspeech2025TTShybrid2026-07-04
interspeech-2025-2075Segmentation-Variant Codebooks for Preservation of ParaInterspeech2025codec2026-07-04
interspeech-2025-2151FaVC: A Validated, Transcribed, Parallel Farsi Speech DUniversity of TehranInterspeech2025VC, evaluationGAN2026-07-04
interspeech-2025-2159Generating Consistent Prosodic Patterns from Open-SourcInterspeech2025TTS, evaluationflow-matching2026-07-04
interspeech-2025-2189ProMode: A Speech Prosody Model Conditioned on AcousticInterspeech2025TTStransformer-enc-dec2026-07-04
interspeech-2025-2283Pairwise Evaluation of Accent Similarity in Speech SyntInterspeech2025TTS, evaluation2026-07-04
interspeech-2025-2328A Watermark for Auto-Regressive Speech Generation ModelUniversity of MarylandInterspeech2025TTS, evaluationautoregressive-LM2026-07-05
interspeech-2025-2536The Text-to-speech in the Wild (TITW) DatabaseInterspeech2025TTS, evaluation2026-07-05
interspeech-2025-2564Towards a Japanese Full-duplex Spoken Dialogue SystemNagoya UniversityInterspeech2025SCAautoregressive-LM2026-07-05
interspeech-2025-2573SawtArabi: A Benchmark Corpus for Arabic TTS. Standard, Dialectal and Code-SwitchingInterspeech2025TTS, evaluationflow-matching, GAN2026-07-05
interspeech-2025-2586Spotlight-TTS: Spotlighting the Style via Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-SpeechKorea UniversityInterspeech2025TTStransformer-enc-dec, GAN2026-07-05
interspeech-2025-2595Harnessing Text-to-Speech Voice Cloning Models for Improved Audiological Speech AssessmentUniversity of CambridgeInterspeech2025TTS, evaluation2026-07-05
interspeech-2025-2679Can We Reconstruct a Dysarthric Voice with the Large Speech Model Parler TTS?University of EdinburghInterspeech2025TTSautoregressive-LM2026-07-05
interspeech-2025-2684Discl-VC: Disentangled Discrete Tokens and In-Context Learning for Controllable Zero-Shot Voice ConversionInterspeech2025VCflow-matching2026-07-05
interspeech-2025-2726DS-Codec: Dual-Stage Training with Mirror-to-NonMirror Architecture Switching for Speech CodecInterspeech2025codecGAN, hybrid2026-07-05
interspeech-2025-2739AF-Vocoder: Artifact-Free Neural Vocoder with Global Artifact FilterByteDanceInterspeech2025TTSGAN2026-07-05
interspeech-2025-2787Towards Adaptable and Intelligible Speech Synthesis in Noisy EnvironmentsKTH Royal Institute of TechnologyInterspeech2025TTS, evaluationautoregressive-LM2026-07-05
interspeech-2025-2815From Static to Dynamic: Enhancing AAC with Generative Imagery and Zero-Shot TTSKTH Royal Institute of TechnologyInterspeech2025TTS2026-07-05
interspeech-2025-bokkahallisatish25_interspeechHear Me Out: Interactive evaluation and bias discovery platform for speech-to-speech conversational AIKTH Royal Institute of TechnologyInterspeech2025SCA, evaluation2026-07-05
2507.16835Evaluating Speech-to-Text x LLM x Text-to-Speech Combinations for AI Interview SystemsarXiv2025SCA, evaluation2026-07-05
2411.19770Noro: Noise-Robust One-shot Voice Conversion with Hidden Speaker Representation LearningarXiv2025VCdiffusion2026-07-05
2025.clicit-1.27Veras Audire Et Reddere Voces: A Corpus of Prosodically-Correct Latin Poetic Audio from Large-Language-Model TTSworkshop2025TTS, evaluationautoregressive-LM2026-07-05
2506.23367You Sound a Little Tense: L2 Tailored Clear TTS Using Durational Vowel PropertiesarXiv2025TTSflow-matching2026-07-12
2509.05359An Empirical Analysis of Discrete Unit Representations in Speech Language Modeling Pre-trainingarXiv2025SCAautoregressive-LM2026-07-12
2509.04093Open-Source Full-Duplex Conversational Datasets for Natural and Interactive Speech SynthesisarXiv2025TTS, SCAautoregressive-LM, flow-matching2026-07-12
2509.04667DarkStream: real-time speech anonymization with low latencyTexas A&M UniversityarXiv2025VCGAN, hybrid2026-07-12
2509.04685Say More with Less: Variable-Frame-Rate Speech Tokenization via Adaptive Clustering and Implicit Duration CodingarXiv2025codecGAN2026-07-12
2509.04702OleSpeech-IV: A Large-Scale Multispeaker and Multilingual Conversational Speech Dataset with Diverse TopicsOlewavearXiv2025TTS, SCA2026-07-12
2509.05863LatinX: Aligning a Multilingual TTS Model with Direct Preference OptimizationarXiv2025TTSautoregressive-LM2026-07-12
2509.06074Multimodal Fine-grained Context Interaction Graph Modeling for Conversational Speech SynthesisEMNLP2025TTStransformer-enc-dec2026-07-12
2509.06502FireRedChat: A Pluggable, Full-Duplex Voice Interaction System with Cascaded and Semi-Cascaded ImplementationsXiaohongshuarXiv2025SCAhybrid2026-07-12
2509.07038Controllable Singing Voice Synthesis using Phoneme-Level Energy SequenceKorea UniversityarXiv2025singingdiffusion2026-07-12
2509.07376Progressive Facial Granularity Aggregation with Bilateral Attribute-based Enhancement for Face-to-Speech SynthesisPOSTECHEMNLP2025TTSVAE2026-07-12
2509.09716VStyle: A Benchmark for Voice Style Adaptation with Spoken InstructionsarXiv2025TTS, evaluation2026-07-12
2509.08379Flow-Matching ModelsarXiv2025VCdiffusion, flow-matching2026-07-12
2509.08696Accelerating Diffusion Transformer-Based Text-to-Speech with Transformer Layer CachingNational University of SingaporearXiv2025TTSflow-matching2026-07-12
2506.04077A Novel Data Augmentation Approach for Automatic Speaking Assessment on Opinion ExpressionsNational Taiwan Normal UniversityarXiv2025TTSautoregressive-LM2026-07-12
2509.09174EchoX: Towards Mitigating Acoustic-Semantic Gap via Echo Training for Speech-to-Speech LLMsThe Chinese University of Hong Kong, ShenzhenarXiv2025SCAautoregressive-LM2026-07-12
2509.09201DeCodec: Rethinking Audio Codecs as Universal Disentangled Representation LearnersChina MobilearXiv2025codecGAN2026-07-13
2509.09550Finite Scalar Quantization Enables Redundant and Transmission-Robust Neural Audio Compression at Low Bit-ratesNeuphonicarXiv2025codechybrid, GAN2026-07-13
2509.09748DiTReducio: A Training-Free Acceleration for DiT-Based TTS via Progressive CalibrationarXiv2025TTSflow-matching2026-07-13
2509.11084Length-Aware Rotary Position Embedding for Text-Speech AlignmentSupertone, Inc.arXiv2025TTSflow-matching2026-07-13
2509.11425FuseCodec: Semantic-Contextual Fusion and Supervision for Neural CodecsarXiv2025codec, TTSGAN, autoregressive-LM2026-07-13
2508.18240MTalk-Bench: Evaluating Speech-to-Speech Models in Multi-Turn Dialogues via Arena-style and Rubrics ProtocolsThe Chinese University of Hong Kong, ShenzhenarXiv2025SCA, evaluation2026-07-13
2509.12171Preservation of Language Understanding Capabilities in Speech-aware Large Language ModelsarXiv2025SCA, evaluationflow-matching2026-07-13
2509.14270SpeechWeave: Diverse Multilingual Synthetic Text & Audio Data Generation Pipeline for Training Text to Speech ModelsOracle AIACL2025TTS2026-07-13
2509.12831A Lightweight Pipeline for Noisy Speech Voice Cloning and Accurate Lip Sync SynthesisInternational Islamic University, IslamabadarXiv2025TTShybrid, GAN2026-07-13
2509.13068MSR-Codec: A Low-Bitrate Multi-Stream Residual Codec for High-Fidelity Speech Generation with Information DisentanglementLIGHTSPEEDarXiv2025TTS, codec, VCautoregressive-LM, VAE2026-07-13
2412.16846KALL-E: Autoregressive Speech Synthesis with Next-Distribution PredictionNorthwestern Polytechnical UniversityarXiv2025TTSautoregressive-LM, VAE2026-07-13
2504.20581ClonEval: An Open Voice Cloning BenchmarkAdam Mickiewicz UniversityarXiv2025TTS, evaluation2026-07-13
interspeech-2025-cho25c_interspeechUnleashing the Inner Monster: Demonstrating High-Fidelity Human to Non-Human Voice ConversionNC AI Co., LtdInterspeech2025VChybrid2026-07-13
interspeech-2025-gourav25_interspeechCode Mix TTS: An Approach to Infer Human Like Speech for Multi-Lingual Input TextsOracle CorporationInterspeech2025TTSdiffusion, GAN2026-07-13
interspeech-2025-raju25_interspeechEnd-to-End Indian Language Dubbing with Zero-Shot Speaker PreservationHitloopInterspeech2025TTSflow-matching2026-07-13
2509.13667A Distilled Low-Latency Neural Vocoder with Explicit Amplitude and Phase PredictionarXiv2025TTSGAN2026-07-13
2509.13670A High-Quality and Low-Complexity Streamable Neural Speech Codec with Knowledge DistillationUniversity of Science and Technology of ChinaarXiv2025codecGAN2026-07-13
2509.13989Do You Hear What I Mean? Quantifying the Instruction-Perception Gap in Instruction-Guided Expressive Text-To-Speech SystemsarXiv2025TTS, evaluation2026-07-14
2509.14579Cross-Lingual F5-TTS: Towards Language-Agnostic Voice Cloning and Speech SynthesisShanghai Jiao Tong UniversityarXiv2025TTS, VCflow-matching2026-07-14
2509.14684DAIEN-TTS: Disentangled Audio Infilling for Environment-Aware Text-to-Speech SynthesisarXiv2025TTSflow-matching2026-07-14
2509.14784MELA-TTS: Joint transformer-diffusion model with representation alignment for speech synthesisarXiv2025TTShybrid2026-07-14
2509.14946SynParaSpeech: Automated Synthesis of Paralinguistic Datasets for Speech Generation and UnderstandingarXiv2025TTSautoregressive-LM, flow-matching2026-07-14
2509.15085Real-Time Streaming Mel Vocoding with Generative Flow MatchingUniversity of HamburgarXiv2025TTSflow-matching2026-07-14
2509.15253Emotion-Aware Speech Generation with Character-Specific Voices for ComicsQueen Mary University of LondonarXiv2025TTS2026-07-14
2509.15462A Novel Semantic Compression Approach for Ultra-low Bandwidth Voice CommunicationSystems & Technology ResearcharXiv2025codec, VCautoregressive-LM, flow-matching2026-07-14
2505.17093P2VA: Converting Persona Descriptions into Voice Attributes for Fair and Controllable Text-to-SpeecharXiv2025TTS2026-07-14
2509.15626LibriTTS-VI: A Public Corpus and Novel Methods for Efficient Voice Impression ControlSony Group CorporationarXiv2025TTSVAE2026-07-14
2509.15629The Singing Voice Conversion Challenge 2025: From Singer Identity Conversion To Singing Style ConversionarXiv2025VC, singing, evaluation2026-07-14
2509.15845Deep Dubbing: End-to-End Auto-Audiobook System with Text-to-Timbre and Context-Aware Instruct-TTSarXiv2025TTSflow-matching, autoregressive-LM2026-07-14
2509.16010Fed-PISA: Federated Voice Cloning via Personalized Identity-Style AdaptationarXiv2025TTS, VChybrid2026-07-14
2509.16195FocalCodec-Stream: Streaming Low-Bitrate Speech Coding via Causal DistillationarXiv2025codec, VChybrid2026-07-14
2509.16589Benchmarking Contextual and Paralinguistic Reasoning in Speech-LLMs: A Case Study with In-the-Wild DataEMNLP2025SCA, evaluation2026-07-14
2509.20378Beyond Global Emotion: Fine-Grained Emotional Speech Synthesis with Dynamic Word-Level ModulationHarbin Institute of TechnologyarXiv2025TTSautoregressive-LM2026-07-14
2509.17006MBCodec: Thorough Disentangle for High-Fidelity Audio CompressionarXiv2025codecGAN2026-07-14
2509.17021Bridging the gap between training and inference in LM-based TTS modelsarXiv2025TTSautoregressive-LM2026-07-14
2509.17143MaskVCT: Masked Voice Codec Transformer for Zero-Shot Voice Conversion With Increased Controllability via Multiple GuidancesarXiv2025VCautoregressive-LM2026-07-14
2509.14882Llama-Mimi: Exploring the Limits of Flattened Speech Language ModelingarXiv2025SCAautoregressive-LM2026-07-14
2509.17516Audiobook-CC: Controllable Long-context Speech Generation for Multicast AudiobookXimalaya Inc.arXiv2025TTSautoregressive-LM, flow-matching, GAN2026-07-15
2509.17765Qwen3-Omni Technical ReportQwen Team (Alibaba)arXiv2025SCA, TTSautoregressive-LM, hybrid2026-07-15
2509.17988Nord-Parl-TTS: Finnish and Swedish TTS Dataset from Parliament SpeecharXiv2025TTSflow-matching2026-07-15
2509.18060TMD-TTS: A Unified Tibetan Multi-Dialect Text-to-Speech Framework for Ü-Tsang, Amdo and Kham Speech Dataset GenerationUniversity of Electronic Science and Technology of ChinaarXiv2025TTSflow-matching2026-07-15
2509.18470Discrete-Time Diffusion-Like Models for Speech SynthesisUniversity of SheffieldarXiv2025TTSdiffusion, flow-matching2026-07-15
2501.04561OpenOmni: Advancing Open-Source Omnimodal Large Language Models with Progressive Multimodal Alignment and Real-Time Self-Aware Emotional Speech SynthesisShenzhen Institute of Advanced Technology, CAS; Alibaba (Tongyi Lab)arXiv2025SCA, TTSautoregressive-LM, hybrid2026-07-15
2509.18531No Verifiable Reward for Prosody: Toward Preference-Guided Prosody Learning in TTSarXiv2025TTSautoregressive-LM2026-07-15
2509.18806Rethinking the joint estimation of magnitude and phase for time-frequency domain neural vocodersInstitute of Acoustics, Chinese Academy of Sciences; Tencent AI LabarXiv2025TTSGAN2026-07-15
2509.18823Towards Evaluating Generative Audio: Insights from Neural Audio Codec Embedding DistancesDolbyarXiv2025evaluation, codecGAN2026-07-15
2509.18928Direct Preference Optimization for Speech Autoregressive Diffusion ModelsByteDance SeedarXiv2025TTSautoregressive-LM, diffusion, hybrid2026-07-15
2509.19025Enhancing Noise Robustness for Neural Speech Codecs through Resource-Efficient Progressive Quantization Perturbation SimulationUniversity of Science and Technology of ChinaarXiv2025codecGAN2026-07-15
2509.19186Improving Test-Time Performance of RVQ-based Neural CodecsSupertone Inc.arXiv2025codecGAN2026-07-15
2509.19231Finding My Voice: Generative Reconstruction of Disordered Speech for Automated Clinical EvaluationCarnegie Mellon UniversityarXiv2025TTS, VC, evaluationdiffusion, GAN2026-07-15
2509.19592Frame-Stacked Local Transformers For Efficient Multi-Codebook Speech GenerationNVIDIAarXiv2025TTSautoregressive-LM2026-07-15
2509.19812Efficient Speech Watermarking for Speech Synthesis via Progressive Knowledge DistillationMicrosoftarXiv2025TTSGAN2026-07-15
2509.19883CoMelSinger: Discrete Token-Based Zero-Shot Singing Synthesis With Structured Melody Control and GuidanceNational University of SingaporearXiv2025singinghybrid2026-07-15
2509.19928Measuring Prosody Diversity in Zero-Shot TTS: A New Metric, Benchmark, and ExplorationarXiv2025TTS, evaluation2026-07-16
2509.20086OLaPh: Optimal Language PhonemizerHof University of Applied SciencesarXiv2025TTSautoregressive-LM2026-07-16
2509.20321Conversational Speech Reveals Structural Robustness Failures in SpeechLLM BackbonesTexas A&M UniversityarXiv2025SCA, evaluationautoregressive-LM2026-07-16
2509.20410Phoenix-VAD: Streaming Semantic Endpoint Detection for Full-Duplex Speech InteractionXiamen University, DiDi GlobalarXiv2025SCAautoregressive-LM2026-07-16
2509.20485Objective Evaluation of Prosody and Intelligibility in Speech Synthesis via Conditional Prediction of Discrete TokensJohns Hopkins University; National University of SingaporearXiv2025evaluation, TTS2026-07-16
2509.22718PerformSinger: Multimodal Singing Voice Synthesis Leveraging Synchronized Lip Cues from Singing Performance VideosarXiv2025singinghybrid2026-07-16
2505.10599UDDETTS: Unifying Discrete and Dimensional Emotions for Controllable Emotional Text-to-SpeecharXiv2025TTSautoregressive-LM, flow-matching, GAN, hybrid2026-07-16
2509.20802SPADE: Structured Pruning and Adaptive Distillation for Efficient LLM-TTSKAIST, 42dot Inc.arXiv2025TTSautoregressive-LM2026-07-16
2509.22727DiaMoE-TTS: A Unified IPA-Based Dialect TTS Framework with Mixture-of-Experts and Parameter-Efficient Zero-Shot AdaptationTsinghua University, Giant Network AI LabarXiv2025TTSflow-matching2026-07-16
2506.21875WildSpeech-Bench: Benchmarking End-to-End SpeechLLMs in the WildWeChat AI, TencentarXiv2025SCA, evaluation2026-07-16
2509.21968AUV: Teaching Audio Universal Vector Quantization with Single Nested CodebookarXiv2025codecGAN, VAE2026-07-16
2509.22062Comprehend and Talk: Text to Speech Synthesis via Dual Language ModelingAMAP Speech, Tsinghua UniversityarXiv2025TTSautoregressive-LM2026-07-17
2509.22167Semantic-VAE: Semantic-Alignment Latent Representation Shanghai Jiao Tong UniversityarXiv2025TTSVAE2026-07-17
2509.22243FLEXI: Benchmarking Full-duplex Human-LLM Speech InteractionNortheastern UniversityarXiv2025SCA, evaluation2026-07-17
2509.23147BFA: Real-time Multilingual Text-to-speech Forced AlignmentBournemouth UniversityarXiv2025TTS2026-07-17
2510.02352Evaluating Bias in Spoken Dialogue LLMs for Real-World arXiv2025SCA2026-07-17
2509.23938Easy Turn: Integrating Acoustic and Linguistic ModalitiarXiv2025SCAhybrid2026-07-17
2509.24457Assessing speech quality metrics for evaluation of neurCisco SystemsarXiv2025codec, evaluation2026-07-17
2509.24570ISSE: An Instruction-Guided Speech Style Editing Dataset And BenchmarkarXiv2025TTS, VC, evaluationautoregressive-LM2026-07-17
2509.24650VoxCPM: Tokenizer-Free TTS for Context-Aware Speech Generation and True-to-Life Voice CloningarXiv2025TTS, VCautoregressive-LM, diffusion2026-07-17
2509.24773VSSFlow: Unifying Video-conditioned Sound and Speech GearXiv2025TTSflow-matching2026-07-17
2509.25131MGM-Omni: Scaling Omni LLMs to Personalized Long-HorizoCUHK, HKUST, SmartMorearXiv2025SCA, TTSautoregressive-LM, flow-matching, hybrid2026-07-17
2509.25416Emotion-Aligned Generation in Diffusion Text to Speech Models via Preference-Guided OptimizationarXiv2025TTSdiffusion2026-07-17
2509.26276Optimizing Speech Language Models for Acoustic ConsistencyUniversity of ZuricharXiv2025TTS, SCAautoregressive-LM2026-07-17
2509.26514BatonVoice: An Operationalist Framework for Enhancing CTencentarXiv2025TTSautoregressive-LM, flow-matching, GAN2026-07-17
2509.26542Voice Evaluation of Reasoning Ability: Diagnosing the Modality-Induced Performance GaparXiv2025SCA, evaluation2026-07-17
2510.00264Baseline Systems For The 2025 Low-Resource Audio Codec ChallengeCisco Systems (Collaboration AI)arXiv2025codec, evaluationGAN2026-07-17
2510.00499MOSS-Speech: Towards True Speech-to-Speech Models Without Text GuidanceShanghai Innovation Institute, Fudan University, MOSIarXiv2025SCAautoregressive-LM, flow-matching2026-07-17
2510.00743From Scores to Preferences: Redefining MOS Benchmarking for Speech Quality Reward ModelingFudan UniversityarXiv2025evaluationautoregressive-LM2026-07-17
2025.vlsp-1.15Twinkle-VC: A Robust and High-Quality Zero-Shot Voice Conversion System for the VLSP 2025 Shared TaskVLSP 20252025VCdiffusion2026-07-17
2025.vlsp-1.14ViettelRoar: Voice conversion approach for VLSP 2025ViettelAI, Viettel GroupVLSP 20252025VCflow-matching2026-07-17
2025.vlsp-1.13The 2025 VLSP Task on Vietnamese Voice Conversion: Overview and Preliminary ResultsHanoi University of Science and TechnologyVLSP 20252025VC, evaluation2026-07-17
2510.05150Chronological Thinking in Full-Duplex Spoken Dialogue Language ModelsNanyang Technological University, StepFun, MilaarXiv2025SCAautoregressive-LM2026-07-17
2510.02066Chain-of-Thought Reasoning in Streaming Full-Duplex End-to-End Spoken Dialogue SystemsCarnegie Mellon University, Sony Group CorporationarXiv2025SCAautoregressive-LM2026-07-17
2510.01722Emotional Text-To-Speech Based on Mutual-Information-Guided Emotion-Timbre DisentanglementThe University of Tokyo / Institute of Science TokyoarXiv2025TTStransformer-enc-dec2026-07-17
2510.01903MelTok: 2D Tokenization for Single-Codebook Audio CompressionInternational Digital Economy Academy (IDEA)arXiv2025codecVAE, GAN2026-07-17
2510.02044Stream RAG: Instant and Accurate Spoken Dialogue Systems with Streaming Tool UsageMeta, Carnegie Mellon UniversityarXiv2025SCAautoregressive-LM2026-07-17
2510.03111Evaluation of preprocessing pipelines in the creation of in-the-wild TTS datasetsUniversidad Nacional de Tres de FebreroarXiv2025TTS, evaluation2026-07-17
2510.03735Soft Disentanglement in Frequency Bands for Neural Audio CodecsTélécom ParisarXiv2025codecGAN2026-07-17
2510.04738Speak, Edit, Repeat: High-Fidelity Voice Editing and Zero-Shot TTS with Cross-Attentive MambaMTS AI, ITMO UniversityarXiv2025TTSautoregressive-LM, hybrid2026-07-17
2510.05619Teaching Machines to Speak Using Articulatory ControlUC BerkeleyarXiv2025TTShybrid2026-07-17
2510.05984ECTSpeech: Enhancing Efficient Speech Synthesis via Easy Consistency TuningXinjiang UniversityarXiv2025TTSdiffusion, transformer-enc-dec2026-07-17
2510.05799Data-efficient Targeted Token-level Preference Optimization for LLM-based Text-to-SpeechSpiralAI Inc.arXiv2025TTSautoregressive-LM2026-07-17
2506.15556PredGen: Accelerated Inference of Large Language Models through Input-Time Speculation for Real-Time Speech InteractionUniversity of California, Los AngelesarXiv2025SCA, TTSautoregressive-LM2026-07-18
2510.07096Modeling Sarcastic Speech: Semantic and Prosodic Cues in a Speech Synthesis FrameworkUniversity of GroningenarXiv2025TTSGAN, VAE2026-07-18
2510.06917SHANKS: Simultaneous Hearing and Thinking for Spoken Language ModelsNational Taiwan University, MicrosoftarXiv2025SCAautoregressive-LM2026-07-18
2510.06927Position: Towards Responsible Evaluation for Text-to-SpeecharXiv2025TTS, evaluation2026-07-18
2510.07881CS3-Bench: Evaluating and Enhancing Speech-to-Speech LLMs for Mandarin-English Code-SwitchingShanghai Jiao Tong University, Ant GrouparXiv2025SCA, evaluationautoregressive-LM2026-07-18
2510.08373DialoSpeech: Dual-Speaker Dialogue Generation with LLM and Flow MatchingNorthwestern Polytechnical UniversityarXiv2025TTSautoregressive-LM, flow-matching2026-07-18
2510.08392MeanVC: Lightweight and Streaming Zero-Shot Voice Conversion via Mean FlowsNorthwestern Polytechnical University (ASLP@NPU)arXiv2025VCflow-matching, hybrid2026-07-18
2510.07978VoiceAgentBench: Are Voice Assistants ready for agentic tasks?Ola Electric / KrutrimarXiv2025SCA, evaluationautoregressive-LM2026-07-18
2510.09061O_O-VC: Synthetic Data-Driven One-to-One Alignment for Any-to-Any Voice ConversionVNPT AI / Hanoi University of Science and Technology / National Economics UniversityEMNLP2025VCVAE, GAN2026-07-18
2510.09016DiTSinger: Scaling Singing Voice Synthesis with Diffusion Transformer and Implicit AlignmentMigu Music, China Mobile Communications CorporationarXiv2025singingdiffusion2026-07-18
2506.12311Phonikud: Hebrew Grapheme-to-Phoneme Conversion for Real-Time Text-to-SpeechIndependent Researcher; Reichman University; Tel Aviv UniversityarXiv2025TTSGAN, diffusion2026-07-18
2510.09424The Speech-LLM Takes It All: A Truly Fully End-to-End Spoken Dialogue State Tracking ApproachOrange ResearcharXiv2025SCAautoregressive-LM2026-07-18
2510.09592Mind-Paced Speaking: A Dual-Brain Approach to Real-Time Reasoning in Spoken Language ModelsStepFunarXiv2025SCAautoregressive-LM2026-07-18
2510.09245SynthVC: Leveraging Synthetic Data for End-to-End Low Latency Streaming Voice ConversionNorthwestern Polytechnical University (ASLP@NPU)arXiv2025VCGAN2026-07-18
2510.10003MTP-S2UT: Enhancing Speech-to-Speech Translation Quality with Multi-token PredictionNortheastern University; NiuTrans Research; Kunming University of Science and TechnologyarXiv2025TTStransformer-enc-dec2026-07-18
2510.10774ParsVoice: A Large-Scale Multi-Speaker Persian Speech Corpus for Text-to-Speech SynthesisUniversity of TehranarXiv2025TTSautoregressive-LM, GAN2026-07-18
2510.11646BridgeCode: A Dual Speech Representation Paradigm for Autoregressive Zero-Shot Text-to-Speech SynthesisSouth China University of TechnologyarXiv2025TTSautoregressive-LM2026-07-18
2510.11124Perturbation Self-Supervised Representations for Cross-Lingual Emotion TTS: Stage-Wise Modeling of Emotion and SpeakerTianjin UniversityarXiv2025TTStransformer-enc-dec, GAN2026-07-18
2510.12964VCTR: A Transformer-Based Model for Non-parallel Voice ConversionIndependent ResearcherarXiv2025VCGAN, hybrid2026-07-18
2510.12995Continuous-Token Diffusion for Speaker-Referenced TTS in Multimodal LLMsAmazon AGIarXiv2025TTSautoregressive-LM, diffusion2026-07-18
2510.13221Acoustic Teleportation via Disentangled Neural Audio Codec RepresentationsFraunhofer IIS / International Audio Laboratories ErlangenarXiv2025codecGAN2026-07-18
2510.13293Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS ModelsAlibaba, Nanyang Technological UniversityarXiv2025TTSautoregressive-LM2026-07-18
2510.13194StressTransfer: Stress-Aware Speech-to-Speech Translation with Emphasis PreservationThe Chinese University of Hong Kong; Nara Institute of Science and TechnologyarXiv2025TTSautoregressive-LM2026-07-18
2510.15364LDCodec: A high quality neural audio codec with low-complexity decoderByteDancearXiv2025codecGAN2026-07-18
2510.15227LongCat-Audio-Codec: An Audio Tokenizer and Detokenizer Solution Designed for Speech Large Language ModelsMeituan (LongCat Team)arXiv2025codecGAN, hybrid2026-07-18
2510.16841SAC: Neural Speech Codec with Semantic-Acoustic Dual-Stream QuantizationShanghai Jiao Tong University (X-LANCE Lab), Soul AI LabarXiv2025codecGAN, VAE, autoregressive-LM2026-07-18
2510.16718U-Codec: Ultra Low Frame-rate Neural Speech Codec for Fast High-fidelity Speech GenerationPeking University / Tencent AI Lab / Tencent HunyuanarXiv2025codec, TTSautoregressive-LM, GAN2026-07-18
2503.06211Late Fusion and Multi-Level Fission Amplify Cross-Modal Transfer in Text-Speech LMsUniversité de Toulon (LIS) / University of CambridgearXiv2025SCA, TTSautoregressive-LM2026-07-18
2510.18308ParaStyleTTS: Toward Efficient and Robust Paralinguistic Style Control for Expressive Text-to-Speech GenerationUniversity of New South WalesarXiv2025TTSVAE, GAN2026-07-18
2506.23670Efficient Interleaved Speech Modeling through Knowledge DistillationNlpie Research; University of ZuricharXiv2025TTS, SCAautoregressive-LM2026-07-18
2510.19509Which Evaluation for Which Model? A Taxonomy for Speech Model AssessmentApplearXiv2025evaluation2026-07-18
2510.10785FAC-FACodec: Controllable Zero-Shot Foreign Accent Conversion with Factorized Speech CodecUniversity of Illinois Urbana-ChampaignarXiv2025VCdiffusion2026-07-18
2510.20210Vox-Evaluator: Enhancing Stability and Fidelity for Zero-shot TTS with A Multi-Level EvaluatorTencent AI LabarXiv2025TTS, evaluationtransformer-enc-dec2026-07-26
2510.20513Decoding the Ear: A Framework for Objectifying Expressiveness from Human Preference Through Efficient AlignmentThe Chinese University of Hong Kong, Shenzhen / Li Auto Inc.arXiv2025SCA, evaluation2026-07-26
2510.20677R2-SVC: Towards Real-World Robust and Expressive Zero-shot Singing Voice ConversionarXiv2025singing, VCflow-matching2026-07-26
2510.21209SpecTokenizer: A Lightweight Streaming Codec in the Compressed Spectrum DomainInterspeech2025codecGAN, VAE2026-07-26
2510.21685StylePitcher: Generating Style-Following and Expressive Pitch Curves for Versatile Singing TasksarXiv2025singing, VCflow-matching2026-07-26
2510.22241FOA Tokenizer: Low-Bitrate Neural Codec for First Order Ambisonics with Spatial Consistency LossarXiv2025codecGAN, VAE2026-07-26
2510.22588UltraVoice: Scaling Fine-Grained Style-Controlled Speech Conversations for Spoken Dialogue ModelsarXiv2025TTS, SCAautoregressive-LM, flow-matching2026-07-26
2511.05516Ming-UniAudio: Speech LLM for Joint Understanding, Generation and Editing with Unified RepresentationarXiv2025TTS, codecautoregressive-LM, VAE, flow-matching, hybrid2026-07-26
2506.21864DeepOmni: Towards Seamless and Smart Speech Interaction with Adaptive Modality-Specific MoEarXiv2025SCA, TTSautoregressive-LM2026-07-26
2510.23312Low-Resource Audio Codec (LRAC): 2025 Challenge DescriptionCisco Systems (Collaboration AI)arXiv2025codec, evaluation2026-07-26
2510.23541SoulX-Podcast: Towards Realistic Long-form Podcasts with Dialectal and Paralinguistic DiversityarXiv2025TTSautoregressive-LM, flow-matching, hybrid2026-07-26
2510.25178SFMS-ALR: Script-First Multilingual Speech Synthesis with Adaptive Locale ResolutionarXiv2025TTS2026-07-26
2510.24372Bayesian Speech Synthesizers Can Learn from Multiple TeachersarXiv2025TTSautoregressive-LM2026-07-26
2510.25566PitchFlower: A flow-based neural audio codec with pitch controllabilityIRCAM, CNRS, Sorbonne UniversitearXiv2025codecflow-matching2026-07-26
2510.25577Lost in Phonation: Voice Quality Variation as an Evaluation Dimension for Speech Foundation ModelsarXiv2025evaluation2026-07-26
2510.26190SP-MCQA: Evaluating Intelligibility of TTS Beyond the Word LevelThe Chinese University of Hong Kong, ShenzhenarXiv2025TTS, evaluation2026-07-26
2511.00256NaturalVoices: A Large-Scale, Spontaneous and Emotional Podcast Dataset for Voice ConversionarXiv2025VC2026-07-27
2511.00850MULTI-Bench: A Multi-Turn Interactive Benchmark for Assessing Emotional Intelligence ability of Spoken Dialogue ModelsStepFunarXiv2025SCA, evaluationautoregressive-LM2026-07-27
2511.01056WhisperVC: Decoupled Cross-Domain Alignment and Speech Generation for Low-Resource Whisper-to-Normal ConversionDuke Kunshan UniversityarXiv2025VCVAE, flow-matching, GAN2026-07-27
2511.01261Speech-DRAME: A Framework for Human-Aligned Benchmarks in Speech Role-PlayCarnegie Mellon University / AnuttaconarXiv2025SCA, evaluationautoregressive-LM2026-07-27
2511.02104Toward Objective and Interpretable Prosody Evaluation in Text-to-Speech: A Linguistically Motivated ApproacharXiv2025TTS, evaluation2026-07-27
2025.emnlp-main.1160C3: A Bilingual Benchmark for Spoken Dialogue Models Exploring Challenges in Complex ConversationsEMNLP2025SCA, evaluation2026-07-27
2025.emnlp-main.1447MULTIVOX: A Benchmark for Evaluating Voice Assistants for Multimodal InteractionsEMNLP2025SCA, evaluation2026-07-27
2025.emnlp-main.1492PACHAT: Persona-Aware Speech Assistant for Multi-party DialogueEMNLP2025SCAhybrid2026-07-27
2025.findings-emnlp.1077EZ-VC: Easy Zero-shot Any-to-Any Voice ConversionIndian Institute of Technology MadrasEMNLP2025VCflow-matching2026-07-27
2025.findings-emnlp.1381UniSpeaker: A Unified Approach for Multimodality-driven Speaker GenerationUniversity of Science and Technology of China / Alibaba GroupEMNLP2025TTS, VCautoregressive-LM, flow-matching, hybrid2026-07-27
2025.findings-emnlp.1394DM-Codec: Distilling Multimodal Representations for Speech TokenizationEMNLP2025codecGAN, VAE, autoregressive-LM2026-07-27
2025.findings-emnlp.241Enhancing Speech-to-Speech Dialogue Modeling with End-to-End Retrieval-Augmented GenerationShanghai Jiao Tong UniversityEMNLP2025SCAautoregressive-LM, flow-matching2026-07-27
2025.findings-emnlp.524Dub-S2ST: Textless Speech-to-Speech Translation for Seamless DubbingKorea Advanced Institute of Science and TechnologyEMNLP2025TTSdiffusion, flow-matching, hybrid2026-07-27
2025.findings-emnlp.933URO-Bench: Towards Comprehensive Evaluation for End-to-End Spoken Dialogue ModelsEMNLP2025SCA, evaluation2026-07-27
2511.03080PolyNorm: Few-Shot LLM-Based Text Normalization for Text-to-SpeechAppleEMNLP2025TTS, evaluation2026-07-28
2511.03601Step-Audio-EditX Technical ReportStepFunarXiv2025TTSautoregressive-LM, flow-matching, hybrid2026-07-28
2511.13732Principled Coarse-Grained Acceptance for Speculative Decoding in SpeecharXiv2025TTS, codecautoregressive-LM2026-07-28
2511.14779The Impact of Prosodic Segmentation on Speech Synthesis of Spontaneous SpeecharXiv2025TTStransformer-enc-dec2026-07-28
2025.arabicnlp-main.38DialG2P: Dialectal Grapheme-to-Phoneme. Arabic as a Case Studyworkshop2025TTStransformer-enc-dec2026-07-28
2511.06150BSCodec: A Band-Split Neural Codec for High-Quality Universal Audio ReconstructionarXiv2025codecGAN, VAE2026-07-28
2511.06246IDMap: A Pseudo-Speaker Generator Framework Based on Speaker Identity Index to Vector MappingarXiv2025VCdiffusion, hybrid2026-07-28
2508.20916SageLM: A Multi-aspect and Explainable Large Language Model for Speech JudgementNortheastern University, China / NiuTrans ResearcharXiv2025SCA, evaluationautoregressive-LM2026-07-28
2511.05143Synthesizing speech with selected perceptual voice qualities - A case study with creaky voiceInterspeech2025TTSVAE2026-07-29
2511.07116BridgeVoC: Revitalizing Neural Vocoder from a Restoration PerspectivearXiv2025TTSdiffusion2026-07-29
2511.07135Generating Novel and Realistic Speakers for Voice ConversionUniversity of RochesterarXiv2025VCVAE2026-07-29
2511.08496HQ-SVC: Towards High-Quality Zero-Shot Singing Voice Conversion in Low-Resource ScenariosarXiv2025singing, VCdiffusion, hybrid2026-07-29
2511.09995Time-Layer Adaptive Alignment for Speaker Similarity in Flow-Matching Based Zero-Shot TTSarXiv2025TTSflow-matching2026-07-29
2511.10112FabasedVC: Enhancing Voice Conversion with Text Modality Fusion and Phoneme-Level SSL FeaturesarXiv2025VCVAE, GAN2026-07-29
2511.10262MTR-DuplexBench: Towards a Comprehensive Evaluation of Multi-Round Conversations for Full-Duplex Speech Language ModelsarXiv2025SCA, evaluation2026-07-29
2511.10913Synthetic Voices, Real Threats: Evaluating Large Text-to-Speech Models in Generating Harmful AudioarXiv2025TTS, evaluation2026-07-29
2511.11104CLARITY: Contextual Linguistic Adaptation and Accent Retrieval for Dual-Bias Mitigation in Text-to-Speech GenerationarXiv2025TTS2026-07-29
2511.11124AV-Dialog: Spoken Dialogue Models with Audio-Visual InputarXiv2025SCAautoregressive-LM, hybrid2026-07-29
2511.12074MF-Speech: Achieving Fine-Grained and Compositional Control in Speech Generation via Factor DisentanglementUniversity of Science and Technology of ChinaarXiv2025VCGAN2026-07-29
2511.12690Improving Direct Persian-English Speech-to-Speech Translation with Discrete Units and Synthetic Parallel DataSharif University of TechnologyarXiv2025TTStransformer-enc-dec2026-07-29
2511.14249Towards Authentic Movie Dubbing with Retrieve-Augmented Director-Actor Interaction LearningInner Mongolia UniversityarXiv2025TTShybrid2026-07-29
2511.16639Codec2Vec: Self-Supervised Speech Representation Learning Using Neural Speech CodecsarXiv2025codec2026-07-29
2511.18487InstructAudio: Unified speech and music generation with natural language instructionarXiv2025TTSflow-matching2026-07-29
2512.05126SyncVoice: Towards Video Dubbing with Vision-Augmented Pretrained TTS ModelarXiv2025TTSflow-matching2026-07-29
2511.19734Evaluating Objective Speech Quality Metrics for Neural Audio CodecsETH ZuricharXiv2025codec, evaluation2026-07-29
2511.20974RosettaSpeech: Zero-Shot Speech-to-Speech Translation without Parallel SpeechUniversity of Texas at Austin / AmazonarXiv2025TTSautoregressive-LM2026-07-29
2511.21045CartoonSing: Unifying Human and Nonhuman Timbres in Singing GenerationCarnegie Mellon UniversityarXiv2025singing, TTS, VCtransformer-enc-dec, GAN2026-07-29
2511.21229Developing an Open Conversational Speech Corpus for the Isan LanguageSCB 10X (Typhoon Team)arXiv2025TTS2026-07-30
2511.21270Multi-Reward GRPO for Stable and Prosodic Single-Codebook TTS LLMs at ScaleTencentarXiv2025TTSautoregressive-LM2026-07-30
2511.22293GLA-Grad++: An Improved Griffin-Lim Guided Diffusion Model for Speech SynthesisarXiv2025TTSdiffusion2026-07-30
2512.00451STCTS: Generative Semantic Compression for Ultra-Low Bitrate Speech via Explicit Text-Prosody-Timbre DecompositionarXiv2025TTS, codecautoregressive-LM, GAN2026-07-30
2511.22687PURE Codec: Progressive Unfolding of Residual Entropy for Speech Codec LearningCarnegie Mellon University, Shanghai Jiao Tong UniversityarXiv2025codecVAE, GAN2026-07-30
2505.17320Comparative Evaluation of Expressive Japanese Character Text-to-Speech with VITS and Style-BERT-VITS2arXiv2025TTS, evaluationVAE, GAN2026-07-30
2512.00937Arabic TTS with FastPitch: Reproducible Baselines, Adversarial Training, and Oversmoothing AnalysisarXiv2025TTStransformer-enc-dec, GAN2026-07-30
2025.iwclul-1.3The world’s first South Sámi TTS – a revitalisation effort of an endangered language by reviving a legacy voiceworkshop2025TTStransformer-enc-dec, GAN2026-07-30
2512.01537Two-Dimensional Quantization for Geometry-Aware Audio CodingarXiv2025codecGAN, VAE2026-07-30
2512.01865Cross-Lingual Interleaving for Speech Language ModelsUniversity of CambridgearXiv2025SCAautoregressive-LM2026-07-30
2512.02523Generative Multi-modal Feedback for Singing Voice Synthesis EvaluationarXiv2025singing, evaluationautoregressive-LM2026-07-30
2512.03486A Universal Harmonic Discriminator for High-quality GAN-based VocoderAlibaba Digital Media & Entertainment GroupASRU2025TTS, singingGAN2026-07-30
2512.04552RRPO: Robust Reward Policy Optimization for LLM-based Emotional TTSAlibaba Group (Tongyi Lab)arXiv2025TTSautoregressive-LM, flow-matching2026-07-30
2512.04779YingMusic-Singer: Zero-shot Singing Voice Synthesis and Editing with Annotation-free Melody GuidanceGiantNetwork AI LabarXiv2025singingflow-matching2026-07-30
2512.04793YingMusic-SVC: Real-World Robust Zero-Shot Singing Voice Conversion with Flow-GRPO and Singing-Specific Inductive BiasesGiantNetwork AI LabarXiv2025singing, VCflow-matching2026-07-30
2512.06304Degrading Voice: A Comprehensive Overview of Robust Voice Conversion Through Input ManipulationarXiv2025VC2026-08-02
2512.07168JEPA as a Neural Tokenizer: Learning Robust Speech Representations with Density Adaptive AttentionarXiv2025codechybrid2026-08-02
2512.08006Beyond Unified Models: A Service-Oriented Approach to Low Latency, Context Aware Phonemization for Real Time TTSarXiv2025TTShybrid2026-08-02
2512.09504DMP-TTS: Disentangled multi-modal Prompting for Controllable Text-to-Speech with Chained GuidancearXiv2025TTSflow-matching2026-08-02
neurips-2025-1cURNMrieeStreamFlow: Streaming Audio Generation from Discrete Tokens via Streaming Flow MatchingNeurIPS2025TTS, codecflow-matching, GAN2026-08-02
neurips-2025-4iehXI36QGOpenOmni: Advancing Open-Source Omnimodal Large Language Models with Progressive Multimodal Alignment and Real-time Emotional Speech SynthesisNeurIPS2025TTS, SCAhybrid2026-08-02
neurips-2025-7Z3wQSu3mHFocalCodec: Low-Bitrate Speech Coding via Focal Modulation NetworksNeurIPS2025codec, VC, TTSVAE, GAN2026-08-02
neurips-2025-AsRB5nmlODSALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex ConversationNeurIPS2025TTS, SCAhybrid2026-08-02
neurips-2025-RTjr4DnS79Metis: A Foundation Speech Generation Model with Masked Generative Pre-trainingNeurIPS2025TTS, VCdiffusion2026-08-02
neurips-2025-SYcggdxX6WWord-Level Emotional Expression Control in Zero-Shot Text-to-Speech SynthesisNeurIPS2025TTSautoregressive-LM, flow-matching2026-08-02
neurips-2025-SoRe80Tg48Shallow Flow Matching for Coarse-to-Fine Text-to-Speech SynthesisNeurIPS2025TTSflow-matching2026-08-02
neurips-2025-pDWwz9F7ZhEfficient Speech Language Modeling via Energy Distance in Continuous Latent SpaceNeurIPS2025TTS, SCAautoregressive-LM2026-08-02
neurips-2025-vhPy3NMsO5OmniResponse: Online Multimodal Conversational Response Generation in Dyadic InteractionsNeurIPS2025TTS, SCAautoregressive-LM2026-08-02
2512.12129A comparative study of generative models for child voice conversionarXiv2025VCGAN, VAE, diffusion, flow-matching2026-08-02
2512.12297F5-TTS-RO: Extending F5-TTS to Romanian TTS via Lightweight Input AdaptationarXiv2025TTSflow-matching2026-08-02
2512.14653Robust Training of Singing Voice Synthesis Using Prior and Posterior UncertaintyASRU2025singingVAE, GAN2026-08-02
2512.14865Audio MultiChallenge: A Multi-Turn Evaluation of Spoken Dialogue Systems on Natural Human InteractionarXiv2025SCA, evaluation2026-08-02
2512.16519Pseudo-Cepstrum: Pitch Modification for Mel-Based Neural VocodersarXiv2025TTSGAN, flow-matching2026-08-02
2512.16832What Do Prosody and Text Convey? Characterizing How Meaningful Information is Distributed Across Multiple ChannelsarXiv2025evaluation2026-08-02
2512.17293Robust TTS Training via Self-Purifying Flow Matching for the WildSpoof 2026 TTS TrackarXiv2025TTSflow-matching2026-08-02
2512.17356Training Text-to-Speech Model with Purely Synthetic Data: Feasibility, Sensitivity, and Generalization CapabilityNCMMSC2025TTSflow-matching2026-08-02
2601.13910Synthetic Singers: A Review of Deep-Learning-based SingIJCNLP-AACL2025singing2026-08-02
2512.18699Task Vector in TTS: Toward Emotionally Expressive DialearXiv2025TTSflow-matching2026-08-02
2512.18706X-Talk: On the Underestimated Potential of Modular SpeearXiv2025SCA2026-08-02
2512.19090JoyVoice: Long-Context Conditioning for AnthropomorphicarXiv2025TTSautoregressive-LM, flow-matching, hybrid2026-08-02
2512.20156Fun-Audio-Chat Technical ReportarXiv2025SCA, TTSautoregressive-LM, flow-matching, GAN, hybrid2026-08-02
2512.20211Aliasing-Free Neural Audio SynthesisarXiv2025TTS, codecGAN2026-08-02
2512.20296TAVID: Text-Driven Audio-Visual Interactive Dialogue GearXiv2025TTS, SCAautoregressive-LM, flow-matching, hybrid2026-08-02
2512.20944SACodec: Asymmetric Quantization with Semantic AnchorinarXiv2025codecGAN, VAE2026-08-02
2512.21653Semantic Codebooks as Effective Priors for Neural SpeecarXiv2025codecGAN, VAE2026-08-02
2512.21706Enabling Conversational Behavior Reasoning CapabilitiesarXiv2025SCAautoregressive-LM2026-08-02
2512.22491ManchuTTS: Towards High-Quality Manchu Speech SynthesisarXiv2025TTSflow-matching2026-08-02
2512.23808MiMo-Audio: Audio Language Models are Few-Shot LearnersarXiv2025TTS, SCAautoregressive-LM, GAN, hybrid2026-08-02
2601.00217Mitigating Latent Mismatch in cVAE-Based Singing Voice Kwangwoon UniversityarXiv2026singingflow-matching, VAE, GAN2026-08-03
2601.00303DepFlow: Disentangled Speech Generation to Mitigate SemNanyang Technological UniversityarXiv2026TTSflow-matching2026-08-03
2601.01459OV-InstructTTS: Towards Open-Vocabulary Instruct Text-tInstitute of Automation, Chinese Academy of SciencesarXiv2026TTSautoregressive-LM2026-08-03
2601.01568MM-Sonate: Multimodal Controllable Audio-Video GeneratiKuaishou TechnologyarXiv2026TTS, VCflow-matching2026-08-03
2601.04233LEMAS: Large A 150K-Hour Large-scale Extensible MultiliInternational Digital Economy Academy (IDEA)arXiv2026TTSflow-matching, autoregressive-LM2026-08-03
2601.02073Towards Prosodically Informed Mizo TTS without ExplicitIndian Institute of Technology GuwahatiarXiv2026TTSVAE, GAN2026-08-03
2601.02753Vclip: Face-based Speaker Generation by Face-voice AssoOPPOarXiv2026TTSVAE, GAN2026-08-03
2601.02776UniSRCodec: Unified and Low-Bitrate Single Codebook CodTsinghua UniversityarXiv2026codecVAE, GAN2026-08-03
2601.03170TED-TTS: Training-Free Intra-Utterance Emotion and DuraNational University of SingaporearXiv2026TTSautoregressive-LM2026-08-03
2601.03632ReStyle-TTS: Relative and Continuous Style Control for Ant GrouparXiv2026TTSflow-matching2026-08-03
2601.03892Lightweight and perceptually-guided voice conversion foGraz University of TechnologyarXiv2026VCGAN2026-08-03
2601.05329CosyEdit: Unlocking End-to-End Speech Editing CapabilitNankai UniversityarXiv2026TTSautoregressive-LM, flow-matching2026-08-03
2601.05554SPAM: Style Prompt Adherence Metric for Prompt-based TTChung-Ang UniversityarXiv2026TTS, evaluation2026-08-03
2601.05564The ICASSP 2026 HumDial Challenge: Benchmarking Human-lNorthwestern Polytechnical University (ASLP@NPU)arXiv2026SCA, evaluation2026-08-03
2505.15727VocalBench: Benchmarking the Vocal Conversational AbiliShanghai Jiao Tong UniversityarXiv2026SCA, evaluation2026-08-03
2601.08450Decoding Order Matters in Autoregressive Speech SynthesUniversity of SheffieldarXiv2026TTSautoregressive-LM, diffusion, hybrid2026-08-03
2601.09239DSA-Tokenizer: Disentangled Semantic-Acoustic TokenizatCity University of Hong KongarXiv2026codec, TTSflow-matching2026-08-03
2602.06053PersonaPlex: Voice and Role Control for Full Duplex ConNVIDIAarXiv2026SCAautoregressive-LM2026-08-03
2601.10629VoiceSculptor: Your Voice, Designed By YouNorthwestern Polytechnical University (ASLP@NPU)arXiv2026TTSautoregressive-LM2026-08-03
2506.12537What Makes a Good Speech Tokenizer for LLM-Centric SpeearXiv2026TTS, codecautoregressive-LM2026-08-03
2601.11141FlashLabs Chroma 1.0: A Real-Time End-to-End Spoken DiaarXiv2026SCA, TTSautoregressive-LM2026-08-03
2601.16225ES4R: Speech Encoding Based on Prepositive Affective MoarXiv2026SCAautoregressive-LM, hybrid2026-08-03
2601.12205Do Neural Codecs Generalize? A Controlled Study Across arXiv2026codec, evaluationGAN, VAE2026-08-03
2601.12289ParaMETA: Towards Learning Disentangled Paralinguistic arXiv2026TTS2026-08-03
2601.12480A Unified Neural Codec Language Model for Selective Editable Text to Speech GenerationMicrosoftarXiv2026TTS, VCautoregressive-LM, hybrid2026-08-11
2409.16681Emotional Dimension Control in Language Model-Based Text-to-Speech: Spanning a Broad Spectrum of Human EmotionsAlibaba GrouparXiv2026TTSautoregressive-LM, flow-matching, hybrid2026-08-11
2601.12966Lombard Speech Synthesis for Any Voice with Controllable Style EmbeddingsKIT / CMUarXiv2026TTSflow-matching2026-08-11
2601.13055VoCodec: An Efficient Lightweight Low-Bitrate Speech CodecNanjing University / Horizon RoboticsarXiv2026codecGAN2026-08-11
2602.11172Synthesizing the Virtual Advocate: A Multi-Persona Speech Generation Framework for Diverse Linguistic Jurisdictions in Indic LanguagesIIT DelhiarXiv2026TTS, evaluation2026-08-11
2601.13758GOMPSNR: Reflourish the Signal-to-Noise Ratio Metric for Audio Generation TasksInstitute of Acoustics, Chinese Academy of SciencesarXiv2026evaluation, codecGAN2026-08-11
2601.13802Habibi: Laying the Open-Source Foundation of Unified-Dialectal Arabic Speech SynthesisSJTUarXiv2026TTS, evaluationflow-matching2026-08-11
2601.13835The Role of Prosodic and Lexical Cues in Turn-Taking with Self-Supervised Speech RepresentationsTrinity College DublinarXiv2026SCAhybrid2026-08-11
2601.13948Stream-Voice-Anon: Enhancing Utility of Real-Time Speaker Anonymization via Neural Audio Codec and Language ModelsNTU / A*STAR / PolyU (HK)arXiv2026VCautoregressive-LM, hybrid2026-08-11
2601.14472Prosody-Guided Harmonic Attention for Phase-Coherent Neural Vocoding in the Complex SpectrumBMEarXiv2026TTSGAN2026-08-11
2601.14960VCNAC: A Variable-Channel Neural Audio Codec for Mono, Stereo, and Surround SoundAmazon AGIarXiv2026codecGAN2026-08-11
2601.15596DeepASMR: LLM-Based Zero-Shot ASMR Speech Generation for Anyone of Any VoiceShanghai Jiao Tong University / VUI LabsarXiv2026TTSautoregressive-LM, flow-matching2026-08-11
2601.16023Timbre-Aware LLM-based Direct Speech-to-Speech TranslatIIT JammuarXiv2026TTSautoregressive-LM, flow-matching, GAN, hybrid2026-08-12
2601.16618PROST-LLM: Progressively Enhancing the Speech-to-SpeechThe Chinese University of Hong Kong / Huawei Artificial Intelligence Laboratory (Leibniz)arXiv2026TTSautoregressive-LM2026-08-12
2601.17086SonoEdit: Null-Space Constrained Knowledge Editing for TU Darmstadt / UMD / Smallest AIarXiv2026TTSautoregressive-LM2026-08-12
2601.13742Hearing Between the Lines: Unlocking the Reasoning PoweBoston University / Amazon AGIarXiv2026SCA, evaluation2026-08-12
2601.18094OneVoice: One Model, Triple Scenarios-Towards Unified ZChina Mobile (JIUTIAN Research)arXiv2026VC, singingautoregressive-LM, flow-matching, hybrid2026-08-12
2601.18281Reflecting Twice before Speaking with Empathy: Self-RefNankai University / Meituan LongCat Interaction TeamarXiv2026SCAautoregressive-LM2026-08-12
2601.18438UrgentMOS: Unified Multi-Metric and Preference LearningShanghai Jiao Tong University / Carnegie Mellon University / TU Braunschweig / Meta / Waseda University / VUI LabsarXiv2026evaluation2026-08-12
2601.18694Neural Multi-Speaker Voice Cloning for Nepali in Low-ReIOE, Thapathali Campus (Institute of Engineering, Nepal)arXiv2026TTStransformer-enc-dec2026-08-12
2601.17761AR-Omni: A Unified Autoregressive Model for Any-to-Any The Hong Kong Polytechnic University / University of Science and Technology of China / Harbin Institute of Technology (Shenzhen)arXiv2026TTS, SCAautoregressive-LM2026-08-12
2601.19952LTS-VoiceAgent: A Listen-Think-Speak Framework for EffiMeituan / University of Chinese Academy of SciencesarXiv2026SCA, TTSautoregressive-LM2026-08-12
2601.19063Optimizing Conversational Quality in Spoken Dialogue SyCarnegie Mellon University / Sony Group CorporationarXiv2026SCAautoregressive-LM2026-08-12
2601.19781Phonological Tokenizer: Prosody-Aware Phonetic Token viThe University of Tokyo / Sony Group Corporation / Carnegie Mellon UniversityarXiv2026codec, TTSGAN2026-08-12
2601.19786Rethinking Discrete Speech Representation Tokens for AcUniversity of Edinburgh (Centre for Speech Technology Research)arXiv2026TTS, VCGAN2026-08-12
2601.20094T-Mimi: A Transformer-based Mimi Decoder for Real-Time MetaarXiv2026TTS, codecGAN2026-08-12
2601.14417Quantifying Speaker Embedding Phonological Rule InteracUniversity of Southern CaliforniaarXiv2026TTS2026-08-12
2601.20481Erasing Your Voice Before It’s Heard: Training-free SpeEwha Womans UniversityarXiv2026TTSflow-matching2026-08-12
2601.20230Unit-Based Agent for Semi-Cascaded Full-Duplex DialogueHunan UniversityarXiv2026SCAautoregressive-LM2026-08-12
2601.21886Speech Quality-Based Localization of Low-Quality SpeechPaderborn UniversityarXiv2026TTS, evaluation2026-08-12
2601.22661Evaluating and Rewarding LALMs for Expressive Role-PlayarXiv2026TTS, evaluationautoregressive-LM2026-08-12
2601.22873EmoShift: Lightweight Activation Steering for Enhanced The Chinese University of Hong Kong, Shenzhen / The Hong Kong Polytechnic University / Tianjin UniversityarXiv2026TTSautoregressive-LM2026-08-12
2601.22889DiffuSpeech: Silent Thought, Spoken Answer via Unified arXiv2026SCA, TTSdiffusion2026-08-12
2601.23174Beyond Fixed Frames: Dynamic Character-Aligned Speech TarXiv2026codec, TTSGAN2026-08-12
2602.00269VoxServe: Streaming-Centric Serving System for Speech LarXiv2026SCA, TTS2026-08-12
2602.00443RVCBench: Benchmarking the Robustness of Voice Cloning arXiv2026VC, evaluation2026-08-12
2602.00594Kanade: A Simple Disentangled Tokenizer for Spoken Language ModelingarXiv2026codec, VC, TTS, SCAVAE, GAN, transformer-enc-dec2026-08-13
2602.02591VividVoice: A Unified Framework for Scene-Aware Visually-Driven Speech SynthesisTsinghua University / Ant GrouparXiv2026TTSdiffusion2026-08-13
2602.03420CoCoEmo: Composable and Controllable Human-Like Emotional TTS via Activation SteeringarXiv2026TTSautoregressive-LM, flow-matching, hybrid2026-08-13
2602.09041DSFlow: Dual Supervision and Step-Aware Architecture for One-Step Flow Matching Speech SynthesisarXiv2026TTSflow-matching2026-08-13
2602.04160PFluxTTS: Hybrid Flow-Matching TTS with Robust Cross-Lingual Voice Cloning and Inference-Time Model FusionRask AIarXiv2026TTSflow-matching, hybrid2026-08-13
2602.04683UniAudio 2.0: A Unified Audio Language Model with Text-Aligned Factorized Audio TokenizationarXiv2026TTS, VC, SCA, codec, singingautoregressive-LM, flow-matching, hybrid2026-08-13
2602.05207ARCHI-TTS: A Flow-Matching-Based TTS Model with Self-Supervised Semantic Aligner and Accelerated InferenceThe Chinese University of Hong KongarXiv2026TTSflow-matching2026-08-13
2602.05443Wave-Trainer-Fit: Neural Vocoder with Trainable Prior and Fixed-Point Iteration towards High-Quality Speech Generation from SSL featuresNara Institute of Science and Technology, LY CorporationarXiv2026TTSdiffusion, GAN, VAE2026-08-13
2602.05770Zero-Shot TTS With Enhanced Audio Prompts: BSc Submission For The 2026 Wildspoof Challenge TTS TrackBarcelona Supercomputing Center, Universitat de Barcelona, DFKI GmbHarXiv2026TTSflow-matching, diffusion, GAN2026-08-13
2602.06180STACodec: Semantic Token Assignment for Balancing Acoustic Fidelity and Semantic Information in Audio CodecsUniversity of California, Los AngelesarXiv2026codecGAN2026-08-13
2602.07803SoulX-Singer: Towards High-Quality Zero-Shot Singing Voice SynthesisSoul AI Lab / Geely Automobile Research Institute (AI Center) / Tianjin University / Northwestern Polytechnical University (ASLP@NPU)arXiv2026singing, VCflow-matching2026-08-13
2602.09823Covo-Audio Technical ReportTencent AI LabarXiv2026TTS, SCAautoregressive-LM, flow-matching, GAN, hybrid2026-08-13
2602.10164Emotion-Coherent Speech Data Augmentation and Self-Supervised Contrastive Style Training for Enhancing Kids’s Story Speech SynthesisLogistics and Supply Chain MultiTech R&D CentrearXiv2026TTStransformer-enc-dec2026-08-13
2506.04518Towards Efficient Speech-Text Jointly Decoding within One Speech Language ModelMicrosoftarXiv2026SCAautoregressive-LM, flow-matching2026-08-13
2602.10735Calliope: A TTS-based Narrated E-book Creator Ensuring Exact Synchronization, Privacy, and Layout FidelityOslo Metropolitan University, SimulaMetarXiv2026TTSautoregressive-LM, flow-matching2026-08-13
2602.10934MOSS-Audio-Tokenizer: Scaling Audio Tokenizers for Future Audio Foundation ModelsMOSI Intelligence, Shanghai Innovation Institute, Fudan UniversityarXiv2026codec, TTSautoregressive-LM, GAN, hybrid2026-08-13
2602.11072Simultaneous Speech-to-Speech Translation Without Aligned DataKyutaiarXiv2026TTSautoregressive-LM, hybrid2026-08-13
2602.11477SLD-L2S: Hierarchical Subspace Latent Diffusion for High-Fidelity Lip to Speech SynthesisInstitute of Acoustics, Chinese Academy of Sciences / Jiangnan University / Zhipu AIarXiv2026TTSflow-matching2026-08-13
2602.12135WavBench: Benchmarking Reasoning, Colloquialism, and Paralinguistics for End-to-End Spoken Dialogue ModelsXiamen University, Zhejiang University, CUHK-ShenzhenarXiv2026SCA, evaluation2026-08-13
2602.13891GSRM: Generative Speech Reward Model for Speech RLHFMeta Superintelligence LabsarXiv2026evaluation, SCAautoregressive-LM2026-08-13
2602.14664Probing Human Articulatory Constraints in End-to-End TTS with Reverse and Mismatched Speech-Text DirectionsTCS ResearcharXiv2026TTShybrid2026-08-13
2602.14686Disentangling Pitch and Creak for Speaker Identity Preservation in Speech SynthesisPaderborn University / Bielefeld UniversityarXiv2026TTSVAE, flow-matching2026-08-13
2602.15491The Equalizer: Introducing Shape-Gain Decomposition in Neural Audio CodecsInria, Univ. Grenoble Alpes, CNRS, LJK / Univ. Grenoble Alpes, Grenoble-INP, GIPSA-labarXiv2026codecGAN2026-08-13
2602.17157CC-G2PnP: Streaming Grapheme-to-Phoneme and Prosody with Conformer-CTC for Unsegmented LanguagesLY CorporationarXiv2026TTShybrid2026-08-13
2602.18104MeanVoiceFlow: One-step Nonparallel Voice Conversion with Mean FlowsNTT, Inc.arXiv2026VCflow-matching2026-08-14
2602.19574CTC-TTS: LLM-based dual-streaming text-to-speech with CTC alignmentXinjiang University, Tsinghua UniversityarXiv2026TTSautoregressive-LM2026-08-14
2602.23068TADA: A Generative Framework for Speech Modeling via Text-Acoustic Dual AlignmentHume AI, Dartmouth CollegearXiv2026TTS, codecautoregressive-LM, flow-matching, GAN, VAE, hybrid2026-08-14
2602.23266Discourse-Aware Dual-Track Streaming Response for Low-Latency Spoken Dialogue SystemsTongji University / The Hong Kong Polytechnic University / Shenzhen University of Advanced Technology / The Chinese University of Hong Kong, ShenzhenarXiv2026SCA, TTShybrid2026-08-14
2602.23765DashengTokenizer: One layer is enough for unified audio understanding and generationXiaomi Inc.arXiv2026codecGAN, hybrid2026-08-14
2603.01467Conversational Speech Naturalness PredictorMetaarXiv2026SCA, evaluation2026-08-14
2603.01476Entropy-Guided GRVQ for Ultra-Low Bitrate Neural Speech CodecWaseda University / NTT, Inc.arXiv2026codecGAN, VAE2026-08-14
2603.02022CodecFlow: Efficient Bandwidth Extension via Conditional Flow Matching in Neural Codec Latent SpaceSingapore Institute of Technology / Nanyang Technological University / National University of SingaporearXiv2026codecflow-matching, GAN2026-08-14
2603.04145VietNormalizer: An Open-Source, Dependency-Free Python Library for Vietnamese Text Normalization in TTS and NLP ApplicationsAustralian Catholic University / ICMS / FPT University / KETEMU / RMIT University Vietnam / NGHI Studio / Phuong Hai JSCarXiv2026TTS2026-08-14
2603.04219ZeSTA: Zero-Shot TTS Augmentation with Domain-ConditionMaum AI Inc., Humelo Inc.arXiv2026TTSVAE, GAN2026-08-15
2603.05299WavSLM: Single-Stream Speech Language Modeling via WavLM DistillationConcordia University / Mila-Quebec AI Institute / Université LavalarXiv2026SCAautoregressive-LM2026-08-15
2603.05373Hierarchical Decoding for Discrete Speech Synthesis with Multi-Resolution Spoof DetectionNational University of SingaporearXiv2026TTSautoregressive-LM2026-08-15
2603.05413Building Enterprise Realtime Voice Agents from Scratch:Salesforce AI ResearcharXiv2026SCA, TTS2026-08-15
2603.05887Reconstruct! Don’t Encode: Self-Supervised RepresentatiarXiv2026codecGAN, VAE2026-08-15
2603.05977Activation Steering for Accent-Neutralized Zero-Shot TeUniversity of Texas at DallasarXiv2026TTSautoregressive-LM2026-08-15
2603.06079StreamVoiceAnon+: Emotion-Preserving Streaming Speaker arXiv2026VCautoregressive-LM, hybrid2026-08-15
2603.06444Prosodic Boundary-Aware Streaming Generation for LLM-BaarXiv2026TTSautoregressive-LM, flow-matching, hybrid2026-08-15
2603.07513Bolbosh: Script-Aware Flow Matching for Kashmiri Text-tKAUST / University of Kashmir / Gaash Lab, NIT SrinagararXiv2026TTSflow-matching2026-08-15
2603.07534Accent Vector: Controllable Accent Manipulation for MulUniversity of Southern CaliforniaarXiv2026TTSautoregressive-LM2026-08-15
2603.07550Learning-free L2-Accented Speech Generation using PhonoUniversity of Southern CaliforniaarXiv2026TTS2026-08-15
2603.07551Targeted Speaker Poisoning Framework in Zero-Shot Text-to-SpeechUniversity of Southern CaliforniaarXiv2026TTSdiffusion, GAN, hybrid2026-08-15
2603.07599StyleBench: Evaluating Speech Language Models on ConverNortheastern University / NiuTrans ResearcharXiv2026SCA, evaluation2026-08-15
2603.08216DualTurn: Learning Turn-Taking from Dual-Channel GeneraarXiv2026SCAautoregressive-LM, hybrid2026-08-15
2603.08574Scalable Neural Vocoder from Range-Null Space DecomposiChinese Academy of Sciences, Tencent AI LabarXiv2026TTSGAN2026-08-15
2603.08977Universal Speech Content FactorizationJohns Hopkins UniversityarXiv2026VC, TTShybrid2026-08-15
2603.09120Emotion-Aware Prefix: Towards Explicit Emotion ControlCenter for Robust Speech Systems (CRSS), UT DallasarXiv2026VCautoregressive-LM, flow-matching, hybrid2026-08-15
2603.09180DuplexCascade: Full-Duplex Speech-to-Speech Dialogue wiSB Intuitions Corp. / The University of TokyoarXiv2026SCA, TTSautoregressive-LM, hybrid2026-08-15
2603.09627Speech-Omni-Lite: Portable Speech Interfaces for VisionarXiv2026TTS, SCAautoregressive-LM, flow-matching2026-08-15
2603.10371Speech Codec Probing from Semantic and Phonetic PerspectivesUniversity of Southern California / Dolby LaboratoriesarXiv2026codec, evaluation2026-08-15
2603.10904When Fine-Tuning Fails and when it Generalises: Role ofSprinklr AIarXiv2026TTSautoregressive-LM2026-08-15
2603.16924SimulU: Training-free Policy for Long-form SimultaneousMBZUAI / FBKarXiv2026TTStransformer-enc-dec, GAN2026-08-15
2603.11678RAF: Relativistic Adversarial Feedback For Universal SpKAISTarXiv2026TTSGAN2026-08-15
2603.11683Causal Prosody Mediation for Text-to-Speech: CounterfaarXiv2026TTStransformer-enc-dec2026-08-15
2603.11947Resurfacing Paralinguistic Awareness in Large Audio LanMonash University / University College LondonarXiv2026SCAautoregressive-LM2026-08-15
2603.12342MamTra: A Hybrid Mamba-Transformer Backbone for Speech KAIST, Chung-Ang UniversityarXiv2026TTShybrid, autoregressive-LM2026-08-15
2603.12565Speech-Worthy Alignment for Japanese SpeechLLMs via DirSB IntuitionsarXiv2026SCA, evaluationhybrid2026-08-19
2603.13518VoXtream2: Full-stream TTS with dynamic speaking rKTH Royal Institute of TechnologyarXiv2026TTSautoregressive-LM2026-08-19
2603.14032Beyond Two-stage Diffusion TTS: Joint Structure and ConarXiv2026TTSdiffusion, transformer-enc-dec2026-08-19
2603.14035Probing neural audio codecs for distinctions among EnglNorthwestern UniversityarXiv2026codec, evaluation2026-08-19
2603.14267DiFlowDubber: Discrete Flow Matching for Automated VideFPT Software AI CenterarXiv2026TTSflow-matching, hybrid2026-08-19
2603.14328CodecMOS-Accent: A MOS Benchmark of Resynthesized and TNagoya University, University of Edinburgh, NICTarXiv2026codec, evaluation2026-08-19
2603.14432Affectron: Emotional Speech Synthesis with Affective anKorea UniversityarXiv2026TTSautoregressive-LM2026-08-19
2603.14853WhispSynth: Scaling Multilingual Whisper Corpus throughNanjing University, Fudan University, ByteDancearXiv2026TTS, VCflow-matching, GAN, hybrid2026-08-19
2603.14877SoulX-Duplug: Plug-and-Play Streaming State Prediction Shanghai Jiao Tong University; Soul AI Lab; Northwestern Polytechnical UniversityarXiv2026SCA, evaluationautoregressive-LM2026-08-19
2603.14889SDiaReward: Modeling and Benchmarking Spoken Dialogue RarXiv2026SCA, evaluationautoregressive-LM2026-08-19
2603.15352NV-Bench: Benchmark of Nonverbal Vocalization SynthesisThe Chinese University of Hong Kong, ShenzhenarXiv2026TTS, evaluationautoregressive-LM, flow-matching, hybrid2026-08-19
2603.15981Aligning Paralinguistic Understanding and Generation inMeta Reality LabsarXiv2026SCAautoregressive-LM2026-08-19
2603.16280CAST-TTS: A Simple Cross-Attention Framework for UnifieShanghai AI Lab / Shanghai Jiao Tong UniversityarXiv2026TTSflow-matching2026-08-19
2603.16483On the Emotion Understanding of Synthesized SpeechNortheastern University / NiuTrans ResearcharXiv2026TTS, evaluation2026-08-19
2603.16783SpokenUS: A Spoken User Simulator for Task-Oriented DiaSeoul National University, Hanyang UniversityarXiv2026SCAautoregressive-LM, flow-matching2026-08-19
2603.17061Collecting Prosody in the Wild: A Content-Controlled, PUniversity of St. Gallen / LMU Munich / University of Mannheim / Charlotte Fresenius HochschulearXiv20262026-08-19
2604.08558WAND: Windowed Attention and Knowledge Distillation forKAIST, Sungkyunkwan UniversityarXiv2026TTSautoregressive-LM2026-08-19
2604.08562Neural networks for Text-to-Speech evaluationarXiv2026TTS, evaluationhybrid2026-08-19
2603.17231Neuron-Level Emotion Control in Speech-Generative LargeJohns Hopkins University / Imperial College LondonarXiv2026VCautoregressive-LM2026-08-19
2603.17837The Silent Thought: Modeling Internal Cognition in FullarXiv2026SCAautoregressive-LM, flow-matching2026-08-19
2603.18359Towards Interpretable Framework for Neural Audio CodecsarXiv2026codec2026-08-19
2603.19798Borderless Long Speech SynthesisXiaomi (MiLM Plus) / Nanjing UniversityarXiv2026TTSautoregressive-LM2026-08-19
2603.19831Gesture2Speech: How Far Can Hand Movements Shape ExpresSony Research IndiaarXiv2026TTSautoregressive-LM, GAN2026-08-19
2603.25750Sommelier: Scalable Open Multi-turn Audio Pre-processinKAIST AI, NAVER CloudarXiv2026SCAautoregressive-LM2026-08-19
2603.20638OmniCodec: Low Frame Rate Universal Audio Codec with SeNorthwestern Polytechnical UniversityarXiv2026codecGAN, VAE2026-08-22
2603.20743The Binding Effect: Analyzing How Multi-Dimensional CueNational Taiwan University, Inventec CorporationarXiv2026TTS, evaluation2026-08-22
2603.21078Assessing the Ability of Neural TTS Systems to Model CoUniversity at Buffalo / Australian National UniversityarXiv2026TTS, evaluation2026-08-22
2603.22252SelfTTS: cross-speaker style transfer through explicit UNICAMP, CPQDarXiv2026TTShybrid2026-08-22
2603.22267TiCo: Time-Controllable Spoken Dialogue ModelMIT, National Taiwan UniversityarXiv2026SCAautoregressive-LM, hybrid2026-08-22
2604.03279Rewriting TTS Inference Economics: Lightning V2 on TensSmallest AIarXiv2026TTSdiffusion2026-08-22
2603.23938OmniACBench: A Benchmark for Evaluating Context-GroundeHanyang University / Seoul National University / KAIST AI / NAVER CloudarXiv2026TTS, SCA, evaluation2026-08-22
2603.24116How Open is Open TTS? A Practical Evaluation of Open SoPOLITEHNICA Bucharest / Technical University of Cluj-NapocaarXiv2026TTS, evaluationtransformer-enc-dec, VAE, GAN, diffusion, flow-matching2026-08-22
2603.24144Semantic-Aware Interruption Detection in Spoken DialoguQwen Team, AlibabaarXiv2026SCA, evaluationhybrid2026-08-22
2603.24430Iterate to Differentiate: Enhancing Discriminability anNanjing University, MiLM Plus (Xiaomi), HKUSTarXiv2026TTS, evaluation2026-08-22
2603.24589YingMusic-Singer-Plus: Controllable Singing Voice SynthNorthwestern Polytechnical University (ASLP@NPU) / GiantNetwork AI LabarXiv2026singingflow-matching, VAE2026-08-22
2604.01247Combining Masked Language Modeling and Cross-Modal ContMTUCI, Moscow, RussiaarXiv2026TTSdiffusion2026-08-22

900 items under this folder.