Full catalog of all ingested papers. Navigate by topic via Concepts. Use search to find a specific paper by title or author.

IDTitleOrgVenueYearTaskArchitectureIngested
1904.02882LibriTTS: A Corpus Derived from LibriSpeech for Text-to-SGoogle AIarXiv2019TTS2026-06-10
2403.03100NaturalSpeech 3: Zero-Shot Speech Synthesis with FactoriMicrosoftarXiv2024TTS, VCdiffusion, hybrid2026-06-10
2504.18425Kimi-Audio Technical ReportMoonshot AIarXiv2025TTS, VC, SCAautoregressive-LM, flow-matching2026-06-10
2204.02152UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 202University of TokyoarXiv2022evaluation2026-06-10
2509.04072Computational Narrative Understanding for Expressive TearXiv2025TTSautoregressive-LM, flow-matching2026-06-04
2508.15827Mini-Omni-Reasoner: Token-Level Thinking-in-Speaking inarXiv2025SCAautoregressive-LM2026-06-03
interspeech-2025-0739FD-Bench: A Full-Duplex Benchmarking Pipeline Designed for Full Duplex Spoken Dialogue SystemsInterspeech2025SCA, evaluation2026-06-03
interspeech-2025-1993Defending Unauthorized Voice Cloning with Watermark-Aware CodecsThe Chinese University of Hong KongInterspeech2025TTS, VCautoregressive-LM2026-06-03
2508.08095Dual Information Speech Language Models for Emotional ConversationsMashang Consumer Finance Co., Ltd.arXiv2025SCAtransformer-enc-dec2026-06-03
interspeech-2025-0948PromptEVC: Controllable Emotional Voice Conversion with Natural Language PromptsSoutheast UniversityInterspeech2025VCVAE, diffusion2026-06-03
interspeech-2025-0203ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and SpeechKyushu University / University of Tokyo / EverestAI XimalayaInterspeech2025VCflow-matching, transformer-enc-dec2026-06-02
interspeech-2025-0196SPCODEC: Split and Prediction for Neural Speech CodecSamsungInterspeech2025codecGAN2026-06-02
2503.04721Full-Duplex-Bench: A Benchmark to Evaluate Full-duplex Spoken Dialogue Models on Turn-taking CapabilitiesarXiv2025SCA, evaluation2026-06-02
2508.08715MultiGen: Child-Friendly Multilingual Speech Generator with LLMsA*STAR Institute for Infocomm ResearcharXiv2025TTSautoregressive-LM, flow-matching, GAN2026-06-02
2508.09767UtterTune: LoRA-Based Target-Language Pronunciation Edit and Control in Multilingual Text-to-SpeecharXiv2025TTSautoregressive-LM2026-06-02
2508.11326MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-ExpertsKunlun Inc.arXiv2025TTSautoregressive-LM, diffusion2026-06-02
2504.12867EmoVoice: LLM-based Emotional Text-To-Speech Model with Freestyle Text PromptingShanghai Jiao Tong University / Tongyi Speech LabarXiv2025TTSautoregressive-LM, flow-matching2026-06-02
2508.08961DualSpeechLM: Towards Unified Speech Understanding and Generation via Dual Speech Token ModelingarXiv2025TTS, SCA, VCautoregressive-LM2026-06-02
2508.08399Exploring Disentangled Neural Speech Codecs from Self-Supervised RepresentationsMERL / Mitsubishi ElectricarXiv2025codec, VCVAE2026-06-02
2508.07711Is GAN Necessary for Mel-Spectrogram-based Neural Vocoder?arXiv2025TTSGAN2026-06-02
2508.07426Scalable Controllable Accented TTSJohns Hopkins UniversityASRU2025TTStransformer-enc-dec, GAN, VAE2026-06-02
2508.07302XEmoRAG: Cross-Lingual Emotion Transfer with Controllable Intensity Using Retrieval-Augmented GenerationNorthwestern Polytechnical UniversityarXiv2025TTS, VCautoregressive-LM, flow-matching2026-06-02
2508.06890Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit ProsodyPOSTECHarXiv2025VCGAN, transformer-enc-dec2026-06-02
2508.06870Text to Speech System for Meitei Mayek ScriptarXiv2025TTStransformer-enc-dec, GAN2026-06-02
2508.05385A Scalable Pipeline for Enabling Non-Verbal Speech Generation and UnderstandingTsinghua UniversityarXiv2025TTS, SCAtransformer-enc-dec2026-06-02
2508.14049MahaTTS: A Unified Framework for Multilingual Text-to-Speech SynthesisDubverse AIarXiv2025TTSautoregressive-LM, flow-matching2026-06-02
2508.04585UniTalker: Conversational Speech-Visual SynthesisInner Mongolia UniversityarXiv2025TTS, SCAautoregressive-LM, flow-matching2026-06-02
2508.04996REF-VC: Robust, Expressive and Fast Zero-Shot Voice Conversion with Diffusion TransformersNorthwestern Polytechnical UniversityarXiv2025VCflow-matching2026-06-02
2508.05207SpectroStream: A Versatile Neural Codec for General AudioGoogle DeepMindarXiv2025codecGAN2026-06-02
2507.20091ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language ModelsarXiv2025SCA, TTSautoregressive-LM2026-06-02
2508.00317Advancing Speech Quality Assessment Through Scientific Challenges and Open-source ActivitiesNagoya UniversityarXiv2025evaluation2026-06-02
2507.22746Next Tokens Denoising for Speech SynthesisMicrosoftarXiv2025TTShybrid2026-06-02
2508.01796Enhancing Spectrogram Realism in Singing Voice Synthesis via Explicit Bandwidth Extension Prior to VocoderTsinghua UniversityarXiv2025singing, TTSdiffusion, GAN2026-06-02
2508.02013SpeechRole: A Large-Scale Dataset and Benchmark for Evaluating Speech Role-Playing AgentsFudan UniversityarXiv2025SCA, evaluation2026-06-02
2508.02849SecoustiCodec: Cross-Modal Aligned Streaming Single-Codebook Speech CodecarXiv2025codecVAE, transformer-enc-dec2026-06-02
2025.naacl-long.110WaveFM: A High-Fidelity and Efficient Vocoder Based on Flow MatchingTsinghua UniversityNAACL2025TTSflow-matching2026-05-30
2025.findings-acl.1051LLMVoX: Autoregressive Streaming Text-to-Speech Model for Any LLMMBZUAIACL2025TTS, SCAautoregressive-LM2026-05-30
2025.emnlp-main.180Scaling Rich Style-Prompted Text-to-Speech DatasetsUT Austin / NYUEMNLP2025TTS, evaluationautoregressive-LM2026-06-01
2507.09318ZipVoice-Dialog: Non-Autoregressive Spoken Dialogue Generation with Flow MatchingXiaomi Corp.arXiv2026TTS, SCAflow-matching2026-05-30
2025.coling-main.518ProsodyFlow: High-fidelity Text-to-Speech through Conditional Flow Matching and Prosody Modeling with Large Speech Language ModelsZhejiang Universityworkshop2025TTSflow-matching, hybrid2026-05-30
interspeech-2025-0469Developing High-Quality TTS for Punjabi and Urdu: Benchmarking against MMS ModelsUniversity of Engineering and Technology, LahoreInterspeech2025TTS, evaluationtransformer-enc-dec2026-05-30
interspeech-2025-0854Bridging the Training–Inference Gap in TTS: Training Strategies for Robust Generative Postprocessing for Low-Resource SpeakersFraunhofer IISInterspeech2025TTSGAN, flow-matching, transformer-enc-dec2026-05-30
interspeech-2025-0973A Dataset for Automatic Assessment of TTS Quality in SpanishInterspeech2025TTS, evaluation2026-05-30
interspeech-2025-0989HiFiTTS-2: A Large-Scale High Bandwidth Speech DatasetNVIDIAInterspeech2025TTS, evaluationautoregressive-LM2026-05-30
interspeech-2025-1034Non-Standard Accent TTS Support via Large Multi-Accent Frontend Pronunciation Knowledge TransferUniversity of EdinburghInterspeech2025TTStransformer-enc-dec2026-05-30
interspeech-2025-0723Counterfactual Activation Editing for Post-hoc Prosody and Mispronunciation Correction in TTS ModelsKAIST / Samsung ElectronicsInterspeech2025TTStransformer-enc-dec, VAE2026-05-30
interspeech-2025-0754EME-TTS: Unlocking the Emphasis and Emotion Link in Speech SynthesisUCAS HangzhouInterspeech2025TTStransformer-enc-dec2026-05-30
interspeech-2025-0762Intrasentential English in Swedish TTS: perceived English-accentednessKTH / MTMInterspeech2025TTSflow-matching2026-05-30
interspeech-2025-0779Intelligibility of Text-to-Speech Systems for Mathematical ExpressionsEricsson R&DInterspeech2025TTS, evaluation2026-05-30
interspeech-2025-0787Gradual modeling of the Lombard effect by modifying speaker embeddings from a Text-To-Speech modelHEAD acousticsInterspeech2025TTSautoregressive-LM2026-05-30
interspeech-2025-0575VoiceMark: Zero-Shot Voice Cloning-Resistant Watermarking Approach Leveraging Speaker-Specific LatentsTsinghua UniversityInterspeech2025VC, TTSVAE2026-05-30
interspeech-2025-0596Facilitating Personalized TTS for Dysarthric Speakers Using Knowledge Anchoring and Curriculum LearningPOSTECHInterspeech2025TTStransformer-enc-dec2026-05-30
interspeech-2025-0648MIKU-PAL: An Automated and Standardized Multimodal Method for Speech Paralinguistic and Affect LabelingFish AudioInterspeech2025TTS, evaluation2026-05-30
interspeech-2025-0669PAST: Phonetic-Acoustic Speech TokenizerHebrew University of JerusalemInterspeech2025codec, TTShybrid2026-05-30
interspeech-2025-0704Differentiable Reward Optimization for LLM based TTS systemAlibaba GroupInterspeech2025TTSautoregressive-LM, flow-matching2026-05-30
interspeech-2025-0406Zero-Shot Mono-to-Binaural Speech SynthesisGoogleInterspeech2025TTSGAN2026-05-30
interspeech-2025-0408Improving User Impression of Spoken Dialogue Systems by Controlling Para-linguistic Expression Based on IntimacyTohoku UniversityInterspeech2025SCA, TTStransformer-enc-dec, GAN2026-05-30
interspeech-2025-0455APTTS: Adversarial Post-training in Latent Flow Matching for Fast and High-fidelity Text-to-SpeechLG AI ResearchInterspeech2025TTSflow-matching, VAE, hybrid2026-05-30
interspeech-2025-0554RapFlow-TTS: Rapid and High-Fidelity Text-to-Speech with Improved Consistency Flow MatchingNAVER CloudInterspeech2025TTSflow-matching, GAN2026-05-30
interspeech-2025-0551Monotonic Attention for Robust Text-to-Speech Synthesis in Large Language Model FrameworksTencentInterspeech2025TTSautoregressive-LM2026-05-30
interspeech-2025-0047Revival with Voice: Multi-modal Controllable Text-to-Speech SynthesisMeta AIInterspeech2025TTSautoregressive-LM2026-05-30
interspeech-2025-0063Fine-Tuning Text-to-Speech Diffusion Models Using Reinforcement Learning with Human FeedbackInterspeech2025TTSdiffusion2026-05-30
interspeech-2025-0143Multimodal Prosody Modeling: A Use Case for Multilingual Sentence Mode PredictionIdiap Research InstituteInterspeech2025TTS, evaluation2026-05-30
interspeech-2025-0310Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language ModelsInterspeech2025TTS, codecautoregressive-LM2026-05-30
interspeech-2025-0319Improving Noise Robustness of LLM-based Zero-shot TTS via Discrete Acoustic Token DenoisingUSTCInterspeech2025TTSautoregressive-LM2026-05-30
2509.02020FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and ChatbotarXiv2025TTS, SCAautoregressive-LM2026-05-26
2507.14534Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice ConversionarXiv2025VCGAN, hybrid2026-05-26
2509.19668Selective Classifier-free Guidance for Zero-shot Text-to-speecharXiv2025TTSflow-matching2026-05-26
2510.00981FlexiCodec: A Dynamic Neural Audio Codec for Low Frame RatesarXiv2025codec, TTShybrid2026-05-26
2412.17048Why Do Speech Language Models Fail to Generate Semantically Coherent Outputs?arXiv2026SCAautoregressive-LM2026-05-26
2025.findings-emnlp.424InteractSpeech: A Speech Dialogue Interaction Corpus for Spoken Dialogue ModelEMNLP2025SCA, evaluationautoregressive-LM2026-05-26
2025.acl-demo.37RT-VC: Real-Time Zero-Shot Voice Conversion with Speech Articulatory CodingUC BerkeleyACL2025VChybrid2026-05-26
2025.acl-industry.42Scaling Under-Resourced TTS: A Data-Optimized Framework with Advanced Acoustic Modeling for ThaiBeijing Logic Intelligence TechnologyACL2025TTSGAN, transformer-enc-dec2026-05-26
2025.acl-long.1043OZSpeech: One-step Zero-shot Speech Synthesis with Learned-Prior-Conditioned Flow MatchingFPT Software AI CenterACL2025TTSflow-matching, hybrid2026-05-26
2025.acl-long.1252Finding A Voice: Exploring the Potential of African American Dialect and Voice Generation for ChatbotsEmory UniversityACL2025TTS, SCAhybrid2026-05-26
2025.acl-long.1471The time scale of redundancy between prosody and linguistic contextMITACL2025evaluationtransformer-enc-dec2026-05-26
2025.acl-long.1498Analyzing and Mitigating Inconsistency in Discrete Speech Tokens for Neural Codec Language ModelsAlibaba Group / Zhejiang UniversityACL2025TTS, codecautoregressive-LM, GAN2026-05-26
2025.acl-long.313F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow MatchingShanghai Jiao Tong UniversityACL2025TTSflow-matching2026-06-01
2025.acl-long.346ControlSpeech: Towards Simultaneous and Independent ZerZhejiang University / Alibaba Tongyi Speech LabACL2025TTStransformer-enc-dec, hybrid2026-05-26
2025.acl-long.388Distilling an End-to-End Voice Assistant Without InstruACL2025SCAtransformer-enc-dec2026-05-26
2025.acl-long.598Advancing Zero-shot Text-to-Speech Intelligibility acroACL2025TTSautoregressive-LM, flow-matching, hybrid2026-05-26
interspeech-2025-0253Long-Context Speech Synthesis with Context-Aware MemorySouth China University of Technology / Alibaba GroupInterspeech2025TTSautoregressive-LM, hybrid2026-05-27
2301.02111Neural Codec Language Models are Zero-Shot Text to Speech SynthesizersMicrosoftarXiv2023TTS, codecautoregressive-LM2026-06-01
interspeech-2025-0902VoiceQualityVC: A Voice Conversion System for Studying the Perceptual Effects of Voice Quality in SpeechKTH Royal Institute of TechnologyInterspeech2025VCVAE, GAN2026-05-27
2025.emnlp-main.989VocalNet: Speech LLMs with Multi-Token Prediction for Faster and High-Quality GenerationSJTU / Ant Group / Wuhan UniversityEMNLP2025TTS, SCAautoregressive-LM, hybrid2026-05-27
2025.acl-long.682Recent Advances in Speech Language Models: A SurveyCUHK / Tencent / NUSACL2025TTS, SCA, evaluationautoregressive-LM, hybrid2026-06-01
2025.americasnlp-1.1Text-to-speech system for low-resource languages: A case study in Shipibo-KoniboPontificia Universidad Católica del Perúworkshop2025TTStransformer-enc-dec, GAN2026-05-27
2025.emnlp-main.1730FillerSpeech: Towards Human-Like Text-to-Speech Synthesis with Filler Insertion and Filler Style ControlKorea University / Samsung ResearchEMNLP2025TTSflow-matching, transformer-enc-dec2026-05-27
2025.findings-naacl.184Continuous Speech Tokenizer in Text To SpeechCUHK / TencentNAACL2025TTS, codecautoregressive-LM, VAE, flow-matching2026-05-27
2025.emnlp-demos.70OpenS2S: Advancing Fully Open-Source End-to-End Empathetic Large Speech Language ModelInstitute of Automation, Chinese Academy of SciencesEMNLP2025SCAautoregressive-LM, hybrid2026-05-27
2406.02430Seed-TTS: A Family of High-Quality Versatile Speech Generation ModelsByteDancearXiv2024TTS, VCautoregressive-LM, diffusion, hybrid2026-05-28
2407.05407CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic TokensAlibaba GrouparXiv2024TTSautoregressive-LM, flow-matching, hybrid2026-05-28
2025.acl-long.65Autoregressive Speech Synthesis without Vector QuantizationMicrosoft / CUHKACL2025TTSautoregressive-LM, VAE, hybrid2026-05-28
2412.10117CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language ModelsAlibaba GrouparXiv2024TTSautoregressive-LM, flow-matching, hybrid2026-05-28
2601.15621Qwen3-TTS Technical ReportAlibaba / Qwen TeamarXiv2026TTSautoregressive-LM, hybrid2026-05-28
2512.14291GLM-TTS Technical ReportZhipu AI / Tsinghua UniversityarXiv2025TTSautoregressive-LM, diffusion, GAN, hybrid2026-05-28
2508.06262Llasa+: Free Lunch for Accelerated and Streaming Llama-Based Speech SynthesisNorthwestern Polytechnical University / HKUSTarXiv2025TTSautoregressive-LM, hybrid2026-05-28
2502.03930DiTAR: Diffusion Transformer Autoregressive Modeling for Speech GenerationByteDance SeedarXiv2025TTSautoregressive-LM, diffusion, hybrid2026-05-28
2504.10352Pseudo-Autoregressive Neural Codec Language Models for Efficient Zero-Shot Text-to-Speech SynthesisMicrosoft / SJTUarXiv2025TTSautoregressive-LM, hybrid2026-05-28
2508.16332Vevo2: A Unified and Controllable Framework for Speech and Singing Voice GenerationCUHK Shenzhen / ByteDance SeedarXiv2025TTS, VC, singingautoregressive-LM, flow-matching, hybrid2026-05-28
2508.02038Marco-Voice Technical ReportAlibaba International Digital CommercearXiv2025TTS, VCautoregressive-LM, flow-matching, hybrid2026-05-28
2604.00688OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language ModelsXiaomi Corp.arXiv2026TTSdiffusion, hybrid2026-05-28
2508.03543EmoSteer-TTS: Fine-Grained and Training-Free Emotion-Controllable Text-to-Speech via Activation SteeringHKUST (Guangzhou) / Tencent AI LabarXiv2025TTSflow-matching2026-05-28
2510.02848Flamed-TTS: Flow Matching Attention-Free Models for Efficient Generating and Dynamic Pacing Zero-shot Text-to-SpeechFPT Software AI CenterarXiv2025TTSflow-matching, hybrid2026-05-28
2506.21619IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-SpeechbilibiliarXiv2025TTSautoregressive-LM, flow-matching, GAN, hybrid2026-05-28
2025.naacl-long.242StyleTTS-ZS: Efficient High-Quality Zero-Shot Text-to-Speech Synthesis with Distilled Time-Varying Style DiffusionColumbia UniversityNAACL2025TTSdiffusion, GAN, VAE, hybrid2026-05-28
2510.12210DiSTAR: Diffusion over a Scalable Token Autoregressive Representation for Speech GenerationSJTU / ByteDancearXiv2025TTSautoregressive-LM, diffusion, hybrid2026-05-28
2025.emnlp-main.40Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic SurveyHKUST-GZ / University of SurreyEMNLP2025TTS, evaluationautoregressive-LM, flow-matching, diffusion, GAN, VAE, transformer-enc-dec, hybrid2026-05-28
2603.08823Fish Audio S2 Technical ReportFish AudioarXiv2026TTSautoregressive-LM, GAN, hybrid2026-05-28
2509.00685MPO: Multidimensional Preference Optimization for Language Model-based Text-to-SpeechNorthwestern Polytechnical UniversityarXiv2025TTSautoregressive-LM2026-05-29
2511.12347VoiceCraft-X: Unifying Multilingual, Voice-Cloning Speech Synthesis and Speech EditingUniversity of Texas at Austin / AmazonEMNLP2025TTSautoregressive-LM2026-05-29
2512.13251DisCo-Speech: Controllable Zero-Shot Speech GenerationChina Mobile Nineverse AI / Peking UniversityarXiv2025TTS, VC, codecautoregressive-LM, GAN2026-05-29
2509.09631DiFlow-TTS: Compact and Low-Latency Zero-Shot Text-to-Speech with Factorized Discrete Flow MatchingFPT Software AI CenterarXiv2025TTSflow-matching, transformer-enc-dec2026-05-29
2512.04720M3-TTS: Multi-modal DiT Alignment & Mel-latent for Zero-shot High-fidelity SpeecharXiv2025TTSdiffusion, VAE2026-05-29
2603.29339LongCat-AudioDiT: High-Fidelity Diffusion Text-to-Speech in the Waveform Latent SpaceMeituanarXiv2026TTSflow-matching, VAE2026-05-29
2508.11273EmoSSLSphere: Multilingual Emotional Speech Synthesis with Spherical Vectors and Discrete Speech TokensarXiv2025TTStransformer-enc-dec2026-05-29
2604.12438An Ultra-Low Latency, End-to-End Streaming Speech Synthesis Architecture via Block-Wise Generation and Depth-Wise Codec DecodingarXiv2026TTStransformer-enc-dec2026-06-01
2604.01760T5Gemma-TTS Technical ReportarXiv2026TTSautoregressive-LM2026-05-29
2508.15442Mitigating Hallucinations in LM-Based TTS Models via Distribution Alignment Using GFlowNetsEMNLP2025TTSautoregressive-LM2026-05-29
2025.acl-long.654Language-Codec: Bridging Discrete Codec Representations and Speech Language ModelsZhejiang UniversityACL2025TTS, codecGAN, VAE2026-05-29
2603.18090MOSS-TTS Technical ReportShanghai Innovation Institute / Fudan UniversityarXiv2026TTSautoregressive-LM, hybrid2026-05-29
2508.04141Parallel GPT: Harmonizing the Independence and Interdependence of Acoustic and Semantic Information for Zero-Shot Text-to-SpeechSouth China University of TechnologyarXiv2025TTSautoregressive-LM, hybrid2026-05-29
2502.11128FELLE: Autoregressive Speech Synthesis with Token-Wise Coarse-to-Fine Flow MatchingarXiv2025TTSautoregressive-LM, flow-matching2026-05-29
2603.26364LLaDA-TTS: Unifying Speech Synthesis and Zero-Shot Editing via Masked Diffusion ModelingBairong, Inc.arXiv2026TTSdiffusion2026-05-29
2508.19098CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech SynthesisarXiv2025TTSautoregressive-LM, flow-matching, VAE2026-05-29
2508.12001FNH-TTS: A Fast, Natural, and Human-Like Speech Synthesis System with advanced prosodic modeling based on Mixture of ExpertsMegatronixarXiv2025TTSVAE, GAN, hybrid2026-05-29
2510.05758EMORL-TTS: Reinforcement Learning for Fine-Grained Emotion Control in LLM-based TTSHangzhou Institute for Advanced Study, UCASICASSP2026TTSautoregressive-LM2026-05-29
2601.03888IndexTTS 2.5 Technical ReportBilibiliarXiv2026TTSautoregressive-LM, flow-matching, hybrid2026-05-29
2509.15969VoXtream: Full-Stream Text-to-Speech with Extremely Low LatencyKTH Royal Institute of TechnologyarXiv2025TTSautoregressive-LM, hybrid2026-05-29
2510.07979IntMeanFlow: Few-step Speech Generation with Integral Velocity DistillationByteDancearXiv2025TTSflow-matching2026-05-29
2025.ccl-1.80Lao-English Code-Switched Speech Synthesis Via Neural Codec Language ModelingKunming University of Science and Technologyworkshop2025TTSautoregressive-LM, hybrid2026-05-29
2025.coling-main.352DiffStyleTTS: Diffusion-based Hierarchical Prosody Modeling for Text-to-Speech with Diverse and Controllable StylesUniversity of Science and Technology of Chinaworkshop2025TTSdiffusion, transformer-enc-dec2026-05-29
2025.acl-long.911DNASpeech: A Contextualized and Situated Text-to-Speech Dataset with Dialogues, Narratives and ActionsRenmin University of ChinaACL2025TTS, evaluationtransformer-enc-dec2026-05-29
2025.acl-short.81Zero-Shot Text-to-Speech for VietnameseMovian AIACL2025TTS, evaluationautoregressive-LM, transformer-enc-dec2026-05-29
2025.acl-long.912LLaMA-Omni 2: LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech SynthesisChinese Academy of SciencesACL2025SCA, TTSautoregressive-LM, flow-matching, hybrid2026-05-29
interspeech-2025-2765The State Of TTS: A Case Study with Human Fooling RatesIIT MadrasInterspeech2025TTS, evaluation2026-06-03
interspeech-2025-0401Enabling the replicability of speech synthesis perceptuInterspeech2025evaluation2026-06-03
interspeech-2025-0115Bringing Interpretability to Neural Audio CodecsInterspeech2025codectransformer-enc-dec2026-06-03
interspeech-2025-0468DualCodec: A Low-Frame-Rate, Semantically-Enhanced NeurCUHK-SZ / BaiduInterspeech2025codecVAE, GAN2026-06-03
interspeech-2025-1641Robust Neural Codec Language Modeling with Phoneme PosiSamsungInterspeech2025TTSautoregressive-LM2026-06-03
interspeech-2025-2447Accelerating Autoregressive Speech Synthesis InferenceTsinghua / TencentInterspeech2025TTSautoregressive-LM2026-06-03
interspeech-2025-1779ReFlow-VC: Zero-shot Voice Conversion Based on RectifieInterspeech2025VCflow-matching2026-06-03
interspeech-2025-0874Efficient and Direct Duplex Modeling for Speech-to-SpeeInterspeech2025SCAautoregressive-LM, hybrid2026-06-03
interspeech-2025-0246DC-Spin: A Speaker-invariant Speech Tokenizer for SpokeInterspeech2025codectransformer-enc-dec2026-06-03
interspeech-2025-1440FreeCodec: A Disentangled Neural Speech Codec with FeweInterspeech2025codecVAE2026-06-03
interspeech-2025-2043Training-Free Voice Conversion with Factorized Optimal Interspeech2025VCtransformer-enc-dec2026-06-03
interspeech-2025-0816Bridging Speech and Singing: Multi-stage Speech-PrompteInterspeech2025singing, VCdiffusion2026-06-03
interspeech-2025-1066Score-Based Training for Energy-Based TTS ModelsInterspeech2025TTSdiffusion2026-06-03
interspeech-2025-1122BitTTS: Highly Compact Text-to-Speech Using 1.58-bit QuInterspeech2025TTSGAN, transformer-enc-dec2026-06-03
2508.20660CodecBench: A Comprehensive Benchmark for Acoustic andFudan UniversityarXiv2025codec, evaluation2026-06-03
interspeech-2025-1344Parameter-Efficient Fine-Tuning for Low-Resource Text-tAjou UniversityInterspeech2025TTSflow-matching2026-06-03
interspeech-2025-2449Accelerating Flow-Matching-Based Text-to-Speech via EmpInterspeech2025TTSflow-matching2026-06-03
interspeech-2025-1595Scheduled Interleaved Speech-Text Training for Speech-tInterspeech2025TTS, SCAautoregressive-LM2026-06-03
interspeech-2025-0815Towards Better Disentanglement in Non-Autoregressive ZeInterspeech2025VCVAE, GAN2026-06-03
interspeech-2025-1101ZSDEVC: Zero-Shot Diffusion-based Emotional Voice ConveInterspeech2025VCdiffusion2026-06-03
interspeech-2025-2660Triadic Multi-party Voice Activity Projection for Turn-Kyoto UniversityInterspeech2025SCAtransformer-enc-dec2026-06-03
2508.07375TurnGuide: Enhancing Meaningful Full Duplex Spoken IntearXiv2025SCAautoregressive-LM2026-06-04
2508.16790TaDiCodec: Text-aware Diffusion Speech Tokenizer for SpCUHK-SZarXiv2025codecdiffusion, transformer-enc-dec, autoregressive-LM2026-06-04
interspeech-2025-1289Unlocking Temporal Flexibility: Neural Speech Codec witInterspeech2025codechybrid2026-06-04
interspeech-2025-0984Benchmarking Neural Speech Codec Intelligibility with SInterspeech2025codec, evaluation2026-06-04
2508.07273Incorporating Contextual Paralinguistic Understanding iarXiv2025SCAtransformer-enc-dec2026-06-04
2508.08957QAMRO: Quality-aware Adaptive Margin Ranking OptimizatiarXiv2025evaluation2026-06-04
2508.09600OSUM-EChat: Enhancing End-to-End Empathetic Spoken ChatNorthwestern Polytechnical UniversityarXiv2025SCAautoregressive-LM, flow-matching2026-06-04
2508.09702M3PDB: A Multimodal, Multi-Label, Multilingual Prompt Database for Speech GenerationarXiv2025TTS, evaluation2026-06-04
2508.11224Benchmarking Prosody Encoding in Discrete Speech TokensThe University of Tokyo / AISTASRU2025evaluation, TTS2026-06-04
2508.13028Integrating Feedback Loss from Bi-modal Sarcasm DetectoarXiv2025TTStransformer-enc-dec2026-06-04
2508.15565Any-to-any Speaker Attribute Perturbation for AsynchronarXiv2025VCGAN2026-06-04
2508.15931QvTAD: Differential Relative Attribute Learning for VoiQifu TechnologyarXiv2025evaluationtransformer-enc-dec2026-06-04
2508.16188Seeing is Believing: Emotion-Aware Audio-Visual LanguagEMNLP2025TTS, SCAautoregressive-LM2026-06-04
2508.17031RephraseTTS: Dynamic Length Text based Speech InsertionIIT KanpurarXiv2025TTS, VCtransformer-enc-dec, GAN2026-06-04
2508.17494Improving French Synthetic Speech Quality via SSML Prosworkshop2025TTShybrid2026-06-04
2508.17623EMO-Reasoning: Benchmarking Emotional Reasoning CapabilarXiv2025SCA, evaluation2026-06-04
2508.18006Unseen Speaker and Language Adaptation for LightweightAmazonarXiv2025TTSGAN2026-06-04
2508.19205VibeVoice Technical ReportMicrosoft ResearcharXiv2025TTShybrid2026-06-04
2509.00503Entropy-based Coarse and Compressed Semantic Speech ReparXiv2025codecautoregressive-LM, transformer-enc-dec2026-06-04
2509.00675Speaker-Conditioned Phrase Break Prediction for Text-toarXiv2025TTStransformer-enc-dec2026-06-04
2509.01391MixedG2P-T5: G2P-free Speech Synthesis for Mixed-scriptarXiv2025TTStransformer-enc-dec2026-06-04
2509.02244Spectrogram Patch Codec: A 2D Block-Quantized VQ-VAE anarXiv2025codecVAE, GAN2026-06-04
2509.03292Improving Perceptual Audio Aesthetic Assessment via TriarXiv2025evaluationhybrid2026-06-04
2509.03940VoxRole: A Comprehensive Benchmark for Evaluating SpeecarXiv2025SCA, evaluation2026-06-04
2410.00037Moshi: a speech-text foundation model for real-time diaKyutaiarXiv2024SCA, TTSautoregressive-LM, hybrid2026-06-09
2411.13577WavChat: A Survey of Spoken Dialogue ModelsarXiv2024SCAautoregressive-LM, transformer-enc-dec, hybrid2026-06-09
2503.20215Qwen2.5-Omni Technical ReportarXiv2025autoregressive-LM, flow-matching, hybrid2026-06-09
2407.21783The Llama 3 Herd of ModelsMetaarXiv2024autoregressive-LM2026-06-09
1912.06670Common Voice: A Massively-Multilingual Speech CorpusarXiv20192026-06-09
2312.15185emotion2vec: Self-Supervised Pre-Training for Speech EmarXiv20232026-06-09
2212.04356Robust Speech Recognition via Large-Scale Weak SupervisarXiv2022transformer-enc-dec2026-06-09
2010.05646HiFi-GAN: Generative Adversarial Networks for EfficientKakao EnterprisearXiv2020TTSGAN2026-06-09
2210.13438High Fidelity Neural Audio CompressionMeta AIarXiv2022codecGAN, VAE2026-06-09
2006.04558FastSpeech 2: Fast and High-Quality End-to-End Text toMicrosoft Research AsiaarXiv2020TTStransformer-enc-dec2026-06-09
2407.10759Qwen2-Audio Technical ReportAlibaba GrouparXiv20242026-06-10
2303.08774GPT-4 Technical ReportOpenAIarXiv2023autoregressive-LM2026-06-10
2412.02612GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken ChatbotTsinghua University / Zhipu.AIarXiv2024TTS, SCAautoregressive-LM, flow-matching2026-06-10
1711.05101Decoupled Weight Decay RegularizationarXiv20172026-06-10
2410.21276GPT-4o System CardarXiv2024autoregressive-LM2026-06-10
2412.15115Qwen2.5 Technical ReportAlibabaarXiv2024autoregressive-LM2026-06-10
2210.02747Flow Matching for Generative ModelingMeta AI (FAIR)arXiv2022flow-matching2026-06-10
2209.03143AudioLM: a Language Modeling Approach to Audio GeneratiarXiv2022TTS, SCAautoregressive-LM2026-06-10
2206.04658BigVGAN: A Universal Neural Vocoder with Large-Scale TrNVIDIAarXiv2022TTSGAN2026-06-10
2207.12598Classifier-Free Diffusion GuidancearXiv2022diffusion2026-06-10
1412.6980Adam: A Method for Stochastic OptimizationarXiv20142026-06-10
2308.16692SpeechTokenizer: Unified Speech Tokenizer for Speech LaFudan UniversityarXiv2023TTS, codecGAN, VAE2026-06-11
2503.01710Spark-TTS: An Efficient LLM-Based Text-to-Speech ModelHKUSTarXiv2025TTSautoregressive-LM2026-06-11
2409.00750MaskGCT: Zero-Shot Text-to-Speech with Masked GenerativCUHK-SZarXiv2024TTS, VCautoregressive-LM2026-06-11
2505.17589CosyVoice 3: Towards In-the-wild Speech Generation viaAlibabaarXiv2025TTSautoregressive-LM, flow-matching2026-06-11
2408.16725Mini-Omni: Language Models Can Hear, Talk While ThinkiInspiraiarXiv2024SCA, TTSautoregressive-LM2026-06-11
2502.04128Llasa: Scaling Train-Time and Inference-Time Compute foarXiv2025TTSautoregressive-LM2026-06-11
2304.09116NaturalSpeech 2: Latent Diffusion Models are Natural anMicrosoft Research AsiaarXiv2023TTS, VC, singingdiffusion, VAE2026-06-11
2406.05370VALL-E 2: Neural Codec Language Models are Human ParityMicrosoftarXiv2024TTSautoregressive-LM2026-06-11
2409.06666LLaMA-Omni: Seamless Speech Interaction with Large LangICT/CASarXiv2024SCAhybrid2026-06-11
2411.00774Freeze-Omni: A Smart and Low Latency Speech-to-speech DTencent Youtu LabarXiv2024SCAautoregressive-LM, hybrid2026-06-11
2408.16532WavTokenizer: an Efficient Acoustic Discrete Codec TokearXiv2024codecGAN, VAE2026-06-11
2407.04051FunAudioLLM: Voice Understanding and Generation FoundatAlibaba GrouparXiv2024TTShybrid2026-06-11
2305.11000SpeechGPT: Empowering Large Language Models with IntrinFudan UniversityarXiv2023SCA, TTSautoregressive-LM2026-06-11
2410.17196VoiceBench: Benchmarking LLM-Based Voice AssistantsNational University of SingaporearXiv2024evaluation, SCA2026-06-11
2409.03283FireRedTTS: A Foundation Text-To-Speech Framework for IXiaohongshuarXiv2024TTSautoregressive-LM, flow-matching, hybrid2026-06-11
2311.07919Qwen-Audio: Advancing Universal Audio Understanding viaarXiv2023transformer-enc-dec2026-06-12
2505.09388Qwen3 Technical ReportarXiv2025autoregressive-LM2026-06-12
2005.07143ECAPA-TDNN: Emphasized Channel Attention, Propagation aarXiv20202026-06-12
2302.13971LLaMA: Open and Efficient Foundation Language ModelsarXiv2023autoregressive-LM2026-06-12
2507.06261Gemini 2.5: Pushing the Frontier with Advanced ReasoninGooglearXiv2025TTS, SCA2026-06-12
2012.03411MLS: A Large-Scale Multilingual Dataset for Speech ResearXiv20202026-06-12
2501.12948DeepSeek-R1: Incentivizing Reasoning Capability in LLMsDeepSeekarXiv2025autoregressive-LM2026-06-12
2312.05187Seamless: Multilingual Expressive and Streaming SpeecharXiv20232026-06-12
2106.06909GigaSpeech: An Evolving, Multi-domain ASR Corpus with 1arXiv20212026-06-12
2309.15505Finite Scalar Quantization: VQ-VAE Made SimplearXiv2023VAE2026-06-12
2306.00814Vocos: Closing the gap between time-domain and Fourier-arXiv2023TTSGAN2026-06-12
2407.05361Emilia: An Extensive, Multilingual, and Diverse SpeecharXiv20242026-06-12
2406.18009E2 TTS: Embarrassingly Easy Fully Non-Autoregressive ZeMicrosoftarXiv2024TTSflow-matching2026-06-12
2406.04904XTTS: a Massively Multilingual Zero-Shot Text-to-SpeechCoqui.ai / NVIDIA / Cantina.aiarXiv2024autoregressive-LM, GAN2026-06-12
2409.05377BigCodec: Pushing the Limits of Low-Bitrate Neural SpeeUniversity of Tokyo, Microsoft, Keio UniversityarXiv2024codecGAN, VAE2026-06-12
2305.02765HiFi-Codec: Group-residual Vector quantization for HighPeking University / Tencent AI LabarXiv2023codecGAN, VAE2026-06-12
2403.16973VoiceCraft: Zero-Shot Speech Editing and Text-to-SpeecharXiv2024TTSautoregressive-LM2026-06-12
2502.11946Step-Audio: Unified Understanding and Generation in IntStepFunarXiv2025TTS, SCAautoregressive-LM, flow-matching, hybrid2026-06-12
2501.06282MinMo: A Multimodal Large Language Model for Seamless VAlibaba GrouparXiv2025TTS, SCAautoregressive-LM, flow-matching, hybrid2026-06-12
2303.03926Speak Foreign Languages with Your Own Voice: Cross-LingMicrosoftarXiv2023TTS, multilingual-ttsautoregressive-LM2026-06-12
2305.09636SoundStorm: Efficient Parallel Audio GenerationGooglearXiv2023TTS, SCAautoregressive-LM2026-06-13
1712.05884Natural TTS Synthesis by Conditioning WaveNet on Mel SpGooglearXiv2017TTStransformer-enc-dec2026-06-13
2402.01912Natural language guidance of high-fidelity text-to-speeStability AIarXiv2024TTSautoregressive-LM2026-06-13
2306.12925AudioPaLM: A Large Language Model That Can Speak and LiGooglearXiv2023TTS, SCAautoregressive-LM2026-06-13
2305.07243Better speech synthesis through scalingarXiv2023TTSautoregressive-LM, diffusion, VAE2026-06-13
1609.03499WaveNet: A Generative Model for Raw AudioGoogle DeepMindarXiv2016TTSautoregressive-LM2026-06-13
2411.19842Scaling Transformers for Low-Bitrate High-Quality SpeecStability AIarXiv2024codectransformer-enc-dec, VAE2026-06-13
2407.08551Autoregressive Speech Synthesis without Vector QuantizaMicrosoftarXiv2024TTSautoregressive-LM2026-06-13
1703.10135Tacotron: Towards End-to-End Speech SynthesisGooglearXiv2017transformer-enc-dec2026-06-13
2502.17239Baichuan-Audio: A Unified Framework for End-to-End SpeeBaichuan Inc.arXiv2025SCA, TTSautoregressive-LM, flow-matching, hybrid2026-06-13
2402.05755Spirit LM: Interleaved Spoken and Written Language ModeMeta AIarXiv2024SCA, TTSautoregressive-LM2026-06-13
2402.08093BASE TTS: Lessons from building a billion-parameter TexAmazon AGIarXiv2024TTSautoregressive-LM2026-06-13
2507.16632Step-Audio 2 Technical ReportStepFunarXiv2025SCA, TTSautoregressive-LM, flow-matching, hybrid2026-06-13
2310.00704UniAudio: An Audio Foundation Model Toward Universal AuarXiv2023TTS, VC, singingautoregressive-LM, hybrid2026-06-13
2106.15561A Survey on Neural Speech SynthesisMicrosoft Research AsiaarXiv2021TTSautoregressive-LM, flow-matching, diffusion, GAN, VAE, transformer-enc-dec2026-06-13
2411.01156Fish-Speech: Leveraging Large Language Models for AdvanFish AudioarXiv2024TTSautoregressive-LM, GAN2026-06-14
2505.07916MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech withMiniMaxarXiv2025TTSautoregressive-LM, flow-matching, VAE2026-06-14
2410.11190Mini-Omni2: Towards Open-source GPT-4o with Vision, SpeInspirai / Tsinghua UniversityarXiv2024SCAautoregressive-LM2026-06-14
2410.03751Recent Advances in Speech Language Models: A SurveyChinese University of Hong KongarXiv2024SCA, TTSautoregressive-LM2026-06-14
2104.00355Speech Resynthesis from Discrete Disentangled Self-SupeFacebook AI ResearcharXiv2021TTS, VCGAN, VAE2026-06-14
2502.06490Recent Advances in Discrete Speech Tokens: A ReviewSJTU / MSRAarXiv2025TTS, VC, SCA, codecautoregressive-LM, transformer-enc-dec, GAN, VAE2026-06-14
2105.06337Grad-TTS: A Diffusion Probabilistic Model for Text-to-SarXiv2021TTSdiffusion, transformer-enc-dec2026-06-14
2412.15649SLAM-Omni: Timbre-Controllable Voice Interaction SystemSJTU / MicrosoftarXiv2024SCAautoregressive-LM, flow-matching2026-06-14
2502.05512IndexTTS: An Industrial-Level Controllable and EfficienbilibiliarXiv2025TTSautoregressive-LM, GAN2026-06-14
2502.07243Vevo: Controllable Zero-Shot Voice Imitation with Self-Meta AIICLR2025TTS, VCautoregressive-LM, flow-matching, hybrid2026-06-14
2406.07855VALL-E R: Robust and Efficient Zero-Shot Text-to-SpeechMicrosoftarXiv2024TTSautoregressive-LM2026-06-14
2410.17799OmniFlatten: An End-to-end GPT Model for Seamless VoiceAlibaba (Tongyi Lab)arXiv2024SCAautoregressive-LM2026-06-14
2504.08528On The Landscape of Spoken Language Models: A ComprehenarXiv2025SCAautoregressive-LM, transformer-enc-dec2026-06-14
2206.08317Paraformer: Fast and Accurate Parallel Transformer forAlibaba GrouparXiv2022transformer-enc-dec2026-06-14
2412.19437DeepSeek-V3 Technical ReportDeepSeek-AIarXiv20242026-06-14
2402.03300DeepSeekMath: Pushing the Limits of Mathematical ReasonarXiv2024autoregressive-LM2026-06-14
2310.13289SALMONN: Towards Generic Hearing Abilities for Large LaTsinghua University / ByteDancearXiv20232026-06-14
1810.04805BERT: Pre-training of Deep Bidirectional Transformers farXiv20182026-06-14
2502.05139Meta Audiobox Aesthetics: Unified Automatic Quality AssMeta (FAIR)arXiv20252026-06-14
2307.09288Llama 2: Open Foundation and Fine-Tuned Chat ModelsMetaarXiv2023autoregressive-LM2026-06-14
2312.11805Gemini: A Family of Highly Capable Multimodal ModelsarXiv20232026-06-14
2005.14165Language Models are Few-Shot LearnersarXiv2020autoregressive-LM2026-06-14
2407.10671Qwen2 Technical ReportAlibaba GrouparXiv20242026-06-14
2106.04624SpeechBrain: A General-Purpose Speech ToolkitarXiv20212026-06-14
2406.14294DASB - Discrete Audio and Speech BenchmarkarXiv20242026-06-14
2301.12503AudioLDM: Text-to-Audio Generation with Latent DiffusioICML2023diffusion, VAE2026-06-15
2301.11325MusicLM: Generating Music From TextGooglearXiv2023SCAautoregressive-LM2026-06-15
2305.15255Spoken Question Answering and Speech Continuation UsingGoogle ResearcharXiv2023SCAautoregressive-LM2026-06-15
2312.01479OpenVoice: Versatile Instant Voice CloningMIT & MyShell.aiarXiv2023TTS, VCGAN, VAE2026-06-15
2312.15821Audiobox: Unified Audio Generation with Natural LanguagMeta FAIRarXiv2023TTSflow-matching2026-06-15
2401.07333ELLA-V: Stable Neural Codec Language Modeling with AligarXiv2024TTSautoregressive-LM2026-06-15
2402.13236Towards audio language modeling — an overviewarXiv2024TTS, SCA, codec2026-06-15
2404.03204RALL-E: Robust Codec Language Modeling with Chain-of-ThMicrosoftarXiv2024TTSautoregressive-LM2026-06-15
2406.00654Enhancing Zero-shot Text-to-Speech Synthesis with HumanNanyang Technological UniversityarXiv2024TTSautoregressive-LM2026-06-15
2406.05551Autoregressive Diffusion Transformer for Text-to-SpeechCUHK ShenzhenarXiv2024TTShybrid2026-06-15
2408.02622Language Model Can Listen While SpeakingShanghai Jiao Tong University / ByteDancearXiv2024SCA, TTSautoregressive-LM2026-06-15
2411.09943Zero-shot Voice Conversion with Diffusion TransformersNanyang Technological UniversityarXiv2024VCdiffusion, transformer-enc-dec2026-06-16
2411.17607Scaling Speech-Text Pre-training with Synthetic InterleTsinghua University / Zhipu.AIarXiv2024SCAautoregressive-LM, flow-matching2026-06-16
2411.18803TS3-Codec: Transformer-Based Simple Streaming Single CoarXiv2024GAN2026-06-16
2412.04724StableVC: Style Controllable Zero-Shot Voice ConversionNorthwestern Polytechnical University / XimalayaarXiv2024VCflow-matching2026-06-16
2506.13053ZipVoice: Fast and High-Quality Zero-Shot Text-to-SpeecXiaomiarXiv2025TTSflow-matching2026-06-16
2505.02625LLaMA-Omni2: LLM-based Real-time Spoken Chatbot with AuICT/CASarXiv2025SCAautoregressive-LM, flow-matching, hybrid2026-06-16
2502.18924MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion TZhejiang University, ByteDancearXiv2025TTSdiffusion, VAE2026-06-16
2507.23159Full-Duplex-Bench v1.5: Evaluating Overlap Handling forarXiv2025SCA, evaluation2026-06-16
2506.16381InstructTTSEval: Benchmarking Complex Natural-LanguageFudan UniversityarXiv2025TTS, evaluation2026-06-16
2506.10274Discrete Audio Tokens: More Than a Survey!arXiv2025codec, TTS, evaluation2026-06-16
2505.13000DualCodec: A Low-Frame-Rate, Semantically-Enhanced NeurCUHK-SZ / BaiduarXiv2025codec, TTSautoregressive-LM2026-06-16
2503.14345MoonCast: High-Quality Zero-Shot Podcast GenerationarXiv2025TTSautoregressive-LM, flow-matching2026-06-16
2508.04195NVSpeech: An Integrated and Scalable Pipeline for HumanarXiv2025autoregressive-LM, flow-matching2026-06-16
2505.09558WavReward: Spoken Dialogue Models With Generalist RewarZhejiang University / Alibaba GrouparXiv2025SCA, evaluationautoregressive-LM2026-06-16
2504.10344ALMTokenizer: A Low-bitrate and Semantic-rich Audio CodarXiv2025codec, TTS, SCAautoregressive-LM, VAE2026-06-16
2504.02407F5R-TTS: Improving Flow-Matching based Text-to-Speech wTencentarXiv2025TTSflow-matching2026-06-16
2511.15848Step-Audio-R1 Technical ReportStepFunarXiv2025SCAautoregressive-LM2026-06-16
2510.07838Full-Duplex-Bench-v2: A Multi-Turn Evaluation FrameworkarXiv2025SCA, evaluation2026-06-16
2505.14648Vox-Profile: A Speech Foundation Model Benchmark for CharXiv2025evaluation2026-06-16
1510.08484MUSAN: A Music, Speech, and Noise CorpusarXiv20152026-06-17
1908.06248JVS corpus: free Japanese multi-speaker voice corpusarXiv2019TTS, VC2026-06-17
1607.06450Layer NormalizationarXiv20162026-06-17
1808.10583AISHELL-2: Transforming Mandarin ASR Research Into InduarXiv20182026-06-17
2002.05202GLU Variants Improve TransformerarXiv20202026-06-17
2302.00482Improving and generalizing flow-based generative modelsarXiv20232026-06-17
2007.10310CoVoST 2 and Massively Multilingual Speech-to-Text TranarXiv20202026-06-17
2001.08361Scaling Laws for Neural Language ModelsarXiv20202026-06-17
2308.10248Steering Language Models With Activation EngineeringarXiv20232026-06-17
2309.16609Qwen Technical ReportarXiv2023autoregressive-LM2026-06-17
2308.05725EXPRESSO: A Benchmark and Analysis of Discrete ExpressiarXiv20232026-06-17
2308.11596SeamlessM4T: Massively Multilingual & Multimodal MachinarXiv2023transformer-enc-dec2026-06-17
2408.05211VITA: Towards Open-Source Interactive Omni Multimodal LarXiv20242026-06-17
2408.01800MiniCPM-V: A GPT-4V Level MLLM on Your PhonearXiv20242026-06-17
2312.10997Retrieval-Augmented Generation for Large Language ModelarXiv20232026-06-17
2402.07729AIR-Bench: Benchmarking Large Audio-Language Models viaarXiv20242026-06-17
2501.07246Audio-CoT: Exploring Chain-of-Thought Reasoning in LargarXiv20252026-06-17
2412.08635Multimodal Latent Language Modeling with Next-Token DifarXiv2024autoregressive-LM, diffusion, VAE, hybrid2026-06-17
2410.19168MMAU: A Massive Multi-Task Audio Understanding and ReasarXiv20242026-06-17
2501.01957VITA-1.5: Towards GPT-4o Level Real-Time Vision and SpearXiv2025autoregressive-LM, transformer-enc-dec2026-06-17
2503.01743Phi-4-Mini Technical Report: Compact yet Powerful MultiarXiv20252026-06-17
2503.19786Gemma 3 Technical ReportGoogle DeepMindarXiv2025autoregressive-LM2026-06-17
2505.03739VITA-Audio: Fast Interleaved Cross-Modal Token GeneratiarXiv2025autoregressive-LM, hybrid2026-06-17
2501.15368Baichuan-Omni-1.5 Technical ReportBaichuan Inc.arXiv2025autoregressive-LM, flow-matching2026-06-17
2506.02863CapSpeech: Enabling Downstream Applications in Style-CaarXiv2025TTS, evaluationautoregressive-LM, flow-matching2026-06-17
2507.12705AudioJudge: Understanding What Works in Large Audio ModarXiv20252026-06-17
2506.07900MiniCPM4: Ultra-Efficient LLMs on End DevicesarXiv2025autoregressive-LM2026-06-17
2507.08128Audio Flamingo 3: Advancing Audio Intelligence with FularXiv2025autoregressive-LM2026-06-17
2510.14664SpeechLLM-as-Judges: Towards General and InterpretablearXiv20252026-06-17
2508.13992MMAU-Pro: A Challenging and Comprehensive Benchmark forarXiv20252026-06-17
2511.09690Omnilingual ASR: Open-Source Multilingual Speech RecognMeta (FAIR)arXiv2025transformer-enc-dec2026-06-17
2509.08753Streaming Sequence-to-Sequence Learning with Delayed StarXiv2025autoregressive-LM2026-06-17
2409.09098AccentBox: Towards High-Fidelity Zero-Shot Accent GenerUniversity of EdinburgharXiv2025TTStransformer-enc-dec2026-06-29
2025.coling-industry.29CarMem: Enhancing Long-Term Memory in LLM Voice AssistaBMW Group / Univ. Augsburg / TUMCOLING2025SCA2026-06-29
2025.chipsal-1.18Impacts of Vocoder Selection on Tacotron-based Nepali TCHiPSAL2025TTS, evaluationGAN, transformer-enc-dec2026-06-29
2025.coling-main.685VoxpopuliTTS: a large-scale multilingual TTS corpus forZhejiang UniversityCOLING2025TTS2026-06-29
2409.20007DeSTA2: Developing Instruction-Following Speech LanguagearXiv2025SCAtransformer-enc-dec2026-06-29
2025.computel-main.6Evaluating Indigenous language speech synthesis for educComputEL2025TTS, evaluation2026-06-29
2025.nodalida-1.32Estonian isolated-word text-to-speech synthesiserInstitute of the Estonian LanguageNoDaLiDa2025TTS2026-06-29
2025.naacl-srw.6Towards Codec-LM Co-design for Neural Codec Language MoCartesia AI / MIT / CMUNAACL2025TTS, codecautoregressive-LM, hybrid2026-06-29
2025.findings-naacl.298Gender Bias in Instruction-Guided Speech Synthesis ModeNAACL2025TTS, evaluation2026-06-29
2025.findings-naacl.471The Role of Prosody in Spoken Question AnsweringNAACL2025evaluation2026-06-29
2025.naacl-long.464ManaTTS Persian: a recipe for creating TTS datasets forSharif University of TechnologyNAACL2025TTS, evaluation2026-06-29
2025.naacl-long.619ProSE: Diffusion Priors for Speech EnhancementUniversity of MarylandNAACL2025TTSdiffusion, transformer-enc-dec2026-06-29
iclr-2025-tQ1PmLfPBLPeriodWave: Multi-Period Flow Matching for High-Fidelity Waveform GenICLR2025TTSflow-matching2026-06-29
iclr-2025-cuFzE8JlvbContinuous Autoregressive Modeling with Stochastic Monotonic AlignmenThe Hong Kong Polytechnic UniversityICLR2025TTSautoregressive-LM, VAE2026-06-29
iclr-2025-dGSOn7sdWgSyllableLM: Learning Coarse Semantic Units for Speech Language ModelsUniversity of Texas at AustinICLR2025SCAautoregressive-LM2026-06-29
iclr-2025-868masI331HALL-E: Hierarchical Neural Codec Language Model for Minute-Long ZeroICLR2025TTSautoregressive-LM2026-06-29
iclr-2025-hQvX9MBowCDiTTo-TTS: Diffusion Transformers for Scalable Text-to-KRAFTONICLR2025TTSdiffusion, transformer-enc-dec2026-06-30
iclr-2025-uxDFlPGRLXFlowDec: A flow-based full-band general audio codec witMetaICLR2025codecflow-matching, GAN2026-06-30
2025.findings-naacl.130DiVISe: Direct Visual-Input Speech Synthesis PreservingShanghai Jiao Tong UniversityNAACL2025TTStransformer-enc-dec, GAN2026-06-30
2025.findings-naacl.279BnTTS: Few-Shot Speaker Adaptation in Low-Resource SettHishab SingaporeNAACL2025TTSautoregressive-LM, GAN, hybrid2026-06-30
2025.findings-naacl.38Prompt-Guided Selective Masking Loss for Context-Aware POSTECHNAACL2025TTStransformer-enc-dec2026-06-30
2025.naacl-demo.12ESPnet-SpeechLM: An Open Speech Language Model ToolkitCarnegie Mellon UniversityNAACL2025TTS, SCAautoregressive-LM2026-06-30
2025.naacl-demo.21ESPnet-SDS: Unified Toolkit and Demo for Spoken DialoguCarnegie Mellon UniversityNAACL2025SCA2026-06-30
2025.naacl-long.484Behavior-SD: Behaviorally Aware Spoken Dialogue GeneratSeoul National UniversityNAACL2025SCAautoregressive-LM2026-06-30
2025.naacl-long.591Robust and Unbounded Length Generalization in AutoregreGoogle DeepMindNAACL2025TTStransformer-enc-dec2026-06-30
2025.naacl-short.65kNN Retrieval for Simple and Effective Zero-Shot Multi-NAACL2025TTShybrid, GAN2026-06-30
2025.naacl-short.69Developing multilingual speech synthesis system for OjiNAACL2025TTSflow-matching2026-06-30
2025.iwsds-1.11Paralinguistic Attitude Recognition for Spoken DialogueFairy Devices Inc.IWSDS2025SCA2026-06-30
2025.iwsds-1.27A Survey of Recent Advances on Turn-taking Modeling in Université Paris-Saclay, CEA, ListIWSDS2025SCA2026-07-01
2505.15772MIKU-PAL: An Automated and Standardized Multi-Modal MetFish Audio; Carnegie Mellon UniversityarXiv2025TTS, evaluation2026-07-01
2507.06235Super Kawaii Vocalics: Amplifying the “Cute” Factor in arXiv2025TTS2026-07-01
2506.23049AURA: Agent for Understanding, Reasoning, and AutomatedarXiv2025SCAautoregressive-LM2026-07-01
2025.acl-long.681SIFT-50M: A Large-Scale Multilingual Dataset for SpeechAmazon AGIACL2025SCAautoregressive-LM2026-07-01
2025.acl-long.790Rhythm Controllable and Efficient Zero-Shot Voice ConveZhejiang UniversityACL2025VCflow-matching2026-07-01
2025.acl-long.817SimulS2S-LLM: Unlocking Simultaneous Inference of SpeecACL2025SCA, TTSautoregressive-LM2026-07-01
2025.acl-long.87Takin-VC: Expressive Zero-Shot Voice Conversion via AdaACL2025VCflow-matching2026-07-01
2025.acl-long.937UniCodec: Unified Audio Codec with Single Domain-AdaptiACL2025codecVAE2026-07-01
2025.acl-long.997Align-SLM: Textless Spoken Language Models with ReinforACL2025SCAautoregressive-LM2026-07-01
2025.conll-1.9A Linguistically Motivated Analysis of Intonational PhrCoNLL2025TTS, evaluation2026-07-01
2025.findings-acl.101Chain-Talker: Chain Understanding and Rendering for EmpACL2025TTSautoregressive-LM, flow-matching2026-07-01
2025.findings-acl.115SLAM-Omni: Timbre-Controllable Voice Interaction SystemACL2025SCA, TTSautoregressive-LM2026-07-01
2025.findings-acl.1226PodAgent: A Comprehensive Framework for Podcast GeneratACL2025TTS, VChybrid2026-07-01
2025.findings-acl.470Does Your Voice Assistant Remember? Analyzing ConversatSeoul National UniversityACL2025SCA, evaluation2026-07-01
2025.findings-acl.534Unlocking Speech Instruction Data Potential with Query ACL2025SCA2026-07-01
2025.findings-acl.631Slamming: Training a Speech Language Model on One GPU iThe Hebrew University of JerusalemACL2025SCAautoregressive-LM2026-07-01
2025.findings-acl.687TCSinger 2: Customizable Multilingual Zero-shot SingingZhejiang UniversityACL2025singing, TTSflow-matching, VAE2026-07-01
2025.findings-acl.71Data-Centric Improvements for Enhancing Multi-Modal UndACL2025SCAautoregressive-LM2026-07-01
2025.findings-acl.75Leveraging Unit Language Guidance to Advance Speech ModACL2025TTS, SCAtransformer-enc-dec2026-07-01
2025.findings-ijcnlp.49Incorporating Dialogue State Tracking into Japanese FulNTT / Nagoya UniversityACL2025SCAautoregressive-LM2026-07-01
2025.iwslt-1.5SSR: Alignment-Aware Modality Connector for Speech LangIWSLT2025SCAautoregressive-LM2026-07-01
2025.unlp-1.11Context-Aware Lexical Stress Prediction and Phonemizatiworkshop2025TTStransformer-enc-dec2026-07-01
2412.18603Long-Form Speech Generation with Spoken Language ModelsICML2025SCAautoregressive-LM2026-07-01
2503.11026MAVFlow: Preserving Paralinguistic Elements with ConditKAISTarXiv2025VC, TTSflow-matching2026-07-01
2505.15670SALM-Duplex: Efficient and Direct Duplex Modeling for SNVIDIAarXiv2025SCAautoregressive-LM2026-07-01
2506.09874UmbraTTS: Adapting Text-to-Speech to Environmental ContarXiv2025TTSflow-matching2026-07-01
2506.18296JIS: A Speech Corpus of Japanese Idol Speakers with VarNTT CorporationInterspeech2025TTS, VC, evaluation2026-07-01
2507.02176Analyzing and Improving Speaker Similarity Assessment farXiv2025evaluation2026-07-01
2507.00808Multi-interaction TTS toward professional recording repNTTarXiv2025TTStransformer-enc-dec2026-07-01
2507.01611QHARMA-GAN: Quasi-Harmonic Neural Vocoder based on AutoarXiv2025TTSGAN2026-07-01
2507.02380JoyTTS: LLM-based Spoken Chatbot With Voice CloningJD Health International Inc.arXiv2025SCA, TTSautoregressive-LM, flow-matching2026-07-01
2507.03887Traceable TTS: Toward Watermark-Free TTS with Strong TrarXiv2025TTSflow-matching2026-07-01
2507.03912Prosody Labeling with Phoneme-BERT and Speech FoundatioCyberAgentarXiv2025TTS2026-07-01
2507.08012RepeaTTS: Towards Feature Discovery through Repeated FiarXiv2025TTStransformer-enc-dec2026-07-01
2507.04349TTS-CtrlNet: Time varying emotion aligned text-to-speecarXiv2025TTSflow-matching2026-07-01
2507.04598Multi-Step Prediction and Control of Hierarchical EmotiarXiv2025TTStransformer-enc-dec2026-07-01
2507.04817Fast-VGAN: Lightweight Voice Conversion with Explicit CarXiv2025VCGAN2026-07-01
2507.01348SpeechAccentLLM: A Unified Framework for Foreign AccentarXiv2025VC, TTSautoregressive-LM, VAE, GAN2026-07-01
2507.06116Speech Quality Assessment Model Based on Mixture of ExpZhejiang UniversityarXiv2025evaluation2026-07-01
2506.23325XY-Tokenizer: Mitigating the Semantic-Acoustic ConflictFudan UniversityarXiv2025codecGAN, hybrid2026-07-01
2507.07799SecureSpeech: Prompt-based Speaker and Content ProtectiarXiv2025TTSautoregressive-LM2026-07-01
2507.08319Active Learning for Text-to-Speech Synthesis with InforarXiv2025TTStransformer-enc-dec2026-07-01
2507.09070SemAlignVC: Enhancing zero-shot timbre conversion usingMeta / KTH Royal Institute of TechnologyarXiv2025VCautoregressive-LM, flow-matching2026-07-01
2507.09282ClaritySpeech: Dementia Obfuscation in SpeechImperial College LondonarXiv2025TTSautoregressive-LM, diffusion, VAE2026-07-02
2507.09310Voice Conversion for Lombard Speaking Style with ImplicAmazon Alexa / Imperial College LondonarXiv2025VC, TTSVAE2026-07-02
2507.10985Pronunciation Deviation Analysis Through Voice Cloning California State University Long BeacharXiv2025TTS2026-07-02
2507.12197Quantize More, Lose Less: Autoregressive Generation froarXiv2025TTS, singingautoregressive-LM, GAN2026-07-02
2507.14988DMOSpeech 2: Reinforcement Learning for Duration PredicColumbia UniversityarXiv2025TTSflow-matching2026-07-02
2507.15272A2TTS: TTS for Low Resource Indian LanguagesarXiv2025TTSdiffusion2026-07-02
2507.16875Technical report: Impact of Duration Prediction on SpeaarXiv2025TTSflow-matching2026-07-02
2507.21138TTS-1 Technical ReportInworld AIarXiv2025TTSautoregressive-LM2026-07-02
2507.18119GOAT-SLM: A Spoken Language Model with Paralinguistic aTeleAI, China TelecomarXiv2025SCAautoregressive-LM, flow-matching2026-07-02
2507.18897HH-Codec: High Compression High-fidelity Discrete NeuraarXiv2025codecGAN2026-07-02
2507.17527Seed LiveInterpret 2.0: End-to-end Simultaneous Speech-ByteDancearXiv2025TTS, VCautoregressive-LM2026-07-02
2507.20140Do Not Mimic My Voice: Speaker Identity Unlearning for arXiv2025TTSflow-matching2026-07-02
2507.20731Learning Neural Vocoder from Range-Null Space DecomposiarXiv2025TTSGAN2026-07-02
2025.ccl-1.77HFSD-V2C: Zero-Shot Visual Voice Cloning Via Hierarchicworkshop2025TTSdiffusion2026-07-02
2025.icnlsp-1.34Beyond Labeled Datasets: Advancing TTS with Direct Prefworkshop2025TTSautoregressive-LM2026-07-02
2025.sigdial-1.21Transition Relevance Point Detection for Spoken Dialoguworkshop2025SCAhybrid2026-07-02
2025.sigdial-1.27EmoNews: A Spoken Dialogue System for Expressive News Cworkshop2025SCA, TTStransformer-enc-dec2026-07-02
2025.sigdial-1.51rrSDS 2.0: Incremental, Modular, Distributed, Multimodaworkshop2025SCA2026-07-02
interspeech-2025-0166Frozen Large Language Models Can Perceive ParalinguistiInterspeech2025SCAtransformer-enc-dec2026-07-02
interspeech-2025-0305DAFMSVC: One-Shot Singing Voice Conversion with Dual AtInterspeech2025singing, VCflow-matching, transformer-enc-dec2026-07-02
interspeech-2025-0347PeriodCodec: A Pitch-Controllable Neural Audio Codec UsInterspeech2025codec, singingGAN, VAE2026-07-02
interspeech-2025-0355Probing the Robustness Properties of Neural Speech CodeInterspeech2025codec, evaluation2026-07-02
interspeech-2025-0383Voice Conversion for Likability Control via Automated RInterspeech2025VCtransformer-enc-dec2026-07-02
interspeech-2025-0433When Humans Growl and Birds Speak: High-Fidelity Voice Interspeech2025VCVAE2026-07-02
interspeech-2025-0438LinearVC: Linear Transformations of Self-Supervised FeaInterspeech2025VChybrid2026-07-02
interspeech-2025-0464Prosody-Adaptable Audio Codecs for Zero-Shot Voice ConvInterspeech2025VC, codecautoregressive-LM2026-07-02
interspeech-2025-0506EnCodecMAE: leveraging neural codecs for universal audiInterspeech2025codectransformer-enc-dec2026-07-02
interspeech-2025-0656EEG-based Voice Conversion : Hearing the Voice of Your Beijing University of Posts and TelecommunicationsInterspeech2025VChybrid2026-07-02
interspeech-2025-0706Contextual Paralinguistic Data Creation for Multi-ModalInterspeech2025SCA2026-07-02
interspeech-2025-0756A-SMiLE: Affective Sparse Mixture-of-Experts Adapter wiInterspeech2025SCAhybrid2026-07-02
interspeech-2025-0998Voice-ENHANCE: Speech Restoration using a Diffusion-basInterspeech2025VCdiffusion, GAN2026-07-02
interspeech-2025-1020Learning Optimal Prosody Embedding Codebook based on F0Interspeech2025TTS, evaluationVAE2026-07-03
interspeech-2025-1081Speaker Normalization and Content Restoration for Zero-Interspeech2025VCGAN2026-07-03
interspeech-2025-1084Efficient Streaming TTS Acoustic Model with Depthwise RInterspeech2025TTSautoregressive-LM2026-07-03
interspeech-2025-1098GST-BERT-TTS: Prosody Prediction Without Accentual LabeInterspeech2025TTStransformer-enc-dec2026-07-03
interspeech-2025-1106LSCodec: Low-Bitrate and Speaker-Decoupled Discrete SpeInterspeech2025codecVAE, GAN2026-07-03
interspeech-2025-1115MPE-TTS: Customized Emotion Zero-Shot Text-To-Speech UsInterspeech2025TTSdiffusion, autoregressive-LM2026-07-03
interspeech-2025-1192Voice Impression Control in Zero-Shot TTSInterspeech2025TTStransformer-enc-dec2026-07-03
interspeech-2025-1210DiffEmotionVC: A Dual-Granularity Disentangled DiffusioInterspeech2025VCdiffusion2026-07-03
interspeech-2025-1229E2E-BPVC: End-to-End Background-Preserving Voice ConverInterspeech2025VCflow-matching2026-07-03
interspeech-2025-1236Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality AlignmentInterspeech2025TTSflow-matching2026-07-03
interspeech-2025-1334MiSTR: Multi-Modal iEEG-to-Speech Synthesis with TransfInterspeech2025TTStransformer-enc-dec2026-07-03
interspeech-2025-1364VS-Singer: Vision-Guided Stereo Singing Voice SynthesisInterspeech2025singing, TTSdiffusion2026-07-03
interspeech-2025-1394DiEmo-TTS: Disentangled Emotion Representations via SelInterspeech2025TTStransformer-enc-dec2026-07-03
interspeech-2025-1397VibE-SVC: Vibrato Extraction with High-frequency F0 ConKorea UniversityInterspeech2025singing, VCdiffusion2026-07-03
interspeech-2025-1434REWIND: Speech Time Reversal for Enhancing Speaker ReprInterspeech2025VCdiffusion2026-07-03
interspeech-2025-1478LightL2S: Ultra-Low Complexity Lip-to-Speech Synthesis Interspeech2025TTShybrid2026-07-03
interspeech-2025-1494VisualSpeech: Enhancing Prosody Modeling in TTS Using VInterspeech2025TTStransformer-enc-dec2026-07-03
interspeech-2025-1531Simple and Effective Content Encoder for Singing Voice Interspeech2025singing, VCVAE, GAN2026-07-03
interspeech-2025-1536Fairness in Dysarthric Speech Synthesis: Understanding Interspeech2025TTS, evaluationflow-matching2026-07-03
interspeech-2025-1538StarVC: A Unified Auto-Regressive Framework for Joint TInterspeech2025VCautoregressive-LM2026-07-03
interspeech-2025-1550ArVoice: A Multi-Speaker Dataset for Arabic Speech SyntMohamed Bin Zayed University of Artificial IntelligenceInterspeech2025TTS, VC, evaluationtransformer-enc-dec, VAE, GAN2026-07-03
interspeech-2025-1625Mimic Blocker: Self-Supervised Adversarial Training forInterspeech2025VCGAN2026-07-03
interspeech-2025-1638EATS-Speech: Emotion-Adaptive Transformation and PrioriInterspeech2025TTShybrid2026-07-03
interspeech-2025-1639LombardTokenizer: Disentanglement and Control of Vocal GIPSA-lab, Univ. Grenoble AlpesInterspeech2025codec, VCGAN2026-07-03
interspeech-2025-1684SA-RAS: Speaker-Aware Style Retrieval Augmented GeneratInterspeech2025TTShybrid2026-07-03
interspeech-2025-1726Voice Reconstruction through Large-Scale TTS Models: CoInterspeech2025TTS, evaluationhybrid2026-07-03
interspeech-2025-1747FasterVoiceGrad: Faster One-step Diffusion-Based Voice NTT, Inc.Interspeech2025VCdiffusion, GAN2026-07-03
interspeech-2025-1763Vocoder-Projected Feature DiscriminatorNTTInterspeech2025VCGAN, diffusion2026-07-03
interspeech-2025-1776SpeechSEC: A Unified Multi-Task Framework for Speech SyInterspeech2025TTShybrid2026-07-04
interspeech-2025-1819Comparative Analysis of Fast and High-Fidelity Neural VInterspeech2025TTSGAN2026-07-04
interspeech-2025-1873Can AI Understand Mandarin Speech Prosody? A FrameworkInterspeech2025SCA, evaluation2026-07-04
interspeech-2025-1940Investigating Stochastic Methods for Prosody Modeling iInterspeech2025TTStransformer-enc-dec, flow-matching2026-07-04
interspeech-2025-2031Kinship in Speech: Leveraging Linguistic Relatedness foInterspeech2025TTStransformer-enc-dec2026-07-04
interspeech-2025-2032ExagTTS: An Approach Towards Controllable Word Stress IIIIT HyderabadInterspeech2025TTShybrid2026-07-04
interspeech-2025-2075Segmentation-Variant Codebooks for Preservation of ParaInterspeech2025codec2026-07-04
interspeech-2025-2151FaVC: A Validated, Transcribed, Parallel Farsi Speech DUniversity of TehranInterspeech2025VC, evaluationGAN2026-07-04
interspeech-2025-2159Generating Consistent Prosodic Patterns from Open-SourcInterspeech2025TTS, evaluationflow-matching2026-07-04
interspeech-2025-2189ProMode: A Speech Prosody Model Conditioned on AcousticInterspeech2025TTStransformer-enc-dec2026-07-04
interspeech-2025-2283Pairwise Evaluation of Accent Similarity in Speech SyntInterspeech2025TTS, evaluation2026-07-04
interspeech-2025-2328A Watermark for Auto-Regressive Speech Generation ModelUniversity of MarylandInterspeech2025TTS, evaluationautoregressive-LM2026-07-05
interspeech-2025-2536The Text-to-speech in the Wild (TITW) DatabaseInterspeech2025TTS, evaluation2026-07-05
interspeech-2025-2564Towards a Japanese Full-duplex Spoken Dialogue SystemNagoya UniversityInterspeech2025SCAautoregressive-LM2026-07-05
interspeech-2025-2573SawtArabi: A Benchmark Corpus for Arabic TTS. Standard, Dialectal and Code-SwitchingInterspeech2025TTS, evaluationflow-matching, GAN2026-07-05
interspeech-2025-2586Spotlight-TTS: Spotlighting the Style via Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-SpeechKorea UniversityInterspeech2025TTStransformer-enc-dec, GAN2026-07-05
interspeech-2025-2595Harnessing Text-to-Speech Voice Cloning Models for Improved Audiological Speech AssessmentUniversity of CambridgeInterspeech2025TTS, evaluation2026-07-05
interspeech-2025-2679Can We Reconstruct a Dysarthric Voice with the Large Speech Model Parler TTS?University of EdinburghInterspeech2025TTSautoregressive-LM2026-07-05
interspeech-2025-2684Discl-VC: Disentangled Discrete Tokens and In-Context Learning for Controllable Zero-Shot Voice ConversionInterspeech2025VCflow-matching2026-07-05
interspeech-2025-2726DS-Codec: Dual-Stage Training with Mirror-to-NonMirror Architecture Switching for Speech CodecInterspeech2025codecGAN, hybrid2026-07-05
interspeech-2025-2739AF-Vocoder: Artifact-Free Neural Vocoder with Global Artifact FilterByteDanceInterspeech2025TTSGAN2026-07-05
interspeech-2025-2787Towards Adaptable and Intelligible Speech Synthesis in Noisy EnvironmentsKTH Royal Institute of TechnologyInterspeech2025TTS, evaluationautoregressive-LM2026-07-05
interspeech-2025-2815From Static to Dynamic: Enhancing AAC with Generative Imagery and Zero-Shot TTSKTH Royal Institute of TechnologyInterspeech2025TTS2026-07-05
interspeech-2025-bokkahallisatish25_interspeechHear Me Out: Interactive evaluation and bias discovery platform for speech-to-speech conversational AIKTH Royal Institute of TechnologyInterspeech2025SCA, evaluation2026-07-05
2507.16835Evaluating Speech-to-Text x LLM x Text-to-Speech Combinations for AI Interview SystemsarXiv2025SCA, evaluation2026-07-05
2411.19770Noro: Noise-Robust One-shot Voice Conversion with Hidden Speaker Representation LearningarXiv2025VCdiffusion2026-07-05
2025.clicit-1.27Veras Audire Et Reddere Voces: A Corpus of Prosodically-Correct Latin Poetic Audio from Large-Language-Model TTSworkshop2025TTS, evaluationautoregressive-LM2026-07-05
2506.23367You Sound a Little Tense: L2 Tailored Clear TTS Using Durational Vowel PropertiesarXiv2025TTSflow-matching2026-07-12
2509.05359An Empirical Analysis of Discrete Unit Representations in Speech Language Modeling Pre-trainingarXiv2025SCAautoregressive-LM2026-07-12
2509.04093Open-Source Full-Duplex Conversational Datasets for Natural and Interactive Speech SynthesisarXiv2025TTS, SCAautoregressive-LM, flow-matching2026-07-12
2509.04667DarkStream: real-time speech anonymization with low latencyTexas A&M UniversityarXiv2025VCGAN, hybrid2026-07-12
2509.04685Say More with Less: Variable-Frame-Rate Speech Tokenization via Adaptive Clustering and Implicit Duration CodingarXiv2025codecGAN2026-07-12
2509.04702OleSpeech-IV: A Large-Scale Multispeaker and Multilingual Conversational Speech Dataset with Diverse TopicsOlewavearXiv2025TTS, SCA2026-07-12
2509.05863LatinX: Aligning a Multilingual TTS Model with Direct Preference OptimizationarXiv2025TTSautoregressive-LM2026-07-12
2509.06074Multimodal Fine-grained Context Interaction Graph Modeling for Conversational Speech SynthesisEMNLP2025TTStransformer-enc-dec2026-07-12
2509.06502FireRedChat: A Pluggable, Full-Duplex Voice Interaction System with Cascaded and Semi-Cascaded ImplementationsXiaohongshuarXiv2025SCAhybrid2026-07-12
2509.07038Controllable Singing Voice Synthesis using Phoneme-Level Energy SequenceKorea UniversityarXiv2025singingdiffusion2026-07-12
2509.07376Progressive Facial Granularity Aggregation with Bilateral Attribute-based Enhancement for Face-to-Speech SynthesisPOSTECHEMNLP2025TTSVAE2026-07-12
2509.09716VStyle: A Benchmark for Voice Style Adaptation with Spoken InstructionsarXiv2025TTS, evaluation2026-07-12
2509.08379Flow-Matching ModelsarXiv2025VCdiffusion, flow-matching2026-07-12
2509.08696Accelerating Diffusion Transformer-Based Text-to-Speech with Transformer Layer CachingNational University of SingaporearXiv2025TTSflow-matching2026-07-12
2506.04077A Novel Data Augmentation Approach for Automatic Speaking Assessment on Opinion ExpressionsNational Taiwan Normal UniversityarXiv2025TTSautoregressive-LM2026-07-12
2509.09174EchoX: Towards Mitigating Acoustic-Semantic Gap via Echo Training for Speech-to-Speech LLMsThe Chinese University of Hong Kong, ShenzhenarXiv2025SCAautoregressive-LM2026-07-12
2509.09201DeCodec: Rethinking Audio Codecs as Universal Disentangled Representation LearnersChina MobilearXiv2025codecGAN2026-07-13
2509.09550Finite Scalar Quantization Enables Redundant and Transmission-Robust Neural Audio Compression at Low Bit-ratesNeuphonicarXiv2025codechybrid, GAN2026-07-13
2509.09748DiTReducio: A Training-Free Acceleration for DiT-Based TTS via Progressive CalibrationarXiv2025TTSflow-matching2026-07-13
2509.11084Length-Aware Rotary Position Embedding for Text-Speech AlignmentSupertone, Inc.arXiv2025TTSflow-matching2026-07-13
2509.11425FuseCodec: Semantic-Contextual Fusion and Supervision for Neural CodecsarXiv2025codec, TTSGAN, autoregressive-LM2026-07-13
2508.18240MTalk-Bench: Evaluating Speech-to-Speech Models in Multi-Turn Dialogues via Arena-style and Rubrics ProtocolsThe Chinese University of Hong Kong, ShenzhenarXiv2025SCA, evaluation2026-07-13
2509.12171Preservation of Language Understanding Capabilities in Speech-aware Large Language ModelsarXiv2025SCA, evaluationflow-matching2026-07-13
2509.14270SpeechWeave: Diverse Multilingual Synthetic Text & Audio Data Generation Pipeline for Training Text to Speech ModelsOracle AIACL2025TTS2026-07-13
2509.12831A Lightweight Pipeline for Noisy Speech Voice Cloning and Accurate Lip Sync SynthesisInternational Islamic University, IslamabadarXiv2025TTShybrid, GAN2026-07-13
2509.13068MSR-Codec: A Low-Bitrate Multi-Stream Residual Codec for High-Fidelity Speech Generation with Information DisentanglementLIGHTSPEEDarXiv2025TTS, codec, VCautoregressive-LM, VAE2026-07-13
2412.16846KALL-E: Autoregressive Speech Synthesis with Next-Distribution PredictionNorthwestern Polytechnical UniversityarXiv2025TTSautoregressive-LM, VAE2026-07-13
2504.20581ClonEval: An Open Voice Cloning BenchmarkAdam Mickiewicz UniversityarXiv2025TTS, evaluation2026-07-13
interspeech-2025-cho25c_interspeechUnleashing the Inner Monster: Demonstrating High-Fidelity Human to Non-Human Voice ConversionNC AI Co., LtdInterspeech2025VChybrid2026-07-13
interspeech-2025-gourav25_interspeechCode Mix TTS: An Approach to Infer Human Like Speech for Multi-Lingual Input TextsOracle CorporationInterspeech2025TTSdiffusion, GAN2026-07-13
interspeech-2025-raju25_interspeechEnd-to-End Indian Language Dubbing with Zero-Shot Speaker PreservationHitloopInterspeech2025TTSflow-matching2026-07-13
2509.13667A Distilled Low-Latency Neural Vocoder with Explicit Amplitude and Phase PredictionarXiv2025TTSGAN2026-07-13
2509.13670A High-Quality and Low-Complexity Streamable Neural Speech Codec with Knowledge DistillationUniversity of Science and Technology of ChinaarXiv2025codecGAN2026-07-13
2509.13989Do You Hear What I Mean? Quantifying the Instruction-Perception Gap in Instruction-Guided Expressive Text-To-Speech SystemsarXiv2025TTS, evaluation2026-07-14
2509.14579Cross-Lingual F5-TTS: Towards Language-Agnostic Voice Cloning and Speech SynthesisShanghai Jiao Tong UniversityarXiv2025TTS, VCflow-matching2026-07-14
2509.14684DAIEN-TTS: Disentangled Audio Infilling for Environment-Aware Text-to-Speech SynthesisarXiv2025TTSflow-matching2026-07-14
2509.14784MELA-TTS: Joint transformer-diffusion model with representation alignment for speech synthesisarXiv2025TTShybrid2026-07-14
2509.14946SynParaSpeech: Automated Synthesis of Paralinguistic Datasets for Speech Generation and UnderstandingarXiv2025TTSautoregressive-LM, flow-matching2026-07-14
2509.15085Real-Time Streaming Mel Vocoding with Generative Flow MatchingUniversity of HamburgarXiv2025TTSflow-matching2026-07-14
2509.15253Emotion-Aware Speech Generation with Character-Specific Voices for ComicsQueen Mary University of LondonarXiv2025TTS2026-07-14
2509.15462A Novel Semantic Compression Approach for Ultra-low Bandwidth Voice CommunicationSystems & Technology ResearcharXiv2025codec, VCautoregressive-LM, flow-matching2026-07-14
2505.17093P2VA: Converting Persona Descriptions into Voice Attributes for Fair and Controllable Text-to-SpeecharXiv2025TTS2026-07-14
2509.15626LibriTTS-VI: A Public Corpus and Novel Methods for Efficient Voice Impression ControlSony Group CorporationarXiv2025TTSVAE2026-07-14
2509.15629The Singing Voice Conversion Challenge 2025: From Singer Identity Conversion To Singing Style ConversionarXiv2025VC, singing, evaluation2026-07-14
2509.15845Deep Dubbing: End-to-End Auto-Audiobook System with Text-to-Timbre and Context-Aware Instruct-TTSarXiv2025TTSflow-matching, autoregressive-LM2026-07-14
2509.16010Fed-PISA: Federated Voice Cloning via Personalized Identity-Style AdaptationarXiv2025TTS, VChybrid2026-07-14
2509.16195FocalCodec-Stream: Streaming Low-Bitrate Speech Coding via Causal DistillationarXiv2025codec, VChybrid2026-07-14
2509.16589Benchmarking Contextual and Paralinguistic Reasoning in Speech-LLMs: A Case Study with In-the-Wild DataEMNLP2025SCA, evaluation2026-07-14
2509.20378Beyond Global Emotion: Fine-Grained Emotional Speech Synthesis with Dynamic Word-Level ModulationHarbin Institute of TechnologyarXiv2025TTSautoregressive-LM2026-07-14
2509.17006MBCodec: Thorough Disentangle for High-Fidelity Audio CompressionarXiv2025codecGAN2026-07-14
2509.17021Bridging the gap between training and inference in LM-based TTS modelsarXiv2025TTSautoregressive-LM2026-07-14
2509.17143MaskVCT: Masked Voice Codec Transformer for Zero-Shot Voice Conversion With Increased Controllability via Multiple GuidancesarXiv2025VCautoregressive-LM2026-07-14
2509.14882Llama-Mimi: Exploring the Limits of Flattened Speech Language ModelingarXiv2025SCAautoregressive-LM2026-07-14
2509.17516Audiobook-CC: Controllable Long-context Speech Generation for Multicast AudiobookXimalaya Inc.arXiv2025TTSautoregressive-LM, flow-matching, GAN2026-07-15
2509.17765Qwen3-Omni Technical ReportQwen Team (Alibaba)arXiv2025SCA, TTSautoregressive-LM, hybrid2026-07-15
2509.17988Nord-Parl-TTS: Finnish and Swedish TTS Dataset from Parliament SpeecharXiv2025TTSflow-matching2026-07-15
2509.18060TMD-TTS: A Unified Tibetan Multi-Dialect Text-to-Speech Framework for Ü-Tsang, Amdo and Kham Speech Dataset GenerationUniversity of Electronic Science and Technology of ChinaarXiv2025TTSflow-matching2026-07-15
2509.18470Discrete-Time Diffusion-Like Models for Speech SynthesisUniversity of SheffieldarXiv2025TTSdiffusion, flow-matching2026-07-15
2501.04561OpenOmni: Advancing Open-Source Omnimodal Large Language Models with Progressive Multimodal Alignment and Real-Time Self-Aware Emotional Speech SynthesisShenzhen Institute of Advanced Technology, CAS; Alibaba (Tongyi Lab)arXiv2025SCA, TTSautoregressive-LM, hybrid2026-07-15
2509.18531No Verifiable Reward for Prosody: Toward Preference-Guided Prosody Learning in TTSarXiv2025TTSautoregressive-LM2026-07-15
2509.18806Rethinking the joint estimation of magnitude and phase for time-frequency domain neural vocodersInstitute of Acoustics, Chinese Academy of Sciences; Tencent AI LabarXiv2025TTSGAN2026-07-15
2509.18823Towards Evaluating Generative Audio: Insights from Neural Audio Codec Embedding DistancesDolbyarXiv2025evaluation, codecGAN2026-07-15
2509.18928Direct Preference Optimization for Speech Autoregressive Diffusion ModelsByteDance SeedarXiv2025TTSautoregressive-LM, diffusion, hybrid2026-07-15
2509.19025Enhancing Noise Robustness for Neural Speech Codecs through Resource-Efficient Progressive Quantization Perturbation SimulationUniversity of Science and Technology of ChinaarXiv2025codecGAN2026-07-15
2509.19186Improving Test-Time Performance of RVQ-based Neural CodecsSupertone Inc.arXiv2025codecGAN2026-07-15
2509.19231Finding My Voice: Generative Reconstruction of Disordered Speech for Automated Clinical EvaluationCarnegie Mellon UniversityarXiv2025TTS, VC, evaluationdiffusion, GAN2026-07-15
2509.19592Frame-Stacked Local Transformers For Efficient Multi-Codebook Speech GenerationNVIDIAarXiv2025TTSautoregressive-LM2026-07-15
2509.19812Efficient Speech Watermarking for Speech Synthesis via Progressive Knowledge DistillationMicrosoftarXiv2025TTSGAN2026-07-15
2509.19883CoMelSinger: Discrete Token-Based Zero-Shot Singing Synthesis With Structured Melody Control and GuidanceNational University of SingaporearXiv2025singinghybrid2026-07-15
2509.19928Measuring Prosody Diversity in Zero-Shot TTS: A New Metric, Benchmark, and ExplorationarXiv2025TTS, evaluation2026-07-16
2509.20086OLaPh: Optimal Language PhonemizerHof University of Applied SciencesarXiv2025TTSautoregressive-LM2026-07-16
2509.20321Conversational Speech Reveals Structural Robustness Failures in SpeechLLM BackbonesTexas A&M UniversityarXiv2025SCA, evaluationautoregressive-LM2026-07-16
2509.20410Phoenix-VAD: Streaming Semantic Endpoint Detection for Full-Duplex Speech InteractionXiamen University, DiDi GlobalarXiv2025SCAautoregressive-LM2026-07-16
2509.20485Objective Evaluation of Prosody and Intelligibility in Speech Synthesis via Conditional Prediction of Discrete TokensJohns Hopkins University; National University of SingaporearXiv2025evaluation, TTS2026-07-16
2509.22718PerformSinger: Multimodal Singing Voice Synthesis Leveraging Synchronized Lip Cues from Singing Performance VideosarXiv2025singinghybrid2026-07-16
2505.10599UDDETTS: Unifying Discrete and Dimensional Emotions for Controllable Emotional Text-to-SpeecharXiv2025TTSautoregressive-LM, flow-matching, GAN, hybrid2026-07-16
2509.20802SPADE: Structured Pruning and Adaptive Distillation for Efficient LLM-TTSKAIST, 42dot Inc.arXiv2025TTSautoregressive-LM2026-07-16
2509.22727DiaMoE-TTS: A Unified IPA-Based Dialect TTS Framework with Mixture-of-Experts and Parameter-Efficient Zero-Shot AdaptationTsinghua University, Giant Network AI LabarXiv2025TTSflow-matching2026-07-16
2506.21875WildSpeech-Bench: Benchmarking End-to-End SpeechLLMs in the WildWeChat AI, TencentarXiv2025SCA, evaluation2026-07-16
2509.21968AUV: Teaching Audio Universal Vector Quantization with Single Nested CodebookarXiv2025codecGAN, VAE2026-07-16
2509.22062Comprehend and Talk: Text to Speech Synthesis via Dual Language ModelingAMAP Speech, Tsinghua UniversityarXiv2025TTSautoregressive-LM2026-07-17
2509.22167Semantic-VAE: Semantic-Alignment Latent Representation Shanghai Jiao Tong UniversityarXiv2025TTSVAE2026-07-17
2509.22243FLEXI: Benchmarking Full-duplex Human-LLM Speech InteractionNortheastern UniversityarXiv2025SCA, evaluation2026-07-17
2509.23147BFA: Real-time Multilingual Text-to-speech Forced AlignmentBournemouth UniversityarXiv2025TTS2026-07-17
2510.02352Evaluating Bias in Spoken Dialogue LLMs for Real-World arXiv2025SCA2026-07-17
2509.23938Easy Turn: Integrating Acoustic and Linguistic ModalitiarXiv2025SCAhybrid2026-07-17
2509.24457Assessing speech quality metrics for evaluation of neurCisco SystemsarXiv2025codec, evaluation2026-07-17
2509.24570ISSE: An Instruction-Guided Speech Style Editing Dataset And BenchmarkarXiv2025TTS, VC, evaluationautoregressive-LM2026-07-17
2509.24650VoxCPM: Tokenizer-Free TTS for Context-Aware Speech Generation and True-to-Life Voice CloningarXiv2025TTS, VCautoregressive-LM, diffusion2026-07-17
2509.24773VSSFlow: Unifying Video-conditioned Sound and Speech GearXiv2025TTSflow-matching2026-07-17
2509.25131MGM-Omni: Scaling Omni LLMs to Personalized Long-HorizoCUHK, HKUST, SmartMorearXiv2025SCA, TTSautoregressive-LM, flow-matching, hybrid2026-07-17
2509.25416Emotion-Aligned Generation in Diffusion Text to Speech Models via Preference-Guided OptimizationarXiv2025TTSdiffusion2026-07-17
2509.26276Optimizing Speech Language Models for Acoustic ConsistencyUniversity of ZuricharXiv2025TTS, SCAautoregressive-LM2026-07-17
2509.26514BatonVoice: An Operationalist Framework for Enhancing CTencentarXiv2025TTSautoregressive-LM, flow-matching, GAN2026-07-17
2509.26542Voice Evaluation of Reasoning Ability: Diagnosing the Modality-Induced Performance GaparXiv2025SCA, evaluation2026-07-17
2510.00264Baseline Systems For The 2025 Low-Resource Audio Codec ChallengeCisco Systems (Collaboration AI)arXiv2025codec, evaluationGAN2026-07-17
2510.00499MOSS-Speech: Towards True Speech-to-Speech Models Without Text GuidanceShanghai Innovation Institute, Fudan University, MOSIarXiv2025SCAautoregressive-LM, flow-matching2026-07-17
2510.00743From Scores to Preferences: Redefining MOS Benchmarking for Speech Quality Reward ModelingFudan UniversityarXiv2025evaluationautoregressive-LM2026-07-17
2025.vlsp-1.15Twinkle-VC: A Robust and High-Quality Zero-Shot Voice Conversion System for the VLSP 2025 Shared TaskVLSP 20252025VCdiffusion2026-07-17
2025.vlsp-1.14ViettelRoar: Voice conversion approach for VLSP 2025ViettelAI, Viettel GroupVLSP 20252025VCflow-matching2026-07-17
2025.vlsp-1.13The 2025 VLSP Task on Vietnamese Voice Conversion: Overview and Preliminary ResultsHanoi University of Science and TechnologyVLSP 20252025VC, evaluation2026-07-17
2510.05150Chronological Thinking in Full-Duplex Spoken Dialogue Language ModelsNanyang Technological University, StepFun, MilaarXiv2025SCAautoregressive-LM2026-07-17
2510.02066Chain-of-Thought Reasoning in Streaming Full-Duplex End-to-End Spoken Dialogue SystemsCarnegie Mellon University, Sony Group CorporationarXiv2025SCAautoregressive-LM2026-07-17
2510.01722Emotional Text-To-Speech Based on Mutual-Information-Guided Emotion-Timbre DisentanglementThe University of Tokyo / Institute of Science TokyoarXiv2025TTStransformer-enc-dec2026-07-17
2510.01903MelTok: 2D Tokenization for Single-Codebook Audio CompressionInternational Digital Economy Academy (IDEA)arXiv2025codecVAE, GAN2026-07-17
2510.02044Stream RAG: Instant and Accurate Spoken Dialogue Systems with Streaming Tool UsageMeta, Carnegie Mellon UniversityarXiv2025SCAautoregressive-LM2026-07-17
2510.03111Evaluation of preprocessing pipelines in the creation of in-the-wild TTS datasetsUniversidad Nacional de Tres de FebreroarXiv2025TTS, evaluation2026-07-17
2510.03735Soft Disentanglement in Frequency Bands for Neural Audio CodecsTélécom ParisarXiv2025codecGAN2026-07-17
2510.04738Speak, Edit, Repeat: High-Fidelity Voice Editing and Zero-Shot TTS with Cross-Attentive MambaMTS AI, ITMO UniversityarXiv2025TTSautoregressive-LM, hybrid2026-07-17
2510.05619Teaching Machines to Speak Using Articulatory ControlUC BerkeleyarXiv2025TTShybrid2026-07-17
2510.05984ECTSpeech: Enhancing Efficient Speech Synthesis via Easy Consistency TuningXinjiang UniversityarXiv2025TTSdiffusion, transformer-enc-dec2026-07-17
2510.05799Data-efficient Targeted Token-level Preference Optimization for LLM-based Text-to-SpeechSpiralAI Inc.arXiv2025TTSautoregressive-LM2026-07-17
2506.15556PredGen: Accelerated Inference of Large Language Models through Input-Time Speculation for Real-Time Speech InteractionUniversity of California, Los AngelesarXiv2025SCA, TTSautoregressive-LM2026-07-18
2510.07096Modeling Sarcastic Speech: Semantic and Prosodic Cues in a Speech Synthesis FrameworkUniversity of GroningenarXiv2025TTSGAN, VAE2026-07-18
2510.06917SHANKS: Simultaneous Hearing and Thinking for Spoken Language ModelsNational Taiwan University, MicrosoftarXiv2025SCAautoregressive-LM2026-07-18
2510.06927Position: Towards Responsible Evaluation for Text-to-SpeecharXiv2025TTS, evaluation2026-07-18
2510.07881CS3-Bench: Evaluating and Enhancing Speech-to-Speech LLMs for Mandarin-English Code-SwitchingShanghai Jiao Tong University, Ant GrouparXiv2025SCA, evaluationautoregressive-LM2026-07-18
2510.08373DialoSpeech: Dual-Speaker Dialogue Generation with LLM and Flow MatchingNorthwestern Polytechnical UniversityarXiv2025TTSautoregressive-LM, flow-matching2026-07-18
2510.08392MeanVC: Lightweight and Streaming Zero-Shot Voice Conversion via Mean FlowsNorthwestern Polytechnical University (ASLP@NPU)arXiv2025VCflow-matching, hybrid2026-07-18
2510.07978VoiceAgentBench: Are Voice Assistants ready for agentic tasks?Ola Electric / KrutrimarXiv2025SCA, evaluationautoregressive-LM2026-07-18
2510.09061O_O-VC: Synthetic Data-Driven One-to-One Alignment for Any-to-Any Voice ConversionVNPT AI / Hanoi University of Science and Technology / National Economics UniversityEMNLP2025VCVAE, GAN2026-07-18
2510.09016DiTSinger: Scaling Singing Voice Synthesis with Diffusion Transformer and Implicit AlignmentMigu Music, China Mobile Communications CorporationarXiv2025singingdiffusion2026-07-18
2506.12311Phonikud: Hebrew Grapheme-to-Phoneme Conversion for Real-Time Text-to-SpeechIndependent Researcher; Reichman University; Tel Aviv UniversityarXiv2025TTSGAN, diffusion2026-07-18
2510.09424The Speech-LLM Takes It All: A Truly Fully End-to-End Spoken Dialogue State Tracking ApproachOrange ResearcharXiv2025SCAautoregressive-LM2026-07-18
2510.09592Mind-Paced Speaking: A Dual-Brain Approach to Real-Time Reasoning in Spoken Language ModelsStepFunarXiv2025SCAautoregressive-LM2026-07-18
2510.09245SynthVC: Leveraging Synthetic Data for End-to-End Low Latency Streaming Voice ConversionNorthwestern Polytechnical University (ASLP@NPU)arXiv2025VCGAN2026-07-18
2510.10003MTP-S2UT: Enhancing Speech-to-Speech Translation Quality with Multi-token PredictionNortheastern University; NiuTrans Research; Kunming University of Science and TechnologyarXiv2025TTStransformer-enc-dec2026-07-18
2510.10774ParsVoice: A Large-Scale Multi-Speaker Persian Speech Corpus for Text-to-Speech SynthesisUniversity of TehranarXiv2025TTSautoregressive-LM, GAN2026-07-18
2510.11646BridgeCode: A Dual Speech Representation Paradigm for Autoregressive Zero-Shot Text-to-Speech SynthesisSouth China University of TechnologyarXiv2025TTSautoregressive-LM2026-07-18
2510.11124Perturbation Self-Supervised Representations for Cross-Lingual Emotion TTS: Stage-Wise Modeling of Emotion and SpeakerTianjin UniversityarXiv2025TTStransformer-enc-dec, GAN2026-07-18
2510.12964VCTR: A Transformer-Based Model for Non-parallel Voice ConversionIndependent ResearcherarXiv2025VCGAN, hybrid2026-07-18
2510.12995Continuous-Token Diffusion for Speaker-Referenced TTS in Multimodal LLMsAmazon AGIarXiv2025TTSautoregressive-LM, diffusion2026-07-18
2510.13221Acoustic Teleportation via Disentangled Neural Audio Codec RepresentationsFraunhofer IIS / International Audio Laboratories ErlangenarXiv2025codecGAN2026-07-18
2510.13293Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS ModelsAlibaba, Nanyang Technological UniversityarXiv2025TTSautoregressive-LM2026-07-18
2510.13194StressTransfer: Stress-Aware Speech-to-Speech Translation with Emphasis PreservationThe Chinese University of Hong Kong; Nara Institute of Science and TechnologyarXiv2025TTSautoregressive-LM2026-07-18
2510.15364LDCodec: A high quality neural audio codec with low-complexity decoderByteDancearXiv2025codecGAN2026-07-18
2510.15227LongCat-Audio-Codec: An Audio Tokenizer and Detokenizer Solution Designed for Speech Large Language ModelsMeituan (LongCat Team)arXiv2025codecGAN, hybrid2026-07-18
2510.16841SAC: Neural Speech Codec with Semantic-Acoustic Dual-Stream QuantizationShanghai Jiao Tong University (X-LANCE Lab), Soul AI LabarXiv2025codecGAN, VAE, autoregressive-LM2026-07-18
2510.16718U-Codec: Ultra Low Frame-rate Neural Speech Codec for Fast High-fidelity Speech GenerationPeking University / Tencent AI Lab / Tencent HunyuanarXiv2025codec, TTSautoregressive-LM, GAN2026-07-18
2503.06211Late Fusion and Multi-Level Fission Amplify Cross-Modal Transfer in Text-Speech LMsUniversité de Toulon (LIS) / University of CambridgearXiv2025SCA, TTSautoregressive-LM2026-07-18
2510.18308ParaStyleTTS: Toward Efficient and Robust Paralinguistic Style Control for Expressive Text-to-Speech GenerationUniversity of New South WalesarXiv2025TTSVAE, GAN2026-07-18
2506.23670Efficient Interleaved Speech Modeling through Knowledge DistillationNlpie Research; University of ZuricharXiv2025TTS, SCAautoregressive-LM2026-07-18
2510.19509Which Evaluation for Which Model? A Taxonomy for Speech Model AssessmentApplearXiv2025evaluation2026-07-18
2510.10785FAC-FACodec: Controllable Zero-Shot Foreign Accent Conversion with Factorized Speech CodecUniversity of Illinois Urbana-ChampaignarXiv2025VCdiffusion2026-07-18

640 items under this folder.