| 1904.02882 | LibriTTS: A Corpus Derived from LibriSpeech for Text-to-S | Google AI | arXiv | 2019 | TTS | | 2026-06-10 |
| 2403.03100 | NaturalSpeech 3: Zero-Shot Speech Synthesis with Factori | Microsoft | arXiv | 2024 | TTS, VC | diffusion, hybrid | 2026-06-10 |
| 2504.18425 | Kimi-Audio Technical Report | Moonshot AI | arXiv | 2025 | TTS, VC, SCA | autoregressive-LM, flow-matching | 2026-06-10 |
| 2204.02152 | UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 202 | University of Tokyo | arXiv | 2022 | evaluation | | 2026-06-10 |
| 2509.04072 | Computational Narrative Understanding for Expressive Te | — | arXiv | 2025 | TTS | autoregressive-LM, flow-matching | 2026-06-04 |
| 2508.15827 | Mini-Omni-Reasoner: Token-Level Thinking-in-Speaking in | — | arXiv | 2025 | SCA | autoregressive-LM | 2026-06-03 |
| interspeech-2025-0739 | FD-Bench: A Full-Duplex Benchmarking Pipeline Designed for Full Duplex Spoken Dialogue Systems | — | Interspeech | 2025 | SCA, evaluation | | 2026-06-03 |
| interspeech-2025-1993 | Defending Unauthorized Voice Cloning with Watermark-Aware Codecs | The Chinese University of Hong Kong | Interspeech | 2025 | TTS, VC | autoregressive-LM | 2026-06-03 |
| 2508.08095 | Dual Information Speech Language Models for Emotional Conversations | Mashang Consumer Finance Co., Ltd. | arXiv | 2025 | SCA | transformer-enc-dec | 2026-06-03 |
| interspeech-2025-0948 | PromptEVC: Controllable Emotional Voice Conversion with Natural Language Prompts | Southeast University | Interspeech | 2025 | VC | VAE, diffusion | 2026-06-03 |
| interspeech-2025-0203 | ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech | Kyushu University / University of Tokyo / EverestAI Ximalaya | Interspeech | 2025 | VC | flow-matching, transformer-enc-dec | 2026-06-02 |
| interspeech-2025-0196 | SPCODEC: Split and Prediction for Neural Speech Codec | Samsung | Interspeech | 2025 | codec | GAN | 2026-06-02 |
| 2503.04721 | Full-Duplex-Bench: A Benchmark to Evaluate Full-duplex Spoken Dialogue Models on Turn-taking Capabilities | — | arXiv | 2025 | SCA, evaluation | | 2026-06-02 |
| 2508.08715 | MultiGen: Child-Friendly Multilingual Speech Generator with LLMs | A*STAR Institute for Infocomm Research | arXiv | 2025 | TTS | autoregressive-LM, flow-matching, GAN | 2026-06-02 |
| 2508.09767 | UtterTune: LoRA-Based Target-Language Pronunciation Edit and Control in Multilingual Text-to-Speech | — | arXiv | 2025 | TTS | autoregressive-LM | 2026-06-02 |
| 2508.11326 | MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts | Kunlun Inc. | arXiv | 2025 | TTS | autoregressive-LM, diffusion | 2026-06-02 |
| 2504.12867 | EmoVoice: LLM-based Emotional Text-To-Speech Model with Freestyle Text Prompting | Shanghai Jiao Tong University / Tongyi Speech Lab | arXiv | 2025 | TTS | autoregressive-LM, flow-matching | 2026-06-02 |
| 2508.08961 | DualSpeechLM: Towards Unified Speech Understanding and Generation via Dual Speech Token Modeling | — | arXiv | 2025 | TTS, SCA, VC | autoregressive-LM | 2026-06-02 |
| 2508.08399 | Exploring Disentangled Neural Speech Codecs from Self-Supervised Representations | MERL / Mitsubishi Electric | arXiv | 2025 | codec, VC | VAE | 2026-06-02 |
| 2508.07711 | Is GAN Necessary for Mel-Spectrogram-based Neural Vocoder? | — | arXiv | 2025 | TTS | GAN | 2026-06-02 |
| 2508.07426 | Scalable Controllable Accented TTS | Johns Hopkins University | ASRU | 2025 | TTS | transformer-enc-dec, GAN, VAE | 2026-06-02 |
| 2508.07302 | XEmoRAG: Cross-Lingual Emotion Transfer with Controllable Intensity Using Retrieval-Augmented Generation | Northwestern Polytechnical University | arXiv | 2025 | TTS, VC | autoregressive-LM, flow-matching | 2026-06-02 |
| 2508.06890 | Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody | POSTECH | arXiv | 2025 | VC | GAN, transformer-enc-dec | 2026-06-02 |
| 2508.06870 | Text to Speech System for Meitei Mayek Script | — | arXiv | 2025 | TTS | transformer-enc-dec, GAN | 2026-06-02 |
| 2508.05385 | A Scalable Pipeline for Enabling Non-Verbal Speech Generation and Understanding | Tsinghua University | arXiv | 2025 | TTS, SCA | transformer-enc-dec | 2026-06-02 |
| 2508.14049 | MahaTTS: A Unified Framework for Multilingual Text-to-Speech Synthesis | Dubverse AI | arXiv | 2025 | TTS | autoregressive-LM, flow-matching | 2026-06-02 |
| 2508.04585 | UniTalker: Conversational Speech-Visual Synthesis | Inner Mongolia University | arXiv | 2025 | TTS, SCA | autoregressive-LM, flow-matching | 2026-06-02 |
| 2508.04996 | REF-VC: Robust, Expressive and Fast Zero-Shot Voice Conversion with Diffusion Transformers | Northwestern Polytechnical University | arXiv | 2025 | VC | flow-matching | 2026-06-02 |
| 2508.05207 | SpectroStream: A Versatile Neural Codec for General Audio | Google DeepMind | arXiv | 2025 | codec | GAN | 2026-06-02 |
| 2507.20091 | ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models | — | arXiv | 2025 | SCA, TTS | autoregressive-LM | 2026-06-02 |
| 2508.00317 | Advancing Speech Quality Assessment Through Scientific Challenges and Open-source Activities | Nagoya University | arXiv | 2025 | evaluation | | 2026-06-02 |
| 2507.22746 | Next Tokens Denoising for Speech Synthesis | Microsoft | arXiv | 2025 | TTS | hybrid | 2026-06-02 |
| 2508.01796 | Enhancing Spectrogram Realism in Singing Voice Synthesis via Explicit Bandwidth Extension Prior to Vocoder | Tsinghua University | arXiv | 2025 | singing, TTS | diffusion, GAN | 2026-06-02 |
| 2508.02013 | SpeechRole: A Large-Scale Dataset and Benchmark for Evaluating Speech Role-Playing Agents | Fudan University | arXiv | 2025 | SCA, evaluation | | 2026-06-02 |
| 2508.02849 | SecoustiCodec: Cross-Modal Aligned Streaming Single-Codebook Speech Codec | — | arXiv | 2025 | codec | VAE, transformer-enc-dec | 2026-06-02 |
| 2025.naacl-long.110 | WaveFM: A High-Fidelity and Efficient Vocoder Based on Flow Matching | Tsinghua University | NAACL | 2025 | TTS | flow-matching | 2026-05-30 |
| 2025.findings-acl.1051 | LLMVoX: Autoregressive Streaming Text-to-Speech Model for Any LLM | MBZUAI | ACL | 2025 | TTS, SCA | autoregressive-LM | 2026-05-30 |
| 2025.emnlp-main.180 | Scaling Rich Style-Prompted Text-to-Speech Datasets | UT Austin / NYU | EMNLP | 2025 | TTS, evaluation | autoregressive-LM | 2026-06-01 |
| 2507.09318 | ZipVoice-Dialog: Non-Autoregressive Spoken Dialogue Generation with Flow Matching | Xiaomi Corp. | arXiv | 2026 | TTS, SCA | flow-matching | 2026-05-30 |
| 2025.coling-main.518 | ProsodyFlow: High-fidelity Text-to-Speech through Conditional Flow Matching and Prosody Modeling with Large Speech Language Models | Zhejiang University | workshop | 2025 | TTS | flow-matching, hybrid | 2026-05-30 |
| interspeech-2025-0469 | Developing High-Quality TTS for Punjabi and Urdu: Benchmarking against MMS Models | University of Engineering and Technology, Lahore | Interspeech | 2025 | TTS, evaluation | transformer-enc-dec | 2026-05-30 |
| interspeech-2025-0854 | Bridging the Training–Inference Gap in TTS: Training Strategies for Robust Generative Postprocessing for Low-Resource Speakers | Fraunhofer IIS | Interspeech | 2025 | TTS | GAN, flow-matching, transformer-enc-dec | 2026-05-30 |
| interspeech-2025-0973 | A Dataset for Automatic Assessment of TTS Quality in Spanish | | Interspeech | 2025 | TTS, evaluation | | 2026-05-30 |
| interspeech-2025-0989 | HiFiTTS-2: A Large-Scale High Bandwidth Speech Dataset | NVIDIA | Interspeech | 2025 | TTS, evaluation | autoregressive-LM | 2026-05-30 |
| interspeech-2025-1034 | Non-Standard Accent TTS Support via Large Multi-Accent Frontend Pronunciation Knowledge Transfer | University of Edinburgh | Interspeech | 2025 | TTS | transformer-enc-dec | 2026-05-30 |
| interspeech-2025-0723 | Counterfactual Activation Editing for Post-hoc Prosody and Mispronunciation Correction in TTS Models | KAIST / Samsung Electronics | Interspeech | 2025 | TTS | transformer-enc-dec, VAE | 2026-05-30 |
| interspeech-2025-0754 | EME-TTS: Unlocking the Emphasis and Emotion Link in Speech Synthesis | UCAS Hangzhou | Interspeech | 2025 | TTS | transformer-enc-dec | 2026-05-30 |
| interspeech-2025-0762 | Intrasentential English in Swedish TTS: perceived English-accentedness | KTH / MTM | Interspeech | 2025 | TTS | flow-matching | 2026-05-30 |
| interspeech-2025-0779 | Intelligibility of Text-to-Speech Systems for Mathematical Expressions | Ericsson R&D | Interspeech | 2025 | TTS, evaluation | | 2026-05-30 |
| interspeech-2025-0787 | Gradual modeling of the Lombard effect by modifying speaker embeddings from a Text-To-Speech model | HEAD acoustics | Interspeech | 2025 | TTS | autoregressive-LM | 2026-05-30 |
| interspeech-2025-0575 | VoiceMark: Zero-Shot Voice Cloning-Resistant Watermarking Approach Leveraging Speaker-Specific Latents | Tsinghua University | Interspeech | 2025 | VC, TTS | VAE | 2026-05-30 |
| interspeech-2025-0596 | Facilitating Personalized TTS for Dysarthric Speakers Using Knowledge Anchoring and Curriculum Learning | POSTECH | Interspeech | 2025 | TTS | transformer-enc-dec | 2026-05-30 |
| interspeech-2025-0648 | MIKU-PAL: An Automated and Standardized Multimodal Method for Speech Paralinguistic and Affect Labeling | Fish Audio | Interspeech | 2025 | TTS, evaluation | | 2026-05-30 |
| interspeech-2025-0669 | PAST: Phonetic-Acoustic Speech Tokenizer | Hebrew University of Jerusalem | Interspeech | 2025 | codec, TTS | hybrid | 2026-05-30 |
| interspeech-2025-0704 | Differentiable Reward Optimization for LLM based TTS system | Alibaba Group | Interspeech | 2025 | TTS | autoregressive-LM, flow-matching | 2026-05-30 |
| interspeech-2025-0406 | Zero-Shot Mono-to-Binaural Speech Synthesis | Google | Interspeech | 2025 | TTS | GAN | 2026-05-30 |
| interspeech-2025-0408 | Improving User Impression of Spoken Dialogue Systems by Controlling Para-linguistic Expression Based on Intimacy | Tohoku University | Interspeech | 2025 | SCA, TTS | transformer-enc-dec, GAN | 2026-05-30 |
| interspeech-2025-0455 | APTTS: Adversarial Post-training in Latent Flow Matching for Fast and High-fidelity Text-to-Speech | LG AI Research | Interspeech | 2025 | TTS | flow-matching, VAE, hybrid | 2026-05-30 |
| interspeech-2025-0554 | RapFlow-TTS: Rapid and High-Fidelity Text-to-Speech with Improved Consistency Flow Matching | NAVER Cloud | Interspeech | 2025 | TTS | flow-matching, GAN | 2026-05-30 |
| interspeech-2025-0551 | Monotonic Attention for Robust Text-to-Speech Synthesis in Large Language Model Frameworks | Tencent | Interspeech | 2025 | TTS | autoregressive-LM | 2026-05-30 |
| interspeech-2025-0047 | Revival with Voice: Multi-modal Controllable Text-to-Speech Synthesis | Meta AI | Interspeech | 2025 | TTS | autoregressive-LM | 2026-05-30 |
| interspeech-2025-0063 | Fine-Tuning Text-to-Speech Diffusion Models Using Reinforcement Learning with Human Feedback | | Interspeech | 2025 | TTS | diffusion | 2026-05-30 |
| interspeech-2025-0143 | Multimodal Prosody Modeling: A Use Case for Multilingual Sentence Mode Prediction | Idiap Research Institute | Interspeech | 2025 | TTS, evaluation | | 2026-05-30 |
| interspeech-2025-0310 | Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models | | Interspeech | 2025 | TTS, codec | autoregressive-LM | 2026-05-30 |
| interspeech-2025-0319 | Improving Noise Robustness of LLM-based Zero-shot TTS via Discrete Acoustic Token Denoising | USTC | Interspeech | 2025 | TTS | autoregressive-LM | 2026-05-30 |
| 2509.02020 | FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot | | arXiv | 2025 | TTS, SCA | autoregressive-LM | 2026-05-26 |
| 2507.14534 | Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion | | arXiv | 2025 | VC | GAN, hybrid | 2026-05-26 |
| 2509.19668 | Selective Classifier-free Guidance for Zero-shot Text-to-speech | | arXiv | 2025 | TTS | flow-matching | 2026-05-26 |
| 2510.00981 | FlexiCodec: A Dynamic Neural Audio Codec for Low Frame Rates | | arXiv | 2025 | codec, TTS | hybrid | 2026-05-26 |
| 2412.17048 | Why Do Speech Language Models Fail to Generate Semantically Coherent Outputs? | | arXiv | 2026 | SCA | autoregressive-LM | 2026-05-26 |
| 2025.findings-emnlp.424 | InteractSpeech: A Speech Dialogue Interaction Corpus for Spoken Dialogue Model | | EMNLP | 2025 | SCA, evaluation | autoregressive-LM | 2026-05-26 |
| 2025.acl-demo.37 | RT-VC: Real-Time Zero-Shot Voice Conversion with Speech Articulatory Coding | UC Berkeley | ACL | 2025 | VC | hybrid | 2026-05-26 |
| 2025.acl-industry.42 | Scaling Under-Resourced TTS: A Data-Optimized Framework with Advanced Acoustic Modeling for Thai | Beijing Logic Intelligence Technology | ACL | 2025 | TTS | GAN, transformer-enc-dec | 2026-05-26 |
| 2025.acl-long.1043 | OZSpeech: One-step Zero-shot Speech Synthesis with Learned-Prior-Conditioned Flow Matching | FPT Software AI Center | ACL | 2025 | TTS | flow-matching, hybrid | 2026-05-26 |
| 2025.acl-long.1252 | Finding A Voice: Exploring the Potential of African American Dialect and Voice Generation for Chatbots | Emory University | ACL | 2025 | TTS, SCA | hybrid | 2026-05-26 |
| 2025.acl-long.1471 | The time scale of redundancy between prosody and linguistic context | MIT | ACL | 2025 | evaluation | transformer-enc-dec | 2026-05-26 |
| 2025.acl-long.1498 | Analyzing and Mitigating Inconsistency in Discrete Speech Tokens for Neural Codec Language Models | Alibaba Group / Zhejiang University | ACL | 2025 | TTS, codec | autoregressive-LM, GAN | 2026-05-26 |
| 2025.acl-long.313 | F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching | Shanghai Jiao Tong University | ACL | 2025 | TTS | flow-matching | 2026-06-01 |
| 2025.acl-long.346 | ControlSpeech: Towards Simultaneous and Independent Zer | Zhejiang University / Alibaba Tongyi Speech Lab | ACL | 2025 | TTS | transformer-enc-dec, hybrid | 2026-05-26 |
| 2025.acl-long.388 | Distilling an End-to-End Voice Assistant Without Instru | | ACL | 2025 | SCA | transformer-enc-dec | 2026-05-26 |
| 2025.acl-long.598 | Advancing Zero-shot Text-to-Speech Intelligibility acro | | ACL | 2025 | TTS | autoregressive-LM, flow-matching, hybrid | 2026-05-26 |
| interspeech-2025-0253 | Long-Context Speech Synthesis with Context-Aware Memory | South China University of Technology / Alibaba Group | Interspeech | 2025 | TTS | autoregressive-LM, hybrid | 2026-05-27 |
| 2301.02111 | Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers | Microsoft | arXiv | 2023 | TTS, codec | autoregressive-LM | 2026-06-01 |
| interspeech-2025-0902 | VoiceQualityVC: A Voice Conversion System for Studying the Perceptual Effects of Voice Quality in Speech | KTH Royal Institute of Technology | Interspeech | 2025 | VC | VAE, GAN | 2026-05-27 |
| 2025.emnlp-main.989 | VocalNet: Speech LLMs with Multi-Token Prediction for Faster and High-Quality Generation | SJTU / Ant Group / Wuhan University | EMNLP | 2025 | TTS, SCA | autoregressive-LM, hybrid | 2026-05-27 |
| 2025.acl-long.682 | Recent Advances in Speech Language Models: A Survey | CUHK / Tencent / NUS | ACL | 2025 | TTS, SCA, evaluation | autoregressive-LM, hybrid | 2026-06-01 |
| 2025.americasnlp-1.1 | Text-to-speech system for low-resource languages: A case study in Shipibo-Konibo | Pontificia Universidad Católica del Perú | workshop | 2025 | TTS | transformer-enc-dec, GAN | 2026-05-27 |
| 2025.emnlp-main.1730 | FillerSpeech: Towards Human-Like Text-to-Speech Synthesis with Filler Insertion and Filler Style Control | Korea University / Samsung Research | EMNLP | 2025 | TTS | flow-matching, transformer-enc-dec | 2026-05-27 |
| 2025.findings-naacl.184 | Continuous Speech Tokenizer in Text To Speech | CUHK / Tencent | NAACL | 2025 | TTS, codec | autoregressive-LM, VAE, flow-matching | 2026-05-27 |
| 2025.emnlp-demos.70 | OpenS2S: Advancing Fully Open-Source End-to-End Empathetic Large Speech Language Model | Institute of Automation, Chinese Academy of Sciences | EMNLP | 2025 | SCA | autoregressive-LM, hybrid | 2026-05-27 |
| 2406.02430 | Seed-TTS: A Family of High-Quality Versatile Speech Generation Models | ByteDance | arXiv | 2024 | TTS, VC | autoregressive-LM, diffusion, hybrid | 2026-05-28 |
| 2407.05407 | CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens | Alibaba Group | arXiv | 2024 | TTS | autoregressive-LM, flow-matching, hybrid | 2026-05-28 |
| 2025.acl-long.65 | Autoregressive Speech Synthesis without Vector Quantization | Microsoft / CUHK | ACL | 2025 | TTS | autoregressive-LM, VAE, hybrid | 2026-05-28 |
| 2412.10117 | CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models | Alibaba Group | arXiv | 2024 | TTS | autoregressive-LM, flow-matching, hybrid | 2026-05-28 |
| 2601.15621 | Qwen3-TTS Technical Report | Alibaba / Qwen Team | arXiv | 2026 | TTS | autoregressive-LM, hybrid | 2026-05-28 |
| 2512.14291 | GLM-TTS Technical Report | Zhipu AI / Tsinghua University | arXiv | 2025 | TTS | autoregressive-LM, diffusion, GAN, hybrid | 2026-05-28 |
| 2508.06262 | Llasa+: Free Lunch for Accelerated and Streaming Llama-Based Speech Synthesis | Northwestern Polytechnical University / HKUST | arXiv | 2025 | TTS | autoregressive-LM, hybrid | 2026-05-28 |
| 2502.03930 | DiTAR: Diffusion Transformer Autoregressive Modeling for Speech Generation | ByteDance Seed | arXiv | 2025 | TTS | autoregressive-LM, diffusion, hybrid | 2026-05-28 |
| 2504.10352 | Pseudo-Autoregressive Neural Codec Language Models for Efficient Zero-Shot Text-to-Speech Synthesis | Microsoft / SJTU | arXiv | 2025 | TTS | autoregressive-LM, hybrid | 2026-05-28 |
| 2508.16332 | Vevo2: A Unified and Controllable Framework for Speech and Singing Voice Generation | CUHK Shenzhen / ByteDance Seed | arXiv | 2025 | TTS, VC, singing | autoregressive-LM, flow-matching, hybrid | 2026-05-28 |
| 2508.02038 | Marco-Voice Technical Report | Alibaba International Digital Commerce | arXiv | 2025 | TTS, VC | autoregressive-LM, flow-matching, hybrid | 2026-05-28 |
| 2604.00688 | OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models | Xiaomi Corp. | arXiv | 2026 | TTS | diffusion, hybrid | 2026-05-28 |
| 2508.03543 | EmoSteer-TTS: Fine-Grained and Training-Free Emotion-Controllable Text-to-Speech via Activation Steering | HKUST (Guangzhou) / Tencent AI Lab | arXiv | 2025 | TTS | flow-matching | 2026-05-28 |
| 2510.02848 | Flamed-TTS: Flow Matching Attention-Free Models for Efficient Generating and Dynamic Pacing Zero-shot Text-to-Speech | FPT Software AI Center | arXiv | 2025 | TTS | flow-matching, hybrid | 2026-05-28 |
| 2506.21619 | IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech | bilibili | arXiv | 2025 | TTS | autoregressive-LM, flow-matching, GAN, hybrid | 2026-05-28 |
| 2025.naacl-long.242 | StyleTTS-ZS: Efficient High-Quality Zero-Shot Text-to-Speech Synthesis with Distilled Time-Varying Style Diffusion | Columbia University | NAACL | 2025 | TTS | diffusion, GAN, VAE, hybrid | 2026-05-28 |
| 2510.12210 | DiSTAR: Diffusion over a Scalable Token Autoregressive Representation for Speech Generation | SJTU / ByteDance | arXiv | 2025 | TTS | autoregressive-LM, diffusion, hybrid | 2026-05-28 |
| 2025.emnlp-main.40 | Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey | HKUST-GZ / University of Surrey | EMNLP | 2025 | TTS, evaluation | autoregressive-LM, flow-matching, diffusion, GAN, VAE, transformer-enc-dec, hybrid | 2026-05-28 |
| 2603.08823 | Fish Audio S2 Technical Report | Fish Audio | arXiv | 2026 | TTS | autoregressive-LM, GAN, hybrid | 2026-05-28 |
| 2509.00685 | MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech | Northwestern Polytechnical University | arXiv | 2025 | TTS | autoregressive-LM | 2026-05-29 |
| 2511.12347 | VoiceCraft-X: Unifying Multilingual, Voice-Cloning Speech Synthesis and Speech Editing | University of Texas at Austin / Amazon | EMNLP | 2025 | TTS | autoregressive-LM | 2026-05-29 |
| 2512.13251 | DisCo-Speech: Controllable Zero-Shot Speech Generation | China Mobile Nineverse AI / Peking University | arXiv | 2025 | TTS, VC, codec | autoregressive-LM, GAN | 2026-05-29 |
| 2509.09631 | DiFlow-TTS: Compact and Low-Latency Zero-Shot Text-to-Speech with Factorized Discrete Flow Matching | FPT Software AI Center | arXiv | 2025 | TTS | flow-matching, transformer-enc-dec | 2026-05-29 |
| 2512.04720 | M3-TTS: Multi-modal DiT Alignment & Mel-latent for Zero-shot High-fidelity Speech | | arXiv | 2025 | TTS | diffusion, VAE | 2026-05-29 |
| 2603.29339 | LongCat-AudioDiT: High-Fidelity Diffusion Text-to-Speech in the Waveform Latent Space | Meituan | arXiv | 2026 | TTS | flow-matching, VAE | 2026-05-29 |
| 2508.11273 | EmoSSLSphere: Multilingual Emotional Speech Synthesis with Spherical Vectors and Discrete Speech Tokens | | arXiv | 2025 | TTS | transformer-enc-dec | 2026-05-29 |
| 2604.12438 | An Ultra-Low Latency, End-to-End Streaming Speech Synthesis Architecture via Block-Wise Generation and Depth-Wise Codec Decoding | | arXiv | 2026 | TTS | transformer-enc-dec | 2026-06-01 |
| 2604.01760 | T5Gemma-TTS Technical Report | | arXiv | 2026 | TTS | autoregressive-LM | 2026-05-29 |
| 2508.15442 | Mitigating Hallucinations in LM-Based TTS Models via Distribution Alignment Using GFlowNets | | EMNLP | 2025 | TTS | autoregressive-LM | 2026-05-29 |
| 2025.acl-long.654 | Language-Codec: Bridging Discrete Codec Representations and Speech Language Models | Zhejiang University | ACL | 2025 | TTS, codec | GAN, VAE | 2026-05-29 |
| 2603.18090 | MOSS-TTS Technical Report | Shanghai Innovation Institute / Fudan University | arXiv | 2026 | TTS | autoregressive-LM, hybrid | 2026-05-29 |
| 2508.04141 | Parallel GPT: Harmonizing the Independence and Interdependence of Acoustic and Semantic Information for Zero-Shot Text-to-Speech | South China University of Technology | arXiv | 2025 | TTS | autoregressive-LM, hybrid | 2026-05-29 |
| 2502.11128 | FELLE: Autoregressive Speech Synthesis with Token-Wise Coarse-to-Fine Flow Matching | | arXiv | 2025 | TTS | autoregressive-LM, flow-matching | 2026-05-29 |
| 2603.26364 | LLaDA-TTS: Unifying Speech Synthesis and Zero-Shot Editing via Masked Diffusion Modeling | Bairong, Inc. | arXiv | 2026 | TTS | diffusion | 2026-05-29 |
| 2508.19098 | CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis | | arXiv | 2025 | TTS | autoregressive-LM, flow-matching, VAE | 2026-05-29 |
| 2508.12001 | FNH-TTS: A Fast, Natural, and Human-Like Speech Synthesis System with advanced prosodic modeling based on Mixture of Experts | Megatronix | arXiv | 2025 | TTS | VAE, GAN, hybrid | 2026-05-29 |
| 2510.05758 | EMORL-TTS: Reinforcement Learning for Fine-Grained Emotion Control in LLM-based TTS | Hangzhou Institute for Advanced Study, UCAS | ICASSP | 2026 | TTS | autoregressive-LM | 2026-05-29 |
| 2601.03888 | IndexTTS 2.5 Technical Report | Bilibili | arXiv | 2026 | TTS | autoregressive-LM, flow-matching, hybrid | 2026-05-29 |
| 2509.15969 | VoXtream: Full-Stream Text-to-Speech with Extremely Low Latency | KTH Royal Institute of Technology | arXiv | 2025 | TTS | autoregressive-LM, hybrid | 2026-05-29 |
| 2510.07979 | IntMeanFlow: Few-step Speech Generation with Integral Velocity Distillation | ByteDance | arXiv | 2025 | TTS | flow-matching | 2026-05-29 |
| 2025.ccl-1.80 | Lao-English Code-Switched Speech Synthesis Via Neural Codec Language Modeling | Kunming University of Science and Technology | workshop | 2025 | TTS | autoregressive-LM, hybrid | 2026-05-29 |
| 2025.coling-main.352 | DiffStyleTTS: Diffusion-based Hierarchical Prosody Modeling for Text-to-Speech with Diverse and Controllable Styles | University of Science and Technology of China | workshop | 2025 | TTS | diffusion, transformer-enc-dec | 2026-05-29 |
| 2025.acl-long.911 | DNASpeech: A Contextualized and Situated Text-to-Speech Dataset with Dialogues, Narratives and Actions | Renmin University of China | ACL | 2025 | TTS, evaluation | transformer-enc-dec | 2026-05-29 |
| 2025.acl-short.81 | Zero-Shot Text-to-Speech for Vietnamese | Movian AI | ACL | 2025 | TTS, evaluation | autoregressive-LM, transformer-enc-dec | 2026-05-29 |
| 2025.acl-long.912 | LLaMA-Omni 2: LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech Synthesis | Chinese Academy of Sciences | ACL | 2025 | SCA, TTS | autoregressive-LM, flow-matching, hybrid | 2026-05-29 |
| interspeech-2025-2765 | The State Of TTS: A Case Study with Human Fooling Rates | IIT Madras | Interspeech | 2025 | TTS, evaluation | | 2026-06-03 |
| interspeech-2025-0401 | Enabling the replicability of speech synthesis perceptu | | Interspeech | 2025 | evaluation | | 2026-06-03 |
| interspeech-2025-0115 | Bringing Interpretability to Neural Audio Codecs | | Interspeech | 2025 | codec | transformer-enc-dec | 2026-06-03 |
| interspeech-2025-0468 | DualCodec: A Low-Frame-Rate, Semantically-Enhanced Neur | CUHK-SZ / Baidu | Interspeech | 2025 | codec | VAE, GAN | 2026-06-03 |
| interspeech-2025-1641 | Robust Neural Codec Language Modeling with Phoneme Posi | Samsung | Interspeech | 2025 | TTS | autoregressive-LM | 2026-06-03 |
| interspeech-2025-2447 | Accelerating Autoregressive Speech Synthesis Inference | Tsinghua / Tencent | Interspeech | 2025 | TTS | autoregressive-LM | 2026-06-03 |
| interspeech-2025-1779 | ReFlow-VC: Zero-shot Voice Conversion Based on Rectifie | | Interspeech | 2025 | VC | flow-matching | 2026-06-03 |
| interspeech-2025-0874 | Efficient and Direct Duplex Modeling for Speech-to-Spee | | Interspeech | 2025 | SCA | autoregressive-LM, hybrid | 2026-06-03 |
| interspeech-2025-0246 | DC-Spin: A Speaker-invariant Speech Tokenizer for Spoke | | Interspeech | 2025 | codec | transformer-enc-dec | 2026-06-03 |
| interspeech-2025-1440 | FreeCodec: A Disentangled Neural Speech Codec with Fewe | | Interspeech | 2025 | codec | VAE | 2026-06-03 |
| interspeech-2025-2043 | Training-Free Voice Conversion with Factorized Optimal | | Interspeech | 2025 | VC | transformer-enc-dec | 2026-06-03 |
| interspeech-2025-0816 | Bridging Speech and Singing: Multi-stage Speech-Prompte | | Interspeech | 2025 | singing, VC | diffusion | 2026-06-03 |
| interspeech-2025-1066 | Score-Based Training for Energy-Based TTS Models | | Interspeech | 2025 | TTS | diffusion | 2026-06-03 |
| interspeech-2025-1122 | BitTTS: Highly Compact Text-to-Speech Using 1.58-bit Qu | | Interspeech | 2025 | TTS | GAN, transformer-enc-dec | 2026-06-03 |
| 2508.20660 | CodecBench: A Comprehensive Benchmark for Acoustic and | Fudan University | arXiv | 2025 | codec, evaluation | | 2026-06-03 |
| interspeech-2025-1344 | Parameter-Efficient Fine-Tuning for Low-Resource Text-t | Ajou University | Interspeech | 2025 | TTS | flow-matching | 2026-06-03 |
| interspeech-2025-2449 | Accelerating Flow-Matching-Based Text-to-Speech via Emp | — | Interspeech | 2025 | TTS | flow-matching | 2026-06-03 |
| interspeech-2025-1595 | Scheduled Interleaved Speech-Text Training for Speech-t | | Interspeech | 2025 | TTS, SCA | autoregressive-LM | 2026-06-03 |
| interspeech-2025-0815 | Towards Better Disentanglement in Non-Autoregressive Ze | | Interspeech | 2025 | VC | VAE, GAN | 2026-06-03 |
| interspeech-2025-1101 | ZSDEVC: Zero-Shot Diffusion-based Emotional Voice Conve | | Interspeech | 2025 | VC | diffusion | 2026-06-03 |
| interspeech-2025-2660 | Triadic Multi-party Voice Activity Projection for Turn- | Kyoto University | Interspeech | 2025 | SCA | transformer-enc-dec | 2026-06-03 |
| 2508.07375 | TurnGuide: Enhancing Meaningful Full Duplex Spoken Inte | — | arXiv | 2025 | SCA | autoregressive-LM | 2026-06-04 |
| 2508.16790 | TaDiCodec: Text-aware Diffusion Speech Tokenizer for Sp | CUHK-SZ | arXiv | 2025 | codec | diffusion, transformer-enc-dec, autoregressive-LM | 2026-06-04 |
| interspeech-2025-1289 | Unlocking Temporal Flexibility: Neural Speech Codec wit | | Interspeech | 2025 | codec | hybrid | 2026-06-04 |
| interspeech-2025-0984 | Benchmarking Neural Speech Codec Intelligibility with S | | Interspeech | 2025 | codec, evaluation | | 2026-06-04 |
| 2508.07273 | Incorporating Contextual Paralinguistic Understanding i | | arXiv | 2025 | SCA | transformer-enc-dec | 2026-06-04 |
| 2508.08957 | QAMRO: Quality-aware Adaptive Margin Ranking Optimizati | | arXiv | 2025 | evaluation | | 2026-06-04 |
| 2508.09600 | OSUM-EChat: Enhancing End-to-End Empathetic Spoken Chat | Northwestern Polytechnical University | arXiv | 2025 | SCA | autoregressive-LM, flow-matching | 2026-06-04 |
| 2508.09702 | M3PDB: A Multimodal, Multi-Label, Multilingual Prompt Database for Speech Generation | | arXiv | 2025 | TTS, evaluation | | 2026-06-04 |
| 2508.11224 | Benchmarking Prosody Encoding in Discrete Speech Tokens | The University of Tokyo / AIST | ASRU | 2025 | evaluation, TTS | | 2026-06-04 |
| 2508.13028 | Integrating Feedback Loss from Bi-modal Sarcasm Detecto | | arXiv | 2025 | TTS | transformer-enc-dec | 2026-06-04 |
| 2508.15565 | Any-to-any Speaker Attribute Perturbation for Asynchron | | arXiv | 2025 | VC | GAN | 2026-06-04 |
| 2508.15931 | QvTAD: Differential Relative Attribute Learning for Voi | Qifu Technology | arXiv | 2025 | evaluation | transformer-enc-dec | 2026-06-04 |
| 2508.16188 | Seeing is Believing: Emotion-Aware Audio-Visual Languag | | EMNLP | 2025 | TTS, SCA | autoregressive-LM | 2026-06-04 |
| 2508.17031 | RephraseTTS: Dynamic Length Text based Speech Insertion | IIT Kanpur | arXiv | 2025 | TTS, VC | transformer-enc-dec, GAN | 2026-06-04 |
| 2508.17494 | Improving French Synthetic Speech Quality via SSML Pros | | workshop | 2025 | TTS | hybrid | 2026-06-04 |
| 2508.17623 | EMO-Reasoning: Benchmarking Emotional Reasoning Capabil | | arXiv | 2025 | SCA, evaluation | | 2026-06-04 |
| 2508.18006 | Unseen Speaker and Language Adaptation for Lightweight | Amazon | arXiv | 2025 | TTS | GAN | 2026-06-04 |
| 2508.19205 | VibeVoice Technical Report | Microsoft Research | arXiv | 2025 | TTS | hybrid | 2026-06-04 |
| 2509.00503 | Entropy-based Coarse and Compressed Semantic Speech Rep | | arXiv | 2025 | codec | autoregressive-LM, transformer-enc-dec | 2026-06-04 |
| 2509.00675 | Speaker-Conditioned Phrase Break Prediction for Text-to | — | arXiv | 2025 | TTS | transformer-enc-dec | 2026-06-04 |
| 2509.01391 | MixedG2P-T5: G2P-free Speech Synthesis for Mixed-script | | arXiv | 2025 | TTS | transformer-enc-dec | 2026-06-04 |
| 2509.02244 | Spectrogram Patch Codec: A 2D Block-Quantized VQ-VAE an | | arXiv | 2025 | codec | VAE, GAN | 2026-06-04 |
| 2509.03292 | Improving Perceptual Audio Aesthetic Assessment via Tri | | arXiv | 2025 | evaluation | hybrid | 2026-06-04 |
| 2509.03940 | VoxRole: A Comprehensive Benchmark for Evaluating Speec | | arXiv | 2025 | SCA, evaluation | | 2026-06-04 |
| 2410.00037 | Moshi: a speech-text foundation model for real-time dia | Kyutai | arXiv | 2024 | SCA, TTS | autoregressive-LM, hybrid | 2026-06-09 |
| 2411.13577 | WavChat: A Survey of Spoken Dialogue Models | | arXiv | 2024 | SCA | autoregressive-LM, transformer-enc-dec, hybrid | 2026-06-09 |
| 2503.20215 | Qwen2.5-Omni Technical Report | | arXiv | 2025 | | autoregressive-LM, flow-matching, hybrid | 2026-06-09 |
| 2407.21783 | The Llama 3 Herd of Models | Meta | arXiv | 2024 | | autoregressive-LM | 2026-06-09 |
| 1912.06670 | Common Voice: A Massively-Multilingual Speech Corpus | | arXiv | 2019 | | | 2026-06-09 |
| 2312.15185 | emotion2vec: Self-Supervised Pre-Training for Speech Em | | arXiv | 2023 | | | 2026-06-09 |
| 2212.04356 | Robust Speech Recognition via Large-Scale Weak Supervis | | arXiv | 2022 | | transformer-enc-dec | 2026-06-09 |
| 2010.05646 | HiFi-GAN: Generative Adversarial Networks for Efficient | Kakao Enterprise | arXiv | 2020 | TTS | GAN | 2026-06-09 |
| 2210.13438 | High Fidelity Neural Audio Compression | Meta AI | arXiv | 2022 | codec | GAN, VAE | 2026-06-09 |
| 2006.04558 | FastSpeech 2: Fast and High-Quality End-to-End Text to | Microsoft Research Asia | arXiv | 2020 | TTS | transformer-enc-dec | 2026-06-09 |
| 2407.10759 | Qwen2-Audio Technical Report | Alibaba Group | arXiv | 2024 | | | 2026-06-10 |
| 2303.08774 | GPT-4 Technical Report | OpenAI | arXiv | 2023 | | autoregressive-LM | 2026-06-10 |
| 2412.02612 | GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot | Tsinghua University / Zhipu.AI | arXiv | 2024 | TTS, SCA | autoregressive-LM, flow-matching | 2026-06-10 |
| 1711.05101 | Decoupled Weight Decay Regularization | | arXiv | 2017 | | | 2026-06-10 |
| 2410.21276 | GPT-4o System Card | | arXiv | 2024 | | autoregressive-LM | 2026-06-10 |
| 2412.15115 | Qwen2.5 Technical Report | Alibaba | arXiv | 2024 | | autoregressive-LM | 2026-06-10 |
| 2210.02747 | Flow Matching for Generative Modeling | Meta AI (FAIR) | arXiv | 2022 | | flow-matching | 2026-06-10 |
| 2209.03143 | AudioLM: a Language Modeling Approach to Audio Generati | | arXiv | 2022 | TTS, SCA | autoregressive-LM | 2026-06-10 |
| 2206.04658 | BigVGAN: A Universal Neural Vocoder with Large-Scale Tr | NVIDIA | arXiv | 2022 | TTS | GAN | 2026-06-10 |
| 2207.12598 | Classifier-Free Diffusion Guidance | | arXiv | 2022 | | diffusion | 2026-06-10 |
| 1412.6980 | Adam: A Method for Stochastic Optimization | | arXiv | 2014 | | | 2026-06-10 |
| 2308.16692 | SpeechTokenizer: Unified Speech Tokenizer for Speech La | Fudan University | arXiv | 2023 | TTS, codec | GAN, VAE | 2026-06-11 |
| 2503.01710 | Spark-TTS: An Efficient LLM-Based Text-to-Speech Model | HKUST | arXiv | 2025 | TTS | autoregressive-LM | 2026-06-11 |
| 2409.00750 | MaskGCT: Zero-Shot Text-to-Speech with Masked Generativ | CUHK-SZ | arXiv | 2024 | TTS, VC | autoregressive-LM | 2026-06-11 |
| 2505.17589 | CosyVoice 3: Towards In-the-wild Speech Generation via | Alibaba | arXiv | 2025 | TTS | autoregressive-LM, flow-matching | 2026-06-11 |
| 2408.16725 | Mini-Omni: Language Models Can Hear, Talk While Thinki | Inspirai | arXiv | 2024 | SCA, TTS | autoregressive-LM | 2026-06-11 |
| 2502.04128 | Llasa: Scaling Train-Time and Inference-Time Compute fo | | arXiv | 2025 | TTS | autoregressive-LM | 2026-06-11 |
| 2304.09116 | NaturalSpeech 2: Latent Diffusion Models are Natural an | Microsoft Research Asia | arXiv | 2023 | TTS, VC, singing | diffusion, VAE | 2026-06-11 |
| 2406.05370 | VALL-E 2: Neural Codec Language Models are Human Parity | Microsoft | arXiv | 2024 | TTS | autoregressive-LM | 2026-06-11 |
| 2409.06666 | LLaMA-Omni: Seamless Speech Interaction with Large Lang | ICT/CAS | arXiv | 2024 | SCA | hybrid | 2026-06-11 |
| 2411.00774 | Freeze-Omni: A Smart and Low Latency Speech-to-speech D | Tencent Youtu Lab | arXiv | 2024 | SCA | autoregressive-LM, hybrid | 2026-06-11 |
| 2408.16532 | WavTokenizer: an Efficient Acoustic Discrete Codec Toke | | arXiv | 2024 | codec | GAN, VAE | 2026-06-11 |
| 2407.04051 | FunAudioLLM: Voice Understanding and Generation Foundat | Alibaba Group | arXiv | 2024 | TTS | hybrid | 2026-06-11 |
| 2305.11000 | SpeechGPT: Empowering Large Language Models with Intrin | Fudan University | arXiv | 2023 | SCA, TTS | autoregressive-LM | 2026-06-11 |
| 2410.17196 | VoiceBench: Benchmarking LLM-Based Voice Assistants | National University of Singapore | arXiv | 2024 | evaluation, SCA | | 2026-06-11 |
| 2409.03283 | FireRedTTS: A Foundation Text-To-Speech Framework for I | Xiaohongshu | arXiv | 2024 | TTS | autoregressive-LM, flow-matching, hybrid | 2026-06-11 |
| 2311.07919 | Qwen-Audio: Advancing Universal Audio Understanding via | | arXiv | 2023 | | transformer-enc-dec | 2026-06-12 |
| 2505.09388 | Qwen3 Technical Report | | arXiv | 2025 | | autoregressive-LM | 2026-06-12 |
| 2005.07143 | ECAPA-TDNN: Emphasized Channel Attention, Propagation a | | arXiv | 2020 | | | 2026-06-12 |
| 2302.13971 | LLaMA: Open and Efficient Foundation Language Models | | arXiv | 2023 | | autoregressive-LM | 2026-06-12 |
| 2507.06261 | Gemini 2.5: Pushing the Frontier with Advanced Reasonin | Google | arXiv | 2025 | TTS, SCA | | 2026-06-12 |
| 2012.03411 | MLS: A Large-Scale Multilingual Dataset for Speech Rese | | arXiv | 2020 | | | 2026-06-12 |
| 2501.12948 | DeepSeek-R1: Incentivizing Reasoning Capability in LLMs | DeepSeek | arXiv | 2025 | | autoregressive-LM | 2026-06-12 |
| 2312.05187 | Seamless: Multilingual Expressive and Streaming Speech | | arXiv | 2023 | | | 2026-06-12 |
| 2106.06909 | GigaSpeech: An Evolving, Multi-domain ASR Corpus with 1 | | arXiv | 2021 | | | 2026-06-12 |
| 2309.15505 | Finite Scalar Quantization: VQ-VAE Made Simple | | arXiv | 2023 | | VAE | 2026-06-12 |
| 2306.00814 | Vocos: Closing the gap between time-domain and Fourier- | | arXiv | 2023 | TTS | GAN | 2026-06-12 |
| 2407.05361 | Emilia: An Extensive, Multilingual, and Diverse Speech | | arXiv | 2024 | | | 2026-06-12 |
| 2406.18009 | E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Ze | Microsoft | arXiv | 2024 | TTS | flow-matching | 2026-06-12 |
| 2406.04904 | XTTS: a Massively Multilingual Zero-Shot Text-to-Speech | Coqui.ai / NVIDIA / Cantina.ai | arXiv | 2024 | | autoregressive-LM, GAN | 2026-06-12 |
| 2409.05377 | BigCodec: Pushing the Limits of Low-Bitrate Neural Spee | University of Tokyo, Microsoft, Keio University | arXiv | 2024 | codec | GAN, VAE | 2026-06-12 |
| 2305.02765 | HiFi-Codec: Group-residual Vector quantization for High | Peking University / Tencent AI Lab | arXiv | 2023 | codec | GAN, VAE | 2026-06-12 |
| 2403.16973 | VoiceCraft: Zero-Shot Speech Editing and Text-to-Speech | | arXiv | 2024 | TTS | autoregressive-LM | 2026-06-12 |
| 2502.11946 | Step-Audio: Unified Understanding and Generation in Int | StepFun | arXiv | 2025 | TTS, SCA | autoregressive-LM, flow-matching, hybrid | 2026-06-12 |
| 2501.06282 | MinMo: A Multimodal Large Language Model for Seamless V | Alibaba Group | arXiv | 2025 | TTS, SCA | autoregressive-LM, flow-matching, hybrid | 2026-06-12 |
| 2303.03926 | Speak Foreign Languages with Your Own Voice: Cross-Ling | Microsoft | arXiv | 2023 | TTS, multilingual-tts | autoregressive-LM | 2026-06-12 |
| 2305.09636 | SoundStorm: Efficient Parallel Audio Generation | Google | arXiv | 2023 | TTS, SCA | autoregressive-LM | 2026-06-13 |
| 1712.05884 | Natural TTS Synthesis by Conditioning WaveNet on Mel Sp | Google | arXiv | 2017 | TTS | transformer-enc-dec | 2026-06-13 |
| 2402.01912 | Natural language guidance of high-fidelity text-to-spee | Stability AI | arXiv | 2024 | TTS | autoregressive-LM | 2026-06-13 |
| 2306.12925 | AudioPaLM: A Large Language Model That Can Speak and Li | Google | arXiv | 2023 | TTS, SCA | autoregressive-LM | 2026-06-13 |
| 2305.07243 | Better speech synthesis through scaling | | arXiv | 2023 | TTS | autoregressive-LM, diffusion, VAE | 2026-06-13 |
| 1609.03499 | WaveNet: A Generative Model for Raw Audio | Google DeepMind | arXiv | 2016 | TTS | autoregressive-LM | 2026-06-13 |
| 2411.19842 | Scaling Transformers for Low-Bitrate High-Quality Speec | Stability AI | arXiv | 2024 | codec | transformer-enc-dec, VAE | 2026-06-13 |
| 2407.08551 | Autoregressive Speech Synthesis without Vector Quantiza | Microsoft | arXiv | 2024 | TTS | autoregressive-LM | 2026-06-13 |
| 1703.10135 | Tacotron: Towards End-to-End Speech Synthesis | Google | arXiv | 2017 | | transformer-enc-dec | 2026-06-13 |
| 2502.17239 | Baichuan-Audio: A Unified Framework for End-to-End Spee | Baichuan Inc. | arXiv | 2025 | SCA, TTS | autoregressive-LM, flow-matching, hybrid | 2026-06-13 |
| 2402.05755 | Spirit LM: Interleaved Spoken and Written Language Mode | Meta AI | arXiv | 2024 | SCA, TTS | autoregressive-LM | 2026-06-13 |
| 2402.08093 | BASE TTS: Lessons from building a billion-parameter Tex | Amazon AGI | arXiv | 2024 | TTS | autoregressive-LM | 2026-06-13 |
| 2507.16632 | Step-Audio 2 Technical Report | StepFun | arXiv | 2025 | SCA, TTS | autoregressive-LM, flow-matching, hybrid | 2026-06-13 |
| 2310.00704 | UniAudio: An Audio Foundation Model Toward Universal Au | | arXiv | 2023 | TTS, VC, singing | autoregressive-LM, hybrid | 2026-06-13 |
| 2106.15561 | A Survey on Neural Speech Synthesis | Microsoft Research Asia | arXiv | 2021 | TTS | autoregressive-LM, flow-matching, diffusion, GAN, VAE, transformer-enc-dec | 2026-06-13 |
| 2411.01156 | Fish-Speech: Leveraging Large Language Models for Advan | Fish Audio | arXiv | 2024 | TTS | autoregressive-LM, GAN | 2026-06-14 |
| 2505.07916 | MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with | MiniMax | arXiv | 2025 | TTS | autoregressive-LM, flow-matching, VAE | 2026-06-14 |
| 2410.11190 | Mini-Omni2: Towards Open-source GPT-4o with Vision, Spe | Inspirai / Tsinghua University | arXiv | 2024 | SCA | autoregressive-LM | 2026-06-14 |
| 2410.03751 | Recent Advances in Speech Language Models: A Survey | Chinese University of Hong Kong | arXiv | 2024 | SCA, TTS | autoregressive-LM | 2026-06-14 |
| 2104.00355 | Speech Resynthesis from Discrete Disentangled Self-Supe | Facebook AI Research | arXiv | 2021 | TTS, VC | GAN, VAE | 2026-06-14 |
| 2502.06490 | Recent Advances in Discrete Speech Tokens: A Review | SJTU / MSRA | arXiv | 2025 | TTS, VC, SCA, codec | autoregressive-LM, transformer-enc-dec, GAN, VAE | 2026-06-14 |
| 2105.06337 | Grad-TTS: A Diffusion Probabilistic Model for Text-to-S | | arXiv | 2021 | TTS | diffusion, transformer-enc-dec | 2026-06-14 |
| 2412.15649 | SLAM-Omni: Timbre-Controllable Voice Interaction System | SJTU / Microsoft | arXiv | 2024 | SCA | autoregressive-LM, flow-matching | 2026-06-14 |
| 2502.05512 | IndexTTS: An Industrial-Level Controllable and Efficien | bilibili | arXiv | 2025 | TTS | autoregressive-LM, GAN | 2026-06-14 |
| 2502.07243 | Vevo: Controllable Zero-Shot Voice Imitation with Self- | Meta AI | ICLR | 2025 | TTS, VC | autoregressive-LM, flow-matching, hybrid | 2026-06-14 |
| 2406.07855 | VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech | Microsoft | arXiv | 2024 | TTS | autoregressive-LM | 2026-06-14 |
| 2410.17799 | OmniFlatten: An End-to-end GPT Model for Seamless Voice | Alibaba (Tongyi Lab) | arXiv | 2024 | SCA | autoregressive-LM | 2026-06-14 |
| 2504.08528 | On The Landscape of Spoken Language Models: A Comprehen | | arXiv | 2025 | SCA | autoregressive-LM, transformer-enc-dec | 2026-06-14 |
| 2206.08317 | Paraformer: Fast and Accurate Parallel Transformer for | Alibaba Group | arXiv | 2022 | | transformer-enc-dec | 2026-06-14 |
| 2412.19437 | DeepSeek-V3 Technical Report | DeepSeek-AI | arXiv | 2024 | | | 2026-06-14 |
| 2402.03300 | DeepSeekMath: Pushing the Limits of Mathematical Reason | | arXiv | 2024 | | autoregressive-LM | 2026-06-14 |
| 2310.13289 | SALMONN: Towards Generic Hearing Abilities for Large La | Tsinghua University / ByteDance | arXiv | 2023 | | | 2026-06-14 |
| 1810.04805 | BERT: Pre-training of Deep Bidirectional Transformers f | | arXiv | 2018 | | | 2026-06-14 |
| 2502.05139 | Meta Audiobox Aesthetics: Unified Automatic Quality Ass | Meta (FAIR) | arXiv | 2025 | | | 2026-06-14 |
| 2307.09288 | Llama 2: Open Foundation and Fine-Tuned Chat Models | Meta | arXiv | 2023 | | autoregressive-LM | 2026-06-14 |
| 2312.11805 | Gemini: A Family of Highly Capable Multimodal Models | | arXiv | 2023 | | | 2026-06-14 |
| 2005.14165 | Language Models are Few-Shot Learners | | arXiv | 2020 | | autoregressive-LM | 2026-06-14 |
| 2407.10671 | Qwen2 Technical Report | Alibaba Group | arXiv | 2024 | | | 2026-06-14 |
| 2106.04624 | SpeechBrain: A General-Purpose Speech Toolkit | | arXiv | 2021 | | | 2026-06-14 |
| 2406.14294 | DASB - Discrete Audio and Speech Benchmark | | arXiv | 2024 | | | 2026-06-14 |
| 2301.12503 | AudioLDM: Text-to-Audio Generation with Latent Diffusio | | ICML | 2023 | | diffusion, VAE | 2026-06-15 |
| 2301.11325 | MusicLM: Generating Music From Text | Google | arXiv | 2023 | SCA | autoregressive-LM | 2026-06-15 |
| 2305.15255 | Spoken Question Answering and Speech Continuation Using | Google Research | arXiv | 2023 | SCA | autoregressive-LM | 2026-06-15 |
| 2312.01479 | OpenVoice: Versatile Instant Voice Cloning | MIT & MyShell.ai | arXiv | 2023 | TTS, VC | GAN, VAE | 2026-06-15 |
| 2312.15821 | Audiobox: Unified Audio Generation with Natural Languag | Meta FAIR | arXiv | 2023 | TTS | flow-matching | 2026-06-15 |
| 2401.07333 | ELLA-V: Stable Neural Codec Language Modeling with Alig | | arXiv | 2024 | TTS | autoregressive-LM | 2026-06-15 |
| 2402.13236 | Towards audio language modeling — an overview | | arXiv | 2024 | TTS, SCA, codec | | 2026-06-15 |
| 2404.03204 | RALL-E: Robust Codec Language Modeling with Chain-of-Th | Microsoft | arXiv | 2024 | TTS | autoregressive-LM | 2026-06-15 |
| 2406.00654 | Enhancing Zero-shot Text-to-Speech Synthesis with Human | Nanyang Technological University | arXiv | 2024 | TTS | autoregressive-LM | 2026-06-15 |
| 2406.05551 | Autoregressive Diffusion Transformer for Text-to-Speech | CUHK Shenzhen | arXiv | 2024 | TTS | hybrid | 2026-06-15 |
| 2408.02622 | Language Model Can Listen While Speaking | Shanghai Jiao Tong University / ByteDance | arXiv | 2024 | SCA, TTS | autoregressive-LM | 2026-06-15 |
| 2411.09943 | Zero-shot Voice Conversion with Diffusion Transformers | Nanyang Technological University | arXiv | 2024 | VC | diffusion, transformer-enc-dec | 2026-06-16 |
| 2411.17607 | Scaling Speech-Text Pre-training with Synthetic Interle | Tsinghua University / Zhipu.AI | arXiv | 2024 | SCA | autoregressive-LM, flow-matching | 2026-06-16 |
| 2411.18803 | TS3-Codec: Transformer-Based Simple Streaming Single Co | | arXiv | 2024 | | GAN | 2026-06-16 |
| 2412.04724 | StableVC: Style Controllable Zero-Shot Voice Conversion | Northwestern Polytechnical University / Ximalaya | arXiv | 2024 | VC | flow-matching | 2026-06-16 |
| 2506.13053 | ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speec | Xiaomi | arXiv | 2025 | TTS | flow-matching | 2026-06-16 |
| 2505.02625 | LLaMA-Omni2: LLM-based Real-time Spoken Chatbot with Au | ICT/CAS | arXiv | 2025 | SCA | autoregressive-LM, flow-matching, hybrid | 2026-06-16 |
| 2502.18924 | MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion T | Zhejiang University, ByteDance | arXiv | 2025 | TTS | diffusion, VAE | 2026-06-16 |
| 2507.23159 | Full-Duplex-Bench v1.5: Evaluating Overlap Handling for | | arXiv | 2025 | SCA, evaluation | | 2026-06-16 |
| 2506.16381 | InstructTTSEval: Benchmarking Complex Natural-Language | Fudan University | arXiv | 2025 | TTS, evaluation | | 2026-06-16 |
| 2506.10274 | Discrete Audio Tokens: More Than a Survey! | | arXiv | 2025 | codec, TTS, evaluation | | 2026-06-16 |
| 2505.13000 | DualCodec: A Low-Frame-Rate, Semantically-Enhanced Neur | CUHK-SZ / Baidu | arXiv | 2025 | codec, TTS | autoregressive-LM | 2026-06-16 |
| 2503.14345 | MoonCast: High-Quality Zero-Shot Podcast Generation | | arXiv | 2025 | TTS | autoregressive-LM, flow-matching | 2026-06-16 |
| 2508.04195 | NVSpeech: An Integrated and Scalable Pipeline for Human | | arXiv | 2025 | | autoregressive-LM, flow-matching | 2026-06-16 |
| 2505.09558 | WavReward: Spoken Dialogue Models With Generalist Rewar | Zhejiang University / Alibaba Group | arXiv | 2025 | SCA, evaluation | autoregressive-LM | 2026-06-16 |
| 2504.10344 | ALMTokenizer: A Low-bitrate and Semantic-rich Audio Cod | | arXiv | 2025 | codec, TTS, SCA | autoregressive-LM, VAE | 2026-06-16 |
| 2504.02407 | F5R-TTS: Improving Flow-Matching based Text-to-Speech w | Tencent | arXiv | 2025 | TTS | flow-matching | 2026-06-16 |
| 2511.15848 | Step-Audio-R1 Technical Report | StepFun | arXiv | 2025 | SCA | autoregressive-LM | 2026-06-16 |
| 2510.07838 | Full-Duplex-Bench-v2: A Multi-Turn Evaluation Framework | | arXiv | 2025 | SCA, evaluation | | 2026-06-16 |
| 2505.14648 | Vox-Profile: A Speech Foundation Model Benchmark for Ch | | arXiv | 2025 | evaluation | | 2026-06-16 |
| 1510.08484 | MUSAN: A Music, Speech, and Noise Corpus | | arXiv | 2015 | | | 2026-06-17 |
| 1908.06248 | JVS corpus: free Japanese multi-speaker voice corpus | | arXiv | 2019 | TTS, VC | | 2026-06-17 |
| 1607.06450 | Layer Normalization | | arXiv | 2016 | | | 2026-06-17 |
| 1808.10583 | AISHELL-2: Transforming Mandarin ASR Research Into Indu | | arXiv | 2018 | | | 2026-06-17 |
| 2002.05202 | GLU Variants Improve Transformer | | arXiv | 2020 | | | 2026-06-17 |
| 2302.00482 | Improving and generalizing flow-based generative models | | arXiv | 2023 | | | 2026-06-17 |
| 2007.10310 | CoVoST 2 and Massively Multilingual Speech-to-Text Tran | | arXiv | 2020 | | | 2026-06-17 |
| 2001.08361 | Scaling Laws for Neural Language Models | | arXiv | 2020 | | | 2026-06-17 |
| 2308.10248 | Steering Language Models With Activation Engineering | | arXiv | 2023 | | | 2026-06-17 |
| 2309.16609 | Qwen Technical Report | | arXiv | 2023 | | autoregressive-LM | 2026-06-17 |
| 2308.05725 | EXPRESSO: A Benchmark and Analysis of Discrete Expressi | | arXiv | 2023 | | | 2026-06-17 |
| 2308.11596 | SeamlessM4T: Massively Multilingual & Multimodal Machin | | arXiv | 2023 | | transformer-enc-dec | 2026-06-17 |
| 2408.05211 | VITA: Towards Open-Source Interactive Omni Multimodal L | | arXiv | 2024 | | | 2026-06-17 |
| 2408.01800 | MiniCPM-V: A GPT-4V Level MLLM on Your Phone | | arXiv | 2024 | | | 2026-06-17 |
| 2312.10997 | Retrieval-Augmented Generation for Large Language Model | | arXiv | 2023 | | | 2026-06-17 |
| 2402.07729 | AIR-Bench: Benchmarking Large Audio-Language Models via | | arXiv | 2024 | | | 2026-06-17 |
| 2501.07246 | Audio-CoT: Exploring Chain-of-Thought Reasoning in Larg | | arXiv | 2025 | | | 2026-06-17 |
| 2412.08635 | Multimodal Latent Language Modeling with Next-Token Dif | | arXiv | 2024 | | autoregressive-LM, diffusion, VAE, hybrid | 2026-06-17 |
| 2410.19168 | MMAU: A Massive Multi-Task Audio Understanding and Reas | | arXiv | 2024 | | | 2026-06-17 |
| 2501.01957 | VITA-1.5: Towards GPT-4o Level Real-Time Vision and Spe | | arXiv | 2025 | | autoregressive-LM, transformer-enc-dec | 2026-06-17 |
| 2503.01743 | Phi-4-Mini Technical Report: Compact yet Powerful Multi | | arXiv | 2025 | | | 2026-06-17 |
| 2503.19786 | Gemma 3 Technical Report | Google DeepMind | arXiv | 2025 | | autoregressive-LM | 2026-06-17 |
| 2505.03739 | VITA-Audio: Fast Interleaved Cross-Modal Token Generati | | arXiv | 2025 | | autoregressive-LM, hybrid | 2026-06-17 |
| 2501.15368 | Baichuan-Omni-1.5 Technical Report | Baichuan Inc. | arXiv | 2025 | | autoregressive-LM, flow-matching | 2026-06-17 |
| 2506.02863 | CapSpeech: Enabling Downstream Applications in Style-Ca | | arXiv | 2025 | TTS, evaluation | autoregressive-LM, flow-matching | 2026-06-17 |
| 2507.12705 | AudioJudge: Understanding What Works in Large Audio Mod | | arXiv | 2025 | | | 2026-06-17 |
| 2506.07900 | MiniCPM4: Ultra-Efficient LLMs on End Devices | | arXiv | 2025 | | autoregressive-LM | 2026-06-17 |
| 2507.08128 | Audio Flamingo 3: Advancing Audio Intelligence with Ful | | arXiv | 2025 | | autoregressive-LM | 2026-06-17 |
| 2510.14664 | SpeechLLM-as-Judges: Towards General and Interpretable | | arXiv | 2025 | | | 2026-06-17 |
| 2508.13992 | MMAU-Pro: A Challenging and Comprehensive Benchmark for | | arXiv | 2025 | | | 2026-06-17 |
| 2511.09690 | Omnilingual ASR: Open-Source Multilingual Speech Recogn | Meta (FAIR) | arXiv | 2025 | | transformer-enc-dec | 2026-06-17 |
| 2509.08753 | Streaming Sequence-to-Sequence Learning with Delayed St | | arXiv | 2025 | | autoregressive-LM | 2026-06-17 |
| 2409.09098 | AccentBox: Towards High-Fidelity Zero-Shot Accent Gener | University of Edinburgh | arXiv | 2025 | TTS | transformer-enc-dec | 2026-06-29 |
| 2025.coling-industry.29 | CarMem: Enhancing Long-Term Memory in LLM Voice Assista | BMW Group / Univ. Augsburg / TUM | COLING | 2025 | SCA | | 2026-06-29 |
| 2025.chipsal-1.18 | Impacts of Vocoder Selection on Tacotron-based Nepali T | | CHiPSAL | 2025 | TTS, evaluation | GAN, transformer-enc-dec | 2026-06-29 |
| 2025.coling-main.685 | VoxpopuliTTS: a large-scale multilingual TTS corpus for | Zhejiang University | COLING | 2025 | TTS | | 2026-06-29 |
| 2409.20007 | DeSTA2: Developing Instruction-Following Speech Language | | arXiv | 2025 | SCA | transformer-enc-dec | 2026-06-29 |
| 2025.computel-main.6 | Evaluating Indigenous language speech synthesis for educ | | ComputEL | 2025 | TTS, evaluation | | 2026-06-29 |
| 2025.nodalida-1.32 | Estonian isolated-word text-to-speech synthesiser | Institute of the Estonian Language | NoDaLiDa | 2025 | TTS | | 2026-06-29 |
| 2025.naacl-srw.6 | Towards Codec-LM Co-design for Neural Codec Language Mo | Cartesia AI / MIT / CMU | NAACL | 2025 | TTS, codec | autoregressive-LM, hybrid | 2026-06-29 |
| 2025.findings-naacl.298 | Gender Bias in Instruction-Guided Speech Synthesis Mode | | NAACL | 2025 | TTS, evaluation | | 2026-06-29 |
| 2025.findings-naacl.471 | The Role of Prosody in Spoken Question Answering | | NAACL | 2025 | evaluation | | 2026-06-29 |
| 2025.naacl-long.464 | ManaTTS Persian: a recipe for creating TTS datasets for | Sharif University of Technology | NAACL | 2025 | TTS, evaluation | | 2026-06-29 |
| 2025.naacl-long.619 | ProSE: Diffusion Priors for Speech Enhancement | University of Maryland | NAACL | 2025 | TTS | diffusion, transformer-enc-dec | 2026-06-29 |
| iclr-2025-tQ1PmLfPBL | PeriodWave: Multi-Period Flow Matching for High-Fidelity Waveform Gen | | ICLR | 2025 | TTS | flow-matching | 2026-06-29 |
| iclr-2025-cuFzE8Jlvb | Continuous Autoregressive Modeling with Stochastic Monotonic Alignmen | The Hong Kong Polytechnic University | ICLR | 2025 | TTS | autoregressive-LM, VAE | 2026-06-29 |
| iclr-2025-dGSOn7sdWg | SyllableLM: Learning Coarse Semantic Units for Speech Language Models | University of Texas at Austin | ICLR | 2025 | SCA | autoregressive-LM | 2026-06-29 |
| iclr-2025-868masI331 | HALL-E: Hierarchical Neural Codec Language Model for Minute-Long Zero | | ICLR | 2025 | TTS | autoregressive-LM | 2026-06-29 |
| iclr-2025-hQvX9MBowC | DiTTo-TTS: Diffusion Transformers for Scalable Text-to- | KRAFTON | ICLR | 2025 | TTS | diffusion, transformer-enc-dec | 2026-06-30 |
| iclr-2025-uxDFlPGRLX | FlowDec: A flow-based full-band general audio codec wit | Meta | ICLR | 2025 | codec | flow-matching, GAN | 2026-06-30 |
| 2025.findings-naacl.130 | DiVISe: Direct Visual-Input Speech Synthesis Preserving | Shanghai Jiao Tong University | NAACL | 2025 | TTS | transformer-enc-dec, GAN | 2026-06-30 |
| 2025.findings-naacl.279 | BnTTS: Few-Shot Speaker Adaptation in Low-Resource Sett | Hishab Singapore | NAACL | 2025 | TTS | autoregressive-LM, GAN, hybrid | 2026-06-30 |
| 2025.findings-naacl.38 | Prompt-Guided Selective Masking Loss for Context-Aware | POSTECH | NAACL | 2025 | TTS | transformer-enc-dec | 2026-06-30 |
| 2025.naacl-demo.12 | ESPnet-SpeechLM: An Open Speech Language Model Toolkit | Carnegie Mellon University | NAACL | 2025 | TTS, SCA | autoregressive-LM | 2026-06-30 |
| 2025.naacl-demo.21 | ESPnet-SDS: Unified Toolkit and Demo for Spoken Dialogu | Carnegie Mellon University | NAACL | 2025 | SCA | | 2026-06-30 |
| 2025.naacl-long.484 | Behavior-SD: Behaviorally Aware Spoken Dialogue Generat | Seoul National University | NAACL | 2025 | SCA | autoregressive-LM | 2026-06-30 |
| 2025.naacl-long.591 | Robust and Unbounded Length Generalization in Autoregre | Google DeepMind | NAACL | 2025 | TTS | transformer-enc-dec | 2026-06-30 |
| 2025.naacl-short.65 | kNN Retrieval for Simple and Effective Zero-Shot Multi- | | NAACL | 2025 | TTS | hybrid, GAN | 2026-06-30 |
| 2025.naacl-short.69 | Developing multilingual speech synthesis system for Oji | | NAACL | 2025 | TTS | flow-matching | 2026-06-30 |
| 2025.iwsds-1.11 | Paralinguistic Attitude Recognition for Spoken Dialogue | Fairy Devices Inc. | IWSDS | 2025 | SCA | | 2026-06-30 |
| 2025.iwsds-1.27 | A Survey of Recent Advances on Turn-taking Modeling in | Université Paris-Saclay, CEA, List | IWSDS | 2025 | SCA | | 2026-07-01 |
| 2505.15772 | MIKU-PAL: An Automated and Standardized Multi-Modal Met | Fish Audio; Carnegie Mellon University | arXiv | 2025 | TTS, evaluation | | 2026-07-01 |
| 2507.06235 | Super Kawaii Vocalics: Amplifying the “Cute” Factor in | | arXiv | 2025 | TTS | | 2026-07-01 |
| 2506.23049 | AURA: Agent for Understanding, Reasoning, and Automated | | arXiv | 2025 | SCA | autoregressive-LM | 2026-07-01 |
| 2025.acl-long.681 | SIFT-50M: A Large-Scale Multilingual Dataset for Speech | Amazon AGI | ACL | 2025 | SCA | autoregressive-LM | 2026-07-01 |
| 2025.acl-long.790 | Rhythm Controllable and Efficient Zero-Shot Voice Conve | Zhejiang University | ACL | 2025 | VC | flow-matching | 2026-07-01 |
| 2025.acl-long.817 | SimulS2S-LLM: Unlocking Simultaneous Inference of Speec | | ACL | 2025 | SCA, TTS | autoregressive-LM | 2026-07-01 |
| 2025.acl-long.87 | Takin-VC: Expressive Zero-Shot Voice Conversion via Ada | | ACL | 2025 | VC | flow-matching | 2026-07-01 |
| 2025.acl-long.937 | UniCodec: Unified Audio Codec with Single Domain-Adapti | | ACL | 2025 | codec | VAE | 2026-07-01 |
| 2025.acl-long.997 | Align-SLM: Textless Spoken Language Models with Reinfor | | ACL | 2025 | SCA | autoregressive-LM | 2026-07-01 |
| 2025.conll-1.9 | A Linguistically Motivated Analysis of Intonational Phr | | CoNLL | 2025 | TTS, evaluation | | 2026-07-01 |
| 2025.findings-acl.101 | Chain-Talker: Chain Understanding and Rendering for Emp | | ACL | 2025 | TTS | autoregressive-LM, flow-matching | 2026-07-01 |
| 2025.findings-acl.115 | SLAM-Omni: Timbre-Controllable Voice Interaction System | | ACL | 2025 | SCA, TTS | autoregressive-LM | 2026-07-01 |
| 2025.findings-acl.1226 | PodAgent: A Comprehensive Framework for Podcast Generat | | ACL | 2025 | TTS, VC | hybrid | 2026-07-01 |
| 2025.findings-acl.470 | Does Your Voice Assistant Remember? Analyzing Conversat | Seoul National University | ACL | 2025 | SCA, evaluation | | 2026-07-01 |
| 2025.findings-acl.534 | Unlocking Speech Instruction Data Potential with Query | | ACL | 2025 | SCA | | 2026-07-01 |
| 2025.findings-acl.631 | Slamming: Training a Speech Language Model on One GPU i | The Hebrew University of Jerusalem | ACL | 2025 | SCA | autoregressive-LM | 2026-07-01 |
| 2025.findings-acl.687 | TCSinger 2: Customizable Multilingual Zero-shot Singing | Zhejiang University | ACL | 2025 | singing, TTS | flow-matching, VAE | 2026-07-01 |
| 2025.findings-acl.71 | Data-Centric Improvements for Enhancing Multi-Modal Und | | ACL | 2025 | SCA | autoregressive-LM | 2026-07-01 |
| 2025.findings-acl.75 | Leveraging Unit Language Guidance to Advance Speech Mod | | ACL | 2025 | TTS, SCA | transformer-enc-dec | 2026-07-01 |
| 2025.findings-ijcnlp.49 | Incorporating Dialogue State Tracking into Japanese Ful | NTT / Nagoya University | ACL | 2025 | SCA | autoregressive-LM | 2026-07-01 |
| 2025.iwslt-1.5 | SSR: Alignment-Aware Modality Connector for Speech Lang | | IWSLT | 2025 | SCA | autoregressive-LM | 2026-07-01 |
| 2025.unlp-1.11 | Context-Aware Lexical Stress Prediction and Phonemizati | | workshop | 2025 | TTS | transformer-enc-dec | 2026-07-01 |
| 2412.18603 | Long-Form Speech Generation with Spoken Language Models | | ICML | 2025 | SCA | autoregressive-LM | 2026-07-01 |
| 2503.11026 | MAVFlow: Preserving Paralinguistic Elements with Condit | KAIST | arXiv | 2025 | VC, TTS | flow-matching | 2026-07-01 |
| 2505.15670 | SALM-Duplex: Efficient and Direct Duplex Modeling for S | NVIDIA | arXiv | 2025 | SCA | autoregressive-LM | 2026-07-01 |
| 2506.09874 | UmbraTTS: Adapting Text-to-Speech to Environmental Cont | | arXiv | 2025 | TTS | flow-matching | 2026-07-01 |
| 2506.18296 | JIS: A Speech Corpus of Japanese Idol Speakers with Var | NTT Corporation | Interspeech | 2025 | TTS, VC, evaluation | | 2026-07-01 |
| 2507.02176 | Analyzing and Improving Speaker Similarity Assessment f | | arXiv | 2025 | evaluation | | 2026-07-01 |
| 2507.00808 | Multi-interaction TTS toward professional recording rep | NTT | arXiv | 2025 | TTS | transformer-enc-dec | 2026-07-01 |
| 2507.01611 | QHARMA-GAN: Quasi-Harmonic Neural Vocoder based on Auto | | arXiv | 2025 | TTS | GAN | 2026-07-01 |
| 2507.02380 | JoyTTS: LLM-based Spoken Chatbot With Voice Cloning | JD Health International Inc. | arXiv | 2025 | SCA, TTS | autoregressive-LM, flow-matching | 2026-07-01 |
| 2507.03887 | Traceable TTS: Toward Watermark-Free TTS with Strong Tr | | arXiv | 2025 | TTS | flow-matching | 2026-07-01 |
| 2507.03912 | Prosody Labeling with Phoneme-BERT and Speech Foundatio | CyberAgent | arXiv | 2025 | TTS | | 2026-07-01 |
| 2507.08012 | RepeaTTS: Towards Feature Discovery through Repeated Fi | | arXiv | 2025 | TTS | transformer-enc-dec | 2026-07-01 |
| 2507.04349 | TTS-CtrlNet: Time varying emotion aligned text-to-speec | | arXiv | 2025 | TTS | flow-matching | 2026-07-01 |
| 2507.04598 | Multi-Step Prediction and Control of Hierarchical Emoti | | arXiv | 2025 | TTS | transformer-enc-dec | 2026-07-01 |
| 2507.04817 | Fast-VGAN: Lightweight Voice Conversion with Explicit C | | arXiv | 2025 | VC | GAN | 2026-07-01 |
| 2507.01348 | SpeechAccentLLM: A Unified Framework for Foreign Accent | | arXiv | 2025 | VC, TTS | autoregressive-LM, VAE, GAN | 2026-07-01 |
| 2507.06116 | Speech Quality Assessment Model Based on Mixture of Exp | Zhejiang University | arXiv | 2025 | evaluation | | 2026-07-01 |
| 2506.23325 | XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict | Fudan University | arXiv | 2025 | codec | GAN, hybrid | 2026-07-01 |
| 2507.07799 | SecureSpeech: Prompt-based Speaker and Content Protecti | | arXiv | 2025 | TTS | autoregressive-LM | 2026-07-01 |
| 2507.08319 | Active Learning for Text-to-Speech Synthesis with Infor | | arXiv | 2025 | TTS | transformer-enc-dec | 2026-07-01 |
| 2507.09070 | SemAlignVC: Enhancing zero-shot timbre conversion using | Meta / KTH Royal Institute of Technology | arXiv | 2025 | VC | autoregressive-LM, flow-matching | 2026-07-01 |
| 2507.09282 | ClaritySpeech: Dementia Obfuscation in Speech | Imperial College London | arXiv | 2025 | TTS | autoregressive-LM, diffusion, VAE | 2026-07-02 |
| 2507.09310 | Voice Conversion for Lombard Speaking Style with Implic | Amazon Alexa / Imperial College London | arXiv | 2025 | VC, TTS | VAE | 2026-07-02 |
| 2507.10985 | Pronunciation Deviation Analysis Through Voice Cloning | California State University Long Beach | arXiv | 2025 | TTS | | 2026-07-02 |
| 2507.12197 | Quantize More, Lose Less: Autoregressive Generation fro | | arXiv | 2025 | TTS, singing | autoregressive-LM, GAN | 2026-07-02 |
| 2507.14988 | DMOSpeech 2: Reinforcement Learning for Duration Predic | Columbia University | arXiv | 2025 | TTS | flow-matching | 2026-07-02 |
| 2507.15272 | A2TTS: TTS for Low Resource Indian Languages | | arXiv | 2025 | TTS | diffusion | 2026-07-02 |
| 2507.16875 | Technical report: Impact of Duration Prediction on Spea | | arXiv | 2025 | TTS | flow-matching | 2026-07-02 |
| 2507.21138 | TTS-1 Technical Report | Inworld AI | arXiv | 2025 | TTS | autoregressive-LM | 2026-07-02 |
| 2507.18119 | GOAT-SLM: A Spoken Language Model with Paralinguistic a | TeleAI, China Telecom | arXiv | 2025 | SCA | autoregressive-LM, flow-matching | 2026-07-02 |
| 2507.18897 | HH-Codec: High Compression High-fidelity Discrete Neura | | arXiv | 2025 | codec | GAN | 2026-07-02 |
| 2507.17527 | Seed LiveInterpret 2.0: End-to-end Simultaneous Speech- | ByteDance | arXiv | 2025 | TTS, VC | autoregressive-LM | 2026-07-02 |
| 2507.20140 | Do Not Mimic My Voice: Speaker Identity Unlearning for | | arXiv | 2025 | TTS | flow-matching | 2026-07-02 |
| 2507.20731 | Learning Neural Vocoder from Range-Null Space Decomposi | | arXiv | 2025 | TTS | GAN | 2026-07-02 |
| 2025.ccl-1.77 | HFSD-V2C: Zero-Shot Visual Voice Cloning Via Hierarchic | | workshop | 2025 | TTS | diffusion | 2026-07-02 |
| 2025.icnlsp-1.34 | Beyond Labeled Datasets: Advancing TTS with Direct Pref | | workshop | 2025 | TTS | autoregressive-LM | 2026-07-02 |
| 2025.sigdial-1.21 | Transition Relevance Point Detection for Spoken Dialogu | | workshop | 2025 | SCA | hybrid | 2026-07-02 |
| 2025.sigdial-1.27 | EmoNews: A Spoken Dialogue System for Expressive News C | | workshop | 2025 | SCA, TTS | transformer-enc-dec | 2026-07-02 |
| 2025.sigdial-1.51 | rrSDS 2.0: Incremental, Modular, Distributed, Multimoda | | workshop | 2025 | SCA | | 2026-07-02 |
| interspeech-2025-0166 | Frozen Large Language Models Can Perceive Paralinguisti | | Interspeech | 2025 | SCA | transformer-enc-dec | 2026-07-02 |
| interspeech-2025-0305 | DAFMSVC: One-Shot Singing Voice Conversion with Dual At | | Interspeech | 2025 | singing, VC | flow-matching, transformer-enc-dec | 2026-07-02 |
| interspeech-2025-0347 | PeriodCodec: A Pitch-Controllable Neural Audio Codec Us | | Interspeech | 2025 | codec, singing | GAN, VAE | 2026-07-02 |
| interspeech-2025-0355 | Probing the Robustness Properties of Neural Speech Code | | Interspeech | 2025 | codec, evaluation | | 2026-07-02 |
| interspeech-2025-0383 | Voice Conversion for Likability Control via Automated R | | Interspeech | 2025 | VC | transformer-enc-dec | 2026-07-02 |
| interspeech-2025-0433 | When Humans Growl and Birds Speak: High-Fidelity Voice | | Interspeech | 2025 | VC | VAE | 2026-07-02 |
| interspeech-2025-0438 | LinearVC: Linear Transformations of Self-Supervised Fea | | Interspeech | 2025 | VC | hybrid | 2026-07-02 |
| interspeech-2025-0464 | Prosody-Adaptable Audio Codecs for Zero-Shot Voice Conv | | Interspeech | 2025 | VC, codec | autoregressive-LM | 2026-07-02 |
| interspeech-2025-0506 | EnCodecMAE: leveraging neural codecs for universal audi | | Interspeech | 2025 | codec | transformer-enc-dec | 2026-07-02 |
| interspeech-2025-0656 | EEG-based Voice Conversion : Hearing the Voice of Your | Beijing University of Posts and Telecommunications | Interspeech | 2025 | VC | hybrid | 2026-07-02 |
| interspeech-2025-0706 | Contextual Paralinguistic Data Creation for Multi-Modal | | Interspeech | 2025 | SCA | | 2026-07-02 |
| interspeech-2025-0756 | A-SMiLE: Affective Sparse Mixture-of-Experts Adapter wi | | Interspeech | 2025 | SCA | hybrid | 2026-07-02 |
| interspeech-2025-0998 | Voice-ENHANCE: Speech Restoration using a Diffusion-bas | | Interspeech | 2025 | VC | diffusion, GAN | 2026-07-02 |
| interspeech-2025-1020 | Learning Optimal Prosody Embedding Codebook based on F0 | | Interspeech | 2025 | TTS, evaluation | VAE | 2026-07-03 |
| interspeech-2025-1081 | Speaker Normalization and Content Restoration for Zero- | | Interspeech | 2025 | VC | GAN | 2026-07-03 |
| interspeech-2025-1084 | Efficient Streaming TTS Acoustic Model with Depthwise R | | Interspeech | 2025 | TTS | autoregressive-LM | 2026-07-03 |
| interspeech-2025-1098 | GST-BERT-TTS: Prosody Prediction Without Accentual Labe | | Interspeech | 2025 | TTS | transformer-enc-dec | 2026-07-03 |
| interspeech-2025-1106 | LSCodec: Low-Bitrate and Speaker-Decoupled Discrete Spe | | Interspeech | 2025 | codec | VAE, GAN | 2026-07-03 |
| interspeech-2025-1115 | MPE-TTS: Customized Emotion Zero-Shot Text-To-Speech Us | | Interspeech | 2025 | TTS | diffusion, autoregressive-LM | 2026-07-03 |
| interspeech-2025-1192 | Voice Impression Control in Zero-Shot TTS | | Interspeech | 2025 | TTS | transformer-enc-dec | 2026-07-03 |
| interspeech-2025-1210 | DiffEmotionVC: A Dual-Granularity Disentangled Diffusio | | Interspeech | 2025 | VC | diffusion | 2026-07-03 |
| interspeech-2025-1229 | E2E-BPVC: End-to-End Background-Preserving Voice Conver | | Interspeech | 2025 | VC | flow-matching | 2026-07-03 |
| interspeech-2025-1236 | Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment | | Interspeech | 2025 | TTS | flow-matching | 2026-07-03 |
| interspeech-2025-1334 | MiSTR: Multi-Modal iEEG-to-Speech Synthesis with Transf | | Interspeech | 2025 | TTS | transformer-enc-dec | 2026-07-03 |
| interspeech-2025-1364 | VS-Singer: Vision-Guided Stereo Singing Voice Synthesis | | Interspeech | 2025 | singing, TTS | diffusion | 2026-07-03 |
| interspeech-2025-1394 | DiEmo-TTS: Disentangled Emotion Representations via Sel | | Interspeech | 2025 | TTS | transformer-enc-dec | 2026-07-03 |
| interspeech-2025-1397 | VibE-SVC: Vibrato Extraction with High-frequency F0 Con | Korea University | Interspeech | 2025 | singing, VC | diffusion | 2026-07-03 |
| interspeech-2025-1434 | REWIND: Speech Time Reversal for Enhancing Speaker Repr | | Interspeech | 2025 | VC | diffusion | 2026-07-03 |
| interspeech-2025-1478 | LightL2S: Ultra-Low Complexity Lip-to-Speech Synthesis | | Interspeech | 2025 | TTS | hybrid | 2026-07-03 |
| interspeech-2025-1494 | VisualSpeech: Enhancing Prosody Modeling in TTS Using V | | Interspeech | 2025 | TTS | transformer-enc-dec | 2026-07-03 |
| interspeech-2025-1531 | Simple and Effective Content Encoder for Singing Voice | | Interspeech | 2025 | singing, VC | VAE, GAN | 2026-07-03 |
| interspeech-2025-1536 | Fairness in Dysarthric Speech Synthesis: Understanding | | Interspeech | 2025 | TTS, evaluation | flow-matching | 2026-07-03 |
| interspeech-2025-1538 | StarVC: A Unified Auto-Regressive Framework for Joint T | | Interspeech | 2025 | VC | autoregressive-LM | 2026-07-03 |
| interspeech-2025-1550 | ArVoice: A Multi-Speaker Dataset for Arabic Speech Synt | Mohamed Bin Zayed University of Artificial Intelligence | Interspeech | 2025 | TTS, VC, evaluation | transformer-enc-dec, VAE, GAN | 2026-07-03 |
| interspeech-2025-1625 | Mimic Blocker: Self-Supervised Adversarial Training for | | Interspeech | 2025 | VC | GAN | 2026-07-03 |
| interspeech-2025-1638 | EATS-Speech: Emotion-Adaptive Transformation and Priori | | Interspeech | 2025 | TTS | hybrid | 2026-07-03 |
| interspeech-2025-1639 | LombardTokenizer: Disentanglement and Control of Vocal | GIPSA-lab, Univ. Grenoble Alpes | Interspeech | 2025 | codec, VC | GAN | 2026-07-03 |
| interspeech-2025-1684 | SA-RAS: Speaker-Aware Style Retrieval Augmented Generat | | Interspeech | 2025 | TTS | hybrid | 2026-07-03 |
| interspeech-2025-1726 | Voice Reconstruction through Large-Scale TTS Models: Co | | Interspeech | 2025 | TTS, evaluation | hybrid | 2026-07-03 |
| interspeech-2025-1747 | FasterVoiceGrad: Faster One-step Diffusion-Based Voice | NTT, Inc. | Interspeech | 2025 | VC | diffusion, GAN | 2026-07-03 |
| interspeech-2025-1763 | Vocoder-Projected Feature Discriminator | NTT | Interspeech | 2025 | VC | GAN, diffusion | 2026-07-03 |
| interspeech-2025-1776 | SpeechSEC: A Unified Multi-Task Framework for Speech Sy | | Interspeech | 2025 | TTS | hybrid | 2026-07-04 |
| interspeech-2025-1819 | Comparative Analysis of Fast and High-Fidelity Neural V | | Interspeech | 2025 | TTS | GAN | 2026-07-04 |
| interspeech-2025-1873 | Can AI Understand Mandarin Speech Prosody? A Framework | | Interspeech | 2025 | SCA, evaluation | | 2026-07-04 |
| interspeech-2025-1940 | Investigating Stochastic Methods for Prosody Modeling i | | Interspeech | 2025 | TTS | transformer-enc-dec, flow-matching | 2026-07-04 |
| interspeech-2025-2031 | Kinship in Speech: Leveraging Linguistic Relatedness fo | | Interspeech | 2025 | TTS | transformer-enc-dec | 2026-07-04 |
| interspeech-2025-2032 | ExagTTS: An Approach Towards Controllable Word Stress I | IIIT Hyderabad | Interspeech | 2025 | TTS | hybrid | 2026-07-04 |
| interspeech-2025-2075 | Segmentation-Variant Codebooks for Preservation of Para | | Interspeech | 2025 | codec | | 2026-07-04 |
| interspeech-2025-2151 | FaVC: A Validated, Transcribed, Parallel Farsi Speech D | University of Tehran | Interspeech | 2025 | VC, evaluation | GAN | 2026-07-04 |
| interspeech-2025-2159 | Generating Consistent Prosodic Patterns from Open-Sourc | | Interspeech | 2025 | TTS, evaluation | flow-matching | 2026-07-04 |
| interspeech-2025-2189 | ProMode: A Speech Prosody Model Conditioned on Acoustic | | Interspeech | 2025 | TTS | transformer-enc-dec | 2026-07-04 |
| interspeech-2025-2283 | Pairwise Evaluation of Accent Similarity in Speech Synt | | Interspeech | 2025 | TTS, evaluation | | 2026-07-04 |
| interspeech-2025-2328 | A Watermark for Auto-Regressive Speech Generation Model | University of Maryland | Interspeech | 2025 | TTS, evaluation | autoregressive-LM | 2026-07-05 |
| interspeech-2025-2536 | The Text-to-speech in the Wild (TITW) Database | | Interspeech | 2025 | TTS, evaluation | | 2026-07-05 |
| interspeech-2025-2564 | Towards a Japanese Full-duplex Spoken Dialogue System | Nagoya University | Interspeech | 2025 | SCA | autoregressive-LM | 2026-07-05 |
| interspeech-2025-2573 | SawtArabi: A Benchmark Corpus for Arabic TTS. Standard, Dialectal and Code-Switching | | Interspeech | 2025 | TTS, evaluation | flow-matching, GAN | 2026-07-05 |
| interspeech-2025-2586 | Spotlight-TTS: Spotlighting the Style via Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech | Korea University | Interspeech | 2025 | TTS | transformer-enc-dec, GAN | 2026-07-05 |
| interspeech-2025-2595 | Harnessing Text-to-Speech Voice Cloning Models for Improved Audiological Speech Assessment | University of Cambridge | Interspeech | 2025 | TTS, evaluation | | 2026-07-05 |
| interspeech-2025-2679 | Can We Reconstruct a Dysarthric Voice with the Large Speech Model Parler TTS? | University of Edinburgh | Interspeech | 2025 | TTS | autoregressive-LM | 2026-07-05 |
| interspeech-2025-2684 | Discl-VC: Disentangled Discrete Tokens and In-Context Learning for Controllable Zero-Shot Voice Conversion | | Interspeech | 2025 | VC | flow-matching | 2026-07-05 |
| interspeech-2025-2726 | DS-Codec: Dual-Stage Training with Mirror-to-NonMirror Architecture Switching for Speech Codec | | Interspeech | 2025 | codec | GAN, hybrid | 2026-07-05 |
| interspeech-2025-2739 | AF-Vocoder: Artifact-Free Neural Vocoder with Global Artifact Filter | ByteDance | Interspeech | 2025 | TTS | GAN | 2026-07-05 |
| interspeech-2025-2787 | Towards Adaptable and Intelligible Speech Synthesis in Noisy Environments | KTH Royal Institute of Technology | Interspeech | 2025 | TTS, evaluation | autoregressive-LM | 2026-07-05 |
| interspeech-2025-2815 | From Static to Dynamic: Enhancing AAC with Generative Imagery and Zero-Shot TTS | KTH Royal Institute of Technology | Interspeech | 2025 | TTS | | 2026-07-05 |
| interspeech-2025-bokkahallisatish25_interspeech | Hear Me Out: Interactive evaluation and bias discovery platform for speech-to-speech conversational AI | KTH Royal Institute of Technology | Interspeech | 2025 | SCA, evaluation | | 2026-07-05 |
| 2507.16835 | Evaluating Speech-to-Text x LLM x Text-to-Speech Combinations for AI Interview Systems | | arXiv | 2025 | SCA, evaluation | | 2026-07-05 |
| 2411.19770 | Noro: Noise-Robust One-shot Voice Conversion with Hidden Speaker Representation Learning | | arXiv | 2025 | VC | diffusion | 2026-07-05 |
| 2025.clicit-1.27 | Veras Audire Et Reddere Voces: A Corpus of Prosodically-Correct Latin Poetic Audio from Large-Language-Model TTS | | workshop | 2025 | TTS, evaluation | autoregressive-LM | 2026-07-05 |
| 2506.23367 | You Sound a Little Tense: L2 Tailored Clear TTS Using Durational Vowel Properties | | arXiv | 2025 | TTS | flow-matching | 2026-07-12 |
| 2509.05359 | An Empirical Analysis of Discrete Unit Representations in Speech Language Modeling Pre-training | | arXiv | 2025 | SCA | autoregressive-LM | 2026-07-12 |
| 2509.04093 | Open-Source Full-Duplex Conversational Datasets for Natural and Interactive Speech Synthesis | | arXiv | 2025 | TTS, SCA | autoregressive-LM, flow-matching | 2026-07-12 |
| 2509.04667 | DarkStream: real-time speech anonymization with low latency | Texas A&M University | arXiv | 2025 | VC | GAN, hybrid | 2026-07-12 |
| 2509.04685 | Say More with Less: Variable-Frame-Rate Speech Tokenization via Adaptive Clustering and Implicit Duration Coding | | arXiv | 2025 | codec | GAN | 2026-07-12 |
| 2509.04702 | OleSpeech-IV: A Large-Scale Multispeaker and Multilingual Conversational Speech Dataset with Diverse Topics | Olewave | arXiv | 2025 | TTS, SCA | | 2026-07-12 |
| 2509.05863 | LatinX: Aligning a Multilingual TTS Model with Direct Preference Optimization | | arXiv | 2025 | TTS | autoregressive-LM | 2026-07-12 |
| 2509.06074 | Multimodal Fine-grained Context Interaction Graph Modeling for Conversational Speech Synthesis | | EMNLP | 2025 | TTS | transformer-enc-dec | 2026-07-12 |
| 2509.06502 | FireRedChat: A Pluggable, Full-Duplex Voice Interaction System with Cascaded and Semi-Cascaded Implementations | Xiaohongshu | arXiv | 2025 | SCA | hybrid | 2026-07-12 |
| 2509.07038 | Controllable Singing Voice Synthesis using Phoneme-Level Energy Sequence | Korea University | arXiv | 2025 | singing | diffusion | 2026-07-12 |
| 2509.07376 | Progressive Facial Granularity Aggregation with Bilateral Attribute-based Enhancement for Face-to-Speech Synthesis | POSTECH | EMNLP | 2025 | TTS | VAE | 2026-07-12 |
| 2509.09716 | VStyle: A Benchmark for Voice Style Adaptation with Spoken Instructions | | arXiv | 2025 | TTS, evaluation | | 2026-07-12 |
| 2509.08379 | Flow-Matching Models | | arXiv | 2025 | VC | diffusion, flow-matching | 2026-07-12 |
| 2509.08696 | Accelerating Diffusion Transformer-Based Text-to-Speech with Transformer Layer Caching | National University of Singapore | arXiv | 2025 | TTS | flow-matching | 2026-07-12 |
| 2506.04077 | A Novel Data Augmentation Approach for Automatic Speaking Assessment on Opinion Expressions | National Taiwan Normal University | arXiv | 2025 | TTS | autoregressive-LM | 2026-07-12 |
| 2509.09174 | EchoX: Towards Mitigating Acoustic-Semantic Gap via Echo Training for Speech-to-Speech LLMs | The Chinese University of Hong Kong, Shenzhen | arXiv | 2025 | SCA | autoregressive-LM | 2026-07-12 |
| 2509.09201 | DeCodec: Rethinking Audio Codecs as Universal Disentangled Representation Learners | China Mobile | arXiv | 2025 | codec | GAN | 2026-07-13 |
| 2509.09550 | Finite Scalar Quantization Enables Redundant and Transmission-Robust Neural Audio Compression at Low Bit-rates | Neuphonic | arXiv | 2025 | codec | hybrid, GAN | 2026-07-13 |
| 2509.09748 | DiTReducio: A Training-Free Acceleration for DiT-Based TTS via Progressive Calibration | | arXiv | 2025 | TTS | flow-matching | 2026-07-13 |
| 2509.11084 | Length-Aware Rotary Position Embedding for Text-Speech Alignment | Supertone, Inc. | arXiv | 2025 | TTS | flow-matching | 2026-07-13 |
| 2509.11425 | FuseCodec: Semantic-Contextual Fusion and Supervision for Neural Codecs | | arXiv | 2025 | codec, TTS | GAN, autoregressive-LM | 2026-07-13 |
| 2508.18240 | MTalk-Bench: Evaluating Speech-to-Speech Models in Multi-Turn Dialogues via Arena-style and Rubrics Protocols | The Chinese University of Hong Kong, Shenzhen | arXiv | 2025 | SCA, evaluation | | 2026-07-13 |
| 2509.12171 | Preservation of Language Understanding Capabilities in Speech-aware Large Language Models | | arXiv | 2025 | SCA, evaluation | flow-matching | 2026-07-13 |
| 2509.14270 | SpeechWeave: Diverse Multilingual Synthetic Text & Audio Data Generation Pipeline for Training Text to Speech Models | Oracle AI | ACL | 2025 | TTS | | 2026-07-13 |
| 2509.12831 | A Lightweight Pipeline for Noisy Speech Voice Cloning and Accurate Lip Sync Synthesis | International Islamic University, Islamabad | arXiv | 2025 | TTS | hybrid, GAN | 2026-07-13 |
| 2509.13068 | MSR-Codec: A Low-Bitrate Multi-Stream Residual Codec for High-Fidelity Speech Generation with Information Disentanglement | LIGHTSPEED | arXiv | 2025 | TTS, codec, VC | autoregressive-LM, VAE | 2026-07-13 |
| 2412.16846 | KALL-E: Autoregressive Speech Synthesis with Next-Distribution Prediction | Northwestern Polytechnical University | arXiv | 2025 | TTS | autoregressive-LM, VAE | 2026-07-13 |
| 2504.20581 | ClonEval: An Open Voice Cloning Benchmark | Adam Mickiewicz University | arXiv | 2025 | TTS, evaluation | | 2026-07-13 |
| interspeech-2025-cho25c_interspeech | Unleashing the Inner Monster: Demonstrating High-Fidelity Human to Non-Human Voice Conversion | NC AI Co., Ltd | Interspeech | 2025 | VC | hybrid | 2026-07-13 |
| interspeech-2025-gourav25_interspeech | Code Mix TTS: An Approach to Infer Human Like Speech for Multi-Lingual Input Texts | Oracle Corporation | Interspeech | 2025 | TTS | diffusion, GAN | 2026-07-13 |
| interspeech-2025-raju25_interspeech | End-to-End Indian Language Dubbing with Zero-Shot Speaker Preservation | Hitloop | Interspeech | 2025 | TTS | flow-matching | 2026-07-13 |
| 2509.13667 | A Distilled Low-Latency Neural Vocoder with Explicit Amplitude and Phase Prediction | | arXiv | 2025 | TTS | GAN | 2026-07-13 |
| 2509.13670 | A High-Quality and Low-Complexity Streamable Neural Speech Codec with Knowledge Distillation | University of Science and Technology of China | arXiv | 2025 | codec | GAN | 2026-07-13 |
| 2509.13989 | Do You Hear What I Mean? Quantifying the Instruction-Perception Gap in Instruction-Guided Expressive Text-To-Speech Systems | | arXiv | 2025 | TTS, evaluation | | 2026-07-14 |
| 2509.14579 | Cross-Lingual F5-TTS: Towards Language-Agnostic Voice Cloning and Speech Synthesis | Shanghai Jiao Tong University | arXiv | 2025 | TTS, VC | flow-matching | 2026-07-14 |
| 2509.14684 | DAIEN-TTS: Disentangled Audio Infilling for Environment-Aware Text-to-Speech Synthesis | | arXiv | 2025 | TTS | flow-matching | 2026-07-14 |
| 2509.14784 | MELA-TTS: Joint transformer-diffusion model with representation alignment for speech synthesis | | arXiv | 2025 | TTS | hybrid | 2026-07-14 |
| 2509.14946 | SynParaSpeech: Automated Synthesis of Paralinguistic Datasets for Speech Generation and Understanding | | arXiv | 2025 | TTS | autoregressive-LM, flow-matching | 2026-07-14 |
| 2509.15085 | Real-Time Streaming Mel Vocoding with Generative Flow Matching | University of Hamburg | arXiv | 2025 | TTS | flow-matching | 2026-07-14 |
| 2509.15253 | Emotion-Aware Speech Generation with Character-Specific Voices for Comics | Queen Mary University of London | arXiv | 2025 | TTS | | 2026-07-14 |
| 2509.15462 | A Novel Semantic Compression Approach for Ultra-low Bandwidth Voice Communication | Systems & Technology Research | arXiv | 2025 | codec, VC | autoregressive-LM, flow-matching | 2026-07-14 |
| 2505.17093 | P2VA: Converting Persona Descriptions into Voice Attributes for Fair and Controllable Text-to-Speech | | arXiv | 2025 | TTS | | 2026-07-14 |
| 2509.15626 | LibriTTS-VI: A Public Corpus and Novel Methods for Efficient Voice Impression Control | Sony Group Corporation | arXiv | 2025 | TTS | VAE | 2026-07-14 |
| 2509.15629 | The Singing Voice Conversion Challenge 2025: From Singer Identity Conversion To Singing Style Conversion | | arXiv | 2025 | VC, singing, evaluation | | 2026-07-14 |
| 2509.15845 | Deep Dubbing: End-to-End Auto-Audiobook System with Text-to-Timbre and Context-Aware Instruct-TTS | | arXiv | 2025 | TTS | flow-matching, autoregressive-LM | 2026-07-14 |
| 2509.16010 | Fed-PISA: Federated Voice Cloning via Personalized Identity-Style Adaptation | | arXiv | 2025 | TTS, VC | hybrid | 2026-07-14 |
| 2509.16195 | FocalCodec-Stream: Streaming Low-Bitrate Speech Coding via Causal Distillation | | arXiv | 2025 | codec, VC | hybrid | 2026-07-14 |
| 2509.16589 | Benchmarking Contextual and Paralinguistic Reasoning in Speech-LLMs: A Case Study with In-the-Wild Data | | EMNLP | 2025 | SCA, evaluation | | 2026-07-14 |
| 2509.20378 | Beyond Global Emotion: Fine-Grained Emotional Speech Synthesis with Dynamic Word-Level Modulation | Harbin Institute of Technology | arXiv | 2025 | TTS | autoregressive-LM | 2026-07-14 |
| 2509.17006 | MBCodec: Thorough Disentangle for High-Fidelity Audio Compression | | arXiv | 2025 | codec | GAN | 2026-07-14 |
| 2509.17021 | Bridging the gap between training and inference in LM-based TTS models | | arXiv | 2025 | TTS | autoregressive-LM | 2026-07-14 |
| 2509.17143 | MaskVCT: Masked Voice Codec Transformer for Zero-Shot Voice Conversion With Increased Controllability via Multiple Guidances | | arXiv | 2025 | VC | autoregressive-LM | 2026-07-14 |
| 2509.14882 | Llama-Mimi: Exploring the Limits of Flattened Speech Language Modeling | | arXiv | 2025 | SCA | autoregressive-LM | 2026-07-14 |
| 2509.17516 | Audiobook-CC: Controllable Long-context Speech Generation for Multicast Audiobook | Ximalaya Inc. | arXiv | 2025 | TTS | autoregressive-LM, flow-matching, GAN | 2026-07-15 |
| 2509.17765 | Qwen3-Omni Technical Report | Qwen Team (Alibaba) | arXiv | 2025 | SCA, TTS | autoregressive-LM, hybrid | 2026-07-15 |
| 2509.17988 | Nord-Parl-TTS: Finnish and Swedish TTS Dataset from Parliament Speech | | arXiv | 2025 | TTS | flow-matching | 2026-07-15 |
| 2509.18060 | TMD-TTS: A Unified Tibetan Multi-Dialect Text-to-Speech Framework for Ü-Tsang, Amdo and Kham Speech Dataset Generation | University of Electronic Science and Technology of China | arXiv | 2025 | TTS | flow-matching | 2026-07-15 |
| 2509.18470 | Discrete-Time Diffusion-Like Models for Speech Synthesis | University of Sheffield | arXiv | 2025 | TTS | diffusion, flow-matching | 2026-07-15 |
| 2501.04561 | OpenOmni: Advancing Open-Source Omnimodal Large Language Models with Progressive Multimodal Alignment and Real-Time Self-Aware Emotional Speech Synthesis | Shenzhen Institute of Advanced Technology, CAS; Alibaba (Tongyi Lab) | arXiv | 2025 | SCA, TTS | autoregressive-LM, hybrid | 2026-07-15 |
| 2509.18531 | No Verifiable Reward for Prosody: Toward Preference-Guided Prosody Learning in TTS | | arXiv | 2025 | TTS | autoregressive-LM | 2026-07-15 |
| 2509.18806 | Rethinking the joint estimation of magnitude and phase for time-frequency domain neural vocoders | Institute of Acoustics, Chinese Academy of Sciences; Tencent AI Lab | arXiv | 2025 | TTS | GAN | 2026-07-15 |
| 2509.18823 | Towards Evaluating Generative Audio: Insights from Neural Audio Codec Embedding Distances | Dolby | arXiv | 2025 | evaluation, codec | GAN | 2026-07-15 |
| 2509.18928 | Direct Preference Optimization for Speech Autoregressive Diffusion Models | ByteDance Seed | arXiv | 2025 | TTS | autoregressive-LM, diffusion, hybrid | 2026-07-15 |
| 2509.19025 | Enhancing Noise Robustness for Neural Speech Codecs through Resource-Efficient Progressive Quantization Perturbation Simulation | University of Science and Technology of China | arXiv | 2025 | codec | GAN | 2026-07-15 |
| 2509.19186 | Improving Test-Time Performance of RVQ-based Neural Codecs | Supertone Inc. | arXiv | 2025 | codec | GAN | 2026-07-15 |
| 2509.19231 | Finding My Voice: Generative Reconstruction of Disordered Speech for Automated Clinical Evaluation | Carnegie Mellon University | arXiv | 2025 | TTS, VC, evaluation | diffusion, GAN | 2026-07-15 |
| 2509.19592 | Frame-Stacked Local Transformers For Efficient Multi-Codebook Speech Generation | NVIDIA | arXiv | 2025 | TTS | autoregressive-LM | 2026-07-15 |
| 2509.19812 | Efficient Speech Watermarking for Speech Synthesis via Progressive Knowledge Distillation | Microsoft | arXiv | 2025 | TTS | GAN | 2026-07-15 |
| 2509.19883 | CoMelSinger: Discrete Token-Based Zero-Shot Singing Synthesis With Structured Melody Control and Guidance | National University of Singapore | arXiv | 2025 | singing | hybrid | 2026-07-15 |
| 2509.19928 | Measuring Prosody Diversity in Zero-Shot TTS: A New Metric, Benchmark, and Exploration | | arXiv | 2025 | TTS, evaluation | | 2026-07-16 |
| 2509.20086 | OLaPh: Optimal Language Phonemizer | Hof University of Applied Sciences | arXiv | 2025 | TTS | autoregressive-LM | 2026-07-16 |
| 2509.20321 | Conversational Speech Reveals Structural Robustness Failures in SpeechLLM Backbones | Texas A&M University | arXiv | 2025 | SCA, evaluation | autoregressive-LM | 2026-07-16 |
| 2509.20410 | Phoenix-VAD: Streaming Semantic Endpoint Detection for Full-Duplex Speech Interaction | Xiamen University, DiDi Global | arXiv | 2025 | SCA | autoregressive-LM | 2026-07-16 |
| 2509.20485 | Objective Evaluation of Prosody and Intelligibility in Speech Synthesis via Conditional Prediction of Discrete Tokens | Johns Hopkins University; National University of Singapore | arXiv | 2025 | evaluation, TTS | | 2026-07-16 |
| 2509.22718 | PerformSinger: Multimodal Singing Voice Synthesis Leveraging Synchronized Lip Cues from Singing Performance Videos | | arXiv | 2025 | singing | hybrid | 2026-07-16 |
| 2505.10599 | UDDETTS: Unifying Discrete and Dimensional Emotions for Controllable Emotional Text-to-Speech | | arXiv | 2025 | TTS | autoregressive-LM, flow-matching, GAN, hybrid | 2026-07-16 |
| 2509.20802 | SPADE: Structured Pruning and Adaptive Distillation for Efficient LLM-TTS | KAIST, 42dot Inc. | arXiv | 2025 | TTS | autoregressive-LM | 2026-07-16 |
| 2509.22727 | DiaMoE-TTS: A Unified IPA-Based Dialect TTS Framework with Mixture-of-Experts and Parameter-Efficient Zero-Shot Adaptation | Tsinghua University, Giant Network AI Lab | arXiv | 2025 | TTS | flow-matching | 2026-07-16 |
| 2506.21875 | WildSpeech-Bench: Benchmarking End-to-End SpeechLLMs in the Wild | WeChat AI, Tencent | arXiv | 2025 | SCA, evaluation | | 2026-07-16 |
| 2509.21968 | AUV: Teaching Audio Universal Vector Quantization with Single Nested Codebook | | arXiv | 2025 | codec | GAN, VAE | 2026-07-16 |
| 2509.22062 | Comprehend and Talk: Text to Speech Synthesis via Dual Language Modeling | AMAP Speech, Tsinghua University | arXiv | 2025 | TTS | autoregressive-LM | 2026-07-17 |
| 2509.22167 | Semantic-VAE: Semantic-Alignment Latent Representation | Shanghai Jiao Tong University | arXiv | 2025 | TTS | VAE | 2026-07-17 |
| 2509.22243 | FLEXI: Benchmarking Full-duplex Human-LLM Speech Interaction | Northeastern University | arXiv | 2025 | SCA, evaluation | | 2026-07-17 |
| 2509.23147 | BFA: Real-time Multilingual Text-to-speech Forced Alignment | Bournemouth University | arXiv | 2025 | TTS | | 2026-07-17 |
| 2510.02352 | Evaluating Bias in Spoken Dialogue LLMs for Real-World | | arXiv | 2025 | SCA | | 2026-07-17 |
| 2509.23938 | Easy Turn: Integrating Acoustic and Linguistic Modaliti | | arXiv | 2025 | SCA | hybrid | 2026-07-17 |
| 2509.24457 | Assessing speech quality metrics for evaluation of neur | Cisco Systems | arXiv | 2025 | codec, evaluation | | 2026-07-17 |
| 2509.24570 | ISSE: An Instruction-Guided Speech Style Editing Dataset And Benchmark | | arXiv | 2025 | TTS, VC, evaluation | autoregressive-LM | 2026-07-17 |
| 2509.24650 | VoxCPM: Tokenizer-Free TTS for Context-Aware Speech Generation and True-to-Life Voice Cloning | | arXiv | 2025 | TTS, VC | autoregressive-LM, diffusion | 2026-07-17 |
| 2509.24773 | VSSFlow: Unifying Video-conditioned Sound and Speech Ge | | arXiv | 2025 | TTS | flow-matching | 2026-07-17 |
| 2509.25131 | MGM-Omni: Scaling Omni LLMs to Personalized Long-Horizo | CUHK, HKUST, SmartMore | arXiv | 2025 | SCA, TTS | autoregressive-LM, flow-matching, hybrid | 2026-07-17 |
| 2509.25416 | Emotion-Aligned Generation in Diffusion Text to Speech Models via Preference-Guided Optimization | | arXiv | 2025 | TTS | diffusion | 2026-07-17 |
| 2509.26276 | Optimizing Speech Language Models for Acoustic Consistency | University of Zurich | arXiv | 2025 | TTS, SCA | autoregressive-LM | 2026-07-17 |
| 2509.26514 | BatonVoice: An Operationalist Framework for Enhancing C | Tencent | arXiv | 2025 | TTS | autoregressive-LM, flow-matching, GAN | 2026-07-17 |
| 2509.26542 | Voice Evaluation of Reasoning Ability: Diagnosing the Modality-Induced Performance Gap | | arXiv | 2025 | SCA, evaluation | | 2026-07-17 |
| 2510.00264 | Baseline Systems For The 2025 Low-Resource Audio Codec Challenge | Cisco Systems (Collaboration AI) | arXiv | 2025 | codec, evaluation | GAN | 2026-07-17 |
| 2510.00499 | MOSS-Speech: Towards True Speech-to-Speech Models Without Text Guidance | Shanghai Innovation Institute, Fudan University, MOSI | arXiv | 2025 | SCA | autoregressive-LM, flow-matching | 2026-07-17 |
| 2510.00743 | From Scores to Preferences: Redefining MOS Benchmarking for Speech Quality Reward Modeling | Fudan University | arXiv | 2025 | evaluation | autoregressive-LM | 2026-07-17 |
| 2025.vlsp-1.15 | Twinkle-VC: A Robust and High-Quality Zero-Shot Voice Conversion System for the VLSP 2025 Shared Task | | VLSP 2025 | 2025 | VC | diffusion | 2026-07-17 |
| 2025.vlsp-1.14 | ViettelRoar: Voice conversion approach for VLSP 2025 | ViettelAI, Viettel Group | VLSP 2025 | 2025 | VC | flow-matching | 2026-07-17 |
| 2025.vlsp-1.13 | The 2025 VLSP Task on Vietnamese Voice Conversion: Overview and Preliminary Results | Hanoi University of Science and Technology | VLSP 2025 | 2025 | VC, evaluation | | 2026-07-17 |
| 2510.05150 | Chronological Thinking in Full-Duplex Spoken Dialogue Language Models | Nanyang Technological University, StepFun, Mila | arXiv | 2025 | SCA | autoregressive-LM | 2026-07-17 |
| 2510.02066 | Chain-of-Thought Reasoning in Streaming Full-Duplex End-to-End Spoken Dialogue Systems | Carnegie Mellon University, Sony Group Corporation | arXiv | 2025 | SCA | autoregressive-LM | 2026-07-17 |
| 2510.01722 | Emotional Text-To-Speech Based on Mutual-Information-Guided Emotion-Timbre Disentanglement | The University of Tokyo / Institute of Science Tokyo | arXiv | 2025 | TTS | transformer-enc-dec | 2026-07-17 |
| 2510.01903 | MelTok: 2D Tokenization for Single-Codebook Audio Compression | International Digital Economy Academy (IDEA) | arXiv | 2025 | codec | VAE, GAN | 2026-07-17 |
| 2510.02044 | Stream RAG: Instant and Accurate Spoken Dialogue Systems with Streaming Tool Usage | Meta, Carnegie Mellon University | arXiv | 2025 | SCA | autoregressive-LM | 2026-07-17 |
| 2510.03111 | Evaluation of preprocessing pipelines in the creation of in-the-wild TTS datasets | Universidad Nacional de Tres de Febrero | arXiv | 2025 | TTS, evaluation | | 2026-07-17 |
| 2510.03735 | Soft Disentanglement in Frequency Bands for Neural Audio Codecs | Télécom Paris | arXiv | 2025 | codec | GAN | 2026-07-17 |
| 2510.04738 | Speak, Edit, Repeat: High-Fidelity Voice Editing and Zero-Shot TTS with Cross-Attentive Mamba | MTS AI, ITMO University | arXiv | 2025 | TTS | autoregressive-LM, hybrid | 2026-07-17 |
| 2510.05619 | Teaching Machines to Speak Using Articulatory Control | UC Berkeley | arXiv | 2025 | TTS | hybrid | 2026-07-17 |
| 2510.05984 | ECTSpeech: Enhancing Efficient Speech Synthesis via Easy Consistency Tuning | Xinjiang University | arXiv | 2025 | TTS | diffusion, transformer-enc-dec | 2026-07-17 |
| 2510.05799 | Data-efficient Targeted Token-level Preference Optimization for LLM-based Text-to-Speech | SpiralAI Inc. | arXiv | 2025 | TTS | autoregressive-LM | 2026-07-17 |
| 2506.15556 | PredGen: Accelerated Inference of Large Language Models through Input-Time Speculation for Real-Time Speech Interaction | University of California, Los Angeles | arXiv | 2025 | SCA, TTS | autoregressive-LM | 2026-07-18 |
| 2510.07096 | Modeling Sarcastic Speech: Semantic and Prosodic Cues in a Speech Synthesis Framework | University of Groningen | arXiv | 2025 | TTS | GAN, VAE | 2026-07-18 |
| 2510.06917 | SHANKS: Simultaneous Hearing and Thinking for Spoken Language Models | National Taiwan University, Microsoft | arXiv | 2025 | SCA | autoregressive-LM | 2026-07-18 |
| 2510.06927 | Position: Towards Responsible Evaluation for Text-to-Speech | | arXiv | 2025 | TTS, evaluation | | 2026-07-18 |
| 2510.07881 | CS3-Bench: Evaluating and Enhancing Speech-to-Speech LLMs for Mandarin-English Code-Switching | Shanghai Jiao Tong University, Ant Group | arXiv | 2025 | SCA, evaluation | autoregressive-LM | 2026-07-18 |
| 2510.08373 | DialoSpeech: Dual-Speaker Dialogue Generation with LLM and Flow Matching | Northwestern Polytechnical University | arXiv | 2025 | TTS | autoregressive-LM, flow-matching | 2026-07-18 |
| 2510.08392 | MeanVC: Lightweight and Streaming Zero-Shot Voice Conversion via Mean Flows | Northwestern Polytechnical University (ASLP@NPU) | arXiv | 2025 | VC | flow-matching, hybrid | 2026-07-18 |
| 2510.07978 | VoiceAgentBench: Are Voice Assistants ready for agentic tasks? | Ola Electric / Krutrim | arXiv | 2025 | SCA, evaluation | autoregressive-LM | 2026-07-18 |
| 2510.09061 | O_O-VC: Synthetic Data-Driven One-to-One Alignment for Any-to-Any Voice Conversion | VNPT AI / Hanoi University of Science and Technology / National Economics University | EMNLP | 2025 | VC | VAE, GAN | 2026-07-18 |
| 2510.09016 | DiTSinger: Scaling Singing Voice Synthesis with Diffusion Transformer and Implicit Alignment | Migu Music, China Mobile Communications Corporation | arXiv | 2025 | singing | diffusion | 2026-07-18 |
| 2506.12311 | Phonikud: Hebrew Grapheme-to-Phoneme Conversion for Real-Time Text-to-Speech | Independent Researcher; Reichman University; Tel Aviv University | arXiv | 2025 | TTS | GAN, diffusion | 2026-07-18 |
| 2510.09424 | The Speech-LLM Takes It All: A Truly Fully End-to-End Spoken Dialogue State Tracking Approach | Orange Research | arXiv | 2025 | SCA | autoregressive-LM | 2026-07-18 |
| 2510.09592 | Mind-Paced Speaking: A Dual-Brain Approach to Real-Time Reasoning in Spoken Language Models | StepFun | arXiv | 2025 | SCA | autoregressive-LM | 2026-07-18 |
| 2510.09245 | SynthVC: Leveraging Synthetic Data for End-to-End Low Latency Streaming Voice Conversion | Northwestern Polytechnical University (ASLP@NPU) | arXiv | 2025 | VC | GAN | 2026-07-18 |
| 2510.10003 | MTP-S2UT: Enhancing Speech-to-Speech Translation Quality with Multi-token Prediction | Northeastern University; NiuTrans Research; Kunming University of Science and Technology | arXiv | 2025 | TTS | transformer-enc-dec | 2026-07-18 |
| 2510.10774 | ParsVoice: A Large-Scale Multi-Speaker Persian Speech Corpus for Text-to-Speech Synthesis | University of Tehran | arXiv | 2025 | TTS | autoregressive-LM, GAN | 2026-07-18 |
| 2510.11646 | BridgeCode: A Dual Speech Representation Paradigm for Autoregressive Zero-Shot Text-to-Speech Synthesis | South China University of Technology | arXiv | 2025 | TTS | autoregressive-LM | 2026-07-18 |
| 2510.11124 | Perturbation Self-Supervised Representations for Cross-Lingual Emotion TTS: Stage-Wise Modeling of Emotion and Speaker | Tianjin University | arXiv | 2025 | TTS | transformer-enc-dec, GAN | 2026-07-18 |
| 2510.12964 | VCTR: A Transformer-Based Model for Non-parallel Voice Conversion | Independent Researcher | arXiv | 2025 | VC | GAN, hybrid | 2026-07-18 |
| 2510.12995 | Continuous-Token Diffusion for Speaker-Referenced TTS in Multimodal LLMs | Amazon AGI | arXiv | 2025 | TTS | autoregressive-LM, diffusion | 2026-07-18 |
| 2510.13221 | Acoustic Teleportation via Disentangled Neural Audio Codec Representations | Fraunhofer IIS / International Audio Laboratories Erlangen | arXiv | 2025 | codec | GAN | 2026-07-18 |
| 2510.13293 | Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS Models | Alibaba, Nanyang Technological University | arXiv | 2025 | TTS | autoregressive-LM | 2026-07-18 |
| 2510.13194 | StressTransfer: Stress-Aware Speech-to-Speech Translation with Emphasis Preservation | The Chinese University of Hong Kong; Nara Institute of Science and Technology | arXiv | 2025 | TTS | autoregressive-LM | 2026-07-18 |
| 2510.15364 | LDCodec: A high quality neural audio codec with low-complexity decoder | ByteDance | arXiv | 2025 | codec | GAN | 2026-07-18 |
| 2510.15227 | LongCat-Audio-Codec: An Audio Tokenizer and Detokenizer Solution Designed for Speech Large Language Models | Meituan (LongCat Team) | arXiv | 2025 | codec | GAN, hybrid | 2026-07-18 |
| 2510.16841 | SAC: Neural Speech Codec with Semantic-Acoustic Dual-Stream Quantization | Shanghai Jiao Tong University (X-LANCE Lab), Soul AI Lab | arXiv | 2025 | codec | GAN, VAE, autoregressive-LM | 2026-07-18 |
| 2510.16718 | U-Codec: Ultra Low Frame-rate Neural Speech Codec for Fast High-fidelity Speech Generation | Peking University / Tencent AI Lab / Tencent Hunyuan | arXiv | 2025 | codec, TTS | autoregressive-LM, GAN | 2026-07-18 |
| 2503.06211 | Late Fusion and Multi-Level Fission Amplify Cross-Modal Transfer in Text-Speech LMs | Université de Toulon (LIS) / University of Cambridge | arXiv | 2025 | SCA, TTS | autoregressive-LM | 2026-07-18 |
| 2510.18308 | ParaStyleTTS: Toward Efficient and Robust Paralinguistic Style Control for Expressive Text-to-Speech Generation | University of New South Wales | arXiv | 2025 | TTS | VAE, GAN | 2026-07-18 |
| 2506.23670 | Efficient Interleaved Speech Modeling through Knowledge Distillation | Nlpie Research; University of Zurich | arXiv | 2025 | TTS, SCA | autoregressive-LM | 2026-07-18 |
| 2510.19509 | Which Evaluation for Which Model? A Taxonomy for Speech Model Assessment | Apple | arXiv | 2025 | evaluation | | 2026-07-18 |
| 2510.10785 | FAC-FACodec: Controllable Zero-Shot Foreign Accent Conversion with Factorized Speech Codec | University of Illinois Urbana-Champaign | arXiv | 2025 | VC | diffusion | 2026-07-18 |