001 Cocktail-Talker: Multi-Speaker Dialog Modeling in Noisy Social Environments with Turn Action GRPO arXiv Paper Xilin Jiang et al. Yesterday 002 Integrating Contextual Embeddings into Evaluation of Expressive MIDI Piano Performances arXiv Paper Dmitrii Gavrilev et al. Yesterday 003 RIPPLE: Generating Multi-Channel Phase, Not Recovering It arXiv Paper Jaehyuk Lee et al. Yesterday 004 Teffic-Audio: Tell Fact from Fiction arXiv Paper Wan Lin et al. Yesterday 005 VocalRender: Score-Native Singing Voice Synthesis for Real-World Composition arXiv Paper Yukun Chen et al. Yesterday 006 Audio-Anchored Fusion of Multi-Ratio DiT Reconstruction Residuals for Cross-Domain Audio Deepfake Detection arXiv Paper Haotian Mo et al. 2 days ago 007 Few-Shot Open-Set Audio Classification via Transductive Prototype Refinement and Class Logit Enhancement arXiv Paper Tianyan Deng et al. 2 days ago 008 MMAC: A Massive Multi-dimensional Benchmark for Audio Captioning arXiv Paper Weijie Wu et al. 2 days ago 009 MPEcho: A Melody and Phoneme-Aware Generative Framework for Controllable Cover Song Generation arXiv Paper Wei-Jaw Lee et al. 2 days ago 010 Prosody-driven Jailbreaks in Audio LLMs: A Controlled Study and Mechanistic Analysis arXiv Paper Jiachen Qian et al. 2 days ago 011 Voice Memory for Agentic Speech Recognition arXiv Paper Chao-Han Huck Yang et al. 2 days ago 012 Zero-Shot Face-to-Speech Synthesis via Latent Space Adaptation of a Style-Diffusion TTS Model arXiv Paper Carlos Muñoz-Romero et al. 2 days ago 013 A large-scale corpus of religious radio broadcast transcripts from webstream recordings in the United States arXiv Paper Samuel Bestvater et al. 3 days ago 014 Device Invariance using Domain Adaptation on Acoustic Scene Classification arXiv Paper A. Dileep, Shubham Sharma et al. 3 days ago 015 Dissecting Sensitivity to Training Language in Self-Supervised Speech Learning Using Neural Audio Codec Tokens arXiv Paper Daigo Takizawa et al. 3 days ago 016 Disentangling Acoustic Cues in Alzheimer's Pathology and Perception: The Roles of Language and Gender arXiv Paper Liu He et al. 4 days ago 017 Enhancing Law-Enforcement Audio Transcription: A LoRA-Based Adaptation of Whisper for BWC Footage arXiv Paper Vivek Senthil et al. 4 days ago 018 MoLGE: Mixture of Language Group Experts for Efficient Scaling of Massively Multilingual Speech Recognition arXiv Paper Sangmin Lee et al. 4 days ago 019 MusiChat: Vibe Composing for Music Creation arXiv Paper Callie C. Liao et al. 4 days ago 020 Memory Efficient Audio Synthesis with Decoupled Temporal Depth Diffusion Transformers arXiv Paper Dongseong Hwang et al. 5 days ago 021 OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation arXiv Paper Jun Zhan et al. 5 days ago 022 Low-Latency Turn-Taking via Context-Aware Preface Generation in a Real-World Dialogue Robot arXiv Paper Yuki Okafuji et al. 6 days ago 023 Music-Source-Separation-Training (MSST): A Unified Framework for Training and Evaluating Music Demixing Models arXiv Paper Roman Solovyev et al. 6 days ago 024 PathRIR: Physics-Guided Acoustic Path Selection and Late-Tail Compensation for Fast Room Impulse Response Simulation arXiv Paper Shaoheng Xu et al. 6 days ago 025 Kutti AI: A Voice-First, Offline-Capable Learning Companion with Real-Time Struggle Detection for Visually-Impaired Children arXiv Paper Kadharmoideen Fadurudeen 7 days ago 026 MemNMF: Memory-Augmented NMF on LPC Spectra for Anomalous Sound Detection arXiv Paper Phurich Saengthong, Takahiro Shinozaki 7 days ago 027 Multimodal Domain Generalization for Depression Detection: An Attention-Based BiLSTM Network with Domain-Adversarial Training arXiv Paper Ali Tabaraei et al. 7 days ago 028 Reflector: Arrangement-Aware Harmonic Retrieval for Sample-Based Composition arXiv Paper Austin Rockman 7 days ago 029 Synthetic Speech, Real Signal: Paralinguistic Preservation and Cross-Lingual Augmentation via Voice Cloning arXiv Paper R. Polle, Owen Parsons et al. 7 days ago 030 Transforming Keystroke Noise to Text: Self-Supervised Acoustic Eavesdropping Attacks on Keyboards arXiv Paper Atsunori Okada, Akira Ito et al. 7 days ago 031 An Evaluation Framework for Structured Audio Captions Validated by Controlled Perturbations arXiv Paper Liang-Yuan Wu, Sripathi Sridhar et al. Jul 23 032 Probing Speaker Identity Sensitivity in Audio Deepfake Detectors arXiv Paper D. Dar, Arun Ross Jul 23 033 VibeVoice-ASR-BitNet Technical Report arXiv Paper Songcheng Xu, Ting Song et al. Jul 23 034 Audio-Zero: Label-Free Self-Evolution for Fine-Grained Audio Reasoning arXiv Paper Siqian Tong, Xuan Li et al. Jul 22 035 Cumsum-Composable Phase Transport for Low-Cost Streaming Keyword Spotting arXiv Paper Mahesh Godavarti Jul 22 036 Efficient Chain-of-Modality Reasoning via Progressive Compression for Spoken Language Models arXiv Paper Pengchao Feng et al. Jul 22 037 Pushing the Frontier of Full-Song Generation: Hierarchical Autoregressive Planning Meets Flow-Matching Rendering arXiv Paper Junyu Dai, Xinyue Fan et al. Jul 22 038 RPPNet: Perceptually-Grouped Rhythm-Pitch Primitives for Long-Term Structure Melody Generation via Boundary-Aware Modeling arXiv Paper Tieyao Zhang, Yuke Liu et al. Jul 22 039 Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio arXiv Paper Abdul Basit Tonmoy, Kazi Fardinul Hoque et al. Jul 21 040 What the Waveform Knows: Transparent-first Speech and Audio Intelligence with Caption Studio arXiv Paper C. Chin, Jianhua Zhang et al. Jul 21 041 Addressing Limited Data in Auditory Attention Decoding with Diffusion Generative Models arXiv Paper David Rannaleet, Victor Gunnarsson et al. Jul 20 042 FlowSonic: Stable Zero-Shot Music Editing via High-Order Trajectory Integration arXiv Paper ali boudaghi, Hadi Zare Jul 20 043 Re-Sonance: A Dysarthric Asynchronous Real-Time Speech Conversion System Based on a Three-Stage Cascaded ASR-LLM-TTS Architecture arXiv Paper Yuxuan Wu, Yifan Xu et al. Jul 20 044 Time-Frequency Consistency Learning for Robust Speech Deepfake Detection arXiv Paper Jun Xue, Zhuolin Yi et al. Jul 20 045 Multi-Level Privacy-Preserving Dementia Detection from Speech via Targeted Adversarial Obfuscation and Representation Learning arXiv Paper H. Kenne, Raphael Anaadumba et al. Jul 19 046 Staged Depth-Pruning Distillation of a Flow-Matching Text-to-Speech Teacher: A Compact Hindi Speech Synthesizer arXiv Paper Sivateja Trikutam Jul 19 047 Do Speech Tokens Leak Voiceprints? Speaker Inversion Attacks Against End-to-End Speech Language Models arXiv Paper Ye Lu, Yihan Yan et al. Jul 18 048 Explainable Lightweight Compact Deep Models for Speech Emotion Recognition arXiv Paper Nelly Elsayed Jul 18 049 RealDESED: A Real-World Domestic Sound Event Detection Benchmark arXiv Paper Florian Schmid, Paul Primus et al. Jul 18 050 A Situational Speech Synthesizer for Yoruba: System Design, Phonological Rule Architecture, and Orthographic Extensions for Contour arXiv Paper Kólá Túbosún, A. Oluokun et al. Jul 17 051 AuEmoChat: Authentic Emotion Understanding and Rendering for Conversational Speech Synthesis arXiv Paper Zhenqi Jia, Yuan Zhao et al. Jul 17 052 Fretiq: Browser-Native Electric Guitar String Classification via Engineered Spectral Features and Held-Out Free-Play Evaluation arXiv Paper Aaditya Garg Jul 17 053 Natural Backdoor Attacks on Speech Recognition Models arXiv Paper Jinwen Xin, X. Lyu et al. Jul 17 054 SpeechGuard: Online Defense against Backdoor Attacks on Speech Recognition Models arXiv Paper Jinwen Xin, Xixiang Lv Jul 17 055 Can Tokens Compete? Token Representations against Supervised CNN Backbones for BirdCLEF+ 2026 arXiv Paper Anthony Miyaguchi, Murilo Gustineli et al. Jul 16 056 Large Audio Language Models for Spoofing-Aware Speaker Verification arXiv Paper Sofya Savelyeva, Mariia Perunova et al. Jul 16 057 MIDI-RAE-JEPA: Hierarchical Representation Learning and Generation for Symbolic Music arXiv Paper Scott H. Hawley Jul 16 058 RW-Voice-EQ Bench: A Real World Benchmark for Evaluating Voice AI Systems arXiv Paper D. Ayllón, Alice Baird et al. Jul 16 059 SceneBind: Binding What and Where Across Vision, Audio and Language arXiv Paper Mingfei Chen, Zijun Cui et al. Jul 16 060 Auditing Protocol-Level Shortcuts in Large Audio Language Model Judges for Speech Evaluation arXiv Paper Joonyong Park, David M. Chan et al. Jul 15