Library

Subject
Tags

13,735 matches · cs.SD

#
001Cocktail-Talker: Multi-Speaker Dialog Modeling in Noisy Social Environments with Turn Action GRPOarXivPaperXilin Jiang et al.Yesterday
002Integrating Contextual Embeddings into Evaluation of Expressive MIDI Piano PerformancesarXivPaperDmitrii Gavrilev et al.Yesterday
003RIPPLE: Generating Multi-Channel Phase, Not Recovering ItarXivPaperJaehyuk Lee et al.Yesterday
004Teffic-Audio: Tell Fact from FictionarXivPaperWan Lin et al.Yesterday
005VocalRender: Score-Native Singing Voice Synthesis for Real-World CompositionarXivPaperYukun Chen et al.Yesterday
006Audio-Anchored Fusion of Multi-Ratio DiT Reconstruction Residuals for Cross-Domain Audio Deepfake DetectionarXivPaperHaotian Mo et al.2 days ago
007Few-Shot Open-Set Audio Classification via Transductive Prototype Refinement and Class Logit EnhancementarXivPaperTianyan Deng et al.2 days ago
008MMAC: A Massive Multi-dimensional Benchmark for Audio CaptioningarXivPaperWeijie Wu et al.2 days ago
009MPEcho: A Melody and Phoneme-Aware Generative Framework for Controllable Cover Song GenerationarXivPaperWei-Jaw Lee et al.2 days ago
010Prosody-driven Jailbreaks in Audio LLMs: A Controlled Study and Mechanistic AnalysisarXivPaperJiachen Qian et al.2 days ago
011Voice Memory for Agentic Speech RecognitionarXivPaperChao-Han Huck Yang et al.2 days ago
012Zero-Shot Face-to-Speech Synthesis via Latent Space Adaptation of a Style-Diffusion TTS ModelarXivPaperCarlos Muñoz-Romero et al.2 days ago
013A large-scale corpus of religious radio broadcast transcripts from webstream recordings in the United StatesarXivPaperSamuel Bestvater et al.3 days ago
014Device Invariance using Domain Adaptation on Acoustic Scene ClassificationarXivPaperA. Dileep, Shubham Sharma et al.3 days ago
015Dissecting Sensitivity to Training Language in Self-Supervised Speech Learning Using Neural Audio Codec TokensarXivPaperDaigo Takizawa et al.3 days ago
016Disentangling Acoustic Cues in Alzheimer's Pathology and Perception: The Roles of Language and GenderarXivPaperLiu He et al.4 days ago
017Enhancing Law-Enforcement Audio Transcription: A LoRA-Based Adaptation of Whisper for BWC FootagearXivPaperVivek Senthil et al.4 days ago
018MoLGE: Mixture of Language Group Experts for Efficient Scaling of Massively Multilingual Speech RecognitionarXivPaperSangmin Lee et al.4 days ago
019MusiChat: Vibe Composing for Music CreationarXivPaperCallie C. Liao et al.4 days ago
020Memory Efficient Audio Synthesis with Decoupled Temporal Depth Diffusion TransformersarXivPaperDongseong Hwang et al.5 days ago
021OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint GenerationarXivPaperJun Zhan et al.5 days ago
022Low-Latency Turn-Taking via Context-Aware Preface Generation in a Real-World Dialogue RobotarXivPaperYuki Okafuji et al.6 days ago
023Music-Source-Separation-Training (MSST): A Unified Framework for Training and Evaluating Music Demixing ModelsarXivPaperRoman Solovyev et al.6 days ago
024PathRIR: Physics-Guided Acoustic Path Selection and Late-Tail Compensation for Fast Room Impulse Response SimulationarXivPaperShaoheng Xu et al.6 days ago
025Kutti AI: A Voice-First, Offline-Capable Learning Companion with Real-Time Struggle Detection for Visually-Impaired ChildrenarXivPaperKadharmoideen Fadurudeen7 days ago
026MemNMF: Memory-Augmented NMF on LPC Spectra for Anomalous Sound DetectionarXivPaperPhurich Saengthong, Takahiro Shinozaki7 days ago
027Multimodal Domain Generalization for Depression Detection: An Attention-Based BiLSTM Network with Domain-Adversarial TrainingarXivPaperAli Tabaraei et al.7 days ago
028Reflector: Arrangement-Aware Harmonic Retrieval for Sample-Based CompositionarXivPaperAustin Rockman7 days ago
029Synthetic Speech, Real Signal: Paralinguistic Preservation and Cross-Lingual Augmentation via Voice CloningarXivPaperR. Polle, Owen Parsons et al.7 days ago
030Transforming Keystroke Noise to Text: Self-Supervised Acoustic Eavesdropping Attacks on KeyboardsarXivPaperAtsunori Okada, Akira Ito et al.7 days ago
031An Evaluation Framework for Structured Audio Captions Validated by Controlled PerturbationsarXivPaperLiang-Yuan Wu, Sripathi Sridhar et al.Jul 23
032Probing Speaker Identity Sensitivity in Audio Deepfake DetectorsarXivPaperD. Dar, Arun RossJul 23
033VibeVoice-ASR-BitNet Technical ReportarXivPaperSongcheng Xu, Ting Song et al.Jul 23
034Audio-Zero: Label-Free Self-Evolution for Fine-Grained Audio ReasoningarXivPaperSiqian Tong, Xuan Li et al.Jul 22
035Cumsum-Composable Phase Transport for Low-Cost Streaming Keyword SpottingarXivPaperMahesh GodavartiJul 22
036Efficient Chain-of-Modality Reasoning via Progressive Compression for Spoken Language ModelsarXivPaperPengchao Feng et al.Jul 22
037Pushing the Frontier of Full-Song Generation: Hierarchical Autoregressive Planning Meets Flow-Matching RenderingarXivPaperJunyu Dai, Xinyue Fan et al.Jul 22
038RPPNet: Perceptually-Grouped Rhythm-Pitch Primitives for Long-Term Structure Melody Generation via Boundary-Aware ModelingarXivPaperTieyao Zhang, Yuke Liu et al.Jul 22
039Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and AudioarXivPaperAbdul Basit Tonmoy, Kazi Fardinul Hoque et al.Jul 21
040What the Waveform Knows: Transparent-first Speech and Audio Intelligence with Caption StudioarXivPaperC. Chin, Jianhua Zhang et al.Jul 21
041Addressing Limited Data in Auditory Attention Decoding with Diffusion Generative ModelsarXivPaperDavid Rannaleet, Victor Gunnarsson et al.Jul 20
042FlowSonic: Stable Zero-Shot Music Editing via High-Order Trajectory IntegrationarXivPaperali boudaghi, Hadi ZareJul 20
043Re-Sonance: A Dysarthric Asynchronous Real-Time Speech Conversion System Based on a Three-Stage Cascaded ASR-LLM-TTS ArchitecturearXivPaperYuxuan Wu, Yifan Xu et al.Jul 20
044Time-Frequency Consistency Learning for Robust Speech Deepfake DetectionarXivPaperJun Xue, Zhuolin Yi et al.Jul 20
045Multi-Level Privacy-Preserving Dementia Detection from Speech via Targeted Adversarial Obfuscation and Representation LearningarXivPaperH. Kenne, Raphael Anaadumba et al.Jul 19
046Staged Depth-Pruning Distillation of a Flow-Matching Text-to-Speech Teacher: A Compact Hindi Speech SynthesizerarXivPaperSivateja TrikutamJul 19
047Do Speech Tokens Leak Voiceprints? Speaker Inversion Attacks Against End-to-End Speech Language ModelsarXivPaperYe Lu, Yihan Yan et al.Jul 18
048Explainable Lightweight Compact Deep Models for Speech Emotion RecognitionarXivPaperNelly ElsayedJul 18
049RealDESED: A Real-World Domestic Sound Event Detection BenchmarkarXivPaperFlorian Schmid, Paul Primus et al.Jul 18
050A Situational Speech Synthesizer for Yoruba: System Design, Phonological Rule Architecture, and Orthographic Extensions for ContourarXivPaperKólá Túbosún, A. Oluokun et al.Jul 17
051AuEmoChat: Authentic Emotion Understanding and Rendering for Conversational Speech SynthesisarXivPaperZhenqi Jia, Yuan Zhao et al.Jul 17
052Fretiq: Browser-Native Electric Guitar String Classification via Engineered Spectral Features and Held-Out Free-Play EvaluationarXivPaperAaditya GargJul 17
053Natural Backdoor Attacks on Speech Recognition ModelsarXivPaperJinwen Xin, X. Lyu et al.Jul 17
054SpeechGuard: Online Defense against Backdoor Attacks on Speech Recognition ModelsarXivPaperJinwen Xin, Xixiang LvJul 17
055Can Tokens Compete? Token Representations against Supervised CNN Backbones for BirdCLEF+ 2026arXivPaperAnthony Miyaguchi, Murilo Gustineli et al.Jul 16
056Large Audio Language Models for Spoofing-Aware Speaker VerificationarXivPaperSofya Savelyeva, Mariia Perunova et al.Jul 16
057MIDI-RAE-JEPA: Hierarchical Representation Learning and Generation for Symbolic MusicarXivPaperScott H. HawleyJul 16
058RW-Voice-EQ Bench: A Real World Benchmark for Evaluating Voice AI SystemsarXivPaperD. Ayllón, Alice Baird et al.Jul 16
059SceneBind: Binding What and Where Across Vision, Audio and LanguagearXivPaperMingfei Chen, Zijun Cui et al.Jul 16
060Auditing Protocol-Level Shortcuts in Large Audio Language Model Judges for Speech EvaluationarXivPaperJoonyong Park, David M. Chan et al.Jul 15

Showing 60 of 13,735 documents · scroll for more