Library

Subject
Tags

28,751 matches · Speech Recognition and Synthesis

#
001TDScore: Learning Synthetic Speech Quality Predictors from TTS Training Dynamics without Human annotationOpenAlexPaperNatacha Miniconi, Meysam Shamsi et al.Sep 27
002Multiresolution Neural Network for One-Class Learning of Machine SoundsOpenAlexPaperXiran Zhang, Vincent Lostanlen et al.Aug 31
003A Framework for the Automation of Preference GrammarOpenAlexPaperC.U.C. Ugorji, Ginikachi Maduako et al.4 days ago
004The Genealogy of Large Language Models: From Auxiliary Tools in ASR to Foundational Transformers and Back AgainOpenAlexPaperJosé Luciano Maldonado4 days ago
005Parameter-Efficient Fine-Tuning: LoRA, QLoRA, and DoRAOpenAlexPaperDheiver Francisco Santos5 days ago
006Parameter-Efficient Fine-Tuning: LoRA, QLoRA, and DoRAOpenAlexPaperDheiver Francisco Santos5 days ago
007Synthetic Data Generation and Self-Improvement: A Technical NoteOpenAlexPaperDheiver Francisco Santos5 days ago
008Synthetic Data Generation and Self-Improvement: A Technical NoteOpenAlexPaperDheiver Francisco Santos5 days ago
009A lightweight hybrid 2D–1D depthwise-separable CNN with frequency-only pooling for efficient small-footprint keyword spottingOpenAlexPaperJaewon Lee, Sangbeom Lee et al.6 days ago
010LiveLingo Voice Translation Benchmarks 2026OpenAlexPaperRon Villomo6 days ago
011A Deep Learning based Approach for Identifying Spoken languages in TurkeyOpenAlexPaperMüge Ertekin, Hamit Erdem7 days ago
012Embedding-Guided Neural Voice Conversion for Indian Regional-Language Speech TransformationOpenAlexPaperBala Raju A, Singh S.P et al.7 days ago
013How Meta-Learning Shapes LoRA Adapter Geometry in Speech Deepfake DetectionOpenAlexPaperIvan Kukanov, Janne Laakkonen et al.7 days ago
014Quality-Targeted, Page-Aligned KV-Cache Allocation for LLM InferenceOpenAlexPaperAthanase Matabaro7 days ago
015Quality-Targeted, Page-Aligned KV-Cache Allocation for LLM InferenceOpenAlexPaperAthanase Matabaro7 days ago
016Run inferenceOpenAlexPaperYahya Saqban7 days ago
017Run inferenceOpenAlexPaperYahya Saqban7 days ago
018hearth-llm: Persistent KV-Cache Snapshots for Interactive Local LLM InferenceOpenAlexPaperForomo Daniel Soromou7 days ago
019hearth-llm: Persistent KV-Cache Snapshots for Interactive Local LLM InferenceOpenAlexPaperForomo Daniel Soromou7 days ago
020From Read Speech to Spoken Digits: A Task-Specific Evaluation of Speech Privacy With Informed AttackersOpenAlexPaperJule Pohlhausen, Anjana Rajasekhar et al.Jul 23
021Toward Interpretable Speech Deepfake Detection using Artifact-Specific Experts and Calibrated Detection ScoresOpenAlexPaperViola Negroni, Xin Wang et al.Jul 23
022Improving the performance of an ASV system using hybrid speech featuresOpenAlexPaperStanisław Ciszkiewicz, Artur JanickiJul 22
023Multimodal Speaker Verification as a Threat to Speaker AnonymizationOpenAlexPaperAshi Garg, Cristina Aggazzotti et al.Jul 22
024Scalable Keyword Spotting via Modular Network ExpansionOpenAlexPaperViktor Khaymonenko, Dzmitry Saladukha et al.Jul 22
025SimulS2ST-Omni: Data-Efficient Streaming Speech-to-Speech Translation via Explicit Trajectory SupervisionOpenAlexPaperRong‐De He, C X Liang et al.Jul 22
026StellarTTS: Sparse Temporal Embedding for Low-Latency and Robust Speech SynthesisOpenAlexPaperKaicheng Luo, X Z Gong et al.Jul 22
027A Multi-Stage Poisoning Detection Pipeline for Embedded Voice-Command SystemsOpenAlexPaperSebastian–Alexandru Drăguşin, Denisa Toma et al.Jul 21
028Edge-based danger detection system using phrase recognition on STM32 microcontroller with FreeRTOS and GSM communicationOpenAlexPaperHo Yen Rou, Hermawan NugrohoJul 21
029Frequency-based Rerpesentation Consistency Regularization for Dense RetrievalOpenAlexPaperHongyeob Kim, Seoyoon Lee et al.Jul 21
030Speech-GAN: A Black-Box Generative Adversarial Network Attack against Automatic Speech Recognition systemsOpenAlexPaperFaraz Mohammad Mushtak Mogal, Ali Osman Topal et al.Jul 21
031Summary of DCASE 2026 Task 5: Audio-Dependent Question AnsweringOpenAlexPaperHaolin He, Renhe Sun et al.Jul 21
032AM-DANet: Additive margin with dynamic augmentation network for speaker identificationOpenAlexPaperJing Xiang Ng, Kian Ming Lim et al.Jul 20
033Enhancing Low-Resource Lampung Speech Recognition through Cross-Lingual XLSR-Wav2Vec 2.0 PretrainingOpenAlexPaperHendra KurniawanJul 20
034Harness TTS: Towards Context-Aware Expressive Speech Synthesis with Harness LayerOpenAlexPaperShengfan Shen, Di Wu et al.Jul 20
035SSTMark: Robust Training-Free Semantic-Level Speech WatermarkingOpenAlexPaperKuan-Lin Chu, Jun-Cheng Chen et al.Jul 20
036The tttAI System for the TSA-ASR Task of the SmartGlasses Challenge 2026OpenAlexPaperXuanji He, Gaoyang Dong et al.Jul 20
037Tibetan language model incorporating grammatical relationships and morphological verbsOpenAlexPaperKuntharrgyal Khysru, Wenjie Tang et al.Jul 20
038X-Translator: A Real-Time Multilingual Speaker-Aware Speech-to-Speech Translation SystemOpenAlexPaperYuxiang Zhao, Yichi Zhang et al.Jul 20
039Choosing the Compute Unit for On-Device Speech Recognition: NPU, GPU and CPU in a Single PipelineOpenAlexPaperAlexander DmitrievJul 19
040Choosing the Compute Unit for On-Device Speech Recognition: NPU, GPU and CPU in a Single PipelineOpenAlexPaperAlexander DmitrievJul 19
041Mixed approach speech-to-text translation for endangered languageOpenAlexPaperBenyamin Langgu Sinaga, ‪Stephanie Pamela Adithama et al.Jul 18
042NABEATs: Noise-Aware Audio Representation LearningOpenAlexPaperTakuya Fujimura, Yoshiki Masuyama et al.Jul 18
043A Geometry-Limited Identification Floor and Its Consequences for Voice-Clone Attribution in Professional Voice ActorsOpenAlexPaperShuhei KatoJul 17
044AMECxSV: Adaptive Metadata-Driven Embedding-Fusion Calibration for X-Lingual Speaker VerificationOpenAlexPaperXin Wei, Shi He et al.Jul 17
045Bsqat: block-wise shared quantization-aware training for large language modelsOpenAlexPaperSenbao Hou, Libin Hou et al.Jul 17
046COMPOSITIONAL RECOGNITION OF AZERBAIJANI SPEECH SIGNALS USING PHONETIC ACOUSTIC COMPONENTSOpenAlexPaperElchin IsmayilovJul 17
047Component-Level Ensemble Fusion for Speech and Environmental Sound Deepfake DetectionOpenAlexPaperAndré Runewicz, Karla Schäfer et al.Jul 17
048Natural Language Processing (NLP) for Speech Recognition and Text-to-SpeechOpenAlexPaperAmrutha Kolhar, Tukkappa K. GundoorJul 17
049Problems of Artificial Intelligence for the Azerbaijani Language: The Impact of Limited Language CorpusOpenAlexPaperRaksana Aliashrafova, Namiq Abdurahmanov et al.Jul 17
050Real-Time Speech-to-Text And Speaker Diarization Using Whisper And Pyannote EmbeddingsOpenAlexPaperShivani Chauhan, Rimmy et al.Jul 17
051Real-Time Speech-to-Text And Speaker Diarization Using Whisper And Pyannote EmbeddingsOpenAlexPaperShivani Chauhan, Rimmy et al.Jul 17
052The Symboken TheoryOpenAlexPaperIHARAJul 17
053The Symboken TheoryOpenAlexPaperIHARA SHUHEI, SHUHEI IHARAJul 17
054A Lightweight Keyword Spotting Method Using a Convolutional Spiking Neural Network with Learnable Synaptic DelaysOpenAlexPaperX W Li, Ying Liu et al.Jul 16
055Audio Deepfake Detection using Cepstral Coefficients and Spectral AnalysisOpenAlexPaperPrapti KapoorJul 16
056Audio Deepfake Detection using Cepstral Coefficients and Spectral AnalysisOpenAlexPaperPrapti KapoorJul 16
057Effects of applying Gaussianising transformations to DNN embeddings before calculating likelihood ratiosOpenAlexPaperRafael O. Ribeiro, Philip Weber et al.Jul 16
058SLT 2026 REAL-TSE Challenge: Real-world Target Speaker Extraction from Conversational RecordingsOpenAlexPaperShuai Wang, Zihan Qian et al.Jul 16
059<p>Cultural and Linguistic Contextualization in Ugandan English: Training Transformer Models for Enhanced Local Relevance</p>OpenAlexPaperDavid Byansi, Bebwa Isingoma et al.Jul 15
060A survey of forensic practitioners on speaker attribution methodsOpenAlexPaperTallulah Buckley, Kirsty McDougall et al.Jul 15

Showing 60 of 28,751 documents · scroll for more