Results 261 to 270 of about 36,891 (314)
Some of the next articles are maybe not open access.
Quantization for OpenAI's Whisper Models: A Comparative Analysis
arXiv.orgAutomated speech recognition (ASR) models have gained prominence for applications such as captioning, speech translation, and live transcription. This paper studies Whisper and two model variants: one optimized for live speech streaming and another for ...
Allison Andreyev
semanticscholar +1 more source
LoRA-Whisper: Parameter-Efficient and Extensible Multilingual ASR
InterspeechRecent years have witnessed significant progress in multilingual automatic speech recognition (ASR), driven by the emergence of end-to-end (E2E) models and the scaling of multilingual datasets.
Zheshu Song +5 more
semanticscholar +1 more source
POWSM: A Phonetic Open Whisper-Style Speech Foundation Model
Annual Meeting of the Association for Computational LinguisticsRecent advances in spoken language processing have led to substantial progress in phonetic tasks such as automatic speech recognition (ASR), phone recognition (PR), grapheme-to-phoneme conversion (G2P), and phoneme-to-grapheme conversion (P2G).
Chin-Jou Li +7 more
semanticscholar +1 more source
Calm-Whisper: Reduce Whisper Hallucination On Non-Speech By Calming Crazy Heads Down
InterspeechOpenAI's Whisper has achieved significant success in Automatic Speech Recognition. However, it has consistently been found to exhibit hallucination issues, particularly in non-speech segments, which limits its broader application in complex industrial ...
Yingzhi Wang +4 more
semanticscholar +1 more source
Whisper-LM: Improving ASR Models with Language Models for Low-Resource Languages
arXiv.orgAutomatic speech recognition systems have undoubtedly advanced with the integration of multilingual and multitask models such as Whisper, which have shown a promising ability to understand and process speech across a wide range of languages.
Xabier de Zuazo +3 more
semanticscholar +1 more source
Whispering in Amharic: Fine-tuning Whisper for Low-resource Language
arXiv.orgThis work explores fine-tuning OpenAI's Whisper automatic speech recognition (ASR) model for Amharic, a low-resource language, to improve transcription accuracy. While the foundational Whisper model struggles with Amharic due to limited representation in
Dawit Ketema Gete +13 more
semanticscholar +1 more source
COLING Workshops
We collect novel data in the public service domain to evaluate the capability of the state-of-the-art automatic speech recognition (ASR) models in capturing regional differences in accents in the United Kingdom (UK), specifically focusing on two accents ...
Melissa Torgbi +3 more
semanticscholar +1 more source
We collect novel data in the public service domain to evaluate the capability of the state-of-the-art automatic speech recognition (ASR) models in capturing regional differences in accents in the United Kingdom (UK), specifically focusing on two accents ...
Melissa Torgbi +3 more
semanticscholar +1 more source
Improving Domain Generalization in Speech Emotion Recognition with Whisper
IEEE International Conference on Acoustics, Speech, and Signal ProcessingTransformers have been used successfully in a variety of settings, including Speech Emotion Recognition (SER). However, use of the latest transformer base models in domain generalization (DG) settings has mostly been unexplored or only weakly explored ...
Erik Goron +3 more
semanticscholar +1 more source
IEEE International Conference on Acoustics, Speech, and Signal Processing
Alzheimer’s disease (AD) is a neurodegenerative disorder that can lead to speech impairments. Early diagnosis is crucial for effective treatment, and speech-based diagnosis is currently a hot research topic.
Jinpeng Li, Wei-Qiang Zhang
semanticscholar +1 more source
Alzheimer’s disease (AD) is a neurodegenerative disorder that can lead to speech impairments. Early diagnosis is crucial for effective treatment, and speech-based diagnosis is currently a hot research topic.
Jinpeng Li, Wei-Qiang Zhang
semanticscholar +1 more source
Muting Whisper: A Universal Acoustic Adversarial Attack on Speech Foundation Models
Conference on Empirical Methods in Natural Language ProcessingRecent developments in large speech foundation models like Whisper have led to their widespread use in many automatic speech recognition (ASR) applications.
Vyas Raina +4 more
semanticscholar +1 more source

