Results 261 to 270 of about 36,891 (314)
Some of the next articles are maybe not open access.

Quantization for OpenAI's Whisper Models: A Comparative Analysis

arXiv.org
Automated speech recognition (ASR) models have gained prominence for applications such as captioning, speech translation, and live transcription. This paper studies Whisper and two model variants: one optimized for live speech streaming and another for ...
Allison Andreyev
semanticscholar   +1 more source

LoRA-Whisper: Parameter-Efficient and Extensible Multilingual ASR

Interspeech
Recent years have witnessed significant progress in multilingual automatic speech recognition (ASR), driven by the emergence of end-to-end (E2E) models and the scaling of multilingual datasets.
Zheshu Song   +5 more
semanticscholar   +1 more source

POWSM: A Phonetic Open Whisper-Style Speech Foundation Model

Annual Meeting of the Association for Computational Linguistics
Recent advances in spoken language processing have led to substantial progress in phonetic tasks such as automatic speech recognition (ASR), phone recognition (PR), grapheme-to-phoneme conversion (G2P), and phoneme-to-grapheme conversion (P2G).
Chin-Jou Li   +7 more
semanticscholar   +1 more source

Calm-Whisper: Reduce Whisper Hallucination On Non-Speech By Calming Crazy Heads Down

Interspeech
OpenAI's Whisper has achieved significant success in Automatic Speech Recognition. However, it has consistently been found to exhibit hallucination issues, particularly in non-speech segments, which limits its broader application in complex industrial ...
Yingzhi Wang   +4 more
semanticscholar   +1 more source

Whisper-LM: Improving ASR Models with Language Models for Low-Resource Languages

arXiv.org
Automatic speech recognition systems have undoubtedly advanced with the integration of multilingual and multitask models such as Whisper, which have shown a promising ability to understand and process speech across a wide range of languages.
Xabier de Zuazo   +3 more
semanticscholar   +1 more source

Whispering in Amharic: Fine-tuning Whisper for Low-resource Language

arXiv.org
This work explores fine-tuning OpenAI's Whisper automatic speech recognition (ASR) model for Amharic, a low-resource language, to improve transcription accuracy. While the foundational Whisper model struggles with Amharic due to limited representation in
Dawit Ketema Gete   +13 more
semanticscholar   +1 more source

Adapting Whisper for Regional Dialects: Enhancing Public Services for Vulnerable Populations in the United Kingdom

COLING Workshops
We collect novel data in the public service domain to evaluate the capability of the state-of-the-art automatic speech recognition (ASR) models in capturing regional differences in accents in the United Kingdom (UK), specifically focusing on two accents ...
Melissa Torgbi   +3 more
semanticscholar   +1 more source

Improving Domain Generalization in Speech Emotion Recognition with Whisper

IEEE International Conference on Acoustics, Speech, and Signal Processing
Transformers have been used successfully in a variety of settings, including Speech Emotion Recognition (SER). However, use of the latest transformer base models in domain generalization (DG) settings has mostly been unexplored or only weakly explored ...
Erik Goron   +3 more
semanticscholar   +1 more source

Whisper-Based Transfer Learning for Alzheimer Disease Classification: Leveraging Speech Segments with Full Transcripts as Prompts

IEEE International Conference on Acoustics, Speech, and Signal Processing
Alzheimer’s disease (AD) is a neurodegenerative disorder that can lead to speech impairments. Early diagnosis is crucial for effective treatment, and speech-based diagnosis is currently a hot research topic.
Jinpeng Li, Wei-Qiang Zhang
semanticscholar   +1 more source

Muting Whisper: A Universal Acoustic Adversarial Attack on Speech Foundation Models

Conference on Empirical Methods in Natural Language Processing
Recent developments in large speech foundation models like Whisper have led to their widespread use in many automatic speech recognition (ASR) applications.
Vyas Raina   +4 more
semanticscholar   +1 more source

Home - About - Disclaimer - Privacy