Results 151 to 160 of about 49,742 (177)
Some of the next articles are maybe not open access.

OWSM v3.1: Better and Faster Open Whisper-Style Speech Models based on E-Branchformer

Interspeech
Recent studies have highlighted the importance of fully open foundation models. The Open Whisper-style Speech Model (OWSM) is an initial step towards reproducing OpenAI Whisper using public data and open-source toolkits.
Yifan Peng   +11 more
semanticscholar   +1 more source

Whisper-SV: Adapting Whisper for low-data-resource speaker verification [PDF]

open access: yesSpeech Communication
Trained on 680,000 hours of massive speech data, Whisper is a multitasking, multilingual speech foundation model demonstrating superior performance in automatic speech recognition, translation, and language identification.
Li Zhang
exaly   +2 more sources

Calm-Whisper: Reduce Whisper Hallucination On Non-Speech By Calming Crazy Heads Down

Interspeech
OpenAI's Whisper has achieved significant success in Automatic Speech Recognition. However, it has consistently been found to exhibit hallucination issues, particularly in non-speech segments, which limits its broader application in complex industrial ...
Yingzhi Wang   +4 more
semanticscholar   +1 more source

Whispering in Amharic: Fine-tuning Whisper for Low-resource Language

arXiv.org
This work explores fine-tuning OpenAI's Whisper automatic speech recognition (ASR) model for Amharic, a low-resource language, to improve transcription accuracy. While the foundational Whisper model struggles with Amharic due to limited representation in
Dawit Ketema Gete   +13 more
semanticscholar   +1 more source

OpenAI Whisper

Review::OpenAI's Whisper is a free, open-source speech recognition model that automatically generates captions and transcripts for audio and video files. Whisper is publicly available to download and use from GitHub, so once downloaded, it can be run locally with command line prompts or in scripting languages like Python.
openaire   +1 more source

Transforming Daily Tasks for Visually Impaired Seniors: Leveraging an OpenAI Model-Enhanced Multimodal Dialogue System with Speech Recognition in POF Smart Homes

SoutheastCon
We proposed a Client Server Based Plastic Fiber Optic Network for an Elderly Support system in a smart home environment. The network will incorporate multimodal dialogue systems (MDS) and Artificial Intelligence (AI) systems.
Jason Zheng   +13 more
semanticscholar   +1 more source

J-j-j-just Stutter: Benchmarking Whisper's Performance Disparities on Different Stuttering Patterns

Interspeech
Despite their prevalence in everyday technologies, automated speech recognition (ASR) systems often struggle with disfluent speech. To diagnose and address these technical challenges, we evaluate OpenAI’s Whisper, a state-of-the-art ASR model, using ...
Charan Sridhar, Shaomei Wu
semanticscholar   +1 more source

Adaptability of ASR Models on Low-Resource Language: A Comparative Study of Whisper and Wav2Vec-BERT on Bangla

2025 2nd International Conference on Next-Generation Computing, IoT and Machine Learning (NCIM)
In recent years, neural models trained on large multilingual text and speech datasets have shown great potential for supporting low-resource languages.
Md Sazzadul Islam Ridoy   +2 more
semanticscholar   +1 more source

Building AI Applications with OpenAI APIs


Unlock the power of AI in your applications with ChatGPT with this practical guide that shows you how to seamlessly integrate OpenAI APIs into your projects, enabling you to navigate complex APIs and ensure seamless functionality with ease.
M. Yanev
semanticscholar   +1 more source

Fine-tuning Whisper on Low-Resource Languages for Real-World Applications

Swiss Text Analytics Conference
This paper presents a new approach to fine-tuning OpenAI's Whisper model for low-resource languages by introducing a novel data generation method that converts sentence-level data into a long-form corpus, using Swiss German as a case study.
Vincenzo Timmel   +4 more
semanticscholar   +1 more source

Home - About - Disclaimer - Privacy