Results 251 to 260 of about 36,891 (314)
Life on Mars? The physiological perspective
Experimental Physiology, EarlyView.
Ronan M. G. Berg, Damian M. Bailey
wiley +1 more source
Audio-Visual Speech Recognition (AVSR) uses lip-based video to improve performance in noise. Since videos are harder to obtain than audio, the video training data of AVSR models is usually limited to a few thousand hours.
Andrew Rouditchenko +6 more
semanticscholar +3 more sources
Whisper Finetuning on Nepali Language
Despite advancements in ASR models, developing robust models for underrepresented languages like Nepali remains a challenge. Our work focuses on making a comprehensive dataset and finetuning OpenAI’s Whisper models to improve transcription accuracy on ...
Sanjay Rijal +4 more
semanticscholar +2 more sources
Some of the next articles are maybe not open access.
Related searches:
Related searches:
Distil-Whisper: Robust Knowledge Distillation via Large-Scale Pseudo Labelling
arXiv.org, 2023As the size of pre-trained speech recognition models increases, running these large models in low-latency or resource-constrained environments becomes challenging.
Sanchit Gandhi +2 more
semanticscholar +1 more source
Investigation of Whisper ASR Hallucinations Induced by Non-Speech Audio
IEEE International Conference on Acoustics, Speech, and Signal ProcessingHallucinations of deep neural models are amongst key challenges in automatic speech recognition (ASR). In this paper, we investigate hallucinations of the Whisper ASR model induced by non-speech audio segments present during inference.
M. Bara'nski +5 more
semanticscholar +1 more source
OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning
InterspeechThe Open Whisper-style Speech Models (OWSM) project has developed a series of fully open speech foundation models using academic-scale resources, but their training data remains insufficient.
Yifan Peng +6 more
semanticscholar +1 more source
OWSM v3.1: Better and Faster Open Whisper-Style Speech Models based on E-Branchformer
InterspeechRecent studies have highlighted the importance of fully open foundation models. The Open Whisper-style Speech Model (OWSM) is an initial step towards reproducing OpenAI Whisper using public data and open-source toolkits.
Yifan Peng +11 more
semanticscholar +1 more source
IEEE Computer Graphics and Applications, 2009
This paper discussed about a software called Apophysis that has known to be a software that can be downloaded by anyone and used by an online fractal-art community.
openaire +2 more sources
This paper discussed about a software called Apophysis that has known to be a software that can be downloaded by anyone and used by an online fractal-art community.
openaire +2 more sources

