Results 251 to 260 of about 36,891 (314)

Life on Mars? The physiological perspective

open access: yes
Experimental Physiology, EarlyView.
Ronan M. G. Berg, Damian M. Bailey
wiley   +1 more source

From hot tub to high altitude: The makings of a human physiology field expedition and the next generation of physiologists

open access: yes
Experimental Physiology, EarlyView.
M. M. Tymko   +4 more
wiley   +1 more source

Whisper-Flamingo: Integrating Visual Features into Whisper for Audio-Visual Speech Recognition and Translation

open access: yesInterspeech
Audio-Visual Speech Recognition (AVSR) uses lip-based video to improve performance in noise. Since videos are harder to obtain than audio, the video training data of AVSR models is usually limited to a few thousand hours.
Andrew Rouditchenko   +6 more
semanticscholar   +3 more sources

Whisper Finetuning on Nepali Language

open access: yes2025 5th International Conference on Electrical, Computer, Communications and Mechatronics Engineering (ICECCME)
Despite advancements in ASR models, developing robust models for underrepresented languages like Nepali remains a challenge. Our work focuses on making a comprehensive dataset and finetuning OpenAI’s Whisper models to improve transcription accuracy on ...
Sanjay Rijal   +4 more
semanticscholar   +2 more sources
Some of the next articles are maybe not open access.

Related searches:

Distil-Whisper: Robust Knowledge Distillation via Large-Scale Pseudo Labelling

arXiv.org, 2023
As the size of pre-trained speech recognition models increases, running these large models in low-latency or resource-constrained environments becomes challenging.
Sanchit Gandhi   +2 more
semanticscholar   +1 more source

Investigation of Whisper ASR Hallucinations Induced by Non-Speech Audio

IEEE International Conference on Acoustics, Speech, and Signal Processing
Hallucinations of deep neural models are amongst key challenges in automatic speech recognition (ASR). In this paper, we investigate hallucinations of the Whisper ASR model induced by non-speech audio segments present during inference.
M. Bara'nski   +5 more
semanticscholar   +1 more source

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning

Interspeech
The Open Whisper-style Speech Models (OWSM) project has developed a series of fully open speech foundation models using academic-scale resources, but their training data remains insufficient.
Yifan Peng   +6 more
semanticscholar   +1 more source

OWSM v3.1: Better and Faster Open Whisper-Style Speech Models based on E-Branchformer

Interspeech
Recent studies have highlighted the importance of fully open foundation models. The Open Whisper-style Speech Model (OWSM) is an initial step towards reproducing OpenAI Whisper using public data and open-source toolkits.
Yifan Peng   +11 more
semanticscholar   +1 more source

Soft as a Whisper

IEEE Computer Graphics and Applications, 2009
This paper discussed about a software called Apophysis that has known to be a software that can be downloaded by anyone and used by an online fractal-art community.
openaire   +2 more sources

Home - About - Disclaimer - Privacy