Results 81 to 90 of about 49,742 (177)

Whisper [PDF]

open access: yes, 2018
The sounds of life can distract us from the ...
Chapman, Tyler
core  

Effectiveness of Whisper's Fine-Tuning for Domain-Specific Use Cases in the Industry

open access: yesInternational Conference on Agents and Artificial Intelligence
: The integration of Speech-to-Text (STT) technology has the potential to enhance the efficiency of industrial workflows. However, standard speech models demonstrate suboptimal performance in domain-specific use cases.
Daniel Pawlowicz   +2 more
semanticscholar   +1 more source

Identity and personality in the social perception of synthesized voices: perceptions of OpenAI’s text-to-speech technology

open access: yesPhonetica: International Journal of Phonetic Science
As the line between human speakers and “AI-generated” voices becomes increasingly blurred, it is important to understand how sociolinguistic knowledge affects human-computer interaction. Human listeners have been shown to rely on real-world biases, along
Eve Fleisig   +6 more
semanticscholar   +1 more source

Whisper in Medusa\u27s Ear: Multi-head Efficient Decoding for Transformer-based ASR [PDF]

open access: yes
Large transformer-based models have significant potential for speech transcription and translation. Their self-attention mechanisms and parallel processing enable them to capture complex patterns and dependencies in audio sequences.
Keshet, Joseph   +4 more
core   +1 more source

ePoster

open access: yes
European Journal of Neurology, Volume 33, Issue S1, June 2026.
wiley   +1 more source

Riconoscimento del parlato mediante OpenAI Whisper

open access: yes
Questa tesi si propone di implementare e analizzare un sistema di riconoscimento vocale in tempo reale in locale utilizzando OpenAI Whisper, un modello avanzato basato su tecniche di deep learning. Whisper rappresenta lo stato dell’arte nella comprensione del parlato umano e si distingue per essere un modello open source.
openaire   +1 more source

Analysis and Categorization of the Rusyn Language Using the Whisper Model: Demographic Influences on Linguistic Convergence [PDF]

open access: yes
The article presents a detailed linguistic analysis of the Rusyn language, focusing on its complex and evolving features, such as pronunciation, as well as individual, regional, and historical variabilities.
Małecki, Paweł
core   +1 more source

mmWave-Whisper: Phone Call Eavesdropping and Transcription Using Millimeter-Wave Radar [PDF]

open access: yes
This paper introduces mmWave-Whisper, a system that demonstrates the feasibility of full-corpus automated speech recognition (ASR) on phone calls eavesdropped remotely using off-the-shelf frequency modulated continuous wave (FMCW) millimeter-wave radars.
Basak, Suryoday   +2 more
core   +1 more source

Whisper Finetuning on Nepali Language

open access: yes
Despite the growing advancements in Automatic Speech Recognition (ASR) models, the development of robust models for underrepresented languages, such as Nepali, remains a challenge.
Rijal, Sanjay   +4 more
core  

Simul-Whisper: Attention-Guided Streaming Whisper with Truncation Detection [PDF]

open access: yes
As a robust and large-scale multilingual speech recognition model, Whisper has demonstrated impressive results in many low-resource and out-of-distribution scenarios.
Wang, Haoyu   +4 more
core   +1 more source

Home - About - Disclaimer - Privacy