Results 161 to 170 of about 49,742 (177)
Some of the next articles are maybe not open access.

Assessing Whisper for Infant Research: Benchmarking ASR Accuracy and Failure Analysis on Caregiver-Infant Interactions

International Conference on Development and Learning
Transcribing naturalistic caregiver-child interactions is a labor-intensive task in developmental research. While automatic speech recognition (ASR) models offer potential solutions, current ASR systems, including OpenAI’s Whisper, were primarily trained
Yueyan Tang   +3 more
semanticscholar   +1 more source

Adapting Whisper for Streaming Speech Recognition via Two-Pass Decoding

Interspeech
OpenAI Whisper is a family of robust Automatic Speech Recognition (ASR) models trained on 680,000 hours of audio. However, its encoder-decoder architecture, trained with a sequence-to-sequence objective, lacks native support for streaming ASR.
Hao Zhou   +11 more
semanticscholar   +1 more source

Voicebridge: An AI-Based Multi-Modal Voice Assistant Using Whisper, GTTS and GPT

INTERNATIONAL JOURNAL OF SCIENTIFIC RESEARCH IN ENGINEERING AND MANAGEMENT
In recent years, voice assistants have emerged as powerful tools for enabling human-machine interaction through natural spoken language. These systems, powered by advances in artificial intelligence and speech processing, offer users the convenience of ...
Dhanimireedi Greeshmanth   +1 more
semanticscholar   +1 more source

Parameter-Efficient Fine-Tuning of Whisper for Multi-Dialectal Arabic ASR

2025 13th International Japan-Africa Conference on Electronics, Communications, and Computations (JAC-ECC)
Arabic automatic speech recognition (ASR) is particularly challenging due to the linguistic divergence of regional dialects from Modern Standard Arabic (MSA).
Zyad Omar   +5 more
semanticscholar   +1 more source

Bidirectional ASL Communication Platform using Whisper Ai and Real-Time Gesture-to-Speech and Speech-to-Gesture Conversion

2025 3rd International Conference on Intelligent Cyber Physical Systems and Internet of Things (ICoICI)
Bridging the communication gap between hearing-impaired individuals and non-signers remains a critical challenge in inclusive technology. This paper presents a real-time, bidirectional communication platform that translates between American Sign Language
K. Sreenivasulu   +5 more
semanticscholar   +1 more source

Fine-Tuning Whisper for Inclusive Prosodic Stress Analysis

arXiv.org
Prosody plays a crucial role in speech perception, influencing both human understanding and automatic speech recognition (ASR) systems. Despite its importance, prosodic stress remains under-studied due to the challenge of efficiently analyzing it.
Samuel S. Sohn   +2 more
semanticscholar   +1 more source

Deploying Multilingual ASR in Digital Twin Systems: A Performance and Efficiency Analysis of Whisper and SeamlessM4T

2025 International Conference on Artificial Intelligence, Computer, Data Sciences and Applications (ACDSA)
This study presents a systematic comparison of two leading automatic speech recognition (ASR) model families—OpenAI's Whisper and Meta's SeamlessM4T—across three typologically diverse languages: English (Germanic), Russian (Slavic), and Kazakh (Turkic ...
B. Amirkhanov   +5 more
semanticscholar   +1 more source

Personality Detection Using Xlm-Roberta and Whisper

2025 IEEE Pune Section International Conference (PuneCon)
This paper describes the development of an innovative and intelligent framework for identifying personality types in real time using voice samples.
Yamjala Ravi Chandra Madhav   +3 more
semanticscholar   +1 more source

Guide méthodologique pour la transcription automatique de l'audio avec Whisper (OpenAI)

Guide méthodologique sous forme de carnet Jupyter exécutable dans Google Colab avec prise en charge GPU, conçu pour transformer les sources sonores en textes exploitables par la recherche. Une part considérable de la production intellectuelle, patrimoniale et testimoniale des sciences humaines et sociales n'existe que sous forme orale — entretiens de ...
openaire   +1 more source

Guía metodológica para la transcripción automática de audio con Whisper (OpenAI)

Guía metodológica en forma de cuaderno Jupyter ejecutable en Google Colab con soporte GPU, concebida para transformar fuentes sonoras en textos explotables por la investigación. Una parte considerable de la producción intelectual, patrimonial y testimonial de las ciencias humanas y sociales existe únicamente en forma oral —entrevistas de campo ...
openaire   +1 more source

Home - About - Disclaimer - Privacy