Results 141 to 150 of about 49,742 (177)
Some of the next articles are maybe not open access.
Leveraging OpenAI Whisper Model to Improve Speech Recognition for Dysarthric Individuals
2024 Asia Pacific Conference on Innovation in Technology (APCIT)Automatic Speech Recognition (ASR) systems are pivotal in facilitating human-technology interactions through voice commands. However, individuals with dysarthria face significant challenges in benefiting from these technologies due to their speech ...
Hepsiba D, L D Vijay Anand
exaly +3 more sources
2025 International Conference on Innovation and Intelligence for Informatics, Computing, and Technologies (3ICT)
Translating Tamazight speech into written text is important to support communication, preserve cultural identity, and improve access to digital services for Tamazight-speaking communities. The goal of this research is to build a system that automatically
Ayman Ait Achour +2 more
exaly +3 more sources
Translating Tamazight speech into written text is important to support communication, preserve cultural identity, and improve access to digital services for Tamazight-speaking communities. The goal of this research is to build a system that automatically
Ayman Ait Achour +2 more
exaly +3 more sources
2024 Asian Conference on Intelligent Technologies (ACOIT)
Automated closed-captioning systems are critical to increasing accessibility and user interaction with digital media. This paper describes a new real-time caption generation system that combines Flask, MoviePy, OpenCV, and the Whisper model.
Henil Shah +2 more
openaire +2 more sources
Automated closed-captioning systems are critical to increasing accessibility and user interaction with digital media. This paper describes a new real-time caption generation system that combines Flask, MoviePy, OpenCV, and the Whisper model.
Henil Shah +2 more
openaire +2 more sources
Evaluating OpenAI Whisper and its variants for increasing the accessibility of audio collections
Original abstract: The development at the Smithsonian Institution of a central Media Asset Delivery Service (MADS) aims to address accessibility compliance by remediating digital content. MADS supports audio and video delivery with captions through Smithsonian's Digital Asset Management System (DAMS).
Trizna, Michael +3 more
openaire +2 more sources
Fine-Tuning OpenAI Whisper and DistilWhisper: An In-Depth Analysis
Smart Innovation, Systems and TechnologiesAnita Shrotriya, Amit Kumar Bairwa
exaly +2 more sources
Maithili Speech Recognition with OpenAI’s Whisper: A Fine-Tuning Approach
Learning and Analytics in Intelligent SystemsRishabh Negi +4 more
exaly +2 more sources
As the field of generative AI evolves, so does the demand for intelligent systems that can understand human speech. Navigating the complexities of automatic speech recognition (ASR) technology is a significant challenge for many professionals.
J. R. Batista
semanticscholar +1 more source

