Results 1 to 10 of about 49,742 (177)

Video Transcripts Summarization using OpenAI Whisper and GPT Model

open access: yesInternational Journal for Research in Applied Science and Engineering Technology
Abstract: In today’s digital age, a vast amount of video content is generated and shared on the internet every minute. However, extracting relevant information from these videos can be time-consuming and challenging. This is where video transcript summarization comes in, providing a concise summary of video content without the need to watch the entire ...
A. Bhargavi
exaly   +3 more sources

Fine-Tuning OpenAI Whisper-Small for Domain-Specific Medical Speech Recognition within a Microservice Architecture

open access: yesInformatica (Slovenia)
We fine-tune Whisper-small (244M parameters) on 8.5 hours of in-domain medical audio and evaluatewith word error rate (WER). Compared to an unadapted Whisper-small baseline, our fine-tuned modelreduces WER from ∼63% to ∼32%. While the relative gain is substantial, this accuracy is not suitablefor unsupervised clinical use; we position the system as a ...
Alaeddine Moussa, Noursene Drine
exaly   +3 more sources

Spoken Language Analysis in Aging Research: The Validity of AI-Generated Speech to Text Using OpenAI's Whisper. [PDF]

open access: yesGerontology
Introduction: Studying what older adults say can provide important insights into cognitive, affective, and social aspects of aging. Available language analysis tools generally require audio-recorded speech to be transcribed into verbatim text, a task that has historically been performed by humans.
Naffah A, Pfeifer VA, Mehl MR.
europepmc   +4 more sources

Instant Transcription and Translation Tool using OpenAI?s Whisper ASR Model

open access: yesInternational Journal of Science and Research (Raipur, India), 2022
Akarsh Ghale, Janaki K, Devaraj Verma C
exaly   +3 more sources

Evaluating OpenAI's Whisper ASR for Punctuation Prediction and Topic Modeling of life histories of the Museum of the Person [PDF]

open access: yesCoRR, 2023
Automatic speech recognition (ASR) systems play a key role in applications involving human-machine interactions. Despite their importance, ASR models for the Portuguese language proposed in the last decade have limitations in relation to the correct identification of punctuation marks in automatic transcriptions, which hinder the use of transcriptions ...
Lucas Rafael Stefanel Gris   +5 more
openaire   +4 more sources

Evaluating OpenAI’s Whisper ASR: Performance Analysis Across Diverse Accents and Speaker Traits

open access: yesJASA Express Letters, 2023
This research explores the performance of the Whisper's ASR system on different native and non-native English accents. The findings indicate better performance on North American vs British and Irish English accents; and on native vs native accents. The analysis also unearths links between speaker traits (sex, L1 typology, and L2 proficiency) and word ...
Graham, Calbert, Roll, Nathan
openaire   +3 more sources

Mi-Go: Test Framework which uses YouTube as Data Source for Evaluating Speech Recognition Models like OpenAI's Whisper [PDF]

open access: yesCoRR, 2023
This article introduces Mi-Go, a novel testing framework aimed at evaluating the performance and adaptability of general-purpose speech recognition machine learning models across diverse real-world scenarios. The framework leverages YouTube as a rich and continuously updated data source, accounting for multiple languages, accents, dialects, speaking ...
Tomasz Wojnar   +2 more
openaire   +3 more sources

AI-based services for inclusive language learning in immersive XR environments: Speech translation, and sign language integration [version 2; peer review: 2 approved] [PDF]

open access: yesOpen Research Europe
Background Extended Reality (XR) technologies offer transformative potential for language education, yet current platforms largely neglect the accessibility needs of deaf and hard-of-hearing individuals.
Evangelos Papatheou   +3 more
doaj   +2 more sources

fAI-BRO: a multimodal AI decision-support system to address diagnostic delay in fibromyalgia syndrome [PDF]

open access: yesRMD Open
Objective To evaluate the diagnostic accuracy and patient acceptability of fAI-BRO (Fibromyalgia AI-Based Rheumatology Observer), a multimodal artificial intelligence (AI) system integrating video-based visual descriptor analysis and psycholinguistic ...
Antonella Celano   +4 more
doaj   +2 more sources

Quantization for OpenAI's Whisper Models: A Comparative Analysis

open access: yesCoRR
Automated speech recognition (ASR) models have gained prominence for applications such as captioning, speech translation, and live transcription. This paper studies Whisper and two model variants: one optimized for live speech streaming and another for offline transcription.
Allison Andreyev
openaire   +4 more sources

Home - About - Disclaimer - Privacy