Нови тенденциї у розвою сучасней линґвистики у Сербї
ANALIZA I KLASYFIKACJA JĘZYKA RUSIŃSKIEGO PRZY UŻYCIU MODELU SZTUCZNEJ SIECI NEURONOWEJ ASR OPENAI WHISPER Artykuł przedstawia analizę lingwistyczną języka rusińskiego, koncentrując się na jego złożonych i zmieniających się aspektach, takich jak wymowa
Paweł Małecki, Magdalena Piotrowska
doaj +2 more sources
— Implementasi Automatic Speech ...
Danny Ferdiansyah +1 more
openaire +2 more sources
Automated Assessment of Word- and Sentence-Level Speech Intelligibility in Developmental Motor Speech Disorders: A Cross-Linguistic Investigation [PDF]
Background/Objectives: Accurate assessment of speech intelligibility is necessary for individuals with motor speech disorders. Transcription or scaled rating methods by naïve listeners are the most reliable tasks for these purposes; however, they are ...
Micalle Carl, Michal Icht
doaj +2 more sources
Enhancing supermarket robot interaction: an equitable multi-level LLM conversational interface for handling diverse customer intents [PDF]
This paper presents the design and evaluation of a comprehensive system to develop voice-based interfaces to support users in supermarkets. These interfaces enable shoppers to convey their needs through both generic and specific queries.
Chandran Nandkumar, Luka Peternel
doaj +2 more sources
Reproducing Whisper-Style Training Using An Open-Source Toolkit And Publicly Available Data [PDF]
Pre-training speech models on large volumes of data has achieved remarkable success. OpenAI Whisper is a multilingual multitask model trained on 680k hours of supervised speech data.
Yifan Peng +15 more
semanticscholar +4 more sources
Speech Emotion Recognition Leveraging OpenAI's Whisper Representations and Attentive Pooling Methods
Speech Emotion Recognition (SER) research has faced limitations due to the lack of standard and sufficiently large datasets. Recent studies have leveraged pre-trained models to extract features for downstream tasks such as SER. This work explores the capabilities of Whisper, a pre-trained ASR system, in speech emotion recognition by proposing two ...
Ali Shendabadi +3 more
openaire +3 more sources
This study tackles language barriers in computer-mediated communication by developing an application that integrates OpenAI’s Whisper ASR model and Google Translate machine translation to enable real-time, continuous speech transcription and translation and the processing of video and audio files.
Dewi Khairani +4 more
openaire +2 more sources
Human-like AI-based auto-field-in-field whole-brain radiotherapy treatment planning with conversation large language model feedback. [PDF]
Abstract Background Whole‐brain radiotherapy (WBRT) is a common treatment due to its simplicity and effectiveness. While automated Field‐in‐Field (Auto‐FiF) functions assist WBRT planning in modern treatment planning systems, it still requires manual approaches for optimal plan generation including patient‐specific hyperparameters definition and plan ...
Jafar A +5 more
europepmc +2 more sources
This paper details the experimental results of adapting the OpenAI's Whisper model for Code-Switch Mandarin-English Speech Recognition (ASR) on the SEAME and ASRU2019 corpora. We conducted 2 experiments: a) using adaptation data from 1 to 100/200 hours to demonstrate effectiveness of adaptation, b) examining different language ID setup on Whisper ...
Xionghu Zhong
exaly +4 more sources
From voice to ink (Vink): development and assessment of an automated, free-of-charge transcription tool [PDF]
Background Verbatim transcription of qualitative audio data is a cornerstone of analytic quality and rigor, yet the time and energy required for such transcription can drain resources, delay analysis, and hinder the timely dissemination of qualitative ...
Hannah Tolle +6 more
doaj +2 more sources

