Results 11 to 20 of about 28,309,773 (304)
Phoneme and Sentence-Level Ensembles for Speech Recognition [PDF]
We address the question of whether and how boosting and bagging can be used for speech recognition. In order to do this, we compare two different boosting schemes, one at the phoneme level and one at the utterance level, with a phoneme-level bagging ...
Samy Bengio, Christos Dimitrakakis
doaj +3 more sources
Pre-trained models for detection and severity level classification of dysarthria from speech
Automatic detection and severity level classification of dysarthria from speech enables noninvasive and effective diagnosis that helps clinical decisions about medication and therapy of patients.
Farhad Javanmardi +2 more
semanticscholar +2 more sources
Despite the apparent egalitarian principles and language of the Qur’an, the use of speech level in the Madurese translation’s dialogue verses appears to reinforce social stratification, which has existed for a long in society.
Masyithah Mardhatillah +4 more
doaj +1 more source
Multimodal Emotion Recognition with High-Level Speech and Text Features [PDF]
Automatic emotion recognition is one of the central concerns of the Human-Computer Interaction field as it can bridge the gap between humans and machines.
M. R. Makiuchi, K. Uto, Koichi Shinoda
semanticscholar +1 more source
WhisperX: Time-Accurate Speech Transcription of Long-Form Audio [PDF]
Large-scale, weakly-supervised speech recognition models, such as Whisper, have demonstrated impressive results on speech recognition across domains and languages.
Max Bain +3 more
semanticscholar +1 more source
Prosodic Prominence and Boundaries in Sequence-to-Sequence Speech Synthesis [PDF]
Recent advances in deep learning methods have elevated synthetic speech quality to human level, and the field is now moving towards addressing prosodic variation in synthetic speech.Despite successes in this effort, the state-of-the-art systems fall ...
Juraj Šimko +7 more
core +1 more source
KBES: A dataset for realistic Bangla speech emotion recognition with intensity level
Speech Emotion Recognition (SER) identifies and categorizes emotional states by analyzing speech signals. SER is an emerging research area using machine learning and deep learning techniques due to its socio-cultural and business importance.
Md. Masum Billah +2 more
doaj +1 more source
Anterior drooling is common in children with cerebral palsy (CP) and poses significant risks to the child's health. Causes of drooling include oro-motor dysfunction, inefficient swallowing and reduced sensation in the orofacial musculature.
Michelle McInerney +3 more
doaj +1 more source
Analysis of speech prosody using WaveNet embeddings : The Lombard effect [PDF]
We present a novel methodology for speech prosody research based on the analysis of embeddings used to condition a convolutional WaveNet speech synthesis system.
Juraj Šimko +5 more
core +1 more source
Speak, Read and Prompt: High-Fidelity Text-to-Speech with Minimal Supervision [PDF]
We introduce SPEAR-TTS, a multi-speaker text-to-speech (TTS) system that can be trained with minimal supervision. By combining two types of discrete speech representations, we cast TTS as a composition of two sequence-to-sequence tasks: from text to high-
E. Kharitonov +8 more
semanticscholar +1 more source

