Results 61 to 70 of about 48,773 (160)
Adversarial joint training with self-attention mechanism for robust end-to-end speech recognition
Lately, the self-attention mechanism has marked a new milestone in the field of automatic speech recognition (ASR). Nevertheless, its performance is susceptible to environmental intrusions as the system predicts the next output symbol depending on the ...
Lujun Li +5 more
doaj +1 more source
NIESR: Nuisance Invariant End-to-End Speech Recognition [PDF]
Deep neural network models for speech recognition have achieved great success recently, but they can learn incorrect associations between the target and nuisance factors of speech (e.g., speaker identities, background noise, etc.), which can lead to overfitting.
I-Hung Hsu +2 more
openaire +4 more sources
At present, automatic speech recognition has become an important bridge for human-computer interaction and is widely applied in multiple fields. The Portuguese speech recognition task is gradually receiving attention due to its unique language stance ...
Yan Li +4 more
doaj +1 more source
End-to-end Music-mixed Speech Recognition
Automatic speech recognition (ASR) in multimedia content is one of the promising applications, but speech data in this kind of content are frequently mixed with background music, which is harmful for the performance of ASR. In this study, we propose a method for improving ASR with background music based on time-domain source separation. We utilize Conv-
Jeongwoo Woo +3 more
openaire +3 more sources
Multilingual Speech Recognition with a Single End-to-End Model [PDF]
Training a conventional automatic speech recognition (ASR) system to support multiple languages is challenging because the sub-word unit, lexicon and word inventories are typically language specific. In contrast, sequence-to-sequence models are well suited for multilingual ASR because they encapsulate an acoustic, pronunciation and language model ...
Shubham Toshniwal +6 more
openaire +3 more sources
End-to-End Speech Recognition from the Raw Waveform [PDF]
State-of-the-art speech recognition systems rely on fixed, hand-crafted features such as mel-filterbanks to preprocess the waveform before the training pipeline. In this paper, we study end-to-end systems trained directly from the raw waveform, building on two alternatives for trainable replacements of mel-filterbanks that use a convolutional ...
Zeghidour, Neil +4 more
openaire +2 more sources
Speech endpoint detection (EPD) benefits from the decoder state features (DSFs) of online automatic speech recognition (ASR) system. However, the DSFs are obtained via the ASR decoding process, which can become prohibitively expensive especially in ...
Inyoung Hwang, Joon-Hyuk Chang
doaj +1 more source
A Waveform-Feature Dual Branch Acoustic Embedding Network for Emotion Recognition
Research in advancing speech emotion recognition (SER) has attracted a lot of attention due to its critical role for better human behaviors understanding scientifically and comprehensive applications commercially.
Jeng-Lin Li +7 more
doaj +1 more source
WaveNet With Cross-Attention for Audiovisual Speech Recognition
In this paper, the WaveNet with cross-attention is proposed for Audio-Visual Automatic Speech Recognition (AV-ASR) to address multimodal feature fusion and frame alignment problems between two data streams.
Hui Wang, Fei Gao, Yue Zhao, Licheng Wu
doaj +1 more source
End-to-end feature fusion for jointly optimized speech enhancement and automatic speech recognition
Speech enhancement (SE) and automatic speech recognition (ASR) in real-time processing involve improving the quality and intelligibility of speech signals on the fly, ensuring accurate transcription as the speech unfolds.
Mohamed Medani +5 more
doaj +1 more source

