Results 61 to 70 of about 48,773 (160)

Adversarial joint training with self-attention mechanism for robust end-to-end speech recognition

open access: yesEURASIP Journal on Audio, Speech, and Music Processing, 2021
Lately, the self-attention mechanism has marked a new milestone in the field of automatic speech recognition (ASR). Nevertheless, its performance is susceptible to environmental intrusions as the system predicts the next output symbol depending on the ...
Lujun Li   +5 more
doaj   +1 more source

NIESR: Nuisance Invariant End-to-End Speech Recognition [PDF]

open access: yesInterspeech 2019, 2019
Deep neural network models for speech recognition have achieved great success recently, but they can learn incorrect associations between the target and nuisance factors of speech (e.g., speaker identities, background noise, etc.), which can lead to overfitting.
I-Hung Hsu   +2 more
openaire   +4 more sources

A review on speech recognition approaches and challenges for Portuguese: exploring the feasibility of fine-tuning large-scale end-to-end models

open access: yesEURASIP Journal on Audio, Speech, and Music Processing
At present, automatic speech recognition has become an important bridge for human-computer interaction and is widely applied in multiple fields. The Portuguese speech recognition task is gradually receiving attention due to its unique language stance ...
Yan Li   +4 more
doaj   +1 more source

End-to-end Music-mixed Speech Recognition

open access: yes, 2020
Automatic speech recognition (ASR) in multimedia content is one of the promising applications, but speech data in this kind of content are frequently mixed with background music, which is harmful for the performance of ASR. In this study, we propose a method for improving ASR with background music based on time-domain source separation. We utilize Conv-
Jeongwoo Woo   +3 more
openaire   +3 more sources

Multilingual Speech Recognition with a Single End-to-End Model [PDF]

open access: yes2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2018
Training a conventional automatic speech recognition (ASR) system to support multiple languages is challenging because the sub-word unit, lexicon and word inventories are typically language specific. In contrast, sequence-to-sequence models are well suited for multilingual ASR because they encapsulate an acoustic, pronunciation and language model ...
Shubham Toshniwal   +6 more
openaire   +3 more sources

End-to-End Speech Recognition from the Raw Waveform [PDF]

open access: yesInterspeech 2018, 2018
State-of-the-art speech recognition systems rely on fixed, hand-crafted features such as mel-filterbanks to preprocess the waveform before the training pipeline. In this paper, we study end-to-end systems trained directly from the raw waveform, building on two alternatives for trainable replacements of mel-filterbanks that use a convolutional ...
Zeghidour, Neil   +4 more
openaire   +2 more sources

End-to-End Speech Endpoint Detection Utilizing Acoustic and Language Modeling Knowledge for Online Low-Latency Speech Recognition

open access: yesIEEE Access, 2020
Speech endpoint detection (EPD) benefits from the decoder state features (DSFs) of online automatic speech recognition (ASR) system. However, the DSFs are obtained via the ASR decoding process, which can become prohibitively expensive especially in ...
Inyoung Hwang, Joon-Hyuk Chang
doaj   +1 more source

A Waveform-Feature Dual Branch Acoustic Embedding Network for Emotion Recognition

open access: yesFrontiers in Computer Science, 2020
Research in advancing speech emotion recognition (SER) has attracted a lot of attention due to its critical role for better human behaviors understanding scientifically and comprehensive applications commercially.
Jeng-Lin Li   +7 more
doaj   +1 more source

WaveNet With Cross-Attention for Audiovisual Speech Recognition

open access: yesIEEE Access, 2020
In this paper, the WaveNet with cross-attention is proposed for Audio-Visual Automatic Speech Recognition (AV-ASR) to address multimodal feature fusion and frame alignment problems between two data streams.
Hui Wang, Fei Gao, Yue Zhao, Licheng Wu
doaj   +1 more source

End-to-end feature fusion for jointly optimized speech enhancement and automatic speech recognition

open access: yesScientific Reports
Speech enhancement (SE) and automatic speech recognition (ASR) in real-time processing involve improving the quality and intelligibility of speech signals on the fly, ensuring accurate transcription as the speech unfolds.
Mohamed Medani   +5 more
doaj   +1 more source

Home - About - Disclaimer - Privacy