Results 21 to 30 of about 48,773 (160)
End-to-End Bengali Speech Recognition
Bengali is a prominent language of the Indian subcontinent. However, while many state-of-the-art acoustic models exist for prominent languages spoken in the region, research and resources for Bengali are few and far between. In this work, we apply CTC based CNN-RNN networks, a prominent deep learning based end-to-end automatic speech recognition ...
Sayan Mandal, Sarthak Yadav, Atul Rai
openaire +3 more sources
Self-Training for End-to-End Speech Recognition [PDF]
We revisit self-training in the context of end-to-end speech recognition. We demonstrate that training with pseudo-labels can substantially improve the accuracy of a baseline model. Key to our approach are a strong baseline acoustic and language model used to generate the pseudo-labels, filtering mechanisms tailored to common errors from sequence-to ...
Jacob Kahn, Ann Lee 0001, Awni Y. Hannun
openaire +3 more sources
Recently, Transformer-based models have shown promising results in automatic speech recognition (ASR), outperforming models based on recurrent neural networks (RNNs) and convolutional neural networks (CNNs).
Pengbin Fu, Daxing Liu, Huirong Yang
doaj +1 more source
End-to-End Speech Emotion Recognition With Gender Information
Many works have focused on speech emotion recognition algorithms. However, most rely on the proper selection of speech acoustic features. In this paper, we propose a novel emotion recognition algorithm that does not rely on any speech acoustic features ...
Ting-Wei Sun
doaj +1 more source
FrameAugment: A Simple Data Augmentation Method for Encoder–Decoder Speech Recognition
As the architecture of deep learning-based speech recognizers has recently changed to the end-to-end style, increasing the effective amount of training data has become an important issue.
Seong-Su Lim, Oh-Wook Kwon
doaj +1 more source
Data Augmentation for Arabic Speech Recognition Based on End-to-End Deep Learning [PDF]
End-to-end deep learning approach has greatly enhanced the performance of speech recognition systems. With deep learning techniques, the overfitting stills the main problem with a little data.
Zaki Taha +3 more
doaj +1 more source
End-to-end visual speech recognition with LSTMS [PDF]
Traditional visual speech recognition systems consist of two stages, feature extraction and classification. Recently, several deep learning approaches have been presented which automatically extract features from the mouth images and aim to replace the feature extraction stage.
Petridis, Stavros +2 more
openaire +3 more sources
Grammar-Supervised End-to-End Speech Recognition with Part-of-Speech Tagging and Dependency Parsing
For most automatic speech recognition systems, many unacceptable hypothesis errors still make the recognition results absurd and difficult to understand.
Genshun Wan +5 more
doaj +1 more source
Two-Pass End-to-End Speech Recognition [PDF]
The requirements for many applications of state-of-the-art speech recognition systems include not only low word error rate (WER) but also low latency. Specifically, for many use-cases, the system must be able to decode utterances in a streaming fashion and faster than real-time.
Tara N. Sainath +11 more
openaire +6 more sources
A new joint CTC-attention-based speech recognition model with multi-level multi-head attention
A method called joint connectionist temporal classification (CTC)-attention-based speech recognition has recently received increasing focus and has achieved impressive performance.
Chu-Xiong Qin, Wen-Lin Zhang, Dan Qu
doaj +1 more source

