Results 21 to 30 of about 48,773 (160)

End-to-End Bengali Speech Recognition

open access: yesCoRR, 2020
Bengali is a prominent language of the Indian subcontinent. However, while many state-of-the-art acoustic models exist for prominent languages spoken in the region, research and resources for Bengali are few and far between. In this work, we apply CTC based CNN-RNN networks, a prominent deep learning based end-to-end automatic speech recognition ...
Sayan Mandal, Sarthak Yadav, Atul Rai
openaire   +3 more sources

Self-Training for End-to-End Speech Recognition [PDF]

open access: yesICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2020
We revisit self-training in the context of end-to-end speech recognition. We demonstrate that training with pseudo-labels can substantially improve the accuracy of a baseline model. Key to our approach are a strong baseline acoustic and language model used to generate the pseudo-labels, filtering mechanisms tailored to common errors from sequence-to ...
Jacob Kahn, Ann Lee 0001, Awni Y. Hannun
openaire   +3 more sources

LAS-Transformer: An Enhanced Transformer Based on the Local Attention Mechanism for Speech Recognition

open access: yesInformation, 2022
Recently, Transformer-based models have shown promising results in automatic speech recognition (ASR), outperforming models based on recurrent neural networks (RNNs) and convolutional neural networks (CNNs).
Pengbin Fu, Daxing Liu, Huirong Yang
doaj   +1 more source

End-to-End Speech Emotion Recognition With Gender Information

open access: yesIEEE Access, 2020
Many works have focused on speech emotion recognition algorithms. However, most rely on the proper selection of speech acoustic features. In this paper, we propose a novel emotion recognition algorithm that does not rely on any speech acoustic features ...
Ting-Wei Sun
doaj   +1 more source

FrameAugment: A Simple Data Augmentation Method for Encoder–Decoder Speech Recognition

open access: yesApplied Sciences, 2022
As the architecture of deep learning-based speech recognizers has recently changed to the end-to-end style, increasing the effective amount of training data has become an important issue.
Seong-Su Lim, Oh-Wook Kwon
doaj   +1 more source

Data Augmentation for Arabic Speech Recognition Based on End-to-End Deep Learning [PDF]

open access: yesInternational Journal of Intelligent Computing and Information Sciences, 2021
End-to-end deep learning approach has greatly enhanced the performance of speech recognition systems. With deep learning techniques, the overfitting stills the main problem with a little data.
Zaki Taha   +3 more
doaj   +1 more source

End-to-end visual speech recognition with LSTMS [PDF]

open access: yes2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2017
Traditional visual speech recognition systems consist of two stages, feature extraction and classification. Recently, several deep learning approaches have been presented which automatically extract features from the mouth images and aim to replace the feature extraction stage.
Petridis, Stavros   +2 more
openaire   +3 more sources

Grammar-Supervised End-to-End Speech Recognition with Part-of-Speech Tagging and Dependency Parsing

open access: yesApplied Sciences, 2023
For most automatic speech recognition systems, many unacceptable hypothesis errors still make the recognition results absurd and difficult to understand.
Genshun Wan   +5 more
doaj   +1 more source

Two-Pass End-to-End Speech Recognition [PDF]

open access: yesInterspeech 2019, 2019
The requirements for many applications of state-of-the-art speech recognition systems include not only low word error rate (WER) but also low latency. Specifically, for many use-cases, the system must be able to decode utterances in a streaming fashion and faster than real-time.
Tara N. Sainath   +11 more
openaire   +6 more sources

A new joint CTC-attention-based speech recognition model with multi-level multi-head attention

open access: yesEURASIP Journal on Audio, Speech, and Music Processing, 2019
A method called joint connectionist temporal classification (CTC)-attention-based speech recognition has recently received increasing focus and has achieved impressive performance.
Chu-Xiong Qin, Wen-Lin Zhang, Dan Qu
doaj   +1 more source

Home - About - Disclaimer - Privacy