Results 11 to 20 of about 47,965 (148)
Study on Keyword Search Framework Based on End-to-End Automatic Speech Recognition [PDF]
In the past decade,end-to-end automatic speech recognition (ASR) frameworks have developed rapidly.End-to-end ASR has shown not only very different characteristics from traditional ASR based on hidden Markov models (HMMs),but also advanced performances ...
YANG Run-yan, CHENG Gao-feng, LIU Jian
doaj +1 more source
KsponSpeech: Korean Spontaneous Speech Corpus for Automatic Speech Recognition
This paper introduces a large-scale spontaneous speech corpus of Korean, named KsponSpeech. This corpus contains 969 h of general open-domain dialog utterances, spoken by about 2000 native Korean speakers in a clean environment. All data were constructed
Jeong-Uk Bang +9 more
doaj +1 more source
Towards multilingual end‐to‐end speech recognition for air traffic control
In this work, an end‐to‐end framework is proposed to achieve multilingual automatic speech recognition (ASR) in air traffic control (ATC) systems. Considering the standard ATC procedure, a recurrent neural network (RNN) based framework is selected to ...
Yi Lin, Bo Yang, Dongyue Guo, Peng Fan
doaj +1 more source
Recently, Transformer-based models have shown promising results in automatic speech recognition (ASR), outperforming models based on recurrent neural networks (RNNs) and convolutional neural networks (CNNs).
Pengbin Fu, Daxing Liu, Huirong Yang
doaj +1 more source
End-to-End Speech Emotion Recognition With Gender Information
Many works have focused on speech emotion recognition algorithms. However, most rely on the proper selection of speech acoustic features. In this paper, we propose a novel emotion recognition algorithm that does not rely on any speech acoustic features ...
Ting-Wei Sun
doaj +1 more source
FrameAugment: A Simple Data Augmentation Method for Encoder–Decoder Speech Recognition
As the architecture of deep learning-based speech recognizers has recently changed to the end-to-end style, increasing the effective amount of training data has become an important issue.
Seong-Su Lim, Oh-Wook Kwon
doaj +1 more source
End-to-end visual speech recognition with LSTMS [PDF]
Traditional visual speech recognition systems consist of two stages, feature extraction and classification. Recently, several deep learning approaches have been presented which automatically extract features from the mouth images and aim to replace the feature extraction stage.
Petridis, Stavros +2 more
openaire +3 more sources
End-to-end Speech-to-Punctuated-Text Recognition
Conventional automatic speech recognition systems do not produce punctuation marks which are important for the readability of the speech recognition results. They are also needed for subsequent natural language processing tasks such as machine translation.
Jumon Nozaki +3 more
openaire +2 more sources
End-To-End Audio-Visual Speech Recognition with Conformers [PDF]
Accepted to ICASSP ...
Pingchuan Ma 0001 +2 more
openaire +2 more sources
End-to-End Bengali Speech Recognition
Bengali is a prominent language of the Indian subcontinent. However, while many state-of-the-art acoustic models exist for prominent languages spoken in the region, research and resources for Bengali are few and far between. In this work, we apply CTC based CNN-RNN networks, a prominent deep learning based end-to-end automatic speech recognition ...
Sayan Mandal, Sarthak Yadav, Atul Rai
openaire +2 more sources

