Results 51 to 60 of about 47,965 (148)
Multilingual Speech Recognition with a Single End-to-End Model [PDF]
Training a conventional automatic speech recognition (ASR) system to support multiple languages is challenging because the sub-word unit, lexicon and word inventories are typically language specific. In contrast, sequence-to-sequence models are well suited for multilingual ASR because they encapsulate an acoustic, pronunciation and language model ...
Shubham Toshniwal +6 more
openaire +2 more sources
Acoustic Model Fusion For End-to-End Speech Recognition
Recent advances in deep learning and automatic speech recognition (ASR) have enabled the end-to-end (E2E) ASR system and boosted the accuracy to a new level. The E2E systems implicitly model all conventional ASR components, such as the acoustic model (AM) and the language model (LM), in a single network trained on audio-text pairs. Despite this simpler
Zhihong Lei +10 more
openaire +2 more sources
A Light-Weight Autoregressive CNN-Based Frame Level Transducer Decoder for End-to-End ASR
A convolutional neural network (CNN) transducer decoder was proposed to reduce the decoding time of an end-to-end automatic speech recognition (ASR) system while maintaining accuracy.
Hyeon-Kyu Noh, Hong-June Park
doaj +1 more source
In this paper, we propose a joint training framework that efficiently combines time-domain speech enhancement (SE) with an end-to-end (E2E) automatic speech recognition (ASR) system utilizing attention-based latent features.
Da-Hee Yang, Joon-Hyuk Chang
doaj +1 more source
Dialog-Context Aware end-to-end Speech Recognition [PDF]
Existing speech recognition systems are typically built at the sentence level, although it is known that dialog context, e.g. higher-level knowledge that spans across sentences or speakers, can help the processing of long conversations. The recent progress in end-to-end speech recognition systems promises to integrate all available information (e.g ...
Suyoun Kim, Florian Metze
openaire +2 more sources
End-to-End Speech Recognition: A Survey
Submitted to IEEE/ACM Transactions on Audio, Speech, and Language ...
Rohit Prabhavalkar +4 more
openaire +3 more sources
Improving End-to-End Models for Children’s Speech Recognition
Children’s Speech Recognition (CSR) is a challenging task due to the high variability in children’s speech patterns and limited amount of available annotated children’s speech data. We aim to improve CSR in the often-occurring scenario that no children’s
Tanvina Patel, Odette Scharenborg
doaj +1 more source
End-to-end Music-mixed Speech Recognition
Automatic speech recognition (ASR) in multimedia content is one of the promising applications, but speech data in this kind of content are frequently mixed with background music, which is harmful for the performance of ASR. In this study, we propose a method for improving ASR with background music based on time-domain source separation. We utilize Conv-
Jeongwoo Woo +3 more
openaire +3 more sources
Comparative Study on End-to-End Speech Recognition Using Pre-trained Models [PDF]
In the field of speech and audio signal processing, pre-trained models (PTMs) are commonly available. Pre-trained models (PTMs) offer a collection of initial weights and biases that may be adjusted for a particular task, which makes them a popular ...
Martha Ghobrial +2 more
doaj +1 more source
Accented Speech Recognition Based on End-to-End Domain Adversarial Training of Neural Networks
The performance of automatic speech recognition (ASR) may be degraded when accented speech is recognized because the speech has some linguistic differences from standard speech.
Hyeong-Ju Na, Jeong-Sik Park
doaj +1 more source

