Results 21 to 30 of about 47,965 (148)
Data Augmentation for Arabic Speech Recognition Based on End-to-End Deep Learning [PDF]
End-to-end deep learning approach has greatly enhanced the performance of speech recognition systems. With deep learning techniques, the overfitting stills the main problem with a little data.
Zaki Taha +3 more
doaj +1 more source
End-to-End Speech Recognition and Disfluency Removal [PDF]
Disfluency detection is usually an intermediate step between an automatic speech recognition (ASR) system and a downstream task. By contrast, this paper aims to investigate the task of end-to-end speech recognition and disfluency removal. We specifically explore whether it is possible to train an ASR model to directly map disfluent speech into fluent ...
Paria Jamshid Lou, Mark Johnson 0001
openaire +2 more sources
Grammar-Supervised End-to-End Speech Recognition with Part-of-Speech Tagging and Dependency Parsing
For most automatic speech recognition systems, many unacceptable hypothesis errors still make the recognition results absurd and difficult to understand.
Genshun Wan +5 more
doaj +1 more source
Two-Pass End-to-End Speech Recognition [PDF]
The requirements for many applications of state-of-the-art speech recognition systems include not only low word error rate (WER) but also low latency. Specifically, for many use-cases, the system must be able to decode utterances in a streaming fashion and faster than real-time.
Tara N. Sainath +11 more
openaire +2 more sources
Towards end-to-end speech recognition with transfer learning
A transfer learning-based end-to-end speech recognition approach is presented in two levels in our framework. Firstly, a feature extraction approach combining multilingual deep neural network (DNN) training with matrix factorization algorithm is ...
Chu-Xiong Qin, Dan Qu, Lian-Hai Zhang
doaj +1 more source
Streaming End-to-End Target-Speaker Automatic Speech Recognition and Activity Detection
Automatic speech recognition of a target speaker in the presence of interfering speakers remains a challenging issue. One approach to tackle this problem is target-speaker speech recognition, which conditions the recognition process on an embedding that ...
Takafumi Moriya +4 more
doaj +1 more source
Multichannel End-to-end Speech Recognition
The field of speech recognition is in the midst of a paradigm shift: end-to-end neural networks are challenging the dominance of hidden Markov models as a core technology. Using an attention mechanism in a recurrent encoder-decoder architecture solves the dynamic time alignment problem, allowing joint end-to-end training of the acoustic and language ...
Tsubasa Ochiai +3 more
openaire +3 more sources
Streaming End-to-End Multi-Talker Speech Recognition [PDF]
5 pages, 3 figures.
Liang Lu 0001 +3 more
openaire +2 more sources
A new joint CTC-attention-based speech recognition model with multi-level multi-head attention
A method called joint connectionist temporal classification (CTC)-attention-based speech recognition has recently received increasing focus and has achieved impressive performance.
Chu-Xiong Qin, Wen-Lin Zhang, Dan Qu
doaj +1 more source
Fast offline transformer-based end-to-end automatic speech recognition for real-world applications
With the recent advances in technology, automatic speech recognition (ASR) has been widely used in real-world applications. The efficiency of converting large amounts of speech into text accurately with limited resources has become more vital than ever ...
Yoo Rhee Oh, Kiyoung Park, Kiyoung Park
doaj +1 more source

