End-to-End Speech Recognition Sequence Training With Reinforcement Learning
End-to-end sequence modeling has become a popular choice for automatic speech recognition (ASR) because of the simpler pipeline compared to the conventional system and its excellent performance.
Andros Tjandra +2 more
doaj +1 more source
Ubranch Conformer: Integrating Up-Down Sampling and Branch Attention for Speech Recognition
Conformer has become one of the most popular models in the field of automatic speech recognition, achieving superior speech recognition performance by integrating a convolutional module into Transformer. However, existing Conformer models face challenges
Yang Yang +3 more
doaj +1 more source
An Improvement to Conformer-Based Model for High-Accuracy Speech Feature Extraction and Learning
Owing to the loss of effective information and incomplete feature extraction caused by the convolution and pooling operations in a convolution subsampling network, the accuracy and speed of current speech processing architectures based on the conformer ...
Mengzhuo Liu, Yangjie Wei
doaj +1 more source
Multi-Stream End-to-End Speech Recognition
submitted to IEEE TASLP (In review).
Ruizhi Li +5 more
openaire +2 more sources
End-to-End Training of a Large Vocabulary End-to-End Speech Recognition System [PDF]
In this paper, we present an end-to-end training framework for building state-of-the-art end-to-end speech recognition systems. Our training system utilizes a cluster of Central Processing Units(CPUs) and Graphics Processing Units (GPUs). The entire data reading, large scale data augmentation, neural network parameter updates are all performed "on-the ...
Chanwoo Kim 0001 +12 more
openaire +2 more sources
End-To-End Multi-Speaker Speech Recognition With Transformer [PDF]
Recently, fully recurrent neural network (RNN) based end-to-end models have been proven to be effective for multi-speaker speech recognition in both the single-channel and multi-channel scenarios. In this work, we explore the use of Transformer models for these tasks by focusing on two aspects.
Xuankai Chang +4 more
openaire +2 more sources
JSUM: A Multitask Learning Speech Recognition Model for Jointly Supervised and Unsupervised Learning
In recent years, the end-to-end speech recognition model has emerged as a popular alternative to the traditional Deep Neural Network—Hidden Markov Model (DNN-HMM).
Nurmemet Yolwas, Weijing Meng
doaj +1 more source
Deep Context: End-to-end Contextual Speech Recognition [PDF]
In automatic speech recognition (ASR) what a user says depends on the particular context she is in. Typically, this context is represented as a set of word n-grams. In this work, we present a novel, all-neural, end-to-end (E2E) ASR sys- tem that utilizes such context.
Golan Pundak +4 more
openaire +2 more sources
Streaming End-to-end Speech Recognition for Mobile Devices [PDF]
End-to-end (E2E) models, which directly predict output character sequences given input speech, are good candidates for on-device speech recognition. E2E models, however, present numerous challenges: In order to be truly useful, such models must decode speech utterances in a streaming fashion, in real time; they must be robust to the long tail of use ...
Yanzhang He +19 more
openaire +2 more sources
Segment boundary detection directed attention for online end-to-end speech recognition
Attention-based encoder-decoder models have recently shown competitive performance for automatic speech recognition (ASR) compared to conventional ASR systems.
Junfeng Hou +3 more
doaj +1 more source

