Results 41 to 50 of about 48,773 (160)
Streaming End-to-End Multi-Talker Speech Recognition [PDF]
5 pages, 3 figures.
Liang Lu 0001 +3 more
openaire +3 more sources
LWMD: A Comprehensive Compression Platform for End-to-End Automatic Speech Recognition Models
Recently end-to-end (E2E) automatic speech recognition (ASR) models have achieved promising performance. However, existing models tend to adopt increasing model sizes and suffer from expensive resource consumption for real-world applications. To compress
Yukun Liu +3 more
doaj +1 more source
A Recurrent Neural Networks (RNN) based attention model has been used in code-switching speech recognition (CSSR). However, due to the sequential computation constraint of RNN, there are stronger short-range dependencies and weaker long-range ...
Zheying Huang +5 more
doaj +1 more source
Multi-Stream End-to-End Speech Recognition
submitted to IEEE TASLP (In review).
Ruizhi Li +5 more
openaire +3 more sources
End-to-End Speech Recognition Sequence Training With Reinforcement Learning
End-to-end sequence modeling has become a popular choice for automatic speech recognition (ASR) because of the simpler pipeline compared to the conventional system and its excellent performance.
Andros Tjandra +2 more
doaj +1 more source
Ubranch Conformer: Integrating Up-Down Sampling and Branch Attention for Speech Recognition
Conformer has become one of the most popular models in the field of automatic speech recognition, achieving superior speech recognition performance by integrating a convolutional module into Transformer. However, existing Conformer models face challenges
Yang Yang +3 more
doaj +1 more source
JSUM: A Multitask Learning Speech Recognition Model for Jointly Supervised and Unsupervised Learning
In recent years, the end-to-end speech recognition model has emerged as a popular alternative to the traditional Deep Neural Network—Hidden Markov Model (DNN-HMM).
Nurmemet Yolwas, Weijing Meng
doaj +1 more source
An Improvement to Conformer-Based Model for High-Accuracy Speech Feature Extraction and Learning
Owing to the loss of effective information and incomplete feature extraction caused by the convolution and pooling operations in a convolution subsampling network, the accuracy and speed of current speech processing architectures based on the conformer ...
Mengzhuo Liu, Yangjie Wei
doaj +1 more source
Segment boundary detection directed attention for online end-to-end speech recognition
Attention-based encoder-decoder models have recently shown competitive performance for automatic speech recognition (ASR) compared to conventional ASR systems.
Junfeng Hou +3 more
doaj +1 more source
In this paper, we propose a joint training framework that efficiently combines time-domain speech enhancement (SE) with an end-to-end (E2E) automatic speech recognition (ASR) system utilizing attention-based latent features.
Da-Hee Yang, Joon-Hyuk Chang
doaj +1 more source

