Results 51 to 60 of about 48,773 (160)
A Light-Weight Autoregressive CNN-Based Frame Level Transducer Decoder for End-to-End ASR
A convolutional neural network (CNN) transducer decoder was proposed to reduce the decoding time of an end-to-end automatic speech recognition (ASR) system while maintaining accuracy.
Hyeon-Kyu Noh, Hong-June Park
doaj +1 more source
End-to-End Training of a Large Vocabulary End-to-End Speech Recognition System [PDF]
In this paper, we present an end-to-end training framework for building state-of-the-art end-to-end speech recognition systems. Our training system utilizes a cluster of Central Processing Units(CPUs) and Graphics Processing Units (GPUs). The entire data reading, large scale data augmentation, neural network parameter updates are all performed "on-the ...
Chanwoo Kim 0001 +12 more
openaire +3 more sources
Acoustic Model Fusion For End-to-End Speech Recognition
Recent advances in deep learning and automatic speech recognition (ASR) have enabled the end-to-end (E2E) ASR system and boosted the accuracy to a new level. The E2E systems implicitly model all conventional ASR components, such as the acoustic model (AM) and the language model (LM), in a single network trained on audio-text pairs. Despite this simpler
Zhihong Lei +10 more
openaire +3 more sources
End-To-End Multi-Speaker Speech Recognition With Transformer [PDF]
Recently, fully recurrent neural network (RNN) based end-to-end models have been proven to be effective for multi-speaker speech recognition in both the single-channel and multi-channel scenarios. In this work, we explore the use of Transformer models for these tasks by focusing on two aspects.
Xuankai Chang +4 more
openaire +3 more sources
Comparative Study on End-to-End Speech Recognition Using Pre-trained Models [PDF]
In the field of speech and audio signal processing, pre-trained models (PTMs) are commonly available. Pre-trained models (PTMs) offer a collection of initial weights and biases that may be adjusted for a particular task, which makes them a popular ...
Martha Ghobrial +2 more
doaj +1 more source
Deep Context: End-to-end Contextual Speech Recognition [PDF]
In automatic speech recognition (ASR) what a user says depends on the particular context she is in. Typically, this context is represented as a set of word n-grams. In this work, we present a novel, all-neural, end-to-end (E2E) ASR sys- tem that utilizes such context.
Golan Pundak +4 more
openaire +4 more sources
Streaming End-to-end Speech Recognition for Mobile Devices [PDF]
End-to-end (E2E) models, which directly predict output character sequences given input speech, are good candidates for on-device speech recognition. E2E models, however, present numerous challenges: In order to be truly useful, such models must decode speech utterances in a streaming fashion, in real time; they must be robust to the long tail of use ...
Yanzhang He +19 more
openaire +4 more sources
A Spelling Correction Model for End-to-end Speech Recognition [PDF]
Attention-based sequence-to-sequence models for speech recognition jointly train an acoustic model, language model (LM), and alignment mechanism using a single neural network and require only parallel audio-text pairs. Thus, the language model component of the end-to-end model is only trained on transcribed audio-text pairs, which leads to performance ...
Jinxi Guo, Tara N. Sainath, Ron J. Weiss
openaire +2 more sources
A Novel Syllable-Level Signal Encryption for Robust Secure Speech Communication System
Speech communication is vital for conveying information and emotions, yet it faces significant security threats. This research presents a novel signal encryption system that operates at the syllable level, preserving the natural flow of speech while ...
Albertus Anugerah Pekerti +3 more
doaj +1 more source
Dialog-Context Aware end-to-end Speech Recognition [PDF]
Existing speech recognition systems are typically built at the sentence level, although it is known that dialog context, e.g. higher-level knowledge that spans across sentences or speakers, can help the processing of long conversations. The recent progress in end-to-end speech recognition systems promises to integrate all available information (e.g ...
Suyoun Kim, Florian Metze
openaire +2 more sources

