Results 1 to 10 of about 47,965 (148)
A study of transformer-based end-to-end speech recognition system for Kazakh language [PDF]
Today, the Transformer model, which allows parallelization and also has its own internal attention, has been widely used in the field of speech recognition.
Mamyrbayev Orken +4 more
doaj +2 more sources
End-to-End Audiovisual Speech Recognition [PDF]
Several end-to-end deep learning approaches have been recently presented which extract either audio or visual features from the input images or audio signals and perform speech recognition. However, research on end-to-end audiovisual models is very limited.
Stavros Petridis +5 more
openaire +2 more sources
End-to-End Speech-to-Dialog-Act Recognition [PDF]
Spoken language understanding, which extracts intents and/or semantic concepts in utterances, is conventionally formulated as a post-processing of automatic speech recognition. It is usually trained with oracle transcripts, but needs to deal with errors by ASR.
Viet-Trung Dang +4 more
openaire +2 more sources
Arabic speech recognition using end‐to‐end deep learning
Arabic automatic speech recognition (ASR) methods with diacritics have the ability to be integrated with other systems better than Arabic ASR methods without diacritics.
Hamzah A. Alsayadi +3 more
doaj +1 more source
End-to-end Multimodal Speech Recognition [PDF]
5 pages, 5 figures, Accepted at IEEE International Conference on Acoustics, Speech and Signal Processing 2018 (ICASSP 2018)
Shruti Palaskar +2 more
openaire +2 more sources
Synchronous Transformers for end-to-end Speech Recognition [PDF]
For most of the attention-based sequence-to-sequence models, the decoder predicts the output sequence conditioned on the entire input sequence processed by the encoder. The asynchronous problem between the encoding and decoding makes these models difficult to be applied for online speech recognition.
Zhengkun Tian +5 more
openaire +2 more sources
End-to-end Anchored Speech Recognition [PDF]
Voice-controlled house-hold devices, like Amazon Echo or Google Home, face the problem of performing speech recognition of device-directed speech in the presence of interfering background speech, i.e., background noise and interfering speech from another person or media device in proximity need to be ignored.
Yiming Wang 0006 +5 more
openaire +2 more sources
Towards End-to-End Unsupervised Speech Recognition
Preprint
Alexander H. Liu +3 more
openaire +2 more sources
Self-Training for End-to-End Speech Recognition [PDF]
We revisit self-training in the context of end-to-end speech recognition. We demonstrate that training with pseudo-labels can substantially improve the accuracy of a baseline model. Key to our approach are a strong baseline acoustic and language model used to generate the pseudo-labels, filtering mechanisms tailored to common errors from sequence-to ...
Jacob Kahn, Ann Lee 0001, Awni Y. Hannun
openaire +2 more sources
Research Status and Prospect of Transformer in Speech Recognition
As a new deep learning algorithm framework, Transformer has attracted more and more researchers?? attention and has become a current research hotspot. Inspired by humans focusing on important things only, the self-attention mechanism in the Transformer ...
ZHANG Xiaoxu, MA Zhiqiang, LIU Zhiqiang, ZHU Fangyuan, WANG Chunyu
doaj +1 more source

