Results 11 to 20 of about 48,773 (160)

End-to-End Audiovisual Speech Recognition [PDF]

open access: yes2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2018
Several end-to-end deep learning approaches have been recently presented which extract either audio or visual features from the input images or audio signals and perform speech recognition. However, research on end-to-end audiovisual models is very limited.
Stavros Petridis   +5 more
openaire   +3 more sources

End-to-end Multimodal Speech Recognition [PDF]

open access: yes2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2018
5 pages, 5 figures, Accepted at IEEE International Conference on Acoustics, Speech and Signal Processing 2018 (ICASSP 2018)
Shruti Palaskar   +2 more
openaire   +3 more sources

End-to-End Speech-to-Dialog-Act Recognition [PDF]

open access: yesInterspeech 2020, 2020
Spoken language understanding, which extracts intents and/or semantic concepts in utterances, is conventionally formulated as a post-processing of automatic speech recognition. It is usually trained with oracle transcripts, but needs to deal with errors by ASR.
Viet-Trung Dang   +4 more
openaire   +3 more sources

End-to-end Anchored Speech Recognition [PDF]

open access: yesICASSP 2019 - 2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2019
Voice-controlled house-hold devices, like Amazon Echo or Google Home, face the problem of performing speech recognition of device-directed speech in the presence of interfering background speech, i.e., background noise and interfering speech from another person or media device in proximity need to be ignored.
Yiming Wang 0006   +5 more
openaire   +2 more sources

Synchronous Transformers for end-to-end Speech Recognition [PDF]

open access: yesICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2020
For most of the attention-based sequence-to-sequence models, the decoder predicts the output sequence conditioned on the entire input sequence processed by the encoder. The asynchronous problem between the encoding and decoding makes these models difficult to be applied for online speech recognition.
Zhengkun Tian   +5 more
openaire   +2 more sources

Towards End-to-End Unsupervised Speech Recognition

open access: yes2022 IEEE Spoken Language Technology Workshop (SLT), 2023
Preprint
Alexander H. Liu   +3 more
openaire   +4 more sources

Research Status and Prospect of Transformer in Speech Recognition

open access: yesJisuanji kexue yu tansuo, 2021
As a new deep learning algorithm framework, Transformer has attracted more and more researchers?? attention and has become a current research hotspot. Inspired by humans focusing on important things only, the self-attention mechanism in the Transformer ...
ZHANG Xiaoxu, MA Zhiqiang, LIU Zhiqiang, ZHU Fangyuan, WANG Chunyu
doaj   +1 more source

Study on Keyword Search Framework Based on End-to-End Automatic Speech Recognition [PDF]

open access: yesJisuanji kexue, 2022
In the past decade,end-to-end automatic speech recognition (ASR) frameworks have developed rapidly.End-to-end ASR has shown not only very different characteristics from traditional ASR based on hidden Markov models (HMMs),but also advanced performances ...
YANG Run-yan, CHENG Gao-feng, LIU Jian
doaj   +1 more source

KsponSpeech: Korean Spontaneous Speech Corpus for Automatic Speech Recognition

open access: yesApplied Sciences, 2020
This paper introduces a large-scale spontaneous speech corpus of Korean, named KsponSpeech. This corpus contains 969 h of general open-domain dialog utterances, spoken by about 2000 native Korean speakers in a clean environment. All data were constructed
Jeong-Uk Bang   +9 more
doaj   +1 more source

Towards multilingual end‐to‐end speech recognition for air traffic control

open access: yesIET Intelligent Transport Systems, 2021
In this work, an end‐to‐end framework is proposed to achieve multilingual automatic speech recognition (ASR) in air traffic control (ATC) systems. Considering the standard ATC procedure, a recurrent neural network (RNN) based framework is selected to ...
Yi Lin, Bo Yang, Dongyue Guo, Peng Fan
doaj   +1 more source

Home - About - Disclaimer - Privacy