Results 31 to 40 of about 48,773 (160)

Streaming End-to-End Target-Speaker Automatic Speech Recognition and Activity Detection

open access: yesIEEE Access, 2023
Automatic speech recognition of a target speaker in the presence of interfering speakers remains a challenging issue. One approach to tackle this problem is target-speaker speech recognition, which conditions the recognition process on an embedding that ...
Takafumi Moriya   +4 more
doaj   +1 more source

End-To-End Silent Speech Recognition with Acoustic Sensing [PDF]

open access: yes2021 IEEE Spoken Language Technology Workshop (SLT), 2021
Silent speech interfaces (SSI) has been an exciting area of recent interest. In this paper, we present a non-invasive silent speech interface that uses inaudible acoustic signals to capture people's lip movements when they speak. We exploit the speaker and microphone of the smartphone to emit signals and listen to their reflections, respectively.
Jian Luo 0007   +4 more
openaire   +3 more sources

Fast offline transformer-based end-to-end automatic speech recognition for real-world applications

open access: yesETRI Journal, 2022
With the recent advances in technology, automatic speech recognition (ASR) has been widely used in real-world applications. The efficiency of converting large amounts of speech into text accurately with limited resources has become more vital than ever ...
Yoo Rhee Oh, Kiyoung Park, Kiyoung Park
doaj   +1 more source

Real-Time End-to-End Speech Emotion Recognition with Cross-Domain Adaptation

open access: yesBig Data and Cognitive Computing, 2022
Language resources are the main factor in speech-emotion-recognition (SER)-based deep learning models. Thai is a low-resource language that has a smaller data size than high-resource languages such as German. This paper describes the framework of using a
Konlakorn Wongpatikaseree   +3 more
doaj   +1 more source

End-To-End Audio-Visual Speech Recognition with Conformers [PDF]

open access: yesICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2021
Accepted to ICASSP ...
Pingchuan Ma 0001   +2 more
openaire   +2 more sources

Performance Monitoring for End-to-End Speech Recognition [PDF]

open access: yesInterspeech 2019, 2019
Submitted to Interspeech ...
Ruizhi Li, Gregory Sell, Hynek Hermansky
openaire   +3 more sources

Multichannel End-to-end Speech Recognition

open access: yesCoRR, 2017
The field of speech recognition is in the midst of a paradigm shift: end-to-end neural networks are challenging the dominance of hidden Markov models as a core technology. Using an attention mechanism in a recurrent encoder-decoder architecture solves the dynamic time alignment problem, allowing joint end-to-end training of the acoustic and language ...
Tsubasa Ochiai   +3 more
openaire   +4 more sources

End-to-end Speech-to-Punctuated-Text Recognition

open access: yesInterspeech 2022, 2022
Conventional automatic speech recognition systems do not produce punctuation marks which are important for the readability of the speech recognition results. They are also needed for subsequent natural language processing tasks such as machine translation.
Jumon Nozaki   +3 more
openaire   +4 more sources

End-to-End Speech Recognition and Disfluency Removal [PDF]

open access: yesFindings of the Association for Computational Linguistics: EMNLP 2020, 2020
Disfluency detection is usually an intermediate step between an automatic speech recognition (ASR) system and a downstream task. By contrast, this paper aims to investigate the task of end-to-end speech recognition and disfluency removal. We specifically explore whether it is possible to train an ASR model to directly map disfluent speech into fluent ...
Paria Jamshid Lou, Mark Johnson 0001
openaire   +3 more sources

End-to-End Amdo-Tibetan Speech Recognition Based on Knowledge Transfer

open access: yesIEEE Access, 2020
The end-to-end speech recognition technology solves the problem that each component is independent and models cannot be jointly optimized in the traditional speech recognition model.
Xiaojun Zhu, Heming Huang
doaj   +1 more source

Home - About - Disclaimer - Privacy