Results 31 to 40 of about 48,773 (160)
Streaming End-to-End Target-Speaker Automatic Speech Recognition and Activity Detection
Automatic speech recognition of a target speaker in the presence of interfering speakers remains a challenging issue. One approach to tackle this problem is target-speaker speech recognition, which conditions the recognition process on an embedding that ...
Takafumi Moriya +4 more
doaj +1 more source
End-To-End Silent Speech Recognition with Acoustic Sensing [PDF]
Silent speech interfaces (SSI) has been an exciting area of recent interest. In this paper, we present a non-invasive silent speech interface that uses inaudible acoustic signals to capture people's lip movements when they speak. We exploit the speaker and microphone of the smartphone to emit signals and listen to their reflections, respectively.
Jian Luo 0007 +4 more
openaire +3 more sources
Fast offline transformer-based end-to-end automatic speech recognition for real-world applications
With the recent advances in technology, automatic speech recognition (ASR) has been widely used in real-world applications. The efficiency of converting large amounts of speech into text accurately with limited resources has become more vital than ever ...
Yoo Rhee Oh, Kiyoung Park, Kiyoung Park
doaj +1 more source
Real-Time End-to-End Speech Emotion Recognition with Cross-Domain Adaptation
Language resources are the main factor in speech-emotion-recognition (SER)-based deep learning models. Thai is a low-resource language that has a smaller data size than high-resource languages such as German. This paper describes the framework of using a
Konlakorn Wongpatikaseree +3 more
doaj +1 more source
End-To-End Audio-Visual Speech Recognition with Conformers [PDF]
Accepted to ICASSP ...
Pingchuan Ma 0001 +2 more
openaire +2 more sources
Performance Monitoring for End-to-End Speech Recognition [PDF]
Submitted to Interspeech ...
Ruizhi Li, Gregory Sell, Hynek Hermansky
openaire +3 more sources
Multichannel End-to-end Speech Recognition
The field of speech recognition is in the midst of a paradigm shift: end-to-end neural networks are challenging the dominance of hidden Markov models as a core technology. Using an attention mechanism in a recurrent encoder-decoder architecture solves the dynamic time alignment problem, allowing joint end-to-end training of the acoustic and language ...
Tsubasa Ochiai +3 more
openaire +4 more sources
End-to-end Speech-to-Punctuated-Text Recognition
Conventional automatic speech recognition systems do not produce punctuation marks which are important for the readability of the speech recognition results. They are also needed for subsequent natural language processing tasks such as machine translation.
Jumon Nozaki +3 more
openaire +4 more sources
End-to-End Speech Recognition and Disfluency Removal [PDF]
Disfluency detection is usually an intermediate step between an automatic speech recognition (ASR) system and a downstream task. By contrast, this paper aims to investigate the task of end-to-end speech recognition and disfluency removal. We specifically explore whether it is possible to train an ASR model to directly map disfluent speech into fluent ...
Paria Jamshid Lou, Mark Johnson 0001
openaire +3 more sources
End-to-End Amdo-Tibetan Speech Recognition Based on Knowledge Transfer
The end-to-end speech recognition technology solves the problem that each component is independent and models cannot be jointly optimized in the traditional speech recognition model.
Xiaojun Zhu, Heming Huang
doaj +1 more source

