Results 31 to 40 of about 714 (168)

Attention-enhanced connectionist temporal classification for discrete speech emotion recognition [PDF]

open access: yes, 2019
Discrete speech emotion recognition (SER), the assignment of a single emotion label to an entire speech utterance, is typically performed as a sequence-to-label task. This approach, however, is limited, in that it can result in models that do not capture
Zhang, Zixing   +11 more
core   +1 more source

A Hardware-Oriented and Memory-Efficient Method for CTC Decoding

open access: yesIEEE Access, 2019
The Connectionist Temporal Classification (CTC) has achieved great success in sequence to sequence analysis tasks such as automatic speech recognition (ASR) and scene text recognition (STR).
Siyuan Lu   +3 more
doaj   +1 more source

Modeling Intra-label Dynamics and Analyzing the Role of Blank in Connectionist Temporal Classification [PDF]

open access: yesComputer and Knowledge Engineering, 2018
The goal of many tasks in the realm of sequence processing is to map a sequence of input data to a sequence of output labels. Long short-term memory (LSTM), a type of recurrent neural network (RNN), equipped with connectionist temporal classification ...
Ashkan Sadeghi Lotfabadi   +2 more
doaj   +1 more source

A Deep Diacritics-Based Recognition Model for Arabic Speech: Quranic Verses as Case Study

open access: yesIEEE Access, 2023
Arabic is the language of more than 422 million of the world’s population. Although classic Arabic is the Quran language that 1.9 billion Muslims are required to recite, limited Arabic speech recognition exists.
Sarah S. Alrumiah, Amal A. Al-Shargabi
doaj   +1 more source

Word Beam Search: A Connectionist Temporal Classification Decoding Algorithm [PDF]

open access: yes, 2018
Recurrent Neural Networks (RNNs) are used for sequence recognition tasks such as Handwritten Text Recognition (HTR) or speech recognition. If trained with the Connectionist Temporal Classification (CTC) loss function, the output of such a RNN is a matrix
Harald Scheidl   +5 more
core   +1 more source

Sound Event Detection with Sequentially Labelled Data Based on Connectionist Temporal Classification and Unsupervised Clustering [PDF]

open access: yes, 2019
Sound event detection (SED) methods typically rely on either strongly labelled data or weakly labelled data. As an alternative, sequentially labelled data (SLD) was proposed.
Li, Shengchen   +7 more
core   +4 more sources

Speech GAU: A Single Head Attention for Mandarin Speech Recognition for Air Traffic Control

open access: yesAerospace, 2022
The rise of end-to-end (E2E) speech recognition technology in recent years has overturned the design pattern of cascading multiple subtasks in classical speech recognition and achieved direct mapping of speech input signals to text labels. In this study,
Shiyu Zhang   +4 more
doaj   +1 more source

Self-distillation Regularized Connectionist Temporal Classification Loss for Text Recognition: A Simple Yet Effective Approach [PDF]

open access: yes, 2023
Text recognition methods are gaining rapid development. Some advanced techniques, e.g., powerful modules, language models, and un- and semi-supervised learning schemes, consecutively push the performance on public benchmarks forward. However, the problem
Li, Cheng   +6 more
core   +1 more source

Fast offline transformer-based end-to-end automatic speech recognition for real-world applications

open access: yesETRI Journal, 2022
With the recent advances in technology, automatic speech recognition (ASR) has been widely used in real-world applications. The efficiency of converting large amounts of speech into text accurately with limited resources has become more vital than ever ...
Yoo Rhee Oh, Kiyoung Park, Kiyoung Park
doaj   +1 more source

Improving CTC-AED model with integrated-CTC and auxiliary loss regularization [PDF]

open access: yes, 2023
Connectionist temporal classification (CTC) and attention-based encoder decoder (AED) joint training has been widely applied in automatic speech recognition (ASR).
Su, Xiangdong   +2 more
core   +1 more source

Home - About - Disclaimer - Privacy