Results 21 to 30 of about 423 (132)
A Hardware-Oriented and Memory-Efficient Method for CTC Decoding
The Connectionist Temporal Classification (CTC) has achieved great success in sequence to sequence analysis tasks such as automatic speech recognition (ASR) and scene text recognition (STR).
Siyuan Lu +3 more
doaj +1 more source
Focal CTC Loss for Chinese Optical Character Recognition on Unbalanced Datasets
In this paper, we propose a novel deep model for unbalanced distribution Character Recognition by employing focal loss based connectionist temporal classification (CTC) function.
Xinjie Feng +2 more
doaj +1 more source
Modeling Intra-label Dynamics and Analyzing the Role of Blank in Connectionist Temporal Classification [PDF]
The goal of many tasks in the realm of sequence processing is to map a sequence of input data to a sequence of output labels. Long short-term memory (LSTM), a type of recurrent neural network (RNN), equipped with connectionist temporal classification ...
Ashkan Sadeghi Lotfabadi +2 more
doaj +1 more source
A Deep Diacritics-Based Recognition Model for Arabic Speech: Quranic Verses as Case Study
Arabic is the language of more than 422 million of the world’s population. Although classic Arabic is the Quran language that 1.9 billion Muslims are required to recite, limited Arabic speech recognition exists.
Sarah S. Alrumiah, Amal A. Al-Shargabi
doaj +1 more source
Speech GAU: A Single Head Attention for Mandarin Speech Recognition for Air Traffic Control
The rise of end-to-end (E2E) speech recognition technology in recent years has overturned the design pattern of cascading multiple subtasks in classical speech recognition and achieved direct mapping of speech input signals to text labels. In this study,
Shiyu Zhang +4 more
doaj +1 more source
As demonstrated in hybrid connectionist temporal classification (CTC)/Attention architecture, joint training with a CTC objective is very effective to solve the misalignment problem existing in the attention-based end-to-end automatic speech recognition (
Long Wu, Ta Li, Li Wang, Yonghong Yan
doaj +1 more source
Fast offline transformer-based end-to-end automatic speech recognition for real-world applications
With the recent advances in technology, automatic speech recognition (ASR) has been widely used in real-world applications. The efficiency of converting large amounts of speech into text accurately with limited resources has become more vital than ever ...
Yoo Rhee Oh, Kiyoung Park, Kiyoung Park
doaj +1 more source
Handwritten Text Recognition (HTR) models often exhibit optimization difficulty when trained with small batch sizes. In this study, we investigate the temporal behavior of CTC logits in Connectionist Temporal Classification (CTC)-based HTR models ...
Pham Doan Tinh, Ha Huu An
doaj +1 more source
To solve the problem of the low recognition rate of continuous dynamic gestures in Chinese sign language, a non-invasive end-to-end continuous dynamic gesture recognition system combining Inertial Measurement Unit (IMU) signal and surface ...
Jinquan Li +3 more
doaj +1 more source
Inter‐Model Feature Fusion for Robust Low‐Resource Speech Recognition
Our Self‐Supervised Feature Fusion (SSF‐FT) method enhances low‐resource speech recognition by adaptively combining features from self‐supervised models trained with Contrastive, Predictive, and Reconstruction objectives. This attention‐weighted ensemble delivers robust performance, particularly in acoustically challenging conditions, extending current
Ussen Kimanuka +2 more
wiley +1 more source

