Results 41 to 50 of about 714 (168)
GTC: Guided Training of CTC towards Efficient and Accurate Scene Text Recognition [PDF]
Connectionist Temporal Classification (CTC) and attention mechanism are two main approaches used in recent scene text recognition works. Compared with attention-based methods, CTC decoder has a much shorter inference time, yet a lower accuracy. To design
Yi, Shuai +4 more
core +1 more source
To solve the problem of the low recognition rate of continuous dynamic gestures in Chinese sign language, a non-invasive end-to-end continuous dynamic gesture recognition system combining Inertial Measurement Unit (IMU) signal and surface ...
Jinquan Li +3 more
doaj +1 more source
We solve the problem of how to densely align actions in videos at frame level, with only the order of occurring actions available, in order to save the time-consuming efforts to accurately annotate the temporal boundaries of each action. We propose three
Lin Wang +3 more
doaj +1 more source
Pengenalan Karakter Optikal Aksara Jawa Menggunakan Connectionist Temporal Classification [PDF]
Aksara Jawa memiliki sejarah panjang dan penting di Pulau Jawa. Sampai saat ini aksara Jawa banyak digunakan untuk obyek penelitian. Salah satunya menggunakan aplikasi OCR.
Saputro, Wahju Tjahjo +2 more
core +2 more sources
Inter‐Model Feature Fusion for Robust Low‐Resource Speech Recognition
Our Self‐Supervised Feature Fusion (SSF‐FT) method enhances low‐resource speech recognition by adaptively combining features from self‐supervised models trained with Contrastive, Predictive, and Reconstruction objectives. This attention‐weighted ensemble delivers robust performance, particularly in acoustically challenging conditions, extending current
Ussen Kimanuka +2 more
wiley +1 more source
Tibetan Data Augmentation via GAN‐Based Handwritten Text Generation
ABSTRACT Increased awareness of Tibetan cultural preservation, along with technological advancements, has led to significant efforts in academic research on Tibetan. However, the structural complexity of the Tibetan language and limited labeled handwriting data impede advancements in Optical Character Recognition (OCR) and other applications.
Dorje Tashi +9 more
wiley +1 more source
A deep neural network-based automatic mispronunciation detection in Bengali accented English speech
Learning a second language, especially English, became necessary as globalisation started. One crucial component of language learning resources is computer-assisted pronunciation training, or CAPT.
Puja Bharati +5 more
doaj +1 more source
Bidirectional Representations for Low-Resource Spoken Language Understanding
Speech representation models lack the ability to efficiently store semantic information and require fine tuning to deliver decent performance. In this research, we introduce a transformer encoder–decoder framework with a multiobjective training strategy,
Quentin Meeus +2 more
doaj +1 more source
Delay-penalized CTC implemented based on Finite State Transducer [PDF]
Connectionist Temporal Classification (CTC) suffers from the latency problem when applied to streaming models. We argue that in CTC lattice, the alignments that can access more future context are preferred during training, thereby leading to higher ...
Kang, Wei +7 more
core +1 more source
This survey systematically reviews state‐of‐the‐art license plate recognition methods, with a focus on hybrid CNN‐transformer frameworks and the joint optimisation of detection and recognition for real‐world deployment. It further analyses existing datasets, highlights persistent challenges, such as domain generalisation, and outlines pathways towards ...
SanXing Deng +3 more
wiley +1 more source

