Results 71 to 80 of about 714 (168)
This study presents an advanced OCR model for complex Chinese character recognition, integrating EfficientNetV2 with a Transformer‐based architecture. Key innovations include a DCCSA module for enhanced feature emphasis and RoPE for precise spatial relationship modeling.
Wei Deng +4 more
wiley +1 more source
SCUT-EPT: New Dataset and Benchmark for Offline Chinese Text Recognition in Examination Paper
Most existing studies and public datasets for handwritten Chinese text recognition are based on the regular documents with clean and blank background, lacking research reports for handwritten text recognition on challenging areas such as educational ...
Yuanzhi Zhu +5 more
doaj +1 more source
Inter-KD: Intermediate Knowledge Distillation for CTC-Based Automatic Speech Recognition [PDF]
Recently, the advance in deep learning has brought a considerable improvement in the end-to-end speech recognition field, simplifying the traditional pipeline while producing promising results.
Lee, Hyeonseung +4 more
core +1 more source
Persian/Arabic Scene Text Recognition With Convolutional Recurrent Neural Network
Advancements in technology have made natural scene text recognition (STR) crucial and challenging due to various factors such as font, colour, texture, illumination, and complex backgrounds. This research study explores optical character recognition (OCR) for Iranian signposts, traffic signs, and licence plates by combining preprocessing techniques ...
Alireza Akoushideh +2 more
wiley +1 more source
Framewise and CTC Training of Neural Networks for Handwriting Recognition [PDF]
-In recent years, Long Short-Term Memory Recurrent Neural Networks (LSTM-RNNs) trained with the Connectionist Temporal Classification (CTC) objective won many international handwriting recognition evaluations.
Théodore Bluche +3 more
core
Abstract Research Summary Multimodal data, comprising interdependent unstructured text, image, and audio data that collectively characterize the same source, with video being a prominent example, offer a wealth of information for strategy researchers.
Xueming Luo +3 more
wiley +1 more source
Relay Protection Setting Sheet Detection and Recognition Approach Based on YOLOv8-CRNN-CTC
To address the challenging recognition of relay protection setting characters in industrial scenarios due to their small size, dense arrangement, complex backgrounds, and susceptibility to lighting effects, a two-stage automatic recognition method ...
Xiaohao Lv +3 more
doaj +1 more source
Cross-modal Alignment with Optimal Transport for CTC-based ASR [PDF]
Temporal connectionist temporal classification (CTC)-based automatic speech recognition (ASR) is one of the most successful end to end (E2E) ASR frameworks. However, due to the token independence assumption in decoding, an external language model (LM) is
Lu, Xugang +3 more
core +1 more source
Cantonese sentence dataset for lip‐reading
Lip‐reading deciphers speech without audio data, and deep learning advancements have improved lip‐reading in English and Chinese. Cantonese lip‐reading sentences, a Cantonese lip‐reading dataset, and a novel visual frontend, 3D‐visual attention net, which achieves comparable performance on Chinese Mandarin lip reading dataset, lip reading sentences 2 ...
Yewei Xiao +5 more
wiley +1 more source
Sequence labeling is a common machine-learning task which not only needs the most likely prediction of label for a local input but also seeks the most suitable annotation for the whole input sequence.
Xiaohui Huang +4 more
doaj +1 more source

