Explainable Connectionist-Temporal-Classification-Based Scene Text Recognition [PDF]
Connectionist temporal classification (CTC) is a favored decoder in scene text recognition (STR) for its simplicity and efficiency. However, most CTC-based methods utilize one-dimensional (1D) vector sequences, usually derived from a recurrent neural ...
Rina Buoy +3 more
doaj +6 more sources
MPSA-Conformer-CTC/Attention: A High-Accuracy, Low-Complexity End-to-End Approach for Tibetan Speech Recognition [PDF]
This study addresses the challenges of low accuracy and high computational demands in Tibetan speech recognition by investigating the application of end-to-end networks. We propose a decoding strategy that integrates Connectionist Temporal Classification
Changlin Wu +3 more
doaj +4 more sources
Integrating international Chinese visualization teaching and vocational skills training: leveraging attention-connectionist temporal classification models [PDF]
The teaching of Chinese as a second language has become increasingly crucial for promoting cross-cultural exchange and mutual learning worldwide. However, traditional approaches to international Chinese language teaching have limitations that hinder ...
Yuan Yao, Zhujun Dai, Muhammad Shahbaz
doaj +5 more sources
Spoken term detection with Connectionist Temporal Classification: A novel hybrid CTC-DBN decoder [PDF]
This paper proposes a novel system for robust keyword detection in continuous speech. Our decoder is composed of a bidirectional Long Short-Term Memory recurrent neural network using a Connectionist Temporal Classification (CTC) output layer, and a Dynamic Bayesian Network (DBN).
Martin Wöllmer +3 more
openaire +4 more sources
A CTC-Based Speech Recognition Network Fusing Local Convolution and Global Attention [PDF]
Integrating wav2vec 2.0 with Connectionist Temporal Classification (CTC) for automatic speech recognition (ASR) often involves a trade-off between capturing global semantic consistency and maintaining local feature discriminability.
Huijuan Hu +3 more
doaj +2 more sources
In Automatic Speech Recognition (ASR) systems, a recurring obstacle is the generation of narrowly focused output distributions. This phenomenon emerges as a side effect of Connectionist Temporal Classification (CTC), a robust sequence learning tool that utilizes dynamic programming for sequence mapping.
SooHwan Eom +5 more
openaire +4 more sources
SqueezeCall: nanopore basecalling using a Squeezeformer network [PDF]
Nanopore sequencing, a third-generation sequencing technique, enables direct RNA sequencing, real-time analysis, and long-read length. Nanopore sequencers measure electrical current changes as nucleotides pass through nanopores; a basecaller identifies ...
Zhongxu Zhu
doaj +2 more sources
A zero-shot LLM framework for multimodal grievance classification, urgency scoring, and abuse detection in civic feedback systems [PDF]
A unified model is presented for civic grievance redressal, integrating multimodal complaint intake, zero-shot semantic routing, sentiment-derived urgency estimation, and behavior-sensitive abuse detection within a scalable microservice architecture. The
S. C. Rajkumar +3 more
doaj +2 more sources
An End-To-End Speech Recognition Model for the North Shaanxi Dialect: Design and Evaluation [PDF]
The coal mining industry in Northern Shaanxi is robust, with a prevalent use of the local dialect, known as “Shapu”, characterized by a distinct Northern Shaanxi accent. This study addresses the practical need for speech recognition in this dialect.
Yi Qin, Feifan Yu
doaj +2 more sources
SVTRv2X: Enhanced scene text recognition via self-distilled mixture-of-experts. [PDF]
Scene Text Recognition (STR) is a fundamental component of intelligent perception systems and plays a crucial role in a wide range of real-world applications such as autonomous driving, document understanding, and human-computer interaction.
Jian Guo +5 more
doaj +2 more sources

