An Input-synchronous Blockwise Decoding Algorithm for CTC-AED Speech Recognition
Automatic speech recognition (ASR) systems for real-life scenarios are required to process audio streams of arbitrary length with stable accuracy under limited computational resources.
Iurii Lezhenin, Natalia Bogach
doaj +1 more source
Digital display instrument identification is a crucial approach for automating the collection of digital display data. In this study, we propose a digital display area detection CTPNpro algorithm to address the problem of recognizing multiclass digital ...
Xuanzhang Wen +5 more
doaj +1 more source
Uncertainty‐Aware Processing for OCR Using Dictionary Routing and Candidate Selection
This paper proposes an Uncertainty‐Aware Processing (UAP) framework that reduces the Character Error Rate (CER) without additional training or inference. The method calculates two metrics using the results of Connectionist Temporal classification (CTC). These two metrics are used to estimate a character‐level uncertainty.
Gi Hoon Kim, JiU Bak, Hyunguk Choi
wiley +1 more source
Audio–Visual Speech Recognition Based on Dual Cross-Modality Attentions with the Transformer Model
Since attention mechanism was introduced in neural machine translation, attention has been combined with the long short-term memory (LSTM) or replaced the LSTM in a transformer model to overcome the sequence-to-sequence (seq2seq) problems with the LSTM ...
Yong-Hyeok Lee +4 more
doaj +1 more source
Investigating Sequence-Level Normalisation for CTC-Like End-To-End ASR [PDF]
End-to-end Automatic Speech Recognition (E2E ASR) significantly simplifies the training process of an ASR model. Connectionist Temporal Classification (CTC) is one of the most popular methods for E2E ASR training.
Zhao, Zeyu, Bell, Peter
core +1 more source
Advances in Detecting RNA Modifications Using Direct RNA Nanopore Sequencing
This review examines recent advances in Oxford Nanopore Technologies direct RNA sequencing, highlighting its expanding capacity to detect RNA modifications beyond m6A. It discusses computational frameworks and basecalling innovations that enable single‐nucleotide and single‐molecule resolution, explores co‐occurring modifications and their regulatory ...
Yaran Liu, Yang Li, Qiang Sun
wiley +1 more source
Freight rail activity inventory system using a vision‐based deep learning framework
Abstract Rail freight serves as a reliable cost‐effective and fuel‐efficient mode for long‐distance ground freight transportation. Existing rail data sources rely heavily on aggregate reports that lead to significant spatiotemporal data gaps for infrastructure planning and regulatory evaluation.
Guoliang Feng +3 more
wiley +1 more source
Terminal strip detection and recognition based on improved YOLOv7-tiny and MAH-CRNN+CTC models
For substation secondary circuit terminal strip wiring, low efficiency, less easy fault detection and inspection, and a variety of other issues, this study proposes a text detection and identification model based on improved YOLOv7-tiny and MAH-CRNN+CTC ...
Guo Zhijun +3 more
doaj +1 more source
Decoding Handwriting Trajectories from Intracortical Brain Signals for Brain‐to‐Text Communication
By developing a novel framework that optimizes both shape and temporal loss during decoder training, the authors successfully reconstruct human‐recognizable handwriting trajectories from intracortical neural signals for both Chinese characters and English letters, effectively resolving the temporal misalignment problem in clinical BCIs, thereby ...
Guangxiang Xu +6 more
wiley +1 more source
Blank Collapse: Compressing CTC emission for the faster decoding [PDF]
Connectionist Temporal Classification (CTC) model is a very efficient method for modeling sequences, especially for speech data. In order to use CTC model as an Automatic Speech Recognition (ASR) task, the beam search decoding with an external language ...
Kwon, Ohhyeok +3 more
core +1 more source

