Results 1 to 10 of about 423 (132)

Explainable Connectionist-Temporal-Classification-Based Scene Text Recognition [PDF]

open access: yesJournal of Imaging, 2023
Connectionist temporal classification (CTC) is a favored decoder in scene text recognition (STR) for its simplicity and efficiency. However, most CTC-based methods utilize one-dimensional (1D) vector sequences, usually derived from a recurrent neural ...
Rina Buoy   +3 more
doaj   +4 more sources

Spoken term detection with Connectionist Temporal Classification: A novel hybrid CTC-DBN decoder [PDF]

open access: yes2010 IEEE International Conference on Acoustics, Speech and Signal Processing, 2010
This paper proposes a novel system for robust keyword detection in continuous speech. Our decoder is composed of a bidirectional Long Short-Term Memory recurrent neural network using a Connectionist Temporal Classification (CTC) output layer, and a Dynamic Bayesian Network (DBN).
Martin Wöllmer   +3 more
openaire   +4 more sources

A CTC-Based Speech Recognition Network Fusing Local Convolution and Global Attention [PDF]

open access: yesSensors
Integrating wav2vec 2.0 with Connectionist Temporal Classification (CTC) for automatic speech recognition (ASR) often involves a trade-off between capturing global semantic consistency and maintaining local feature discriminability.
Huijuan Hu   +3 more
doaj   +2 more sources

MPSA-Conformer-CTC/Attention: A High-Accuracy, Low-Complexity End-to-End Approach for Tibetan Speech Recognition [PDF]

open access: yesSensors
This study addresses the challenges of low accuracy and high computational demands in Tibetan speech recognition by investigating the application of end-to-end networks. We propose a decoding strategy that integrates Connectionist Temporal Classification
Changlin Wu   +3 more
doaj   +2 more sources

AdaMER-CTC: Connectionist Temporal Classification with Adaptive Maximum Entropy Regularization for Automatic Speech Recognition

open access: yesICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
In Automatic Speech Recognition (ASR) systems, a recurring obstacle is the generation of narrowly focused output distributions. This phenomenon emerges as a side effect of Connectionist Temporal Classification (CTC), a robust sequence learning tool that utilizes dynamic programming for sequence mapping.
SooHwan Eom   +5 more
openaire   +4 more sources

SqueezeCall: nanopore basecalling using a Squeezeformer network [PDF]

open access: yesGigaByte
Nanopore sequencing, a third-generation sequencing technique, enables direct RNA sequencing, real-time analysis, and long-read length. Nanopore sequencers measure electrical current changes as nucleotides pass through nanopores; a basecaller identifies ...
Zhongxu Zhu
doaj   +2 more sources

A zero-shot LLM framework for multimodal grievance classification, urgency scoring, and abuse detection in civic feedback systems [PDF]

open access: yesScientific Reports
A unified model is presented for civic grievance redressal, integrating multimodal complaint intake, zero-shot semantic routing, sentiment-derived urgency estimation, and behavior-sensitive abuse detection within a scalable microservice architecture. The
S. C. Rajkumar   +3 more
doaj   +2 more sources

Context Conditioning via Surrounding Predictions for Non-Recurrent CTC Models

open access: yesIEEE Access, 2023
Connectionist Temporal Classification (CTC) loss has become widely used in sequence modeling tasks such as Automatic Speech Recognition (ASR) and Handwritten Text Recognition (HTR) due to its ease of use.
Burin Naowarat   +2 more
doaj   +1 more source

Recognition of English speech – using a deep learning algorithm

open access: yesJournal of Intelligent Systems, 2023
The accurate recognition of speech is beneficial to the fields of machine translation and intelligent human–computer interaction. After briefly introducing speech recognition algorithms, this study proposed to recognize speech with a recurrent neural ...
Wang Shuyan
doaj   +1 more source

End-to-End Automatic Pronunciation Error Detection Based on Improved Hybrid CTC/Attention Architecture

open access: yesSensors, 2020
Advanced automatic pronunciation error detection (APED) algorithms are usually based on state-of-the-art automatic speech recognition (ASR) techniques. With the development of deep learning technology, end-to-end ASR technology has gradually matured and ...
Long Zhang   +7 more
doaj   +1 more source

Home - About - Disclaimer - Privacy