End-to-end spoken language understanding using joint CTC loss and self-supervised, pretrained acoustic encoders [PDF]
It is challenging to extract semantic meanings directly from audio signals in spoken language understanding (SLU), due to the lack of textual information. Popular end-to-end (E2E) SLU models utilize sequence-to-sequence automatic speech recognition (ASR)
Wei, Kai +3 more
core +1 more source
A person re‐identification method for sports event scenes incorporating textual information mining
This paper designs a multi‐source information mutual gain mechanism to improve the accuracy of the person re‐identification task. The contribution is threefold: (1) A mutual gain mechanism between a person's visual features and high‐level semantic information is designed to enhance the performance of person re‐identification methods. (2) To the best of
Runmin Wang +7 more
wiley +1 more source
Deep learning based 3D residual convolutional and Multi-Head Attention (3D-RMA) for lip-reading
Lip reading, an essential yet intricate facet of communication, has seen notable progress through the application of advanced deep learning techniques. This research introduces a deep learning-based lip-reading model that integrates Conv3D layers, Multi ...
Archana Chaudhari +6 more
doaj +1 more source
Advancing CTC models for better speech alignment:A topological approach [PDF]
Automatic Speech Recognition (ASR) systems often face challenges in alignment quality, particularly with the Connectionist Temporal Classification (CTC) approach, which frequently results in a high number of blank frames, known as the “peaky” issue.
Zhao, Zeyu, Bell, Peter
core +1 more source
Bypass Temporal Classification: Weakly Supervised Automatic Speech Recognition with Imperfect Transcripts [PDF]
This paper presents a novel algorithm for building an automatic speech recognition (ASR) model with imperfect training data. Imperfectly transcribed speech is a prevalent issue in human-annotated speech corpora, which degrades the performance of ASR ...
Xu, Hainan +5 more
core +1 more source
A CRNN‐based method for Chinese ship license plate recognition
A novel CRNN‐based method for Chinese ship license plate recognition under harsh marine environment is reported. This method can achieve significant recognition accuracy. Abstract Existing deep learning methods cannot achieve satisfactory ship license plate (SLP) recognition due to the harsh marine environment, such as foggy weather, unstable ship ...
Fan Xu +4 more
wiley +1 more source
InterMPL: Momentum Pseudo-Labeling with Intermediate CTC Loss [PDF]
This paper presents InterMPL, a semi-supervised learning method of end-to-end automatic speech recognition (ASR) that performs pseudo-labeling (PL) with intermediate supervision.
Higuchi, Yosuke +3 more
core +1 more source
CTC-based Non-autoregressive Speech Translation [PDF]
Combining end-to-end speech translation (ST) and non-autoregressive (NAR) generation is promising in language and speech processing for their advantages of less error propagation and low latency.
Zhu, Jingbo +11 more
core +1 more source
Monitoring digital television broadcasting stations requires accurate extraction of technical measurement parameters from screenshot images generated by television analyzer devices.
Henrian Robby Fakhriannur +1 more
doaj +1 more source
Char+CV-CTC: Combining Graphemes and Consonant/Vowel Units for CTC-Based ASR Using Multitask Learning [PDF]
Previous work has shown that end-to-end neural-based speech recognition systems can be improved by adding auxiliary tasks at intermediate layers. In this paper, we report multitask learning (MTL) experiments in the context of connectionist temporal ...
Abdelwahab Heba +7 more
core +1 more source

