Results 81 to 90 of about 714 (168)

End-to-end spoken language understanding using joint CTC loss and self-supervised, pretrained acoustic encoders [PDF]

open access: yes, 2023
It is challenging to extract semantic meanings directly from audio signals in spoken language understanding (SLU), due to the lack of textual information. Popular end-to-end (E2E) SLU models utilize sequence-to-sequence automatic speech recognition (ASR)
Wei, Kai   +3 more
core   +1 more source

A person re‐identification method for sports event scenes incorporating textual information mining

open access: yesIET Image Processing, Volume 18, Issue 7, Page 1681-1693, 29 May 2024.
This paper designs a multi‐source information mutual gain mechanism to improve the accuracy of the person re‐identification task. The contribution is threefold: (1) A mutual gain mechanism between a person's visual features and high‐level semantic information is designed to enhance the performance of person re‐identification methods. (2) To the best of
Runmin Wang   +7 more
wiley   +1 more source

Deep learning based 3D residual convolutional and Multi-Head Attention (3D-RMA) for lip-reading

open access: yesResults in Control and Optimization
Lip reading, an essential yet intricate facet of communication, has seen notable progress through the application of advanced deep learning techniques. This research introduces a deep learning-based lip-reading model that integrates Conv3D layers, Multi ...
Archana Chaudhari   +6 more
doaj   +1 more source

Advancing CTC models for better speech alignment:A topological approach [PDF]

open access: yes
Automatic Speech Recognition (ASR) systems often face challenges in alignment quality, particularly with the Connectionist Temporal Classification (CTC) approach, which frequently results in a high number of blank frames, known as the “peaky” issue.
Zhao, Zeyu, Bell, Peter
core   +1 more source

Bypass Temporal Classification: Weakly Supervised Automatic Speech Recognition with Imperfect Transcripts [PDF]

open access: yes, 2023
This paper presents a novel algorithm for building an automatic speech recognition (ASR) model with imperfect training data. Imperfectly transcribed speech is a prevalent issue in human-annotated speech corpora, which degrades the performance of ASR ...
Xu, Hainan   +5 more
core   +1 more source

A CRNN‐based method for Chinese ship license plate recognition

open access: yesIET Image Processing, Volume 18, Issue 2, Page 298-311, 7 February 2024.
A novel CRNN‐based method for Chinese ship license plate recognition under harsh marine environment is reported. This method can achieve significant recognition accuracy. Abstract Existing deep learning methods cannot achieve satisfactory ship license plate (SLP) recognition due to the harsh marine environment, such as foggy weather, unstable ship ...
Fan Xu   +4 more
wiley   +1 more source

InterMPL: Momentum Pseudo-Labeling with Intermediate CTC Loss [PDF]

open access: yes, 2023
This paper presents InterMPL, a semi-supervised learning method of end-to-end automatic speech recognition (ASR) that performs pseudo-labeling (PL) with intermediate supervision.
Higuchi, Yosuke   +3 more
core   +1 more source

CTC-based Non-autoregressive Speech Translation [PDF]

open access: yes, 2023
Combining end-to-end speech translation (ST) and non-autoregressive (NAR) generation is promising in language and speech processing for their advantages of less error propagation and low latency.
Zhu, Jingbo   +11 more
core   +1 more source

ResNet18-BLSTM-CTC-Based OCR Model for Extracting Technical Parameter Data from Digital Television Analyzer Screenshots

open access: yesTeknika
Monitoring digital television broadcasting stations requires accurate extraction of technical measurement parameters from screenshot images generated by television analyzer devices.
Henrian Robby Fakhriannur   +1 more
doaj   +1 more source

Char+CV-CTC: Combining Graphemes and Consonant/Vowel Units for CTC-Based ASR Using Multitask Learning [PDF]

open access: yes, 2019
Previous work has shown that end-to-end neural-based speech recognition systems can be improved by adding auxiliary tasks at intermediate layers. In this paper, we report multitask learning (MTL) experiments in the context of connectionist temporal ...
Abdelwahab Heba   +7 more
core   +1 more source

Home - About - Disclaimer - Privacy