Results 61 to 70 of about 714 (168)

The Rise of Large Language Models: Evolution, Applications, and Future Directions

open access: yesEngineering Reports, Volume 7, Issue 9, September 2025.
This paper provides a comprehensive Systematic Literature Review (SLR) on Large Language Models (LLMs), covering their evolution, applications, evaluation metrics, and challenges. It identifies key research gaps and future directions, offering a structured taxonomy and analysis of performance environments, datasets, and open issues in LLM research ...
Amir Masoud Rahmani   +2 more
wiley   +1 more source

Self-Distillation Regularized Connectionist Temporal Classification Loss for Text Recognition: A Simple Yet Effective Approach [PDF]

open access: yes
Text recognition methods are gaining rapid development. Some advanced techniques, e.g., powerful modules, language models, and un- and semi-supervised learning schemes, consecutively push the performance on public benchmarks forward. However, the problem
Li, Cheng   +6 more
core   +1 more source

Efficient CTC Regularization via Coarse Labels for End-to-End Speech Translation [PDF]

open access: yes, 2023
For end-to-end speech translation, regularizing the encoder with the Connectionist Temporal Classification (CTC) objective using the source transcript or target translation as labels can greatly improve quality.
Zhang, Biao   +2 more
core   +1 more source

Non‐Autoregressive Translation Algorithm Based on LLM Knowledge Distillation in English Corpus

open access: yesEngineering Reports, Volume 7, Issue 1, January 2025.
This research introduces an English corpus‐based machine translation algorithm that leverages knowledge distillation from large language model, with the goal of enhancing translation quality and reducing the computational demands of the model. ABSTRACT Although significant advancements have been made in the quality of machine translation by large‐scale
Fang Ju, Weihui Wang
wiley   +1 more source

A new joint CTC-attention-based speech recognition model with multi-level multi-head attention

open access: yesEURASIP Journal on Audio, Speech, and Music Processing, 2019
A method called joint connectionist temporal classification (CTC)-attention-based speech recognition has recently received increasing focus and has achieved impressive performance.
Chu-Xiong Qin, Wen-Lin Zhang, Dan Qu
doaj   +1 more source

BERT Meets CTC: New Formulation of End-to-End Speech Recognition with Pre-trained Masked Language Model [PDF]

open access: yes, 2023
This paper presents BERT-CTC, a novel formulation of end-to-end speech recognition that adapts BERT for connectionist temporal classification (CTC).
Higuchi, Yosuke   +5 more
core   +1 more source

Multi global context‐aware transformer for ship name recognition in IoT

open access: yesIET Communications, Volume 19, Issue 1, January/December 2025.
In this work, we propose a novel multi global context‐aware transformer (MG‐CAT) to address the problem of ship name recognition. Abstract Scene text recognition has gained increasing attention in recent years, as it can connect products without an open interface in IoT.
Yunting Xian   +3 more
wiley   +1 more source

MVDT: Multiview Distillation Transformer for View‐Invariant Sign Language Translation

open access: yesIET Computer Vision, Volume 19, Issue 1, January/December 2025.
Sign language translation (SLT) is a critical technology for bridging communication gaps between deaf and hearing communities. However, existing single‐view SLT models suffer from performance degradation due to viewpoint dependency, occlusion and limited observation angles.
Zhong Guan   +4 more
wiley   +1 more source

Speech Recognition for Air Traffic Control Utilizing a Multi-Head State-Space Model and Transfer Learning

open access: yesAerospace
In the present study, a novel end-to-end automatic speech recognition (ASR) framework, namely, ResNeXt-Mssm-CTC, has been developed for air traffic control (ATC) systems.
Haijun Liang, Hanwen Chang, Jianguo Kong
doaj   +1 more source

A Comprehensive Survey of Advancement in Lip Reading Models: Techniques and Future Directions

open access: yesIET Image Processing, Volume 19, Issue 1, January/December 2025.
Efficient and accurate lip reading models increase information processing and decision‐making by understanding massive quantities of text. Lip reading can make communication more inclusive, especially for hearing‐impaired people, according to this study. From 2020 to 2024, researchers track lip‐reading algorithm progress.
Sampada Deshpande   +5 more
wiley   +1 more source

Home - About - Disclaimer - Privacy