The Rise of Large Language Models: Evolution, Applications, and Future Directions
This paper provides a comprehensive Systematic Literature Review (SLR) on Large Language Models (LLMs), covering their evolution, applications, evaluation metrics, and challenges. It identifies key research gaps and future directions, offering a structured taxonomy and analysis of performance environments, datasets, and open issues in LLM research ...
Amir Masoud Rahmani +2 more
wiley +1 more source
Self-Distillation Regularized Connectionist Temporal Classification Loss for Text Recognition: A Simple Yet Effective Approach [PDF]
Text recognition methods are gaining rapid development. Some advanced techniques, e.g., powerful modules, language models, and un- and semi-supervised learning schemes, consecutively push the performance on public benchmarks forward. However, the problem
Li, Cheng +6 more
core +1 more source
Efficient CTC Regularization via Coarse Labels for End-to-End Speech Translation [PDF]
For end-to-end speech translation, regularizing the encoder with the Connectionist Temporal Classification (CTC) objective using the source transcript or target translation as labels can greatly improve quality.
Zhang, Biao +2 more
core +1 more source
Non‐Autoregressive Translation Algorithm Based on LLM Knowledge Distillation in English Corpus
This research introduces an English corpus‐based machine translation algorithm that leverages knowledge distillation from large language model, with the goal of enhancing translation quality and reducing the computational demands of the model. ABSTRACT Although significant advancements have been made in the quality of machine translation by large‐scale
Fang Ju, Weihui Wang
wiley +1 more source
A new joint CTC-attention-based speech recognition model with multi-level multi-head attention
A method called joint connectionist temporal classification (CTC)-attention-based speech recognition has recently received increasing focus and has achieved impressive performance.
Chu-Xiong Qin, Wen-Lin Zhang, Dan Qu
doaj +1 more source
BERT Meets CTC: New Formulation of End-to-End Speech Recognition with Pre-trained Masked Language Model [PDF]
This paper presents BERT-CTC, a novel formulation of end-to-end speech recognition that adapts BERT for connectionist temporal classification (CTC).
Higuchi, Yosuke +5 more
core +1 more source
Multi global context‐aware transformer for ship name recognition in IoT
In this work, we propose a novel multi global context‐aware transformer (MG‐CAT) to address the problem of ship name recognition. Abstract Scene text recognition has gained increasing attention in recent years, as it can connect products without an open interface in IoT.
Yunting Xian +3 more
wiley +1 more source
MVDT: Multiview Distillation Transformer for View‐Invariant Sign Language Translation
Sign language translation (SLT) is a critical technology for bridging communication gaps between deaf and hearing communities. However, existing single‐view SLT models suffer from performance degradation due to viewpoint dependency, occlusion and limited observation angles.
Zhong Guan +4 more
wiley +1 more source
In the present study, a novel end-to-end automatic speech recognition (ASR) framework, namely, ResNeXt-Mssm-CTC, has been developed for air traffic control (ATC) systems.
Haijun Liang, Hanwen Chang, Jianguo Kong
doaj +1 more source
A Comprehensive Survey of Advancement in Lip Reading Models: Techniques and Future Directions
Efficient and accurate lip reading models increase information processing and decision‐making by understanding massive quantities of text. Lip reading can make communication more inclusive, especially for hearing‐impaired people, according to this study. From 2020 to 2024, researchers track lip‐reading algorithm progress.
Sampada Deshpande +5 more
wiley +1 more source

