Audio–Visual Speech Recognition Based on Dual Cross-Modality Attentions with the Transformer Model
Since attention mechanism was introduced in neural machine translation, attention has been combined with the long short-term memory (LSTM) or replaced the LSTM in a transformer model to overcome the sequence-to-sequence (seq2seq) problems with the LSTM ...
Yong-Hyeok Lee +4 more
doaj +1 more source
Digital display instrument identification is a crucial approach for automating the collection of digital display data. In this study, we propose a digital display area detection CTPNpro algorithm to address the problem of recognizing multiclass digital ...
Xuanzhang Wen +5 more
doaj +1 more source
Decoding Handwriting Trajectories from Intracortical Brain Signals for Brain‐to‐Text Communication
By developing a novel framework that optimizes both shape and temporal loss during decoder training, the authors successfully reconstruct human‐recognizable handwriting trajectories from intracortical neural signals for both Chinese characters and English letters, effectively resolving the temporal misalignment problem in clinical BCIs, thereby ...
Guangxiang Xu +6 more
wiley +1 more source
Terminal strip detection and recognition based on improved YOLOv7-tiny and MAH-CRNN+CTC models
For substation secondary circuit terminal strip wiring, low efficiency, less easy fault detection and inspection, and a variety of other issues, this study proposes a text detection and identification model based on improved YOLOv7-tiny and MAH-CRNN+CTC ...
Guo Zhijun +3 more
doaj +1 more source
The Rise of Large Language Models: Evolution, Applications, and Future Directions
This paper provides a comprehensive Systematic Literature Review (SLR) on Large Language Models (LLMs), covering their evolution, applications, evaluation metrics, and challenges. It identifies key research gaps and future directions, offering a structured taxonomy and analysis of performance environments, datasets, and open issues in LLM research ...
Amir Masoud Rahmani +2 more
wiley +1 more source
Non‐Autoregressive Translation Algorithm Based on LLM Knowledge Distillation in English Corpus
This research introduces an English corpus‐based machine translation algorithm that leverages knowledge distillation from large language model, with the goal of enhancing translation quality and reducing the computational demands of the model. ABSTRACT Although significant advancements have been made in the quality of machine translation by large‐scale
Fang Ju, Weihui Wang
wiley +1 more source
Multi global context‐aware transformer for ship name recognition in IoT
In this work, we propose a novel multi global context‐aware transformer (MG‐CAT) to address the problem of ship name recognition. Abstract Scene text recognition has gained increasing attention in recent years, as it can connect products without an open interface in IoT.
Yunting Xian +3 more
wiley +1 more source
MVDT: Multiview Distillation Transformer for View‐Invariant Sign Language Translation
Sign language translation (SLT) is a critical technology for bridging communication gaps between deaf and hearing communities. However, existing single‐view SLT models suffer from performance degradation due to viewpoint dependency, occlusion and limited observation angles.
Zhong Guan +4 more
wiley +1 more source
A new joint CTC-attention-based speech recognition model with multi-level multi-head attention
A method called joint connectionist temporal classification (CTC)-attention-based speech recognition has recently received increasing focus and has achieved impressive performance.
Chu-Xiong Qin, Wen-Lin Zhang, Dan Qu
doaj +1 more source
A Comprehensive Survey of Advancement in Lip Reading Models: Techniques and Future Directions
Efficient and accurate lip reading models increase information processing and decision‐making by understanding massive quantities of text. Lip reading can make communication more inclusive, especially for hearing‐impaired people, according to this study. From 2020 to 2024, researchers track lip‐reading algorithm progress.
Sampada Deshpande +5 more
wiley +1 more source

