Results 41 to 50 of about 423 (132)

Audio–Visual Speech Recognition Based on Dual Cross-Modality Attentions with the Transformer Model

open access: yesApplied Sciences, 2020
Since attention mechanism was introduced in neural machine translation, attention has been combined with the long short-term memory (LSTM) or replaced the LSTM in a transformer model to overcome the sequence-to-sequence (seq2seq) problems with the LSTM ...
Yong-Hyeok Lee   +4 more
doaj   +1 more source

Design of recognition algorithm for multiclass digital display instrument based on convolution neural network

open access: yesBiomimetic Intelligence and Robotics, 2023
Digital display instrument identification is a crucial approach for automating the collection of digital display data. In this study, we propose a digital display area detection CTPNpro algorithm to address the problem of recognizing multiclass digital ...
Xuanzhang Wen   +5 more
doaj   +1 more source

Decoding Handwriting Trajectories from Intracortical Brain Signals for Brain‐to‐Text Communication

open access: yesAdvanced Science, Volume 12, Issue 40, October 27, 2025.
By developing a novel framework that optimizes both shape and temporal loss during decoder training, the authors successfully reconstruct human‐recognizable handwriting trajectories from intracortical neural signals for both Chinese characters and English letters, effectively resolving the temporal misalignment problem in clinical BCIs, thereby ...
Guangxiang Xu   +6 more
wiley   +1 more source

Terminal strip detection and recognition based on improved YOLOv7-tiny and MAH-CRNN+CTC models

open access: yesFrontiers in Energy Research
For substation secondary circuit terminal strip wiring, low efficiency, less easy fault detection and inspection, and a variety of other issues, this study proposes a text detection and identification model based on improved YOLOv7-tiny and MAH-CRNN+CTC ...
Guo Zhijun   +3 more
doaj   +1 more source

The Rise of Large Language Models: Evolution, Applications, and Future Directions

open access: yesEngineering Reports, Volume 7, Issue 9, September 2025.
This paper provides a comprehensive Systematic Literature Review (SLR) on Large Language Models (LLMs), covering their evolution, applications, evaluation metrics, and challenges. It identifies key research gaps and future directions, offering a structured taxonomy and analysis of performance environments, datasets, and open issues in LLM research ...
Amir Masoud Rahmani   +2 more
wiley   +1 more source

Non‐Autoregressive Translation Algorithm Based on LLM Knowledge Distillation in English Corpus

open access: yesEngineering Reports, Volume 7, Issue 1, January 2025.
This research introduces an English corpus‐based machine translation algorithm that leverages knowledge distillation from large language model, with the goal of enhancing translation quality and reducing the computational demands of the model. ABSTRACT Although significant advancements have been made in the quality of machine translation by large‐scale
Fang Ju, Weihui Wang
wiley   +1 more source

Multi global context‐aware transformer for ship name recognition in IoT

open access: yesIET Communications, Volume 19, Issue 1, January/December 2025.
In this work, we propose a novel multi global context‐aware transformer (MG‐CAT) to address the problem of ship name recognition. Abstract Scene text recognition has gained increasing attention in recent years, as it can connect products without an open interface in IoT.
Yunting Xian   +3 more
wiley   +1 more source

MVDT: Multiview Distillation Transformer for View‐Invariant Sign Language Translation

open access: yesIET Computer Vision, Volume 19, Issue 1, January/December 2025.
Sign language translation (SLT) is a critical technology for bridging communication gaps between deaf and hearing communities. However, existing single‐view SLT models suffer from performance degradation due to viewpoint dependency, occlusion and limited observation angles.
Zhong Guan   +4 more
wiley   +1 more source

A new joint CTC-attention-based speech recognition model with multi-level multi-head attention

open access: yesEURASIP Journal on Audio, Speech, and Music Processing, 2019
A method called joint connectionist temporal classification (CTC)-attention-based speech recognition has recently received increasing focus and has achieved impressive performance.
Chu-Xiong Qin, Wen-Lin Zhang, Dan Qu
doaj   +1 more source

A Comprehensive Survey of Advancement in Lip Reading Models: Techniques and Future Directions

open access: yesIET Image Processing, Volume 19, Issue 1, January/December 2025.
Efficient and accurate lip reading models increase information processing and decision‐making by understanding massive quantities of text. Lip reading can make communication more inclusive, especially for hearing‐impaired people, according to this study. From 2020 to 2024, researchers track lip‐reading algorithm progress.
Sampada Deshpande   +5 more
wiley   +1 more source

Home - About - Disclaimer - Privacy