Results 61 to 70 of about 423 (132)

Deep learning based 3D residual convolutional and Multi-Head Attention (3D-RMA) for lip-reading

open access: yesResults in Control and Optimization
Lip reading, an essential yet intricate facet of communication, has seen notable progress through the application of advanced deep learning techniques. This research introduces a deep learning-based lip-reading model that integrates Conv3D layers, Multi ...
Archana Chaudhari   +6 more
doaj   +1 more source

ResNet18-BLSTM-CTC-Based OCR Model for Extracting Technical Parameter Data from Digital Television Analyzer Screenshots

open access: yesTeknika
Monitoring digital television broadcasting stations requires accurate extraction of technical measurement parameters from screenshot images generated by television analyzer devices.
Henrian Robby Fakhriannur   +1 more
doaj   +1 more source

Text-Independent Phone-to-Audio Alignment Leveraging SSL (TIPAA-SSL) Pre-Trained Model Latent Representation and Knowledge Transfer

open access: yesAcoustics
In this paper, we present a novel approach for text-independent phone-to-audio alignment based on phoneme recognition, representation learning and knowledge transfer.
Noé Tits   +2 more
doaj   +1 more source

A Statistically Validated and Decoding-Aware CNN–Transformer–CTC Framework for Multi-Font Printed Arabic Word Recognition

open access: yesApplied Sciences
Printed Arabic Optical Character Recognition (OCR) remains challenging due to complex glyph morphology, typographic variability, and sensitivity to Unicode-preserved evaluation protocols. This work introduces a methodology that explicitly treats decoding
Abderrahime Tabzaoui, Loqman Chakir
doaj   +1 more source

Voice Conversion Combining Vector Quantization and CTC Introducing Pre-Trained Representation [PDF]

open access: yesJisuanji gongcheng
Pre-trained models have achieved significant breakthroughs in nonparallel corpus Voice Conversion (VC) via Self-Supervised Pre-trained Representation (SSPR).
WANG Lin, HUANG Hao
doaj   +1 more source

Lightweight End-to-End Diacritical Arabic Speech Recognition Using CTC-Transformer with Relative Positional Encoding

open access: yesMathematics
Arabic automatic speech recognition (ASR) faces distinct challenges due to its complex morphology, dialectal variations, and the presence of diacritical marks that strongly influence pronunciation and meaning. This study introduces a lightweight approach
Haifa Alaqel, Khalil El Hindi
doaj   +1 more source

Low-resource Lingao dialect speech recognition method based on multi-feature transfer learning

open access: yesTongxin xuebao
To address the challenges of data scarcity and high character error rate in low-resource Lingao dialect automatic speech recognition, an end-to-end speech recognition method based on multi-feature transfer learning was proposed.
WANG Zhong   +5 more
doaj  

HCVEA: Personalized Residential Layout Generation via an Improved Conditional Variational Autoencoder with Reinforcement Learning

open access: yesDesigns
With the growing demand for personalized design, residential layout generation has become a key research area in architecture and artificial intelligence.
Hongting He, Zunyue Liu, Fei Xiao
doaj   +1 more source

A textual-guided compact visual attention features for continuous sign language recognition

open access: yesDiscover Artificial Intelligence
Continuous Sign Language Recognition (CSLR) datasets contain RGB videos annotated with similar sentence glosses representing sign sequences. However, the performance of CSLR systems is often hindered by issues such as redundant video frames, sparse text ...
V. Prathyusha, P. V. V. Kishore
doaj   +1 more source

Home - About - Disclaimer - Privacy