Results 41 to 50 of about 2,146 (182)
WaveNet With Cross-Attention for Audiovisual Speech Recognition
In this paper, the WaveNet with cross-attention is proposed for Audio-Visual Automatic Speech Recognition (AV-ASR) to address multimodal feature fusion and frame alignment problems between two data streams.
Hui Wang, Fei Gao, Yue Zhao, Licheng Wu
doaj +1 more source
Documentation comprises a large proportion of Child and Adolescent Mental Health Service (CAMHS) clinicians/practitioners' work burden and is often completed outside contracted working hours, either requiring overtime pay and/or contributing to burnout.
Judith Dineley +9 more
wiley +1 more source
ASR-Dictation Accuracy for Second Language Speech: Programs, Tasks, and Language Backgrounds
Over the last 20 years, automatic speech recognition (ASR) has improved substantially and spread into people’s everyday lives through hands-free texting and intelligent personal assistants.
Shannon McCrocklin +2 more
doaj +1 more source
UrbanClipAtlas: A Visual Analytics Framework for Event and Scene Retrieval in Urban Videos
Abstract Extracting actionable insights from long‐duration urban videos is often labor‐intensive: analysts must manually sift through raw footage to pinpoint target events or uncover broader behavioral trends. In this work, we present UrbanClipAtlas, a visual analytics system for exploring long urban videos recorded at street intersections ...
Joel Perca +5 more
wiley +1 more source
Arabic Automatic Speech Recognition: A Systematic Literature Review
Automatic Speech Recognition (ASR), also known as Speech-To-Text (STT) or computer speech recognition, has been an active field of research recently. This study aims to chart this field by performing a Systematic Literature Review (SLR) to give insight ...
Amira Dhouib +4 more
doaj +1 more source
SiGnature: Explicit Motion Diffusion for Stylized Semantic Gesture Generation
Abstract While recent advances in co‐speech gesture generation have achieved impressive rhythmic synchronization, synthesizing gestures that are both semantically meaningful and faithful to a speaker's unique non‐verbal style remains an open challenge.
Adi Rosenthal +4 more
wiley +1 more source
There is a large interest in the annotation of speech addressed to infants. Infant-directed speech (IDS) has acoustic properties that might pose a challenge to automatic speech recognition (ASR) tools developed for adult-directed speech (ADS).
Anika van der Klis +3 more
doaj +1 more source
FaceMamba: Geometry‐Aware Mamba for Efficient Speech‐Driven 3D Facial Animation
Abstract Speech‐driven 3D facial animation plays a pivotal role in immersive digital human applications. Recent works have explored Mamba‐based sequence modeling as an efficient alternative to Transformer, but they often suffer from limited cross‐modal alignment and insufficient control over fine‐grained facial deformations.
Yifan Ge +5 more
wiley +1 more source
Unrealistic Feedback in Socially Prescriptive Speech Technologies
ABSTRACT Socially prescriptive speech technologies (SPSTs) are technologies that claim to provide feedback to speakers about how other humans perceive their speech and communication style, based on algorithms trained on a set of static, prescriptive standards.
Nicole Holliday
wiley +1 more source
Speech corpus for Medina dialect
Automatic Speech Recognition (ASR) has standard rules which must be followed and considered carefully. Some difficulties that lead to less ASR performance is variations in pronunciation and small words misrecognition.
Haneen Bahjat Khalafallah +2 more
doaj +1 more source

