Results 41 to 50 of about 2,146 (182)

WaveNet With Cross-Attention for Audiovisual Speech Recognition

open access: yesIEEE Access, 2020
In this paper, the WaveNet with cross-attention is proposed for Audio-Visual Automatic Speech Recognition (AV-ASR) to address multimodal feature fusion and frame alignment problems between two data streams.
Hui Wang, Fei Gao, Yue Zhao, Licheng Wu
doaj   +1 more source

Technology Matters: Ambient voice technology in mental health and neurodevelopmental settings – considerations for responsible implementation and evidence gaps

open access: yesChild and Adolescent Mental Health, EarlyView.
Documentation comprises a large proportion of Child and Adolescent Mental Health Service (CAMHS) clinicians/practitioners' work burden and is often completed outside contracted working hours, either requiring overtime pay and/or contributing to burnout.
Judith Dineley   +9 more
wiley   +1 more source

ASR-Dictation Accuracy for Second Language Speech: Programs, Tasks, and Language Backgrounds

open access: yesHuman Behavior and Emerging Technologies
Over the last 20 years, automatic speech recognition (ASR) has improved substantially and spread into people’s everyday lives through hands-free texting and intelligent personal assistants.
Shannon McCrocklin   +2 more
doaj   +1 more source

UrbanClipAtlas: A Visual Analytics Framework for Event and Scene Retrieval in Urban Videos

open access: yesComputer Graphics Forum, EarlyView.
Abstract Extracting actionable insights from long‐duration urban videos is often labor‐intensive: analysts must manually sift through raw footage to pinpoint target events or uncover broader behavioral trends. In this work, we present UrbanClipAtlas, a visual analytics system for exploring long urban videos recorded at street intersections ...
Joel Perca   +5 more
wiley   +1 more source

Arabic Automatic Speech Recognition: A Systematic Literature Review

open access: yesApplied Sciences, 2022
Automatic Speech Recognition (ASR), also known as Speech-To-Text (STT) or computer speech recognition, has been an active field of research recently. This study aims to chart this field by performing a Systematic Literature Review (SLR) to give insight ...
Amira Dhouib   +4 more
doaj   +1 more source

SiGnature: Explicit Motion Diffusion for Stylized Semantic Gesture Generation

open access: yesComputer Graphics Forum, EarlyView.
Abstract While recent advances in co‐speech gesture generation have achieved impressive rhythmic synchronization, synthesizing gestures that are both semantically meaningful and faithful to a speaker's unique non‐verbal style remains an open challenge.
Adi Rosenthal   +4 more
wiley   +1 more source

Using Open-Source Automatic Speech Recognition Tools for the Annotation of Dutch Infant-Directed Speech

open access: yesMultimodal Technologies and Interaction, 2023
There is a large interest in the annotation of speech addressed to infants. Infant-directed speech (IDS) has acoustic properties that might pose a challenge to automatic speech recognition (ASR) tools developed for adult-directed speech (ADS).
Anika van der Klis   +3 more
doaj   +1 more source

FaceMamba: Geometry‐Aware Mamba for Efficient Speech‐Driven 3D Facial Animation

open access: yesComputer Graphics Forum, EarlyView.
Abstract Speech‐driven 3D facial animation plays a pivotal role in immersive digital human applications. Recent works have explored Mamba‐based sequence modeling as an efficient alternative to Transformer, but they often suffer from limited cross‐modal alignment and insufficient control over fine‐grained facial deformations.
Yifan Ge   +5 more
wiley   +1 more source

Unrealistic Feedback in Socially Prescriptive Speech Technologies

open access: yesInternational Journal of Applied Linguistics, EarlyView.
ABSTRACT Socially prescriptive speech technologies (SPSTs) are technologies that claim to provide feedback to speakers about how other humans perceive their speech and communication style, based on algorithms trained on a set of static, prescriptive standards.
Nicole Holliday
wiley   +1 more source

Speech corpus for Medina dialect

open access: yesJournal of King Saud University: Computer and Information Sciences
Automatic Speech Recognition (ASR) has standard rules which must be followed and considered carefully. Some difficulties that lead to less ASR performance is variations in pronunciation and small words misrecognition.
Haneen Bahjat Khalafallah   +2 more
doaj   +1 more source

Home - About - Disclaimer - Privacy