Results 81 to 90 of about 457 (166)

Metrics for Polyphonic Sound Event Detection

open access: yesApplied Sciences, 2016
This paper presents and discusses various metrics proposed for evaluation of polyphonic sound event detection systems used in realistic situations where there are typically multiple sound sources active simultaneously.
Annamaria Mesaros   +2 more
doaj   +1 more source

diaLogic: A Multi-Modal Framework for Automated Team Behavior Modeling Based on Speech Acquisition

open access: yesMultimodal Technologies and Interaction
This paper presents diaLogic, a humans-in-the-loop system for modeling the behavior of teams during collective problem solving. Team behavior is modeled using multi-modal data about cognition, social interactions, and emotions acquired from speech inputs.
Ryan Duke, Alex Doboli
doaj   +1 more source

Unsupervised adaptation of PLDA models for broadcast diarization

open access: yesEURASIP Journal on Audio, Speech, and Music Processing, 2019
We present a novel model adaptation approach to deal with data variability for speaker diarization in a broadcast environment. Expensive human annotated data can be used to mitigate the domain mismatch by means of supervised model adaptation approaches ...
Ignacio Viñals   +4 more
doaj   +1 more source

Whisper Automatic Speech Recognition and GPT Large Language Models as Best Practice for Assessing Communication Progress in Autism Spectrum Disorder

open access: yesJurnal Teknologi Pendidikan
Autism Spectrum Disorder (ASD) is a developmental disorder that affects communication, social interaction, and behavior. Communication assessments for children with ASD are often conducted manually, making the process time-consuming, which can lead to ...
Naela Fauzul Muna   +1 more
doaj   +1 more source

Exploring AI Techniques for Generalizable Teaching Practice Identification

open access: yesIEEE Access
Using automated models to analyze classroom discourse is a valuable tool for educators to improve their teaching methods. In this paper, we focus on exploring alternatives to ensure the generalizability of models for identifying teaching practices across
Federico Pardo Garcia   +2 more
doaj   +1 more source

Relative Applicability of Diverse Automatic Speech Recognition Platforms for Transcription of Psychiatric Treatment Sessions

open access: yesIEEE Access
Service delivery in mental healthcare involves documentation of sensitive patient-clinician conversations that require serious caution. Conventionally, clinicians take handwritten notes, which causes low readability and lack of database which hinders ...
Rana Zeeshan   +2 more
doaj   +1 more source

Active Speaker Detection Using Audio, Visual, and Depth Modalities: A Survey

open access: yesIEEE Access
The rapid progress of multimodal signal processing in recent years has cleared the way for novel applications in human-computer interaction, surveillance, and telecommunication.
Siti Nur Aisyah Mohd Robi   +4 more
doaj   +1 more source

Echo: A crowd-sourced Romanian speech dataset.

open access: yesInteraction Design and Architecture(s)
Romanian is the seventh most popular European language, with around 30 million speakers worldwide. Despite its popularity, the available speech resources are limited.
Remus-Dan Ungureanu, Mihai Dascalu
doaj   +1 more source

USED: Universal Speaker Extraction and Diarization

open access: yesIEEE Transactions on Audio, Speech and Language Processing
Speaker extraction and diarization are two enabling techniques for real-world speech applications. Speaker extraction aims to extract a target speaker's voice from a speech mixture, while speaker diarization demarcates speech segments by speaker, annotating `who spoke when'. Previous studies have typically treated the two tasks independently.
Junyi Ao   +8 more
openaire   +2 more sources

Home - About - Disclaimer - Privacy