Results 61 to 70 of about 457 (166)

Dynamic harmony: Unveiling therapeutic attunement in emotionally focused couples therapy via machine learning

open access: yesFamily Relations, Volume 74, Issue 3, Page 1323-1340, July 2025.
Abstract Objective The goals of the study were to examine therapists' and clients' emotional states and expressions in an emotionally focused therapy (EFT) couple session, to assess therapeutic attunement between the clients and the therapist, and to explore its alignment with EFT techniques. Background Therapeutic attunement is crucial for fostering a
Gökçenay Başer   +3 more
wiley   +1 more source

Maleo-Short: An "In-the-Wild" Indonesian Dataset for Speaker Diarization

open access: yesJOIN: Jurnal Online Informatika
Speaker diarization (SD), the task of partitioning an audio stream into speaker-homogenous segments, is fundamental for analyzing multi-speaker recordings.
Ardi Mardiana   +3 more
doaj   +1 more source

Silent Suffering: Using Machine Learning to Measure CEO Depression

open access: yesJournal of Accounting Research, Volume 63, Issue 2, Page 689-767, May 2025.
ABSTRACT We introduce a novel measure of CEO depression by applying machine learning models that analyze vocal acoustic features from CEOs' conference call recordings. Our research was preregistered via the Journal of Accounting Research's registration‐based editorial process. In this study, we validate this measure and examine associated factors.
SUNG‐YUAN (MARK) CHENG   +1 more
wiley   +1 more source

ATC-SD Net: Radiotelephone Communications Speaker Diarization Network

open access: yesAerospace
This study addresses the challenges that high-noise environments and complex multi-speaker scenarios present in civil aviation radio communications. A novel radiotelephone communications speaker diffraction network is developed specifically for these ...
Weijun Pan   +3 more
doaj   +1 more source

Obfuscation via pitch‐shifting for balancing privacy and diagnostic utility in voice‐based cognitive assessment

open access: yesAlzheimer's &Dementia, Volume 21, Issue 3, March 2025.
Abstract INTRODUCTION Digital voice analysis is an emerging tool for differentiating cognitive states, but it poses privacy risks as automated systems may inadvertently identify speakers. METHODS We developed a computational framework to evaluate the trade‐off between voice obfuscation and cognitive assessment accuracy, using pitch‐shifting as a ...
Meysam Ahangaran   +5 more
wiley   +1 more source

A Script and Tutorial for Using Rev AI's Automatic Speech Transcription

open access: yesInfant and Child Development, Volume 34, Issue 2, March/April 2025.
ABSTRACT We introduce Speech Transcriber with Rev AI (STR) ‐ a Python script that allows for easy interfacing with the Rev AI speech transcription service. Recent advancements in technology have led to increased accuracy and affordability of automatic transcription services, making them preferable over the laborious and time‐consuming process of manual
Margaret Broeren   +3 more
wiley   +1 more source

A multimodal approach to support teacher, researcher and AI collaboration in STEM+C learning environments

open access: yesBritish Journal of Educational Technology, Volume 56, Issue 2, Page 595-620, March 2025.
Abstract Recent advances in generative artificial intelligence (AI) and multimodal learning analytics (MMLA) have allowed for new and creative ways of leveraging AI to support K12 students' collaborative learning in STEM+C domains. To date, there is little evidence of AI methods supporting students' collaboration in complex, open‐ended environments. AI
Clayton Cohn   +5 more
wiley   +1 more source

Empowering Speaker Segmentation With Self‐Supervised Learning

open access: yesElectronics Letters, Volume 61, Issue 1, January/December 2025.
In this paper, we aim to enhance speaker diarization framework with a novel segmentation model leveraging self‐supervised learning. A lightweight backend for WavLM, comprising a cross‐layer extractor, multi‐head factorized attentive module and a classification head, is proposed to further improve the diarization performance.
Jie Yi, Yunfei Gu, Yinfei Xu
wiley   +1 more source

A Novel Sentence‐Level Visual Speech Recognition System for Vietnamese Language Using ResNet3D and Zipformer

open access: yesModelling and Simulation in Engineering, Volume 2025, Issue 1, 2025.
This paper presents the first sentence‐level visual speech recognition (VSR) system specifically designed for the Vietnamese language. We have developed a unique dataset comprising 115 h of video recordings from over 100 speakers, focusing on single‐speaker scenarios.
Phat Nguyen Huu   +2 more
wiley   +1 more source

Retrieval of TV Talk-Show Speakers by Associating Audio Transcript to Visual Clusters

open access: yesIEEE Access, 2017
Retrieval of TV talk-show speakers based on solely visual face recognition is hard because of the significant visual variation caused by illumination, pose, size, and expression, which can exceed those due to identity.
Yina Han, Shanghuan Song, Weikang Zhao
doaj   +1 more source

Home - About - Disclaimer - Privacy