Results 51 to 60 of about 457 (166)
Latent class model with application to speaker diarization
In this paper, we apply a latent class model (LCM) to the task of speaker diarization. LCM is similar to Patrick Kenny’s variational Bayes (VB) method in that it uses soft information and avoids premature hard decisions in its iterations.
Liang He +5 more
doaj +1 more source
Robust speaker diarization for meetings
Aquesta tesi doctoral mostra la recerca feta en l'àrea de la diarització de locutor per a sales de reunions. En la present s'estudien els algorismes i la implementació d'un sistema en diferit de segmentació i aglomerat de locutor per a grabacions de reunions a on normalment es té accés a més d'un micròfon per al processat.
openaire +3 more sources
Speaker embedding loss for end-to-end speaker diarization without external embedding networks
This paper introduces a novel speaker embedding loss function designed to improve the performance of end-to-end neural diarization (EEND) systems by enhancing speaker discrimination.
Jaehee Jung, Wooil Kim
doaj +1 more source
Throughout the history of automated personal identification, now called “biometrics” in some communities, there has been controversy over its implications for personal privacy and human dignity. This controversy has been deepened by equivocation regarding the philosophical concept of personal identity, the social concept of recognition of persons, and ...
Emilio Mordini +2 more
wiley +1 more source
Multimodal Diarization Systems by Training Enrollment Models as Identity Representations
This paper describes a post-evaluation analysis of the system developed by ViVoLAB research group for the IberSPEECH-RTVE 2020 Multimodal Diarization (MD) Challenge.
Victoria Mingote +5 more
doaj +1 more source
Speaker diarization is the task of determining "who spoke when?" in an audio or video recording that contains an unknown amount of speech and an unknown number of speakers. It is a challenging task due to the variability of human speech, the presence of overlapping speech, and the lack of prior information about the speakers.
openaire +1 more source
Improving Speaker Diarization for Overlapped Speech with Texture-Aware Feature Fusion
Speaker diarization (SD), which aims to address the “who spoke when” problem, is a key technology in speech processing. Although end-to-end neural speaker diarization methods have simplified the traditional multi-stage pipeline, their capability to ...
Chengli Sun, Miao Sun, Wenrui Wei
doaj +1 more source
Multistage speaker diarization of broadcast news [PDF]
This paper describes recent advances in speaker diarization with a multistage segmentation and clustering system, which incorporates a speaker identification step. This system builds upon the baseline audio partitioner used in the LIMSI broadcast news transcription system.
Barras, Claude +3 more
openaire +2 more sources
Neural Speaker Diarization with Speaker-Wise Chain Rule
Speaker diarization is an essential step for processing multi-speaker audio. Although an end-to-end neural diarization (EEND) method achieved state-of-the-art performance, it is limited to a fixed number of speakers. In this paper, we solve this fixed number of speaker issue by a novel speaker-wise conditional inference method based on the ...
Yusuke Fujita +5 more
openaire +2 more sources
Preschool Language Environments and Children's School Readiness Skills
ABSTRACT Early language environments are considered to support children's language development; however, it is unclear to what extent early language environments relate to skills other than language abilities. We examined (1) whether the preschool language environment (measured as adult words heard and conversational turns) is associated with children ...
Kirsten L. Anderson +5 more
wiley +1 more source

