Results 21 to 30 of about 457 (166)
Speaker diarization is the task of automatically identifying speaker identities and detecting their speaking times in an audio recording. Several algorithms have shown improvements in the performance of this task during the past years. However, it still
Alejandro Chacón-Vargas +2 more
doaj +1 more source
Speaker diarization refers to methods for identifying speakers from audio recordings. An important application comes from the need to assess student interactions in collaborative learning environments.
Antonio Gomez +2 more
doaj +1 more source
ECAPA-TDNN Embeddings for Speaker Diarization [PDF]
Learning robust speaker embeddings is a crucial step in speaker diarization. Deep neural networks can accurately capture speaker discriminative characteristics and popular deep embeddings such as x-vectors are nowadays a fundamental component of modern diarization systems.
Nauman Dawalatabad +5 more
openaire +3 more sources
The Domain Mismatch Problem in the Broadcast Speaker Attribution Task
The demand of high-quality metadata for the available multimedia content requires the development of new techniques able to correctly identify more and more information, including the speaker information.
Ignacio Viñals +3 more
doaj +1 more source
Combining speaker identification and BIC for speaker diarization [PDF]
This paper describes recent advances in speaker diarization by incorporating a speaker identification step. This system builds upon the LIMSI baseline data partitioner used in the broadcast news transcription system. This partitioner provides a high cluster purity but has a tendency to split the data from a speaker into several clusters, when there is ...
Zhu, Xuan +3 more
openaire +2 more sources
Speaker Diarization with Lexical Information [PDF]
This work presents a novel approach for speaker diarization to leverage lexical information provided by automatic speech recognition. We propose a speaker diarization system that can incorporate word-level speaker turn probabilities with speaker embeddings into a speaker clustering process to improve the overall diarization accuracy.
Tae Jin Park +6 more
openaire +2 more sources
Automatic segmentation and classification of audio streams is a challenging problem, with many applications, such as indexing multi – media digital libraries, information retrieving, and the building of speech corpus or spoken corpus) for particular ...
Roberto Sánchez Cárdenas +1 more
doaj +1 more source
Self-supervised Speaker Diarization
Submitted to Interspeech ...
Yehoshua Dissen +2 more
openaire +2 more sources
Fully Supervised Speaker Diarization [PDF]
In this paper, we propose a fully supervised speaker diarization approach, named unbounded interleaved-state recurrent neural networks (UIS-RNN). Given extracted speaker-discriminative embeddings (a.k.a. d-vectors) from input utterances, each individual speaker is modeled by a parameter-sharing RNN, while the RNN states for different speakers ...
Aonan Zhang +4 more
openaire +2 more sources
Privacy-Preserving Automatic Speaker Diarization
Automatic Speaker Diarization (ASD) is an enabling technology with numerous applications, which deals with recordings of multiple speakers, raising special concerns in terms of privacy. In fact, in remote settings, where recordings are shared with a server, clients relinquish not only the privacy of their conversation, but also of all the information ...
Francisco Teixeira +3 more
openaire +2 more sources

