Results 21 to 30 of about 4,664 (261)
Self-Supervised Speaker Embeddings [PDF]
Preprint. Submitted to Interspeech 2019.
Themos Stafylakis +4 more
openaire +2 more sources
Ordered and Binary Speaker Embedding
to be published in INTERSPEECH ...
Jiaying Wang +4 more
openaire +2 more sources
Compact Speaker Embedding: lrx-Vector [PDF]
Accepted to INTERSPEECH ...
Munir Georges +2 more
openaire +3 more sources
ECAPA-TDNN Embeddings for Speaker Diarization [PDF]
Learning robust speaker embeddings is a crucial step in speaker diarization. Deep neural networks can accurately capture speaker discriminative characteristics and popular deep embeddings such as x-vectors are nowadays a fundamental component of modern diarization systems.
Nauman Dawalatabad +5 more
openaire +3 more sources
In this paper, we propose self-supervised speaker representation learning strategies, which comprise of a bootstrap equilibrium speaker representation learning in the front-end and an uncertainty-aware probabilistic speaker embedding training in the back-
Sung Hwan Mun +4 more
doaj +1 more source
Streaming End-to-End Target-Speaker Automatic Speech Recognition and Activity Detection
Automatic speech recognition of a target speaker in the presence of interfering speakers remains a challenging issue. One approach to tackle this problem is target-speaker speech recognition, which conditions the recognition process on an embedding that ...
Takafumi Moriya +4 more
doaj +1 more source
Emotion-Aware Speaker Identification With Transfer Learning
Speech is a natural communication method used by humans. Speaker identification (SI) technology based on human speech has been used as an entry point for many human–computer-interaction applications.
Kyoungju Noh, Hyuntae Jeong
doaj +1 more source
Magnitude-Aware Probabilistic Speaker Embeddings
Accepted to Odyssey 2022: The Speaker and Language Recognition Workshop, camera-ready ...
Nikita Kuzmin +2 more
openaire +2 more sources
A Survey on Text-Dependent and Text-Independent Speaker Verification
Speaker verification (SV) aims to detect an individual’s identity from his/her voice. SV has been successfully applied in various areas such as access control, remote service customization, financial transactions, etc.
Youzhi Tu, Weiwei Lin, Man-Wai Mak
doaj +1 more source
Design of a Multi-Condition Emotional Speech Synthesizer
Recently, researchers have developed text-to-speech models based on deep learning, which have produced results superior to those of previous approaches.
Sung-Woo Byun, Seok-Pil Lee
doaj +1 more source

