Results 31 to 40 of about 4,664 (261)
DNN Speaker Tracking with Embeddings
In multi-speaker applications is common to have pre-computed models from enrolled speakers. Using these models to identify the instances in which these speakers intervene in a recording is the task of speaker tracking. In this paper, we propose a novel embedding-based speaker tracking method.
Carlos Rodrigo Castillo-Sanchez +2 more
openaire +2 more sources
Embeddings for DNN Speaker Adaptive Training [PDF]
Accepted at ASRU ...
Joanna Rownicka +2 more
openaire +2 more sources
Introducing phonetic information to speaker embedding for speaker verification
Phonetic information is one of the most essential components of a speech signal, playing an important role for many speech processing tasks. However, it is difficult to integrate phonetic information into speaker verification systems since it occurs ...
Yi Liu +3 more
doaj +1 more source
Automatic speaker verification (ASV) is an emerging biometric verification technique with more and more applications. However, both verification accuracy and anti-spoofing should be considered carefully before putting ASV into practice, where anti ...
Jiakang Li +3 more
doaj +1 more source
Xi-Vector Embedding for Speaker Recognition [PDF]
We present a Bayesian formulation for deep speaker embedding, wherein the xi-vector is the Bayesian counterpart of the x-vector, taking into account the uncertainty estimate. On the technology front, we offer a simple and straightforward extension to the now widely used x-vector.
Kong Aik Lee +2 more
openaire +2 more sources
Global–Local Self-Attention Based Transformer for Speaker Verification
Transformer models are now widely used for speech processing tasks due to their powerful sequence modeling capabilities. Previous work determined an efficient way to model speaker embeddings using the Transformer model by combining transformers with ...
Fei Xie, Dalong Zhang, Chengming Liu
doaj +1 more source
On deep speaker embeddings for text-independent speaker recognition [PDF]
We investigate deep neural network performance in the textindependent speaker recognition task. We demonstrate that using angular softmax activation at the last classification layer of a classification neural network instead of a simple softmax activation allows to train a more generalized discriminative speaker embedding extractor.
Sergey Novoselov +4 more
openaire +2 more sources
Sequence-to-Sequence Emotional Voice Conversion With Strength Control
This paper proposes an improved emotional voice conversion (EVC) method with emotional strength and duration controllability. EVC methods without duration mapping generate emotional speech with identical duration to that of the neutral input speech.
Heejin Choi, Minsoo Hahn
doaj +1 more source
Development of Supervised Speaker Diarization System Based on the PyAnnote Audio Processing Library
Diarization is an important task when work with audiodata is executed, as it provides a solution to the problem related to the need of dividing one analyzed call recording into several speech recordings, each of which belongs to one speaker.
Volodymyr Khoma +3 more
doaj +1 more source
Accepted to ICASSP ...
Shota Horiguchi +6 more
openaire +2 more sources

