Results 11 to 20 of about 4,664 (261)

Content-Aware Speaker Embeddings for Speaker Diarisation [PDF]

open access: yesICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2021
Recent speaker diarisation systems often convert variable length speech segments into fixed-length vector representations for speaker clustering, which are known as speaker embeddings. In this paper, the content-aware speaker embeddings (CASE) approach is proposed, which extends the input of the speaker classifier to include not only acoustic features ...
Guangzhi Sun   +3 more
openaire   +2 more sources

Probabilistic Embeddings for Speaker Diarization [PDF]

open access: yesThe Speaker and Language Recognition Workshop (Odyssey 2020), 2020
Awarded: Jack Godfrey Best Student Paper Award, at Odyssey 2020: The Speaker and Language Recognition Workshop ...
Silnova, Anna   +4 more
openaire   +2 more sources

Binary speaker embedding [PDF]

open access: yes2016 10th International Symposium on Chinese Spoken Language Processing (ISCSLP), 2016
The popular i-vector model represents speakers as low-dimensional continuous vectors (i-vectors), and hence it is a way of continuous speaker embedding. In this paper, we investigate binary speaker embedding, which transforms i-vectors to binary vectors (codes) by a hash function.
Lantian Li   +4 more
openaire   +2 more sources

Speaker diarization through speaker embeddings [PDF]

open access: yes2015 23rd European Signal Processing Conference (EUSIPCO), 2015
Publication in the conference proceedings of EUSIPCO, Nice, France ...
Mickael Rouvier   +2 more
openaire   +2 more sources

Zero-shot style transfer for gesture animation driven by text and speech using adversarial disentanglement of multimodal style encoding

open access: yesFrontiers in Artificial Intelligence, 2023
Modeling virtual agents with behavior style is one factor for personalizing human-agent interaction. We propose an efficient yet effective machine learning approach to synthesize gestures driven by prosodic features and text in the style of different ...
Mireille Fares   +2 more
doaj   +1 more source

Combination of deep speaker embeddings for diarisation [PDF]

open access: yesNeural Networks, 2021
Significant progress has recently been made in speaker diarisation after the introduction of d-vectors as speaker embeddings extracted from neural network (NN) speaker classifiers for clustering speech segments. To extract better-performing and more robust speaker embeddings, this paper proposes a c-vector method by combining multiple sets of ...
Guangzhi Sun   +2 more
openaire   +3 more sources

Residual Information in Deep Speaker Embedding Architectures

open access: yesMathematics, 2022
Speaker embeddings represent a means to extract representative vectorial representations from a speech signal such that the representation pertains to the speaker identity alone.
Adriana Stan
doaj   +1 more source

Speaker/Style-Dependent Neural Network Speech Synthesis Based on Speaker/Style Embedding [PDF]

open access: yesJournal of Universal Computer Science, 2020
The paper presents a novel architecture and method for training neural networks to produce synthesized speech in a particular voice and speaking style, based on a small quantity of target speaker/style training data. The method is based on neural network
Milan Sečujski   +4 more
doaj   +3 more sources

Cross-Lingual Voice Conversion With Controllable Speaker Individuality Using Variational Autoencoder and Star Generative Adversarial Network

open access: yesIEEE Access, 2021
This paper proposes a non-parallel cross-lingual voice conversion (CLVC) model that can mimic voice while continuously controlling speaker individuality on the basis of the variational autoencoder (VAE) and star generative adversarial network (StarGAN ...
Tuan Vu Ho, Masato Akagi
doaj   +1 more source

U-Vectors: Generating Clusterable Speaker Embedding from Unlabeled Data

open access: yesApplied Sciences, 2021
Speaker recognition deals with recognizing speakers by their speech. Most speaker recognition systems are built upon two stages, the first stage extracts low dimensional correlation embeddings from speech, and the second performs the classification task.
Muhammad Firoz Mridha   +5 more
doaj   +1 more source

Home - About - Disclaimer - Privacy