Results 21 to 30 of about 981,169 (268)
A Two-Stage Approach to Note-Level Transcription of a Specific Piano
This paper presents a two-stage transcription framework for a specific piano, which combines deep learning and spectrogram factorization techniques. In the first stage, two convolutional neural networks (CNNs) are adopted to recognize the notes of the ...
Qi Wang, Ruohua Zhou, Yonghong Yan
doaj +1 more source
Dysarthric speech classification from coded telephone speech using glottal features
This paper proposes a new dysarthric speech classification method from coded telephone speech using glottal features. The proposed method utilizes glottal features, which are efficiently estimated from coded telephone speech using a recently proposed ...
Alku, Paavo +1 more
core +1 more source
Chinese Dialogue Intention Classification Based on Multi-Model Ensemble
In dialogue systems, understanding the user utterances is crucial for providing appropriate responses. A traditional dialogue act classification (DA) task is to classify each user reply into “ACCEPT, REJECT, PROPOSE, and others”.
Manshu Tu, Bing Wang, Xuemin Zhao
doaj +1 more source
Spectral envelope modeling is an instrumental part of speech and audio codecs, which can be used to enable efficient entropy coding of spectral components.
Bäckström, Tom +3 more
core +1 more source
Unsupervised cross-lingual speaker adaptation for HMM-based speech synthesis using two-pass decision tree construction [PDF]
This paper demonstrates how unsupervised cross-lingual adaptation of HMM-based speech synthesis models may be performed without explicit knowledge of the adaptation data language.
Teemu Hirsimaki +5 more
core +2 more sources
Common sound source localization algorithms focus on localizing all the active sources in the environment. While the source identities are generally unknown, retrieving the location of a speaker of interest requires extra effort. This paper addresses the
Ziteng Wang, Junfeng Li, Yonghong Yan
doaj +1 more source
Relevancy between Objects Based on Common Sense for Semantic Segmentation
Research on image classification sparked the latest deep-learning boom. Many downstream tasks, including semantic segmentation, benefit from it. The state-of-the-art semantic segmentation models are all based on deep learning, and they sometimes make ...
Jun Zhou, Xing Bai, Qin Zhang
doaj +1 more source
Lombard speech synthesis using transfer learning in a Tacotron text-to-speech system
Currently, there is increasing interest to use sequence-to-sequence models in text-to-speech (TTS) synthesis with attention like that in Tacotron models.
Bajibabu Bollepalli +5 more
core +1 more source
The Use of Audio Fingerprints for Authentication of Speakers on Speech Operated Interfaces
In a multi-speaker and multi-device environment, we need acoustic fingerprint information for authentication between devices. Thus, in these kinds of environments, it is crucial to continuously check the authenticity of speakers and devices within a ...
Zewoudie, Abraham +5 more
core +1 more source
Polyphonic Piano Transcription with a Note-Based Music Language Model
This paper proposes a note-based music language model (MLM) for improving note-level polyphonic piano transcription. The MLM is based on the recurrent structure, which could model the temporal correlations between notes in music sequences. To combine the
Qi Wang, Ruohua Zhou, Yonghong Yan
doaj +1 more source

