Results 21 to 30 of about 981,169 (268)

A Two-Stage Approach to Note-Level Transcription of a Specific Piano

open access: yesApplied Sciences, 2017
This paper presents a two-stage transcription framework for a specific piano, which combines deep learning and spectrogram factorization techniques. In the first stage, two convolutional neural networks (CNNs) are adopted to recognize the notes of the ...
Qi Wang, Ruohua Zhou, Yonghong Yan
doaj   +1 more source

Dysarthric speech classification from coded telephone speech using glottal features

open access: yes, 2021
This paper proposes a new dysarthric speech classification method from coded telephone speech using glottal features. The proposed method utilizes glottal features, which are efficiently estimated from coded telephone speech using a recently proposed ...
Alku, Paavo   +1 more
core   +1 more source

Chinese Dialogue Intention Classification Based on Multi-Model Ensemble

open access: yesIEEE Access, 2019
In dialogue systems, understanding the user utterances is crucial for providing appropriate responses. A traditional dialogue act classification (DA) task is to classify each user reply into “ACCEPT, REJECT, PROPOSE, and others”.
Manshu Tu, Bing Wang, Xuemin Zhao
doaj   +1 more source

End-to-End Optimized Multi-Stage Vector Quantization of Spectral Envelopes for Speech and Audio Coding

open access: yes, 2021
Spectral envelope modeling is an instrumental part of speech and audio codecs, which can be used to enable efficient entropy coding of spectral components.
Bäckström, Tom   +3 more
core   +1 more source

Unsupervised cross-lingual speaker adaptation for HMM-based speech synthesis using two-pass decision tree construction [PDF]

open access: yes, 2010
This paper demonstrates how unsupervised cross-lingual adaptation of HMM-based speech synthesis models may be performed without explicit knowledge of the adaptation data language.
Teemu Hirsimaki   +5 more
core   +2 more sources

Target Speaker Localization Based on the Complex Watson Mixture Model and Time-Frequency Selection Neural Network

open access: yesApplied Sciences, 2018
Common sound source localization algorithms focus on localizing all the active sources in the environment. While the source identities are generally unknown, retrieving the location of a speaker of interest requires extra effort. This paper addresses the
Ziteng Wang, Junfeng Li, Yonghong Yan
doaj   +1 more source

Relevancy between Objects Based on Common Sense for Semantic Segmentation

open access: yesApplied Sciences, 2022
Research on image classification sparked the latest deep-learning boom. Many downstream tasks, including semantic segmentation, benefit from it. The state-of-the-art semantic segmentation models are all based on deep learning, and they sometimes make ...
Jun Zhou, Xing Bai, Qin Zhang
doaj   +1 more source

Lombard speech synthesis using transfer learning in a Tacotron text-to-speech system

open access: yes, 2019
Currently, there is increasing interest to use sequence-to-sequence models in text-to-speech (TTS) synthesis with attention like that in Tacotron models.
Bajibabu Bollepalli   +5 more
core   +1 more source

The Use of Audio Fingerprints for Authentication of Speakers on Speech Operated Interfaces

open access: yes, 2021
In a multi-speaker and multi-device environment, we need acoustic fingerprint information for authentication between devices. Thus, in these kinds of environments, it is crucial to continuously check the authenticity of speakers and devices within a ...
Zewoudie, Abraham   +5 more
core   +1 more source

Polyphonic Piano Transcription with a Note-Based Music Language Model

open access: yesApplied Sciences, 2018
This paper proposes a note-based music language model (MLM) for improving note-level polyphonic piano transcription. The MLM is based on the recurrent structure, which could model the temporal correlations between notes in music sequences. To combine the
Qi Wang, Ruohua Zhou, Yonghong Yan
doaj   +1 more source

Home - About - Disclaimer - Privacy