Results 101 to 110 of about 3,633,480 (246)
We propose a method for minimum mean-square error (MMSE) estimation of mel-frequency cepstral features for noise robust automatic speech recognition (ASR).
Tan, Zheng-Hua; id_orcid +3 more
core +1 more source
Existing studies in classification of phonation types in singing use voice source features and Mel-frequency cepstral coefficients (MFCCs) showing poor performance due to high pitch in singing.
semanticscholar +1 more source
ABSTRACT Human newborns are able to discriminate between certain languages but not others. This ability has long been attributed to sensitivity to rhythm—the temporal regularities in speech of different languages. Here, we demonstrate through a series of computational simulations that this discrimination behavior can be achieved using no temporal ...
Ruolan Leslie Famularo +3 more
wiley +1 more source
On the Optimal Selection of Mel‐Frequency Cepstral Coefficients for Voice Deepfake Detection
ABSTRACT The continuous evolution of techniques for generating manipulated audio, known as voice deepfakes, and the widespread availability of tools that produce convincing forgeries have created an urgent need for reliable detection methods.
Sergio A. Falcón-López +3 more
openaire +2 more sources
Inter‐Model Feature Fusion for Robust Low‐Resource Speech Recognition
Our Self‐Supervised Feature Fusion (SSF‐FT) method enhances low‐resource speech recognition by adaptively combining features from self‐supervised models trained with Contrastive, Predictive, and Reconstruction objectives. This attention‐weighted ensemble delivers robust performance, particularly in acoustically challenging conditions, extending current
Ussen Kimanuka +2 more
wiley +1 more source
Feature Extracting in the Presence of Environmental Noise, using Subband Adaptive Filtering [PDF]
In this work, a new feature extracting method in noisy environments is proposed. The approach is based on subband decomposition of speech signals followed by adaptive filtering in the noisiest subbbands of speech.
Samad, Salina Abdul
core
Voice source characteristics in different phonation types vary due to the tension of laryngeal muscles along with the respiratory effort. This study investigates the use of mel-frequency cepstral coefficients (MFCCs) derived from voice source waveforms ...
Kadiri, Sudarsana Reddy +3 more
core +1 more source
Automatic Speaker Recognition Based on Mel-Frequency Cepstral Coefficients and Gaussian Mixture Models [PDF]
This paper investigates the task of SR (Speaker Recognition) for the state-of-the-art techniques. The paper initially presents the technical description of automatic SR, followed by the comparative analysis of a number of methods available for feature ...
Sheeraz Memon +2 more
doaj
Application of Music Data Visualization Technology in Music Appreciation Teaching
The simulation environment is used to simulate real‐world music appreciation scenarios. DL is employed to preprocess music data, extract features, and identify rhythm information, which is then associated with visual design parameters to construct a parametric model.
Xiaowei Chen
wiley +1 more source
On Compensating the Mel-Frequency Cepstral Coefficients
This paper describes a novel noise-robust automatic speech recognition (ASR) front-end that employs a combination of Mel-filterbank output compensation and cumulative distribution mapping of cepstral coefficients with truncated Gaussian distribution ...
For Noisy Speech, Eric H. C. Choi
core

