Results 91 to 100 of about 3,610,992 (210)

Inter‐Model Feature Fusion for Robust Low‐Resource Speech Recognition

open access: yesApplied AI Letters, Volume 7, Issue 2, June 2026.
Our Self‐Supervised Feature Fusion (SSF‐FT) method enhances low‐resource speech recognition by adaptively combining features from self‐supervised models trained with Contrastive, Predictive, and Reconstruction objectives. This attention‐weighted ensemble delivers robust performance, particularly in acoustically challenging conditions, extending current
Ussen Kimanuka   +2 more
wiley   +1 more source

On Compensating the Mel-Frequency Cepstral Coefficients

open access: yes, 2008
This paper describes a novel noise-robust automatic speech recognition (ASR) front-end that employs a combination of Mel-filterbank output compensation and cumulative distribution mapping of cepstral coefficients with truncated Gaussian distribution ...
For Noisy Speech, Eric H. C. Choi
core  

Application of Music Data Visualization Technology in Music Appreciation Teaching

open access: yesEngineering Reports, Volume 8, Issue 6, June 2026.
The simulation environment is used to simulate real‐world music appreciation scenarios. DL is employed to preprocess music data, extract features, and identify rhythm information, which is then associated with visual design parameters to construct a parametric model.
Xiaowei Chen
wiley   +1 more source

A dual-stream CNN-based ConvMixer for durian ripeness classification using magnitude and phase features from knocking sounds

open access: yesResults in Engineering
Knocking-sound analysis provides a non-destructive method for assessing durian ripeness; however, most convolutional neural network methods mainly emphasize magnitude information while disregarding phase information that reflects the overall signal ...
Khomdet Phapatanaburi   +7 more
doaj   +1 more source

Automated classification of vowel category and speaker type in the high-frequency spectrum

open access: yesAudiology Research, 2016
The high-frequency region of vowel signals (above the third formant or F3) has received little research attention. Recent evidence, however, has documented the perceptual utility of high-frequency information in the speech signal above the traditional ...
Jeremy J. Donai   +2 more
doaj   +1 more source

Comparison of cepstral coefficients to other voice evaluation parameters in patients with occupational dysphonia

open access: yesMedycyna Pracy, 2013
Background: Special consideration has recently been given to cepstral analysis with mel-frequency cepstral coefficients (MFCCs). The aim of this study was to assess the applicability of MFCCs in acoustic analysis for diagnosing occupational dysphonia in ...
Ewa Niebudek-Bogusz   +3 more
doaj   +1 more source

Use of Mel Frequency Cepstral Coefficients for Automatic Pathology Detection on Sustained Vowel Phonations: Mathematical and Statistical Justification [PDF]

open access: yes, 2008
This paper presents a justification for the use of MFCC parameters in automatic pathology detection on speech. While such an application has produced good results up to now, only partial explanations to this good performance had been given before.
Sáenz Lechón, Nicolas   +4 more
core  

Cepstral Coefficients Effectiveness for Gunshot Classifying [PDF]

open access: yes
This paper analyses the efficiency of various frequency cepstral coefficients (fCC) in a non-speech application, specifically in classifying acoustic impulse events - gunshots. There are various methods for such event identification available.
Holub J., Svatoš J.
core   +1 more source

Identifikasi Suara Pengontrol Lampu Menggunakan Mel-Frequency Cepstral Coefficients dan Hidden Markov Model

open access: yes, 2017
Identification of voice signals can be used to command a computer system. Identification can be made of voice owner and spoken word. Identification of voice owner is used for security, while spoken word identification is often used to execute command on ...
Munggaran, Angga Kersana   +2 more
core  

Multilingual Speaker Identification by Combining Evidence from LPR and Multitaper MFCC

open access: yesJournal of Intelligent Systems, 2013
In this work, the significance of combining the evidence from multitaper mel-frequency cepstral coefficients (MFCC), linear prediction residual (LPR), and linear prediction residual phase (LPRP) features for multilingual speaker identification with the ...
Nagaraja B.G., Jayanna H.S.
doaj   +1 more source

Home - About - Disclaimer - Privacy