Results 71 to 80 of about 3,612,587 (208)
The Capacity of Mel Frequency Cepstral Coefficients for Speech Recognition
Speech recognition is of an important contribution in promoting new technologies in human computer interaction. Today, there is a growing need to employ speech technology in daily life and business activities. However, speech recognition is a challenging
Dia AbuZeina, Fawaz S. Al-Anzi
core +1 more source
Inter‐Model Feature Fusion for Robust Low‐Resource Speech Recognition
Our Self‐Supervised Feature Fusion (SSF‐FT) method enhances low‐resource speech recognition by adaptively combining features from self‐supervised models trained with Contrastive, Predictive, and Reconstruction objectives. This attention‐weighted ensemble delivers robust performance, particularly in acoustically challenging conditions, extending current
Ussen Kimanuka +2 more
wiley +1 more source
Application of Music Data Visualization Technology in Music Appreciation Teaching
The simulation environment is used to simulate real‐world music appreciation scenarios. DL is employed to preprocess music data, extract features, and identify rhythm information, which is then associated with visual design parameters to construct a parametric model.
Xiaowei Chen
wiley +1 more source
En las agrupaciones musicales sinfónicas, existen ocasiones donde no se consigue un buen balance auditivo. Esto es generalmente ocasionado por enmascaramientos de unos instrumentos con otros. Asimismo, a la hora de llevar a cabo un acto musical,no se consideran diferentes factores como el nivel sonoro, que por lo general es un inconveniente a nivel ...
openaire +1 more source
Minimum Mean-Squared Error Estimation of Mel-Frequency Cepstral Coefficients Using a Novel Distortion Model [PDF]
In this paper, a new method for statistical estimation of Mel-frequency cepstral coefficients (MFCCs) in noisy speech signals is proposed. Previous research has shown that model-based feature domain enhancement of speech signals for use in robust speech ...
R.J. Povinelli +5 more
core +1 more source
Construction and Application of GAN Enhanced Virtual Interpretation Model for Sports Communication
This model is an end‐to‐end framework. Firstly, the style‐based generation network Style‐based GAN2 (StyleGAN2) is used to generate a highly realistic and adjustable static narrator portrait. Then, Bi‐directional Long Short‐Term Memory (Bi‐LSTM) is used to encode the Mel‐frequency Cepstral Coefficients (MFCCs), phonemes, and prosodic features of the ...
Li Zhang +3 more
wiley +1 more source
Optoelectronic control of redox‐active polyoxometalate clusters in polymer matrices yields hybrid memristors with switchable volatile and non‐volatile modes, enabling reservoir‐type in‐sensor optical preprocessing and stable multilevel synapses for multimodal neuromorphic computing, including noise‐tolerant audiovisual keyword recognition and hardware ...
Xiangyu Ma +13 more
wiley +1 more source
Statistically Significant Duration-Independent-based Noise-Robust Speaker Verification [PDF]
A speaker verification system models individual speakers using different speech features to improve their robustness. However, redundant features degrade the system's performance.
Asmita Nirmal +2 more
doaj +1 more source
Digital processing of speech signal and voice recognition algorithm is very important for fast and accurate automatic voice recognition technology. The voice is a signal of infinite information. A direct analysis and synthesizing the complex voice signal is due to too much information contained in the signal. Therefore the digital signal processes such
Lindasalwa Muda +2 more
openaire +3 more sources
Mandarin speech‐based early detection of SCD: a feature‐fusion residual network method
Abstract INTRODUCTION Alzheimer's disease (AD) poses a global health challenge. Early intervention during the stage of subjective cognitive decline (SCD) – a potential window for delaying disease progression – is crucial. This study aims to assess an exploratory speech‐based model for rapid SCD screening.
Zhou Liu +6 more
wiley +1 more source

