Results 61 to 70 of about 3,813 (182)
Inter‐Model Feature Fusion for Robust Low‐Resource Speech Recognition
Our Self‐Supervised Feature Fusion (SSF‐FT) method enhances low‐resource speech recognition by adaptively combining features from self‐supervised models trained with Contrastive, Predictive, and Reconstruction objectives. This attention‐weighted ensemble delivers robust performance, particularly in acoustically challenging conditions, extending current
Ussen Kimanuka +2 more
wiley +1 more source
This study investigates the classification of dangdut music sub-genres using Mel-Frequency Cepstral Coefficients (MFCC) and machine learning approaches.
I Nyoman Surya Jaya +5 more
doaj +1 more source
SPEAKER IDENTIFICATION SYSTEM USING AUDIO SIGNAL AND DEEP LEARNING METHOD [PDF]
Automatic Speaker Identification (ASI) does not result in high accuracy, so it is essential to develop a highly accurate Speaker Identification (SI) system. Artificial Intelligence has shown remarkable improvement in the development of such systems using
Neelam Nehra +2 more
doaj +1 more source
Overview of the proposed Gate‐Align‐SED, including two stages of training: (1) Mean‐Teacher SSL Training; and (2) Enhancer Model Training. In complex real‐world environments such as disaster monitoring, effective sound event detection (SED) is often hindered by the presence of noise and limited labeled data.
Jieli Chen +4 more
wiley +1 more source
To support the preservation of the Sundanese language, speech recognition systems based on machine learning canbe developed. This study aims to evaluate and compare the classification performance of Support Vector Machine, Random Forest, and K-Nearest ...
Laela Nur Rohmah +2 more
doaj +1 more source
Application of Music Data Visualization Technology in Music Appreciation Teaching
The simulation environment is used to simulate real‐world music appreciation scenarios. DL is employed to preprocess music data, extract features, and identify rhythm information, which is then associated with visual design parameters to construct a parametric model.
Xiaowei Chen
wiley +1 more source
Construction and Application of GAN Enhanced Virtual Interpretation Model for Sports Communication
This model is an end‐to‐end framework. Firstly, the style‐based generation network Style‐based GAN2 (StyleGAN2) is used to generate a highly realistic and adjustable static narrator portrait. Then, Bi‐directional Long Short‐Term Memory (Bi‐LSTM) is used to encode the Mel‐frequency Cepstral Coefficients (MFCCs), phonemes, and prosodic features of the ...
Li Zhang +3 more
wiley +1 more source
Optoelectronic control of redox‐active polyoxometalate clusters in polymer matrices yields hybrid memristors with switchable volatile and non‐volatile modes, enabling reservoir‐type in‐sensor optical preprocessing and stable multilevel synapses for multimodal neuromorphic computing, including noise‐tolerant audiovisual keyword recognition and hardware ...
Xiangyu Ma +13 more
wiley +1 more source
Mandarin speech‐based early detection of SCD: a feature‐fusion residual network method
Abstract INTRODUCTION Alzheimer's disease (AD) poses a global health challenge. Early intervention during the stage of subjective cognitive decline (SCD) – a potential window for delaying disease progression – is crucial. This study aims to assess an exploratory speech‐based model for rapid SCD screening.
Zhou Liu +6 more
wiley +1 more source
Flow chart of the base sound trace extraction of speech. ABSTRACT With the acceleration of internationalization, the deficiency of traditional English classroom in English spoken teaching is becoming more obvious, especially the lack of effectiveness and immediate feedback. Therefore, a spoken English assisted training model is proposed.
Lubing Shang, Lu Ma
wiley +1 more source

