Spoken term detection with Connectionist Temporal Classification: A novel hybrid CTC-DBN decoder [PDF]
This paper proposes a novel system for robust keyword detection in continuous speech. Our decoder is composed of a bidirectional Long Short-Term Memory recurrent neural network using a Connectionist Temporal Classification (CTC) output layer, and a Dynamic Bayesian Network (DBN).
Martin Wöllmer +3 more
openaire +2 more sources
Inter‐Model Feature Fusion for Robust Low‐Resource Speech Recognition
Our Self‐Supervised Feature Fusion (SSF‐FT) method enhances low‐resource speech recognition by adaptively combining features from self‐supervised models trained with Contrastive, Predictive, and Reconstruction objectives. This attention‐weighted ensemble delivers robust performance, particularly in acoustically challenging conditions, extending current
Ussen Kimanuka +2 more
wiley +1 more source
A deep neural network-based automatic mispronunciation detection in Bengali accented English speech
Learning a second language, especially English, became necessary as globalisation started. One crucial component of language learning resources is computer-assisted pronunciation training, or CAPT.
Puja Bharati +5 more
doaj +1 more source
A Linear Memory CTC-Based Algorithm for Text-to-Voice Alignment of Very Long Audio Recordings
Synchronisation of a voice recording with the corresponding text is a common task in speech and music processing, and is used in many practical applications (automatic subtitling, audio indexing, etc.).
Guillaume Doras +2 more
doaj +1 more source
Online Sequence Training of Recurrent Neural Networks with Connectionist Temporal Classification
Final version: Kyuyeon Hwang and Wonyong Sung, "Sequence to Sequence Training of CTC-RNNs with Partial Windowing," Proceedings of The 33rd International Conference on Machine Learning, pp. 2178-2187, 2016.
Kyuyeon Hwang, Wonyong Sung
openaire +2 more sources
A Survey for Deep Reinforcement Learning Based Network Intrusion Detection
This paper surveys deep reinforcement learning (DRL) for network intrusion detection, evaluating model efficiency, minority attack detection, and dataset imbalance. Findings show DRL achieves state‐of‐the‐art results on public datasets, sometimes surpassing traditional deep learning.
Wanrong Yang +3 more
wiley +1 more source
Accented Speech Recognition Based on End-to-End Domain Adversarial Training of Neural Networks
The performance of automatic speech recognition (ASR) may be degraded when accented speech is recognized because the speech has some linguistic differences from standard speech.
Hyeong-Ju Na, Jeong-Sik Park
doaj +1 more source
Probabilistic asr feature extraction applying context-sensitive connectionist temporal classification networks [PDF]
This paper proposes a novel automatic speech recognition (ASR) front-end that unites the principles of bidirectional Long Short-Term Memory (BLSTM), Connectionist Temporal Classification (CTC), and Bottleneck (BN) feature generation. BLSTM networks are known to produce better probabilistic ASR features than conventional multilayer perceptrons since ...
Martin Wöllmer +2 more
openaire +1 more source
This review explores how quantum activation functions can contribute to the evolution of neural networks toward quantum computing. The results show that classical‐quantum hybrid architectures are being tested in some practical applications, while fully quantum models are still in the development phase. These functions represent an important step toward
Petterson Pina dos Santos +2 more
wiley +1 more source
We introduce BERT-NAR-BERT (BnB) – a pre-trained non-autoregressive sequence-to-sequence model, which employs BERT as the backbone for the encoder and decoder for natural language understanding and generation tasks.
Mohammad Golam Sohrab +3 more
doaj +1 more source

