Results 61 to 70 of about 47,965 (148)
NIESR: Nuisance Invariant End-to-End Speech Recognition [PDF]
Deep neural network models for speech recognition have achieved great success recently, but they can learn incorrect associations between the target and nuisance factors of speech (e.g., speaker identities, background noise, etc.), which can lead to overfitting.
I-Hung Hsu +2 more
openaire +2 more sources
A Spelling Correction Model for End-to-end Speech Recognition [PDF]
Attention-based sequence-to-sequence models for speech recognition jointly train an acoustic model, language model (LM), and alignment mechanism using a single neural network and require only parallel audio-text pairs. Thus, the language model component of the end-to-end model is only trained on transcribed audio-text pairs, which leads to performance ...
Jinxi Guo, Tara N. Sainath, Ron J. Weiss
openaire +2 more sources
A Novel Syllable-Level Signal Encryption for Robust Secure Speech Communication System
Speech communication is vital for conveying information and emotions, yet it faces significant security threats. This research presents a novel signal encryption system that operates at the syllable level, preserving the natural flow of speech while ...
Albertus Anugerah Pekerti +3 more
doaj +1 more source
End-to-End Speech Recognition from the Raw Waveform [PDF]
State-of-the-art speech recognition systems rely on fixed, hand-crafted features such as mel-filterbanks to preprocess the waveform before the training pipeline. In this paper, we study end-to-end systems trained directly from the raw waveform, building on two alternatives for trainable replacements of mel-filterbanks that use a convolutional ...
Zeghidour, Neil +4 more
openaire +2 more sources
Deep Speaker Recognition: Process, Progress, and Challenges
Speaker recognition is related to human biometrics dealing with the identification of speakers from their speech. Speaker recognition is an active research area and being widely investigated using artificially intelligent mechanisms.
Abu Quwsar Ohi +3 more
doaj +1 more source
Cycle-consistency Training for End-to-end Speech Recognition [PDF]
This paper presents a method to train end-to-end automatic speech recognition (ASR) models using unpaired data. Although the end-to-end approach can eliminate the need for expert knowledge such as pronunciation dictionaries to build ASR systems, it still requires a large amount of paired data, i.e., speech utterances and their transcriptions.
Takaaki Hori +5 more
openaire +2 more sources
Adversarial joint training with self-attention mechanism for robust end-to-end speech recognition
Lately, the self-attention mechanism has marked a new milestone in the field of automatic speech recognition (ASR). Nevertheless, its performance is susceptible to environmental intrusions as the system predicts the next output symbol depending on the ...
Lujun Li +5 more
doaj +1 more source
At present, automatic speech recognition has become an important bridge for human-computer interaction and is widely applied in multiple fields. The Portuguese speech recognition task is gradually receiving attention due to its unique language stance ...
Yan Li +4 more
doaj +1 more source
End-To-End deep neural models for Automatic Speech Recognition for Polish Language [PDF]
This article concerns research on deep learning models (DNN) used for automatic speech recognition (ASR). In such systems, recognition is based on Mel Frequency Cepstral Coefficients (MFCC) acoustic features and spectrograms.
Karolina Pondel-Sycz +2 more
doaj +1 more source
End-to-End Neural Segmental Models for Speech Recognition [PDF]
Segmental models are an alternative to frame-based models for sequence prediction, where hypothesized path weights are based on entire segment scores rather than a single frame at a time. Neural segmental models are segmental models that use neural network-based weight functions.
Hao Tang 0002 +7 more
openaire +3 more sources

