Results 131 to 140 of about 48,773 (160)
Some of the next articles are maybe not open access.

End-to-End Speech Recognition in Russian

2018
End-to-end speech recognition systems incorporating deep neural networks (DNNs) have achieved good results. We propose applying CTC (Connectionist Temporal Classification) models and attention-based encoder-decoder in automatic recognition of the Russian continuous speech.
Nikita Markovnikov   +2 more
openaire   +1 more source

Parameter Uncertainty for End-to-end Speech Recognition

ICASSP 2019 - 2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2019
Recent work on neural networks with probabilistic parameters has shown that parameter uncertainty improves network regularization. Parameter-specific signal-to-noise ratio (SNR) levels derived from parameter distributions were further found to have high correlations with task importance.
Stefan Braun 0005, Shih-Chii Liu
openaire   +2 more sources

Triggered Attention for End-to-end Speech Recognition

ICASSP 2019 - 2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2019
A new system architecture for end-to-end automatic speech recognition (ASR) is proposed that combines the alignment capabilities of the connectionist temporal classification (CTC) approach and the modeling strength of the attention mechanism. The proposed system architecture, named triggered attention (TA), uses a CTC-based classifier to control the ...
Niko Moritz   +2 more
openaire   +2 more sources

End-to-End Speech Recognition

2019
In Chap. 8, we aimed to create an ASR system by dividing the fundamental equation $$\displaystyle W^* = \operatorname *{argmax}_{W \in V^*} P(W|X) $$ into an acoustic model, lexicon model, and language model by using Bayes’ theorem. This approach relies heavily on the use of the conditional independence assumption and separate optimization ...
Uday Kamath, John Liu, James Whitaker
openaire   +1 more source

End-to-end Korean Digits Speech Recognition

2019 International Conference on Information and Communication Technology Convergence (ICTC), 2019
The traditional speech recognition model consisting of an acoustic model and a language model is mainly used. Recently, an end-to-end speech recognition model consisting of a single integrated neural network model is being studied. This model has the advantage that it does not require a lot of training and it is easy to understand the structure of the ...
Jong-Hyuk Roh   +3 more
openaire   +2 more sources

End-to-End Speech Recognition in Agglutinative Languages

2020
This paper considers end-to-end speech recognition systems based on deep neural networks (DNN). The studies used different types of neural networks, CTC model and attention-based encoder-decoder models. As a result of the study, it was proved that the CTC model works without language models directly for agglutinative languages, but the best is ResNet ...
Orken Mamyrbayev   +4 more
openaire   +2 more sources

An End-to-End Model for Vietnamese Speech Recognition

2019 IEEE-RIVF International Conference on Computing and Communication Technologies (RIVF), 2019
This paper presents an approach of End-to-End model based on Long Short-Term Memory (LSTM) and Time Delay Deep Neural Network (TDNN) models for Vietnamese speech recognition. Two Vietnamese End-to-End architectures using Connectionist Temporal Classification (CTC) as the loss function are proposed.
openaire   +1 more source

End-to-End Multi-Speaker Speech Recognition

2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2018
Current advances in deep learning have resulted in a convergence of methods across a wide range of tasks, opening the door for tighter integration of modules that were previously developed and optimized in isolation. Recent ground-breaking works have produced end-to-end deep network methods for both speech separation and end-to-end automatic speech ...
Shane Settle   +4 more
openaire   +2 more sources

End-to-End Speech Recognition Models

2016
For the past few decades, the bane of Automatic Speech Recognition (ASR) systems have been phonemes and Hidden Markov Models (HMMs). HMMs assume conditional indepen-dence between observations, and the reliance on explicit phonetic representations requires expensive handcrafted pronunciation dictionaries.
openaire   +1 more source

Home - About - Disclaimer - Privacy