Mandarin recognition and improvement based on CTC criterion [PDF]
The cross-entropy criterion of mainstream neural network training classifies and optimizes each frame of acoustic data, while the continuous speech recognition uses the sequence-level transcription accuracy as the performance measurement.For this ...
ZHANG Limin,WANG Yanzhe,ZHANG Bingqiang,ZHU Nianbin
doaj +1 more source
Star Temporal Classification: Sequence Classification with Partially Labeled Data [PDF]
We develop an algorithm which can learn from partially labeled and unsegmented sequential data. Most sequential loss functions, such as Connectionist Temporal Classification (CTC), break down when many labels are missing.
Ronan Collobert +3 more
core +1 more source
Attention-based CNN-ConvLSTM for Handwritten Arabic Word Extraction
Word extraction is one of the most critical steps in handwritten recognition systems. It is challenging for many reasons, such as the variability of handwritten writing styles, touching and overlapping characters, skewness problems, diacritics ...
takwa Ben Aicha, Afef Kacem Echi
doaj +1 more source
CTCModel: a Keras Model for Connectionist Temporal Classification [PDF]
We report an extension of a Keras Model, called CTCModel, to perform the Connection-ist Temporal Classification (CTC) in a transparent way. Combined with Recurrent Neural Networks, the Connectionist Temporal Classification is the reference method for ...
Soullard, Yann +2 more
core +5 more sources
A Recurrent Neural Networks (RNN) based attention model has been used in code-switching speech recognition (CSSR). However, due to the sequential computation constraint of RNN, there are stronger short-range dependencies and weaker long-range ...
Zheying Huang +5 more
doaj +1 more source
Speech Recognition Transformer Decoding Acceleration Method with Discarding Redundant Blocks [PDF]
Transformer and its variants have become mainstream models in the field of speech recognition owing to their excellent contextual modeling capabilities. Although they can achieve good recognition results, the decoding speed is limited because the decoder
Dechun ZHAO, Yang SHU, Ling LI, Huan CHEN, Zihao ZHANG
doaj +1 more source
Speechformer-CTC: Sequential Modeling of Depression Detection with Speech Temporal Classification. [PDF]
Speech-based automatic depression detection systems have been extensively explored over the past few years. Typically, each speaker is assigned a single label (Depressive or Non-depressive), and most approaches formulate depression detection as a speech ...
Wang J, Ravi V, Flint J, Alwan A.
europepmc +3 more sources
A Linear Memory CTC-Based Algorithm for Text-to-Voice Alignment of Very Long Audio Recordings
Synchronisation of a voice recording with the corresponding text is a common task in speech and music processing, and is used in many practical applications (automatic subtitling, audio indexing, etc.).
Guillaume Doras +2 more
doaj +1 more source
Variational Connectionist Temporal Classification for Order-Preserving Sequence Modeling [PDF]
Connectionist temporal classification (CTC) is commonly adopted for sequence modeling tasks like speech recognition, where it is necessary to preserve order between the input and target sequences.
Ahmed, Beena +3 more
core +1 more source
Towards end-to-end speech recognition with transfer learning
A transfer learning-based end-to-end speech recognition approach is presented in two levels in our framework. Firstly, a feature extraction approach combining multilingual deep neural network (DNN) training with matrix factorization algorithm is ...
Chu-Xiong Qin, Dan Qu, Lian-Hai Zhang
doaj +1 more source

