Results 151 to 160 of about 76,886 (181)
Some of the next articles are maybe not open access.
Asia-Pacific Signal and Information Processing Association Annual Summit and Conference, 2023
This paper describes end-to-end automatic drum transcription for directly estimating a drum score from an audio signal of popular music using non-aligned paired data.
Daichi Kamakura +2 more
semanticscholar +1 more source
This paper describes end-to-end automatic drum transcription for directly estimating a drum score from an audio signal of popular music using non-aligned paired data.
Daichi Kamakura +2 more
semanticscholar +1 more source
Chinese Audio Transcription Using Connectionist Temporal Classification
Proceedings of the 8th International Conference on Computer and Communications Management, 2020Mandarin is one of the global languages that have large users and speakers. There are several important factors for learners to be an expert in Mandarin. To be able to communicate properly, mastery in Chinese character (hanzi) and pīnyīn are required. We develop an Android-based app to help students who learn Mandarin.
Jansen +2 more
openaire +1 more source
International Conference on the Software Process, 2021
Keyword spotting (KWS) is an essential feature for speech-based applications on mobile devices. For the sake of reducing power consumption and improving robustness on substandard pronunciations of KWS systems, this paper proposes a query-by-example on ...
Meirong Li +3 more
semanticscholar +1 more source
Keyword spotting (KWS) is an essential feature for speech-based applications on mobile devices. For the sake of reducing power consumption and improving robustness on substandard pronunciations of KWS systems, this paper proposes a query-by-example on ...
Meirong Li +3 more
semanticscholar +1 more source
Connectionist temporal classification
Proceedings of the 23rd international conference on Machine learning - ICML '06, 2006Many real-world sequence learning tasks require the prediction of sequences of labels from noisy, unsegmented input data. In speech recognition, for example, an acoustic signal is transcribed into words or sub-word units. Recurrent neural networks (RNNs) are powerful sequence learners that would seem well suited to such tasks.
Alex Graves +3 more
openaire +1 more source
Interspeech, 2019
The state-of-the-art neural network architecture named Trans-former has been used successfully for many sequence-to-sequence transformation tasks. The advantage of this architecture is that it has a fast iteration speed in the training stage because ...
Shigeki Karita +5 more
semanticscholar +1 more source
The state-of-the-art neural network architecture named Trans-former has been used successfully for many sequence-to-sequence transformation tasks. The advantage of this architecture is that it has a fast iteration speed in the training stage because ...
Shigeki Karita +5 more
semanticscholar +1 more source
International Conference on Language Resources and Evaluation
Previous studies employ the autoregressive translation (AT) paradigm in the document-to-document neural machine translation. These methods extend the translation unit from a single sentence to a pseudo-document and encodes the full pseudo-document ...
Hao Yu +4 more
semanticscholar +1 more source
Previous studies employ the autoregressive translation (AT) paradigm in the document-to-document neural machine translation. These methods extend the translation unit from a single sentence to a pseudo-document and encodes the full pseudo-document ...
Hao Yu +4 more
semanticscholar +1 more source
Deep Learning Model Efficient Lip Reading by Using Connectionist Temporal Classification Algorithm
Students Conference on Engineering and SystemsThe technique of translating text from a person's mouth movements is said to be lipreading. The main work of lip reading is the process that is divided into two stages using old traditional techniques: prediction of a speaker's lips and producing or ...
B. Vidhya, R. O, K. Yazhini
semanticscholar +1 more source
Modeling intra-label dynamics in connectionist temporal classification
2017 7th International Conference on Computer and Knowledge Engineering (ICCKE), 2017Most sequence processing tasks can be cast as a problem of mapping a sequence of observations into a sequence of labels. This is a very difficult problem since the association between input data sequences and output label sequences is not given at the frame level. Recurrent neural networks (RNNs) equipped with connectionist temporal classification (CTC)
Ashkan Sadeghi Lotfabadi +2 more
openaire +1 more source
Interspeech, 2020
We propose a two-stage sound event detection (SED) model to deal with sound events overlapping in time-frequency. In the first stage which consists of a faster R-CNN and an attention-LSTM, each log-mel spectrogram segment is divided into one or more ...
Inyoung Park, Hong Kook Kim
semanticscholar +1 more source
We propose a two-stage sound event detection (SED) model to deal with sound events overlapping in time-frequency. In the first stage which consists of a faster R-CNN and an attention-LSTM, each log-mel spectrogram segment is divided into one or more ...
Inyoung Park, Hong Kook Kim
semanticscholar +1 more source

