Results 11 to 20 of about 2,645,934 (333)
Differential Gaze Patterns on Eyes and Mouth During Audiovisual Speech Segmentation. [PDF]
Speech is inextricably multisensory: both auditory and visual components provide critical information for all aspects of speech processing, including speech segmentation, the visual components of which have been the target of a growing number of studies.
Lusk LG, Mitchel AD.
europepmc +2 more sources
In order to acquire their native languages, children must learn richly structured systems with regularities at multiple levels. While structure at different levels could be learned serially, e.g.
Daniel eYurovsky +2 more
doaj +3 more sources
Stress placement and word segmentation by Spanish speakers [PDF]
Several studies have shown that the stress pattern of one's native language is applied to new linguistic stimuli. Regarding the segmentation of artificial synthesized speech, this idea has been supported by experiments with languages where the stress ...
Rodríguez Fornells, Antoni +2 more
core +6 more sources
HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units [PDF]
Self-supervised approaches for speech representation learning are challenged by three unique problems: (1) there are multiple sound units in each input utterance, (2) there is no lexicon of input sound units during the pre-training phase, and (3) sound ...
Wei-Ning Hsu +5 more
semanticscholar +1 more source
Hippocampal and auditory contributions to speech segmentation. [PDF]
Statistical learning has been proposed as a mechanism to structure and segment the continuous flow of information in several sensory modalities. Previous studies proposed that the medial temporal lobe, and in particular the hippocampus, may be crucial to
Neus Ramos-Escobar +5 more
semanticscholar +1 more source
Selection of acoustic modeling unit for Tibetan speech recognition based on deep learning [PDF]
The selection of the speech recognition modeling unit is the primary problem of acoustic modeling in speech recognition, and different acoustic modeling units will directly affect the overall performance of speech recognition.
Gong Baojia +4 more
doaj +1 more source
End-to-End Simultaneous Speech Translation with Differentiable Segmentation [PDF]
End-to-end simultaneous speech translation (SimulST) outputs translation while receiving the streaming speech inputs (a.k.a. streaming speech translation), and hence needs to segment the speech inputs and then translate based on the current received ...
Shaolei Zhang, Yang Feng
semanticscholar +1 more source
A digital neural network approach to speech recognition [PDF]
This thesis was submitted for the degree of Doctor of Philosophy and awarded by Brunel University.This thesis presents two novel methods for isolated word speech recognition based on sub-word components.
Haider, Najmi Ghani
core +7 more sources
SHAS: Approaching optimal Segmentation for End-to-End Speech Translation [PDF]
Speech translation models are unable to directly process long audios, like TED talks, which have to be split into shorter segments. Speech translation datasets provide manual segmentations of the audios, which are not available in real-world scenarios ...
Yiannis (Ioannis) Tsiamas +3 more
semanticscholar +1 more source
Neural oscillations constitute an intrinsic property of functional brain organization that facilitates the tracking of linguistic units at multiple time scales through brain-to-stimulus alignment.
S. Elmer +2 more
semanticscholar +1 more source

