Results 21 to 30 of about 2,681 (242)
Survey of Visual Question Answering Based on Deep Learning [PDF]
Visual question answering(VQA) is an interdisciplinary research paradigm that involves computer vision and natural language processing.VQA generally requires both image and text data to be encoded,their mappings learned,and their features fused,before ...
LI Xiang, FAN Zhiguang, LI Xuexiang, ZHANG Weixing, YANG Cong, CAO Yangjie
doaj +1 more source
The effects of cross-modal feature and location mappings on visual performance
Janna Wennberg, Viola S. Stoermer
exaly +2 more sources
Cross-modal Map Learning for Vision and Language Navigation
We consider the problem of Vision-and-Language Navigation (VLN). The majority of current methods for VLN are trained end-to-end using either unstructured memory such as LSTM, or using cross-modal attention over the egocentric observations of the agent.
Georgios Georgakis +6 more
openaire +2 more sources
Deep Perceptual Mapping for Cross-Modal Face Recognition [PDF]
This is the extended version (invited IJCV submission) with new results of our previous submission (arXiv:1507.02879)
M. Saquib Sarfraz, Rainer Stiefelhagen
openaire +3 more sources
Cross Modal Facial Image Synthesis Using a Collaborative Bidirectional Style Transfer Network
In this paper, we present a novel collaborative bidirectional style transfer network based on generative adversarial network (GAN) for cross modal facial image synthesis, possibly with large modality gap.
Nizam Ud Din +4 more
doaj +1 more source
Statistical learning of cross-modal correspondence with non-linear mappings
Kazuhiko Yokosawa +2 more
exaly +2 more sources
Cross-modal re-mapping influences the Simon effect [PDF]
Tagliabue, Zorzi, Umiltà, and Bassignani (2000) showed that one's practicing of a spatially incompatible task influences performance in a Simon task even when the interval between the two tasks is as long as 1 week. In the present study, three experiments were conducted to investigate whether such an effect could be found in a cross-modal paradigm ...
TAGLIABUE, MARIAELENA +2 more
openaire +3 more sources
People conceptualize auditory pitch as vertical space: low and high pitch correspond to low and high space, respectively. The strength of this cross-modal correspondence, however, seems to vary across different cultural contexts and a debate on the ...
Valentijn Prové
doaj +1 more source
Pix2Map: Cross-Modal Retrieval for Inferring Street Maps from Images
12 pages, 8 ...
Xindi Wu +4 more
openaire +2 more sources
Iconicity correlated with vowel harmony in Korean ideophones
This paper aims to establish connections between the following phenomena pertaining to Korean ideophonic vowel harmony: A set of vowel patterns classified (phonologically) as ‘harmonic,’ ‘neutral,’ and ‘disharmonic’; a set of ideophones classified ...
Nahyun Kwon
doaj +2 more sources

