Results 11 to 20 of about 4,895,571 (286)
Localized Questions in Medical Visual Question Answering [PDF]
Visual Question Answering (VQA) models aim to answer natural language questions about given images. Due to its ability to ask questions that differ from those used when training the model, medical VQA has received substantial attention in recent years.
Sergio Tascon-Morales +2 more
openaire +6 more sources
Counterfactual Mix-Up for Visual Question Answering
Counterfactuals have been shown to be a powerful method in Visual Question Answering in the alleviation of Visual Question Answering’s unimodal bias. However, existing counterfactual methods tend to generate samples that are not diverse or require
Jae Won Cho +3 more
doaj +2 more sources
Intra-assessor consistency in question answering [PDF]
In this paper we investigate the consistency of answer assessment in a complex question answering task examining features of assessor consistency, types of answers and question ...
Murat Yakici +20 more
core +6 more sources
Review of Visual Question Answering Technology [PDF]
Visual question answering (VQA) is a popular cross-modal task that combines natural language pro-cessing and computer vision techniques. The main objective of this task is to enable computers to intelligently recognize and retrieve visual content and ...
WANG Yu, SUN Haichun
doaj +1 more source
Vision-language models for medical report generation and visual question answering: a review [PDF]
Ghulam Rasool
exaly +2 more sources
SBVQA 2.0: Robust End-to-End Speech-Based Visual Question Answering for Open-Ended Questions
Speech-based Visual Question Answering (SBVQA) is a challenging task that aims to answer spoken questions about images. The challenges of this task involve the variability of speakers, the different recording environments, as well as the various objects ...
Faris Alasmary, Saad Al-Ahmadi
doaj +1 more source
Multi-Module Co-Attention Model for Visual Question Answering [PDF]
Visual Question Answering(VQA) is a typical multi-modal problem in computer vision and natural language processing.Most of the existing VQA models ignore the dynamic relationships of semantic information between two modes and the rich spatial structure ...
ZOU Pinrong, XIAO Feng, ZHANG Wenjuan, ZHANG Wanyu, WANG Chenyang
doaj +1 more source
A Comprehensive Review and Open Challenges on Visual Question Answering Models
Users are now able to actively interact with images and pose different questions based on images, thanks to recent developments in artificial intelligence. In turn, a response in a natural language answer is expected.
Fasi Ahamad Shaik +4 more
doaj +1 more source
VQA: Visual Question Answering [PDF]
The first three authors contributed equally.
Stanislaw Antol +6 more
openaire +4 more sources
TASTA: Text‐Assisted Spatial and Temporal Attention Network for Video Question Answering
Video question answering (VideoQA) is a typical task that integrates language and vision. The key for VideoQA is to extract relevant and effective visual information for answering a specific question. Information selection is believed to be necessary for
Tian Wang +5 more
doaj +1 more source

