Results 11 to 20 of about 4,486 (211)
Advancing surgical VQA with scene graph knowledge [PDF]
Abstract Purpose The modern operating room is becoming increasingly complex, requiring innovative intra-operative support systems. While the focus of surgical data science has largely been on video analysis, integrating surgical computer vision with natural language capabilities is emerging as a necessity.
Yuan, Kun +5 more
core +13 more sources
Towards Reasoning-Aware Explainable VQA
The domain of joint vision-language understanding, especially in the context of reasoning in Visual Question Answering (VQA) models, has garnered significant attention in the recent past. While most of the existing VQA models focus on improving the accuracy of VQA, the way models arrive at an answer is oftentimes a black box.
Rakesh Vaideeswaran +3 more
openaire +3 more sources
Making the V in Text-VQA Matter
Text-based VQA aims at answering questions by reading the text present in the images. It requires a large amount of scene-text relationship understanding compared to the VQA task. Recent studies have shown that the question-answer pairs in the dataset are more focused on the text present in the image but less importance is given to visual features and ...
Shamanthak Hegde +2 more
openaire +3 more sources
On the Role of Visual Grounding in VQA [PDF]
Visual Grounding (VG) in VQA refers to a model's proclivity to infer answers based on question-relevant image regions. Conceptually, VG identifies as an axiomatic requirement of the VQA task. In practice, however, DNN-based VQA models are notorious for bypassing VG by way of shortcut (SC) learning without suffering obvious performance losses in ...
Daniel Reich, Tanja Schultz
core +4 more sources
Semi-Supervised Implicit Augmentation for Data-Scarce VQA
Vision-language models (VLMs) have demonstrated increasing potency in solving complex vision-language tasks in the recent past. Visual question answering (VQA) is one of the primary downstream tasks for assessing the capability of VLMs, as it helps in ...
Bhargav Dodla +2 more
doaj +2 more sources
Adversarial VQA: A New Benchmark for Evaluating the Robustness of VQA Models [PDF]
To appear in ICCV 2021; Website: https://adversarialvqa.github.io/
Linjie Li +3 more
openaire +3 more sources
VQA With No Questions-Answers Training [PDF]
Methods for teaching machines to answer visual questions have made significant progress in recent years, but current methods still lack important human capabilities, including integrating new visual classes and concepts in a modular manner, providing explanations for the answers and handling new domains without explicit examples.
Ben Zion Vatashsky, Shimon Ullman
openaire +2 more sources
SQuINTing at VQA Models: Introspecting VQA Models With Sub-Questions [PDF]
Existing VQA datasets contain questions with varying levels of complexity. While the majority of questions in these datasets require perception for recognizing existence, properties, and spatial relationships of entities, a significant portion of questions pose challenges that correspond to reasoning tasks - tasks that can only be answered through a ...
Ramprasaath R. Selvaraju +6 more
openaire +3 more sources
An Experimental Study of the Vision-Bottleneck in Vqa [PDF]
As in many tasks combining vision and language, both modalities play a crucial role in Visual Question Answering (VQA). To properly solve the task, a given model should both understand the content of the proposed image and the nature of the question. While the fusion between modalities, which is another obviously important part of the problem, has been
Pierre Marza +4 more
openaire +2 more sources

