Results 11 to 20 of about 4,486 (211)

Advancing surgical VQA with scene graph knowledge [PDF]

open access: yesInternational Journal of Computer Assisted Radiology and Surgery
Abstract Purpose The modern operating room is becoming increasingly complex, requiring innovative intra-operative support systems. While the focus of surgical data science has largely been on video analysis, integrating surgical computer vision with natural language capabilities is emerging as a necessity.
Yuan, Kun   +5 more
core   +13 more sources

Towards Reasoning-Aware Explainable VQA

open access: yesCoRR, 2022
The domain of joint vision-language understanding, especially in the context of reasoning in Visual Question Answering (VQA) models, has garnered significant attention in the recent past. While most of the existing VQA models focus on improving the accuracy of VQA, the way models arrive at an answer is oftentimes a black box.
Rakesh Vaideeswaran   +3 more
openaire   +3 more sources

Making the V in Text-VQA Matter

open access: yes2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), 2023
Text-based VQA aims at answering questions by reading the text present in the images. It requires a large amount of scene-text relationship understanding compared to the VQA task. Recent studies have shown that the question-answer pairs in the dataset are more focused on the text present in the image but less importance is given to visual features and ...
Shamanthak Hegde   +2 more
openaire   +3 more sources

On the Role of Visual Grounding in VQA [PDF]

open access: yesCoRR
Visual Grounding (VG) in VQA refers to a model's proclivity to infer answers based on question-relevant image regions. Conceptually, VG identifies as an axiomatic requirement of the VQA task. In practice, however, DNN-based VQA models are notorious for bypassing VG by way of shortcut (SC) learning without suffering obvious performance losses in ...
Daniel Reich, Tanja Schultz
core   +4 more sources

Semi-Supervised Implicit Augmentation for Data-Scarce VQA

open access: yesComputer Sciences & Mathematics Forum
Vision-language models (VLMs) have demonstrated increasing potency in solving complex vision-language tasks in the recent past. Visual question answering (VQA) is one of the primary downstream tasks for assessing the capability of VLMs, as it helps in ...
Bhargav Dodla   +2 more
doaj   +2 more sources

Adversarial VQA: A New Benchmark for Evaluating the Robustness of VQA Models [PDF]

open access: yes2021 IEEE/CVF International Conference on Computer Vision (ICCV), 2021
To appear in ICCV 2021; Website: https://adversarialvqa.github.io/
Linjie Li   +3 more
openaire   +3 more sources

VQA With No Questions-Answers Training [PDF]

open access: yes2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020
Methods for teaching machines to answer visual questions have made significant progress in recent years, but current methods still lack important human capabilities, including integrating new visual classes and concepts in a modular manner, providing explanations for the answers and handling new domains without explicit examples.
Ben Zion Vatashsky, Shimon Ullman
openaire   +2 more sources

January 2020 VQA Sale

open access: yes, 2020
Sale results of VQA Calf ...
Overbay, Andrew
openaire   +2 more sources

SQuINTing at VQA Models: Introspecting VQA Models With Sub-Questions [PDF]

open access: yes2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020
Existing VQA datasets contain questions with varying levels of complexity. While the majority of questions in these datasets require perception for recognizing existence, properties, and spatial relationships of entities, a significant portion of questions pose challenges that correspond to reasoning tasks - tasks that can only be answered through a ...
Ramprasaath R. Selvaraju   +6 more
openaire   +3 more sources

An Experimental Study of the Vision-Bottleneck in Vqa [PDF]

open access: yesSSRN Electronic Journal, 2022
As in many tasks combining vision and language, both modalities play a crucial role in Visual Question Answering (VQA). To properly solve the task, a given model should both understand the content of the proposed image and the nature of the question. While the fusion between modalities, which is another obviously important part of the problem, has been
Pierre Marza   +4 more
openaire   +2 more sources

Home - About - Disclaimer - Privacy