Results 21 to 30 of about 4,486 (211)

MUST-VQA: MUltilingual Scene-Text VQA

open access: yes, 2023
In this paper, we present a framework for Multilingual Scene Text Visual Question Answering that deals with new languages in a zero-shot fashion. Specifically, we consider the task of Scene Text Visual Question Answering (STVQA) in which the question can be asked in different languages and it is not necessarily aligned to the scene text language. Thus,
Emanuele Vivoli   +4 more
openaire   +2 more sources

VQA: Visual Question Answering [PDF]

open access: yes2015 IEEE International Conference on Computer Vision (ICCV), 2015
The first three authors contributed equally.
Stanislaw Antol   +6 more
openaire   +4 more sources

Object-Based Reasoning in VQA [PDF]

open access: yes2018 IEEE Winter Conference on Applications of Computer Vision (WACV), 2018
10 pages, 15 figures, published as a conference paper at 2018 IEEE Winter Conf. on Applications of Computer Vision (WACV'2018)
Mikyas T. Desta   +2 more
openaire   +2 more sources

How Transferable are Reasoning Patterns in VQA? [PDF]

open access: yes2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021
Since its inception, Visual Question Answering (VQA) is notoriously known as a task, where models are prone to exploit biases in datasets to find shortcuts instead of performing high-level reasoning. Classical methods address this by removing biases from training data, or adding branches to models to detect and remove biases.
Corentin Kervadec   +5 more
openaire   +3 more sources

How (not) to ensemble LVLMs for VQA

open access: yesCoRR, 2023
This paper studies ensembling in the era of Large Vision-Language Models (LVLMs). Ensembling is a classical method to combine different models to get increased performance. In the recent work on Encyclopedic-VQA the authors examine a wide variety of models to solve their task: from vanilla LVLMs, to models including the caption as extra context, to ...
Lisa Alazraki   +5 more
openaire   +4 more sources

DocVQA: A Dataset for VQA on Document Images [PDF]

open access: yes2021 IEEE Winter Conference on Applications of Computer Vision (WACV), 2021
We present a new dataset for Visual Question Answering (VQA) on document images called DocVQA. The dataset consists of 50,000 questions defined on 12,000+ document images. Detailed analysis of the dataset in comparison with similar datasets for VQA and reading comprehension is presented.
Minesh Mathew   +3 more
openaire   +3 more sources

Exploring Question Decomposition for Zero-Shot VQA [PDF]

open access: yes, 2023
Visual question answering (VQA) has traditionally been treated as a single-step task where each question receives the same amount of effort, unlike natural human question-answering strategies.
Khan, Zaid   +4 more
core   +1 more source

DSGEM: Dual scene graph enhancement module‐based visual question answering

open access: yesIET Computer Vision, 2023
Visual Question Answering (VQA) aims to appropriately answer a text question by understanding the image content. Attention‐based VQA models mine the implicit relationships between objects according to the feature similarity, which neglects the explicit ...
Boyue Wang   +5 more
doaj   +1 more source

Building Trustworthy Multimodal AI: A Review of Fairness, Transparency, and Ethics in Vision-Language Tasks [PDF]

open access: yesInternational Journal of Web Research
Objective: This review explores the trustworthiness of multimodal artificial intelligence (AI) systems, specifically focusing on vision-language tasks. It addresses critical challenges related to fairness, transparency, and ethical implications in these ...
Mohammad Saleh, Azadeh Tabatabaei
doaj   +1 more source

MRET: Multi-resolution transformer for video quality assessment

open access: yesFrontiers in Signal Processing, 2023
No-reference video quality assessment (NR-VQA) for user generated content (UGC) is crucial for understanding and improving visual experience. Unlike video recognition tasks, VQA tasks are sensitive to changes in input resolution.
Junjie Ke   +4 more
doaj   +1 more source

Home - About - Disclaimer - Privacy