Results 21 to 30 of about 4,486 (211)
MUST-VQA: MUltilingual Scene-Text VQA
In this paper, we present a framework for Multilingual Scene Text Visual Question Answering that deals with new languages in a zero-shot fashion. Specifically, we consider the task of Scene Text Visual Question Answering (STVQA) in which the question can be asked in different languages and it is not necessarily aligned to the scene text language. Thus,
Emanuele Vivoli +4 more
openaire +2 more sources
VQA: Visual Question Answering [PDF]
The first three authors contributed equally.
Stanislaw Antol +6 more
openaire +4 more sources
Object-Based Reasoning in VQA [PDF]
10 pages, 15 figures, published as a conference paper at 2018 IEEE Winter Conf. on Applications of Computer Vision (WACV'2018)
Mikyas T. Desta +2 more
openaire +2 more sources
How Transferable are Reasoning Patterns in VQA? [PDF]
Since its inception, Visual Question Answering (VQA) is notoriously known as a task, where models are prone to exploit biases in datasets to find shortcuts instead of performing high-level reasoning. Classical methods address this by removing biases from training data, or adding branches to models to detect and remove biases.
Corentin Kervadec +5 more
openaire +3 more sources
How (not) to ensemble LVLMs for VQA
This paper studies ensembling in the era of Large Vision-Language Models (LVLMs). Ensembling is a classical method to combine different models to get increased performance. In the recent work on Encyclopedic-VQA the authors examine a wide variety of models to solve their task: from vanilla LVLMs, to models including the caption as extra context, to ...
Lisa Alazraki +5 more
openaire +4 more sources
DocVQA: A Dataset for VQA on Document Images [PDF]
We present a new dataset for Visual Question Answering (VQA) on document images called DocVQA. The dataset consists of 50,000 questions defined on 12,000+ document images. Detailed analysis of the dataset in comparison with similar datasets for VQA and reading comprehension is presented.
Minesh Mathew +3 more
openaire +3 more sources
Exploring Question Decomposition for Zero-Shot VQA [PDF]
Visual question answering (VQA) has traditionally been treated as a single-step task where each question receives the same amount of effort, unlike natural human question-answering strategies.
Khan, Zaid +4 more
core +1 more source
DSGEM: Dual scene graph enhancement module‐based visual question answering
Visual Question Answering (VQA) aims to appropriately answer a text question by understanding the image content. Attention‐based VQA models mine the implicit relationships between objects according to the feature similarity, which neglects the explicit ...
Boyue Wang +5 more
doaj +1 more source
Building Trustworthy Multimodal AI: A Review of Fairness, Transparency, and Ethics in Vision-Language Tasks [PDF]
Objective: This review explores the trustworthiness of multimodal artificial intelligence (AI) systems, specifically focusing on vision-language tasks. It addresses critical challenges related to fairness, transparency, and ethical implications in these ...
Mohammad Saleh, Azadeh Tabatabaei
doaj +1 more source
MRET: Multi-resolution transformer for video quality assessment
No-reference video quality assessment (NR-VQA) for user generated content (UGC) is crucial for understanding and improving visual experience. Unlike video recognition tasks, VQA tasks are sensitive to changes in input resolution.
Junjie Ke +4 more
doaj +1 more source

