Results 101 to 110 of about 4,486 (211)

Question Modifiers in VQA: Evaluating Model Sensitivity [PDF]

open access: yes, 2022
Visual Question Answering (VQA) is a challenge problem that can advance AI by integrating several important sub-disciplines including natural language understanding and computer vision.
Britton, William Johnstone
core  

Cross-Attention Mechanism for Medical Visual Question Answering

open access: yesمجلة بغداد للعلوم
Visual Question Answering (VQA) is a machine learning task that aims to create systems capable of answering natural language questions based on given images.
Nada Fadhil Mohammed, Israa H. Ali
doaj   +1 more source

Visually Grounded VQA by Lattice-based Retrieval

open access: yes, 2022
Visual Grounding (VG) in Visual Question Answering (VQA) systems describes how well a system manages to tie a question and its answer to relevant image regions. Systems with strong VG are considered intuitively interpretable and suggest an improved scene
Reich, Daniel   +2 more
core  

Improving Automatic VQA Evaluation Using Large Language Models

open access: yes
8 years after the visual question answering (VQA) task was proposed, accuracy remains the primary metric for automatic evaluation. VQA Accuracy has been effective so far in the IID evaluation setting.
Agrawal, Aishwarya   +2 more
core   +2 more sources

Visual Question Answering in Robotic Surgery: A Comprehensive Review

open access: yesIEEE Access
Visual Question Answering (VQA) in robotic surgery is rapidly becoming a pivotal technology in medical AI, addressing the complex challenge of interpreting multimodal surgical data to support real-time decision-making.
Di Ding   +3 more
doaj   +1 more source

RSMoDM: Multimodal Momentum Distillation Model for Remote Sensing Visual Question Answering

open access: yesIEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing
Remote sensing (RS) visual question answering (VQA) is a task that answers questions about a given RS image by utilizing both image and textual information.
Pengfei Li   +5 more
doaj   +1 more source

PTCR: Knowledge-Based Visual Question Answering Framework Based on Large Language Model [PDF]

open access: yesJisuanji kexue yu tansuo
Aiming at the problems of insufficient model input information and poor reasoning performance in knowledge-based visual question answering (VQA), this paper constructs a PTCR knowledge-based VQA framework based on large language model (LLM), which ...
XUE Di, LI Xin, LIU Mingshuai
doaj   +1 more source

Indian Traffic VQA Dataset

open access: yes
Overview Indian Traffic VQA is a real-world Visual Question Answering (VQA) dataset focusing on Indian road traffic signboards. The dataset is designed for training and evaluating Vision-Language Models (VLMs) and VQA systems in the traffic and ...
Bhuma, Chandra Mohan
core   +1 more source

Estimating semantic structure for the VQA answer space

open access: yes, 2020
Since its appearance, Visual Question Answering (VQA, i.e. answering a question posed over an image), has always been treated as a classification problem over a set of predefined answers.
Baccouche, Moez   +3 more
core   +3 more sources

Analysing human vs. neural attention in VQA [PDF]

open access: yes
Visual Question Answering (VQA) has drawn substantial interest in both academic and industrial research fields in recent years. Driven by Vision Transformers (ViT) and the vision-text co-attention mechanism, these models have shown notable performance ...
Ma, Yingpeng
core   +1 more source

Home - About - Disclaimer - Privacy