Results 101 to 110 of about 4,486 (211)
Question Modifiers in VQA: Evaluating Model Sensitivity [PDF]
Visual Question Answering (VQA) is a challenge problem that can advance AI by integrating several important sub-disciplines including natural language understanding and computer vision.
Britton, William Johnstone
core
Cross-Attention Mechanism for Medical Visual Question Answering
Visual Question Answering (VQA) is a machine learning task that aims to create systems capable of answering natural language questions based on given images.
Nada Fadhil Mohammed, Israa H. Ali
doaj +1 more source
Visually Grounded VQA by Lattice-based Retrieval
Visual Grounding (VG) in Visual Question Answering (VQA) systems describes how well a system manages to tie a question and its answer to relevant image regions. Systems with strong VG are considered intuitively interpretable and suggest an improved scene
Reich, Daniel +2 more
core
Improving Automatic VQA Evaluation Using Large Language Models
8 years after the visual question answering (VQA) task was proposed, accuracy remains the primary metric for automatic evaluation. VQA Accuracy has been effective so far in the IID evaluation setting.
Agrawal, Aishwarya +2 more
core +2 more sources
Visual Question Answering in Robotic Surgery: A Comprehensive Review
Visual Question Answering (VQA) in robotic surgery is rapidly becoming a pivotal technology in medical AI, addressing the complex challenge of interpreting multimodal surgical data to support real-time decision-making.
Di Ding +3 more
doaj +1 more source
RSMoDM: Multimodal Momentum Distillation Model for Remote Sensing Visual Question Answering
Remote sensing (RS) visual question answering (VQA) is a task that answers questions about a given RS image by utilizing both image and textual information.
Pengfei Li +5 more
doaj +1 more source
PTCR: Knowledge-Based Visual Question Answering Framework Based on Large Language Model [PDF]
Aiming at the problems of insufficient model input information and poor reasoning performance in knowledge-based visual question answering (VQA), this paper constructs a PTCR knowledge-based VQA framework based on large language model (LLM), which ...
XUE Di, LI Xin, LIU Mingshuai
doaj +1 more source
Overview Indian Traffic VQA is a real-world Visual Question Answering (VQA) dataset focusing on Indian road traffic signboards. The dataset is designed for training and evaluating Vision-Language Models (VLMs) and VQA systems in the traffic and ...
Bhuma, Chandra Mohan
core +1 more source
Estimating semantic structure for the VQA answer space
Since its appearance, Visual Question Answering (VQA, i.e. answering a question posed over an image), has always been treated as a classification problem over a set of predefined answers.
Baccouche, Moez +3 more
core +3 more sources
Analysing human vs. neural attention in VQA [PDF]
Visual Question Answering (VQA) has drawn substantial interest in both academic and industrial research fields in recent years. Driven by Vision Transformers (ViT) and the vision-text co-attention mechanism, these models have shown notable performance ...
Ma, Yingpeng
core +1 more source

