Results 71 to 80 of about 4,486 (211)

VQA task examples.

open access: yes, 2023
Visual Question Answering (VQA) is a multimodal task that uses natural language to ask and answer questions based on image content. For multimodal tasks, obtaining accurate modality feature information is crucial.
Yangshuyi Xu (16462504)   +2 more
core   +1 more source

The System of Visual Question Answering: Based on The Architectural Perspective [PDF]

open access: yesITM Web of Conferences
With the rapid development of computer vision and natural language processing technology, Visual Question Answering (VQA), as a cross-modal task, has gradually become a research hotspot in the field of artificial intelligence.
Wan Kuangming
doaj   +1 more source

Vision–Language Model for Visual Question Answering in Medical Imagery

open access: yesBioengineering, 2023
In the clinical and healthcare domains, medical images play a critical role. A mature medical visual question answering system (VQA) can improve diagnosis by answering clinical questions presented with a medical image.
Yakoub Bazi   +3 more
doaj   +1 more source

Toward Autonomous Clinics: Human–Robot Collaboration in Clinical Care

open access: yesMedComm, Volume 7, Issue 8, August 2026.
This graphical abstract summarizes the conceptual structure of human–robot collaboration (HRC) for autonomous clinical systems. The framework is organized around a dual‐brain architecture in which a professional brain supports clinical reasoning tasks such as diagnosis, image analysis, treatment planning, and patient monitoring, while a physical brain ...
Xinyuan Wu   +8 more
wiley   +1 more source

BPI-MVQA: a bi-branch model for medical visual question answering

open access: yesBMC Medical Imaging, 2022
Background Visual question answering in medical domain (VQA-Med) exhibits great potential for enhancing confidence in diagnosing diseases and helping patients better understand their medical conditions.
Shengyan Liu   +3 more
doaj   +1 more source

BMPCQA: Bioinspired Metaverse Point Cloud Quality Assessment Based on Large Multimodal Models

open access: yesAdvanced Intelligent Systems, Volume 8, Issue 7, July 2026.
This study presents a bioinspired metaverse point cloud quality assessment metric, which simulates the human visual evaluation process to perform the point cloud quality assessment task. It first extracts rendering projection video features, normal image features, and point cloud patch features, which are then fed into a large multimodal model to ...
Huiyu Duan   +7 more
wiley   +1 more source

MRAN-VQA: Multimodal Recursive Attention Network for Visual Question Answering

open access: yesEngineering Science and Technology, an International Journal
Visual Question Answering (VQA) is a fundamental challenge in multimodal AI, requiring models to integrate and reason over both visual and textual information. Despite advancements in deep learning, existing VQA models struggle with multi-step reasoning,
Mohammad Shariful Islam   +9 more
doaj   +1 more source

VQA-MHUG

open access: yes
We present VQA-MHUG - a novel 49-participant dataset of multimodal human gaze on both images and questions during visual question answering (VQA), collected using a high-speed eye tracker.
Kögel, Fabian   +2 more
core   +1 more source

Generative Artificial Intelligence and Large Language Models in Clinical Oncology

open access: yesMedComm, Volume 7, Issue 7, July 2026.
By integrating multimodal data, including medical imaging, pathology, omics, and electronic health records, generative AI and LLMs support cancer diagnosis, treatment planning, and follow‐up management. These technologies also enhance physician–patient communication, improve personalized treatment strategies, and leverage intelligent agent automation ...
Yunfang Yu   +14 more
wiley   +1 more source

Zero-Shot Transfer VQA Dataset

open access: yesCoRR, 2018
Acquiring a large vocabulary is an important aspect of human intelligence. Onecommon approach for human to populating vocabulary is to learn words duringreading or listening, and then use them in writing or speaking. This ability totransfer from input to output is natural for human, but it is difficult for machines.Human spontaneously performs this ...
Yuanpeng Li 0001   +3 more
openaire   +2 more sources

Home - About - Disclaimer - Privacy