Results 51 to 60 of about 4,486 (211)

K-VQA: A visual question answering method

open access: yesJournal of Hebei University of Science and Technology, 2020
The types of questions answered by the visual question answering of images and texts are roughly divided into two types. The first type is the questions that can get the answers directly from the images, and the second type is the questions that need the
Hongbin GAO, Jinying MAO, Huiyong WANG
doaj   +1 more source

A Review of Artificial Intelligence in Ophthalmology: Key Aspects, Challenges, and Future Directions

open access: yesEye &ENT Research, Volume 3, Issue 3, Page 125-154, September 2026.
ABSTRACT Artificial intelligence (AI) is increasingly reshaping ophthalmology because the specialty depends heavily on structured imaging, quantitative measurements, and repeatable diagnostic workflows. This review provides a clinically grounded and translationally oriented synthesis of AI in ophthalmology, covering methodological foundations ...
Partha Pratim Ray
wiley   +1 more source

Iterated learning for emergent systematicity in VQA

open access: yesCoRR, 2021
Published as a conference paper at ICLR 2021.
Ankit Vani   +4 more
openaire   +3 more sources

MedFuseT: A Transformer-Based Model for Advancing Medical Visual Question Answering

open access: yesIEEE Access
As medicine and healthcare continue to evolve, Visual Question Answering (VQA) has emerged as an important application of artificial intelligence.
Hamza Mbarek   +3 more
doaj   +1 more source

AI Real‐Time Compositing Algorithm for Film and Television Special Effects Based on Latent Space Diffusion Model

open access: yesEngineering Reports, Volume 8, Issue 9, September 2026.
The proposed latent diffusion compositing framework achieves high‐quality, temporally coherent, and efficient real‐time video compositing for practical AI‐driven visual effects applications. ABSTRACT Artificial intelligence–driven visual effects have become increasingly important in film and television production; however, achieving high‐quality real ...
Bowen An, Yue Xie, Yalun Lei
wiley   +1 more source

SURE-VQA: Systematic Understanding of Robustness Evaluation in Medical VQA Tasks

open access: yesCoRR
TMLR 07 ...
Kim-Celine Kahl   +6 more
openaire   +4 more sources

Plug-and-Play VQA: Zero-shot VQA by Conjoining Large Pretrained Models with Zero Training

open access: yesFindings of the Association for Computational Linguistics: EMNLP 2022, 2022
Visual question answering (VQA) is a hallmark of vision and language reasoning and a challenging task under the zero-shot setting. We propose Plug-and-Play VQA (PNP-VQA), a modular framework for zero-shot VQA. In contrast to most existing works, which require substantial adaptation of pretrained language models (PLMs) for the vision modality, PNP-VQA ...
Anthony Meng Huat Tiong   +4 more
openaire   +4 more sources

Prompt-Driven Fuzzing Debiasing Framework for Robust Visual Question Answering

open access: yesMultimodal Technologies and Interaction
Visual Question Answering (VQA) systems have achieved impressive performance with the rise of large-scale vision–language models (VLMs). However, these models remain vulnerable to multiple forms of multimodal bias, severely limiting their robustness and ...
Yali Fan   +3 more
doaj   +1 more source

MedFuseNet: An attention-based multimodal deep learning model for visual question answering in the medical domain

open access: yesScientific Reports, 2021
Medical images are difficult to comprehend for a person without expertise. The scarcity of medical practitioners across the globe often face the issue of physical and mental fatigue due to the high number of cases, inducing human errors during the ...
Dhruv Sharma   +2 more
doaj   +1 more source

Large Language Model Benchmarks in Medical Tasks

open access: yesMedicine Advances, Volume 4, Issue 3, Page 316-341, September 2026.
This survey provides a comprehensive review of benchmark datasets for evaluating large language models in medical tasks, spanning text, image, and multimodal modalities, and covering key resources including MIMIC‐III/IV, BioASQ, PubMedQA, and CheXpert.
Lawrence K. Q. Yan   +18 more
wiley   +1 more source

Home - About - Disclaimer - Privacy