Results 51 to 60 of about 4,486 (211)
K-VQA: A visual question answering method
The types of questions answered by the visual question answering of images and texts are roughly divided into two types. The first type is the questions that can get the answers directly from the images, and the second type is the questions that need the
Hongbin GAO, Jinying MAO, Huiyong WANG
doaj +1 more source
A Review of Artificial Intelligence in Ophthalmology: Key Aspects, Challenges, and Future Directions
ABSTRACT Artificial intelligence (AI) is increasingly reshaping ophthalmology because the specialty depends heavily on structured imaging, quantitative measurements, and repeatable diagnostic workflows. This review provides a clinically grounded and translationally oriented synthesis of AI in ophthalmology, covering methodological foundations ...
Partha Pratim Ray
wiley +1 more source
Iterated learning for emergent systematicity in VQA
Published as a conference paper at ICLR 2021.
Ankit Vani +4 more
openaire +3 more sources
MedFuseT: A Transformer-Based Model for Advancing Medical Visual Question Answering
As medicine and healthcare continue to evolve, Visual Question Answering (VQA) has emerged as an important application of artificial intelligence.
Hamza Mbarek +3 more
doaj +1 more source
The proposed latent diffusion compositing framework achieves high‐quality, temporally coherent, and efficient real‐time video compositing for practical AI‐driven visual effects applications. ABSTRACT Artificial intelligence–driven visual effects have become increasingly important in film and television production; however, achieving high‐quality real ...
Bowen An, Yue Xie, Yalun Lei
wiley +1 more source
SURE-VQA: Systematic Understanding of Robustness Evaluation in Medical VQA Tasks
TMLR 07 ...
Kim-Celine Kahl +6 more
openaire +4 more sources
Plug-and-Play VQA: Zero-shot VQA by Conjoining Large Pretrained Models with Zero Training
Visual question answering (VQA) is a hallmark of vision and language reasoning and a challenging task under the zero-shot setting. We propose Plug-and-Play VQA (PNP-VQA), a modular framework for zero-shot VQA. In contrast to most existing works, which require substantial adaptation of pretrained language models (PLMs) for the vision modality, PNP-VQA ...
Anthony Meng Huat Tiong +4 more
openaire +4 more sources
Prompt-Driven Fuzzing Debiasing Framework for Robust Visual Question Answering
Visual Question Answering (VQA) systems have achieved impressive performance with the rise of large-scale vision–language models (VLMs). However, these models remain vulnerable to multiple forms of multimodal bias, severely limiting their robustness and ...
Yali Fan +3 more
doaj +1 more source
Medical images are difficult to comprehend for a person without expertise. The scarcity of medical practitioners across the globe often face the issue of physical and mental fatigue due to the high number of cases, inducing human errors during the ...
Dhruv Sharma +2 more
doaj +1 more source
Large Language Model Benchmarks in Medical Tasks
This survey provides a comprehensive review of benchmark datasets for evaluating large language models in medical tasks, spanning text, image, and multimodal modalities, and covering key resources including MIMIC‐III/IV, BioASQ, PubMedQA, and CheXpert.
Lawrence K. Q. Yan +18 more
wiley +1 more source

