Visual Question Answering (VQA) on Images with Superimposed Text
Superimposed text annotations have been under-investigated, yet are ubiquitous, useful and important, especially in medical images. Medical images also highlight the challenges posed by low resolution, noise and superimposed textual meta-information ...
Kodali, Venkat, Berleant, Daniel
core
Enabling Multimodal Understanding: Lidar Data Meets VQA
This chapter explores the integration of Light Detection and Ranging (LiDAR) data with multimodal systems such as Visual Question Answering (VQA) to enable robust contextual understanding.
Dhananjay Thiruvady (13066857) +3 more
core
Knowledge Generation for Zero-shot Knowledge-based VQA
Previous solutions to knowledge-based visual question answering~(K-VQA) retrieve knowledge from external knowledge bases and use supervised learning to train the K-VQA model.
Jiang, Jing, Cao, Rui
core
Structured multi-level knowledge augmentation via small-to-large evidence-guided collaboration for knowledge-based VQA. [PDF]
Zhang M, Wu D, Chung W.
europepmc +1 more source
VQA-Diff: Exploiting VQA and Diffusion for Zero-Shot Image-to-3D Vehicle Asset Generation in Autonomous Driving [PDF]
Generating 3D vehicle assets from in-the-wild observations is crucial to autonomous driving. Existing image-to-3D methods cannot well address this problem because they learn generation merely from image RGB information without a deeper understanding of ...
Shan, Jinjun +7 more
core
SurgViVQA: temporally grounded video question answering for surgical scene understanding. [PDF]
Drago MO +9 more
europepmc +1 more source
A Visual Question Answering Dataset for Benchmarking Vision-Language Models in Plant Science. [PDF]
Sakib SN, Haque N, Hossain MZ, Arman SE.
europepmc +1 more source
GME-Init: Gamma-Moment Equalization for LoRA Initialization in Parameter-Efficient Fine-Tuning. [PDF]
Lin Y, Li C, Shen Z, Liu J, Zeng M.
europepmc +1 more source
Scene graph-guided uncertainty decomposition improves confidence calibration in surgical visual question answering. [PDF]
Song J +9 more
europepmc +1 more source
A multimodal vision-language model for comprehensive dental diagnosis and enhanced clinical practice. [PDF]
Meng Z +22 more
europepmc +1 more source

