Results 111 to 120 of about 4,486 (211)
Large Language Models (LLM) are increasingly multimodal, and Zero-Shot Visual Question Answering (VQA) shows promise for image interpretation. If zero-shot VQA can be applied to a 12-lead electrocardiogram (ECG), a prevalent diagnostic tool in the ...
Tomohisa Seki +7 more
doaj +1 more source
Vehicle Feature VQA: Visual Question Answering for Vehicle Feature [PDF]
Visual Question Answering (VQA) can automatically produce the predict answers for questions and real-world images. In this paper, we propose the VQA dataset for Vehicle Feature to know the knowledge of Vehicle. We develop the VQA model using RestNet50 in
Khin Mar Soe, Pa Pa Tun
core
Optimizing feature pooling and prediction models of VQA algorithms
International audienceIn this paper, we propose a strategy to optimize feature pooling and prediction models of video quality assessment (VQA) algorithms with a much smaller number of parameters than methods based on machine learning, such as neural ...
Le Callet, Patrick +4 more
core +1 more source
Diffusion models for text-to-image (T2I) synthesis, e.g. Stable Diffusion, generate visually realistic images, but often fail to capture the fine-grained semantic nuances of complex prompts.
Debashis Bhowmikdebashis Bhowmik +2 more
doaj +1 more source
Hierarchical Modeling for Medical Visual Question Answering with Cross-Attention Fusion
Medical Visual Question Answering (Med-VQA) is designed to accurately answer medical questions by analyzing medical images when given both a medical image and its corresponding clinical question.
Junkai Zhang, Bin Li, Shoujun Zhou
doaj +1 more source
Visual Question Answering for Intelligent Communication Systems: A Systematic Review
Visual Question Answering (VQA) has emerged as a transformative multimodal artificial intelligence paradigm that integrates computer vision and natural language processing to enable intelligent systems to comprehend visual content and respond to natural ...
Merve Gullu, Necaattin Barisci
doaj +1 more source
This paper introduces a novel framework for Visual Question Answering (VQA) that combines Graph Attention Networks (GATs) with Transformers to improve visual-semantic reasoning. The proposed method constructs question-conditioned visual graphs to capture
Hassan Nazeer Chaudhry +2 more
doaj +1 more source
Surgical-VQA: Visual Question Answering in Surgical Scenes using Transformer
Visual question answering (VQA) in surgery is largely unexplored. Expert surgeons are scarce and are often overloaded with clinical and academic workloads.
Islam, Mobarakol +3 more
core
C3-VQA: Cryogenic Counter-Based Coprocessor for Variational Quantum Algorithms
Cryogenic quantum computers play a leading role in demonstrating quantum advantage. Given the severe constraints on the cooling capacity in cryogenic environments, thermal design is crucial for the scalability of these computers.
Yosuke Ueno +7 more
doaj +1 more source
Visual Question Answering Bahasa Indonesia Berbasis Deep Learning untuk Pembelajaran Visual Anak TK
Indonesia semakin gencar melakukan persiapan transformasi digital dalam berbagai sektor, termasuk dalam bidang pendidikan. Salah satu upaya yang dilakukan pemerintah adalah dengan mengimplementasikan platform e-learning dalam kegiatan belajar mengajar ...
Asiyah Hanifah +2 more
doaj +1 more source

