Results 41 to 50 of about 4,486 (211)

Building a Framework for Visual Question Answering Systems

open access: yesSyrian Journal for Science and Innovation
VQA (Visual Question Answering) systems are among the latest advancements in the fields of artificial intelligence and deep learning. They integrate image processing with natural language understanding to enable intelligent systems to answer questions ...
Maya Abu Hamoud, Wasim Safi
doaj   +1 more source

VQA-Levels: A Hierarchical Approach for Classifying Questions in VQA

open access: yesCoRR
Designing datasets for Visual Question Answering (VQA) is a difficult and complex task that requires NLP for parsing and computer vision for analysing the relevant aspects of the image for answering the question asked. Several benchmark datasets have been developed by researchers but there are many issues with using them for methodical performance ...
Madhuri Latha Madaka   +1 more
openaire   +2 more sources

VQA-GEN: A Visual Question Answering Benchmark for Domain Generalization [PDF]

open access: yes, 2023
Visual question answering (VQA) models are designed to demonstrate visual-textual reasoning capabilities. However, their real-world applicability is hindered by a lack of comprehensive benchmark datasets.
Unni, Suraj Jyothi   +2 more
core  

MD-VQA: Multi-Dimensional Quality Assessment for UGC Live Videos [PDF]

open access: yes, 2023
User-generated content (UGC) live videos are often bothered by various distortions during capture procedures and thus exhibit diverse visual qualities.
Sun, Wei   +7 more
core   +1 more source

An Efficient Modern Baseline for FloodNet VQA

open access: yesCoRR, 2022
Designing efficient and reliable VQA systems remains a challenging problem, more so in the case of disaster management and response systems. In this work, we revisit fundamental combination methods like concatenation, addition and element-wise multiplication with modern image and text feature abstraction models.
Aditya Kane, Sahil Khose
openaire   +3 more sources

Unsupervised Keyword Extraction for Full-Sentence VQA [PDF]

open access: yesProceedings of the First International Workshop on Natural Language Processing Beyond Text, 2020
EMNLP 2020 workshop: NLP Beyond Text (NLPBT)
Kohei Uehara, Tatsuya Harada
openaire   +2 more sources

Informed-Learning-Guided Visual Question Answering Model of Crop Disease

open access: yesPlant Phenomics
In contemporary agriculture, experts develop preventative and remedial strategies for various disease stages in diverse crops. Decision-making regarding the stages of disease occurrence exceeds the capabilities of single-image tasks, such as image ...
Yunpeng Zhao   +6 more
doaj   +1 more source

Research on Embodied Intelligence Technology for Electric Power Equipment Based on Large‐Scale Pre‐Trained Models

open access: yesHigh Voltage, EarlyView.
ABSTRACT With the development of electric power artificial intelligence (AI) technology, many complicated challenges emerged during application. Traditional AI algorithms work well in specialised tasks such as detection and classification. However, they are unlikely to solve general problems.
Yuanpeng Tan   +4 more
wiley   +1 more source

Story2Board: A Training‐Free Approach for Expressive Visual Storytelling

open access: yesComputer Graphics Forum, EarlyView.
Abstract We present Story2Board, a training‐free framework for expressive storyboard generation from natural language. Existing methods narrowly focus on subject identity, overlooking key aspects of visual storytelling such as spatial composition, background evolution, and narrative pacing.
D. Dinkevich   +4 more
wiley   +1 more source

C-VQA: A Compositional Split of the Visual Question Answering (VQA) v1.0 Dataset

open access: yesCoRR, 2017
Visual Question Answering (VQA) has received a lot of attention over the past couple of years. A number of deep learning models have been proposed for this task. However, it has been shown that these models are heavily driven by superficial correlations in the training data and lack compositionality -- the ability to answer questions about unseen ...
Aishwarya Agrawal   +3 more
openaire   +2 more sources

Home - About - Disclaimer - Privacy