Results 31 to 40 of about 4,486 (211)
Towards VQA Models That Can Read [PDF]
Studies have shown that a dominant class of questions asked by visually impaired users on images of their surroundings involves reading text in the image. But today's VQA models can not read! Our paper takes a first step towards addressing this problem.
Amanpreet Singh +7 more
openaire +3 more sources
VMAF and variants: towards a unified VQA [PDF]
Some calculational errors have been fixed in this ...
Topiwala, Pankaj +4 more
openaire +2 more sources
Supervising the Transfer of Reasoning Patterns in VQA
Methods for Visual Question Anwering (VQA) are notorious for leveraging dataset biases rather than performing reasoning, hindering generalization. It has been recently shown that better reasoning patterns emerge in attention layers of a state-of-the-art VQA model when they are trained on perfect (oracle) visual inputs.
Corentin Kervadec +4 more
openaire +3 more sources
No-Reference Video Quality Assessment Using Distortion Learning and Temporal Attention
The rapid growth of video consumption and multimedia applications has increased the interest of the academia and industry in building tools that can evaluate perceptual video quality. Since videos might be distorted when they are captured or transmitted,
Koffi Kossi +3 more
doaj +1 more source
RankDVQA: Deep VQA based on Ranking-inspired Hybrid Training [PDF]
In recent years, deep learning techniques have shown significant potential for improving video quality assessment (VQA), achieving higher correlation with subjective opinions compared to conventional approaches.
Bull, David +5 more
core +1 more source
Multiple-Question Multiple-Answer Text-VQA [PDF]
We present Multiple-Question Multiple-Answer (MQMA), a novel approach to do text-VQA in encoder-decoder transformer models. The text-VQA task requires a model to answer a question by understanding multi-modal content: text (typically from OCR) and an ...
Tang, Peng +4 more
core +1 more source
Medical Visual Question Answering Based on Cross-Modal Attention Feature Enhancement [PDF]
Medical Visual Question Answering (Med-VQA) requires an understanding of content related to both medical images and text-based questions. Therefore, designing effective modal representations and cross-modal fusion methods is crucial for performing well ...
LIU Kai, REN Hongyi, LI Ying, JI Yi, LIU Chunping
doaj +1 more source
Continual VQA for Disaster Response Systems
Accepted at Tackling Climate Change with Machine Learning workshop at NeurIPS ...
Aditya Kane, V. Manushree, Sahil Khose
openaire +2 more sources
Video quality assessment using motion-compensated temporal filtering and manifold feature similarity. [PDF]
Well-performed Video quality assessment (VQA) method should be consistent with human visual systems for better prediction accuracy. In this paper, we propose a VQA method using motion-compensated temporal filtering (MCTF) and manifold feature similarity.
Yang Song +4 more
doaj +1 more source
Distraction-free Embeddings for Robust VQA
The generation of effective latent representations and their subsequent refinement to incorporate precise information is an essential prerequisite for Vision-Language Understanding (VLU) tasks such as Video Question Answering (VQA). However, most existing methods for VLU focus on sparsely sampling or fine-graining the input information (e.g., sampling ...
Atharvan Dogra +4 more
openaire +3 more sources

