Results 31 to 40 of about 4,486 (211)

Towards VQA Models That Can Read [PDF]

open access: yes2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019
Studies have shown that a dominant class of questions asked by visually impaired users on images of their surroundings involves reading text in the image. But today's VQA models can not read! Our paper takes a first step towards addressing this problem.
Amanpreet Singh   +7 more
openaire   +3 more sources

VMAF and variants: towards a unified VQA [PDF]

open access: yesApplications of Digital Image Processing XLIV, 2021
Some calculational errors have been fixed in this ...
Topiwala, Pankaj   +4 more
openaire   +2 more sources

Supervising the Transfer of Reasoning Patterns in VQA

open access: yesCoRR, 2021
Methods for Visual Question Anwering (VQA) are notorious for leveraging dataset biases rather than performing reasoning, hindering generalization. It has been recently shown that better reasoning patterns emerge in attention layers of a state-of-the-art VQA model when they are trained on perfect (oracle) visual inputs.
Corentin Kervadec   +4 more
openaire   +3 more sources

No-Reference Video Quality Assessment Using Distortion Learning and Temporal Attention

open access: yesIEEE Access, 2022
The rapid growth of video consumption and multimedia applications has increased the interest of the academia and industry in building tools that can evaluate perceptual video quality. Since videos might be distorted when they are captured or transmitted,
Koffi Kossi   +3 more
doaj   +1 more source

RankDVQA: Deep VQA based on Ranking-inspired Hybrid Training [PDF]

open access: yes, 2023
In recent years, deep learning techniques have shown significant potential for improving video quality assessment (VQA), achieving higher correlation with subjective opinions compared to conventional approaches.
Bull, David   +5 more
core   +1 more source

Multiple-Question Multiple-Answer Text-VQA [PDF]

open access: yes, 2023
We present Multiple-Question Multiple-Answer (MQMA), a novel approach to do text-VQA in encoder-decoder transformer models. The text-VQA task requires a model to answer a question by understanding multi-modal content: text (typically from OCR) and an ...
Tang, Peng   +4 more
core   +1 more source

Medical Visual Question Answering Based on Cross-Modal Attention Feature Enhancement [PDF]

open access: yesJisuanji gongcheng
Medical Visual Question Answering (Med-VQA) requires an understanding of content related to both medical images and text-based questions. Therefore, designing effective modal representations and cross-modal fusion methods is crucial for performing well ...
LIU Kai, REN Hongyi, LI Ying, JI Yi, LIU Chunping
doaj   +1 more source

Continual VQA for Disaster Response Systems

open access: yesCoRR, 2022
Accepted at Tackling Climate Change with Machine Learning workshop at NeurIPS ...
Aditya Kane, V. Manushree, Sahil Khose
openaire   +2 more sources

Video quality assessment using motion-compensated temporal filtering and manifold feature similarity. [PDF]

open access: yesPLoS ONE, 2017
Well-performed Video quality assessment (VQA) method should be consistent with human visual systems for better prediction accuracy. In this paper, we propose a VQA method using motion-compensated temporal filtering (MCTF) and manifold feature similarity.
Yang Song   +4 more
doaj   +1 more source

Distraction-free Embeddings for Robust VQA

open access: yesCoRR, 2023
The generation of effective latent representations and their subsequent refinement to incorporate precise information is an essential prerequisite for Vision-Language Understanding (VLU) tasks such as Video Question Answering (VQA). However, most existing methods for VLU focus on sparsely sampling or fine-graining the input information (e.g., sampling ...
Atharvan Dogra   +4 more
openaire   +3 more sources

Home - About - Disclaimer - Privacy