Results 61 to 70 of about 4,895,571 (286)
Visual Storytelling Based on Planning Learning [PDF]
Visual storytelling is a growing area of interest for scholars in computer vision and natural language processing.Current models concentrate on enhancing image representation,like using external knowledge and scene diagrams.Despite some advancements have
WANG Yuanlong, ZHANG Ningqian, ZHANG Hu
doaj +1 more source
Multimodal Encoder-Decoder Attention Networks for Visual Question Answering
Visual Question Answering (VQA) is a multimodal task involving Computer Vision (CV) and Natural Language Processing (NLP), the goal is to establish a high-efficiency VQA model.
Chongqing Chen, Dezhi Han, Jun Wang
doaj +1 more source
Answer-Type Prediction for Visual Question Answering
Recently, algorithms for object recognition and related tasks have become sufficiently proficient that new vision tasks can now be pursued. In this paper, we build a system capable of answering open-ended text-based questions about images, which is known as Visual Question Answering (VQA).
Kafle, Kushal, Kanan, Christopher
openaire +3 more sources
Visual Question Answering on 360° Images [PDF]
Accepted to WACV ...
Shih-Han Chou +4 more
openaire +3 more sources
Questioning the Stability of Visual Question Answering
Visual Language Models (VLMs) have achieved remarkable progress, yet their reliability under small, meaning-preserving input changes remains poorly understood. We present the first large-scale, systematic study of VLM robustness to benign visual and textual perturbations: pixel-level shifts, light geometric transformations, padded rescaling ...
Amir Rosenfeld +2 more
openaire +2 more sources
Privacy Preserving Visual Question Answering
We introduce a novel privacy-preserving methodology for performing Visual Question Answering on the edge. Our method constructs a symbolic representation of the visual scene, using a low-complexity computer vision model that jointly predicts classes, attributes and predicates. This symbolic representation is non-differentiable, which means it cannot be
Cristian-Paul Bara +5 more
openaire +3 more sources
Aesthetic Visual Question Answering of Photographs
Aesthetic assessment of images can be categorized into two main forms: numerical assessment and language assessment. Aesthetics caption of photographs is the only task of aesthetic language assessment that has been addressed. In this paper, we propose a new task of aesthetic language assessment: aesthetic visual question and answering (AVQA) of images.
Xin Jin 0015 +6 more
openaire +2 more sources
Design and analysis strategies for robust microbiome ageing research
The gut microbiome changes with age and associates with age‐related morbidity and mortality, establishing it as a potential biomarker and intervention target for ageing. Realising this potential requires methodological rigour, yet distinguishing biological signals from methodological artefacts remains challenging across cohorts. This review provides an
Mark Olenik +5 more
wiley +1 more source
Investigating transcription factor dynamics in health and disease using FRAP
FRAP analysis of GFP‐tagged transcription factors reveals how molecular mobility and target engagement change in response to drug treatment. By combining live‐cell imaging, quantitative model fitting, and statistical analysis, this approach uncovers transcription factor dynamics linked to disease mechanisms, providing a powerful framework for ...
Kannan Govindaraj +3 more
wiley +1 more source
Pretrained model parameters and pregenerated evaluation data for our visual analysis system for scene-graph-based visual question answering (https://doi.org/10.18419/darus-3589)
Weiskopf, Daniel +6 more
core +1 more source

