Results 81 to 90 of about 4,486 (211)
The importance of maintaining consistent visual quality has increased significantly due to the rapid advancement of video acquisition technologies, high-speed Internet connectivity, the widespread adoption of social media platforms, and the growing ...
Anish Kumar Vishwakarma +3 more
doaj +1 more source
Crayons × Code: Re‐Exploring Children's Drawings With Multimodal AI and Dynamic Ethics
ABSTRACT Children's drawings are used to study their environmental perception, but AI‐powered analysis of these visual materials remains underexplored. This paper re‐explores the use of children's drawings as a tool for understanding their perceptions of the environment, leveraging multimodal AI for analysis.
Chen Qu
wiley +1 more source
Collaborative Modality Fusion for Mitigating Language Bias in Visual Question Answering
Language bias stands as a noteworthy concern in visual question answering (VQA), wherein models tend to rely on spurious correlations between questions and answers for prediction.
Qiwen Lu, Shengbo Chen, Xiaoke Zhu
doaj +1 more source
Abstract Purpose This study aimed to quantitatively evaluate the efficacy of plan adaptation in stereotactic body proton therapy (SBPT) for pancreatic cancer using daily CT‐based dose evaluation, investigate the appropriate adaptation frequency. Methods This retrospective planning study included 10 patients previously treated with X‐ray stereotactic ...
Yuto Matsuo +8 more
wiley +1 more source
Splattalk: 3D VQA with Gaussian Splatting
Language-guided 3D scene understanding is important for advancing applications in robotics, AR/VR, and human-computer interaction, enabling models to comprehend and interact with 3D environments through natural language. While 2D vision-language models (VLMs) have achieved remarkable success in 2D VQA tasks, progress in the 3D domain has been ...
Anh Thai +4 more
openaire +2 more sources
Artificial intelligence (AI) is reshaping autonomous mobile robot navigation beyond classical pipelines. This review analyzes how AI techniques are integrated into core navigation tasks, including path planning and control, localization and mapping, perception, and context‐aware decision‐making. Learning‐based, probabilistic, and soft‐computing methods
Giovanna Guaragnella +5 more
wiley +1 more source
Previous works employ the Large Language Model (LLM) like GPT-3 for knowledge-based Visual Question Answering (VQA). We argue that the inferential capacity of LLM can be enhanced through knowledge injection. Although methods that utilize knowledge graphs
Zhongjian Hu +4 more
doaj +1 more source
Explainable Artificial Intelligence Through the Lens of Bibliometric Citation Analysis
ABSTRACT This study provides a comprehensive bibliometric analysis of the development of Explainable Artificial Intelligence (XAI) research from 1993 to 2024. The objective is to explore key contributors, thematic trends, and the evolution of methodologies within the field.
Mariateresa Russo, Domenico Vistocco
wiley +1 more source
Deep Exemplar Networks for VQA and VQG
In this paper, we consider the problem of solving semantic tasks such as `Visual Question Answering' (VQA), where one aims to answers related to an image and `Visual Question Generation' (VQG), where one aims to generate a natural question pertaining to an image.
Badri N. Patro, Vinay P. Namboodiri
openaire +3 more sources
A Self-supervised Strategy for the Robustness of VQA Models
Part 6: Game Theory and EmotionInternational audienceIn visual question answering (VQA), most existing models suffer from language biases which make models not robust.
Jing, Chenchen +3 more
core +1 more source

