Results 131 to 140 of about 15,745,009 (280)
Open-vocabulary detection (OVD) aims to detect and classify objects from an unrestricted set of categories, including those unseen during training. Existing open-vocabulary detectors often suffer from visual-textual misalignment and long-tailed category ...
Caixiong Li +5 more
doaj +1 more source
Improving Visual Storytelling with Multimodal Large Language Models
Visual storytelling is an emerging field that combines images and narratives to create engaging and contextually rich stories. Despite its potential, generating coherent and emotionally resonant visual stories remains challenging due to the complexity of aligning visual and textual information.
Xiaochuan Lin, Xiangyong Chen
openaire +3 more sources
Concept Drift in Large Language Models: Challenges of Evolving Language, Contexts, and the Web
Deep learning models including large language models (LLMs) suffer from concept drift for various reasons and causes. A drift can occur either during data collection due to changes in data distribution and data imbalance, or due to external factors such ...
Thierry J. Chaussalet +3 more
core +1 more source
Learning‐Based Soft Robotic Grasping: Recent Progress and Remaining Challenges
This review analyzes learning‐based soft robotic grasping from a pipeline‐oriented perspective, encompassing soft gripper design, multimodal sensing, and learning‐based planning and control. It surveys key neural network architectures and benchmark datasets and identifies critical challenges such as sim‐to‐real transfer, generalization, and continual ...
Arnab Majumder +3 more
wiley +1 more source
Exploring Multimodal Perception in Large Language Models Through Perceptual Strength Ratings
This study investigated whether multimodal large language models can achieve human-like sensory grounding by examining their ability to capture perceptual strength ratings across sensory modalities. We explored how model characteristics (size, multimodal
Jonghyun Lee +4 more
doaj +1 more source
Multimodal Large Language Models for Medicine: A Comprehensive Survey
MLLMs have recently become a focal point in the field of artificial intelligence research. Building on the strong capabilities of LLMs, MLLMs are adept at addressing complex multi-modal tasks. With the release of GPT-4, MLLMs have gained substantial attention from different domains.
Jiarui Ye, Hao Tang 0005
openaire +3 more sources
Probing the Sequential Enumeration Skills of Large Language Models [PDF]
openNumerosity estimation, which we share with many animal species, is a cornerstone in human cognitive development leading to higher mathematical competencies.
DAIBASOGLU, KAAN
core
Can We Edit Multimodal Large Language Models? [PDF]
In this paper, we focus on editing Multimodal Large Language Models (MLLMs). Compared to editing single-modal LLMs, multimodal model editing is more challenging, which demands a higher level of scrutiny and careful consideration in the editing process ...
Tian, Bozhong +6 more
core +1 more source
Flexible Sensors for Robotics Tactile Perception: A Review
Flexible tactile sensing for robotics is reviewed through four interconnected dimensions. Physical mechanisms include piezoresistive, capacitive, piezoelectric, triboelectric, iontronic, and optical sensing. Structural design includes bioinspired, defect‐based, and MEMS‐based tactile systems.
Yu Song, Ying Chen, Yihao Chen, Xue Feng
wiley +1 more source
Automating Steering for Safe Multimodal Large Language Models
EMNLP 2025 Main Conference.
Lyucheng Wu +6 more
openaire +4 more sources

