Results 131 to 140 of about 15,745,009 (280)

MQADet: a plug-and-play paradigm for enhancing open-vocabulary object detection via multimodal question answering

open access: yesScientific Reports
Open-vocabulary detection (OVD) aims to detect and classify objects from an unrestricted set of categories, including those unseen during training. Existing open-vocabulary detectors often suffer from visual-textual misalignment and long-tailed category ...
Caixiong Li   +5 more
doaj   +1 more source

Improving Visual Storytelling with Multimodal Large Language Models

open access: yesCoRR
Visual storytelling is an emerging field that combines images and narratives to create engaging and contextually rich stories. Despite its potential, generating coherent and emotionally resonant visual stories remains challenging due to the complexity of aligning visual and textual information.
Xiaochuan Lin, Xiangyong Chen
openaire   +3 more sources

Concept Drift in Large Language Models: Challenges of Evolving Language, Contexts, and the Web

open access: yes
Deep learning models including large language models (LLMs) suffer from concept drift for various reasons and causes. A drift can occur either during data collection due to changes in data distribution and data imbalance, or due to external factors such ...
Thierry J. Chaussalet   +3 more
core   +1 more source

Learning‐Based Soft Robotic Grasping: Recent Progress and Remaining Challenges

open access: yesAdvanced Robotics Research, EarlyView.
This review analyzes learning‐based soft robotic grasping from a pipeline‐oriented perspective, encompassing soft gripper design, multimodal sensing, and learning‐based planning and control. It surveys key neural network architectures and benchmark datasets and identifies critical challenges such as sim‐to‐real transfer, generalization, and continual ...
Arnab Majumder   +3 more
wiley   +1 more source

Exploring Multimodal Perception in Large Language Models Through Perceptual Strength Ratings

open access: yesIEEE Access
This study investigated whether multimodal large language models can achieve human-like sensory grounding by examining their ability to capture perceptual strength ratings across sensory modalities. We explored how model characteristics (size, multimodal
Jonghyun Lee   +4 more
doaj   +1 more source

Multimodal Large Language Models for Medicine: A Comprehensive Survey

open access: yesCoRR
MLLMs have recently become a focal point in the field of artificial intelligence research. Building on the strong capabilities of LLMs, MLLMs are adept at addressing complex multi-modal tasks. With the release of GPT-4, MLLMs have gained substantial attention from different domains.
Jiarui Ye, Hao Tang 0005
openaire   +3 more sources

Probing the Sequential Enumeration Skills of Large Language Models [PDF]

open access: yes
openNumerosity estimation, which we share with many animal species, is a cornerstone in human cognitive development leading to higher mathematical competencies.
DAIBASOGLU, KAAN
core  

Can We Edit Multimodal Large Language Models? [PDF]

open access: yes
In this paper, we focus on editing Multimodal Large Language Models (MLLMs). Compared to editing single-modal LLMs, multimodal model editing is more challenging, which demands a higher level of scrutiny and careful consideration in the editing process ...
Tian, Bozhong   +6 more
core   +1 more source

Flexible Sensors for Robotics Tactile Perception: A Review

open access: yesAdvanced Robotics Research, EarlyView.
Flexible tactile sensing for robotics is reviewed through four interconnected dimensions. Physical mechanisms include piezoresistive, capacitive, piezoelectric, triboelectric, iontronic, and optical sensing. Structural design includes bioinspired, defect‐based, and MEMS‐based tactile systems.
Yu Song, Ying Chen, Yihao Chen, Xue Feng
wiley   +1 more source

Automating Steering for Safe Multimodal Large Language Models

open access: yesProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing
EMNLP 2025 Main Conference.
Lyucheng Wu   +6 more
openaire   +4 more sources

Home - About - Disclaimer - Privacy