Results 11 to 20 of about 15,745,009 (280)
Survey of Security Research on Multimodal Large Language Models [PDF]
With the rapid development of large language models,multimodal large language models have garnered attention for their outstanding performance across various modalities,such as language and images.These models have not only become valuable assistants in ...
CHEN Jinyin, XI Changkun, ZHENG Haibin, GAO Ming, ZHANG Tianxin
doaj +2 more sources
Harnessing Multimodal Large Language Models for Multimodal Sequential Recommendation
Recent advances in Large Language Models (LLMs) have demonstrated significant potential in the field of Recommendation Systems (RSs). Most existing studies have focused on converting user behavior logs into textual prompts and leveraging techniques such as prompt tuning to enable LLMs for recommendation tasks. Meanwhile, research interest has recently
Yuyang Ye 0002 +8 more
core +10 more sources
Personalized Multimodal Large Language Models: A Survey
Multimodal Large Language Models (MLLMs) have become increasingly important due to their state-of-the-art performance and ability to integrate multiple data modalities, such as text, images, and audio, to perform complex tasks with high accuracy. This paper presents a comprehensive survey on personalized multimodal large language models, focusing on ...
Junda Wu +26 more
core +5 more sources
Leveraging Large Language Models for Multimodal Search
Published at CVPRW ...
Barbany Mayor, Oriol +3 more
core +6 more sources
Attention re-alignment in multimodal large language models via intermediate-layer guidance [PDF]
Multimodal large language models (MLLMs) have achieved impressive performance in understanding and describing visual content, setting new state-of-the-art results on a variety of visual question answering (VQA) benchmarks. However, during decoding, these
Yanming Chen +5 more
doaj +2 more sources
Adaptive diagnostic reasoning framework for pathology with multimodal large language models [PDF]
Background Artificial intelligence enhances pathology screening efficiency, yet clinical adoption remains limited because most systems operate as opaque black boxes.
Yunqi Hong +7 more
doaj +2 more sources
Woodpecker: hallucination correction for multimodal large language models
Hallucination is a big shadow hanging over the rapidly evolving Multimodal Large Language Models (MLLMs), referring to the phenomenon that the generated text is inconsistent with the image content. In order to mitigate hallucinations, existing studies mainly resort to an instruction-tuning manner that requires retraining the models with specific data ...
Xing Sun, Enhong Chen
exaly +5 more sources
Multimodal Relation Extraction (MRE) is a core task for constructing Multimodal Knowledge images (MKGs). Most current research is based on fine-tuning small-scale single-modal image and text pre-trained models, but we find that image-text datasets from ...
Wentao He +5 more
doaj +1 more source
The application of multimodal large language models in medicine
Jianing Qiu, Wu Yuan, Kyle Lam
doaj +3 more sources
Multimodal Large Language Model for Visual Navigation
Recent efforts to enable visual navigation using large language models have mainly focused on developing complex prompt systems. These systems incorporate instructions, observations, and history into massive text prompts, which are then combined with pre-trained large language models to facilitate visual navigation.
Yao-Hung Hubert Tsai +4 more
openaire +3 more sources

