Results 121 to 130 of about 15,745,009 (280)

LITE: Modeling Environmental Ecosystems with Multimodal Large Language Models

open access: yesCoRR
The modeling of environmental ecosystems plays a pivotal role in the sustainable management of our planet. Accurate prediction of key environmental variables over space and time can aid in informed policy and decision-making, thus improving people's livelihood.
Haoran Li 0011   +5 more
openaire   +3 more sources

Multimodal Engagement Assessment in Children During Invented Story Paradigm With a Social Robot

open access: yesAdvanced Robotics Research, EarlyView.
A multimodal framework is proposed to assess children's engagement during storytelling interactions with a social robot. Gaze, physiological, and behavioral data are combined and validated against observer ratings. An automated gaze‐labeling strategy is introduced, and supervised classifiers achieve high accuracy. The study supports scalable engagement
Laura Fiorini   +7 more
wiley   +1 more source

A Deep Hybrid Recommendation Method for Multimodal Information Integrating Content Generated by Large Language Models

open access: yesInformation
Item description information plays a crucial role in helping users understand the basic situation of an item and is also vital auxiliary information in recommendation systems.
Chao Duan   +5 more
doaj   +1 more source

Multimodal Human–Robot Interaction Using Human Pose Estimation and Local Large Language Models

open access: yesAdvanced Robotics Research, EarlyView.
A multimodal human–robot interaction framework integrates human pose estimation (HPE) and a large language model (LLM) for gesture‐ and voice‐based robot control. Speech‐to‐text (STT) enables voice command interpretation, while a safety‐aware arbitration mechanism prioritizes gesture input for rapid intervention.
Nasiru Aboki   +2 more
wiley   +1 more source

Evaluating vision-capable chatbots in interpreting kinematics graphs: a comparative study of free and subscription-based models

open access: yesFrontiers in Education
This study investigates the performance of eight large multimodal model (LMM)-based chatbots on the Test of Understanding Graphs in Kinematics (TUG-K), a research-based concept inventory. Graphs are a widely used representation in STEM and medical fields,
Giulia Polverini, Bor Gregorcic
doaj   +1 more source

A Refer-and-Ground Multimodal Large Language Model for Biomedicine

open access: yes
With the rapid development of multimodal large language models (MLLMs), especially their capabilities in visual chat through refer and ground functionalities, their significance is increasingly recognized. However, the biomedical field currently exhibits a substantial gap in this area, primarily due to the absence of a dedicated refer and ground ...
Xiaoshuang Huang   +6 more
openaire   +3 more sources

LLM‐Integrated Human–Robot Interaction System for Microrobots

open access: yesAdvanced Robotics Research, EarlyView.
This paper proposes an LLM‐based control framework for guiding microrobots using human natural language. This framework can convert the natural human speech into safe and executable command sets for reliable navigation in complex environments. The experimental results show high accuracy and robustness in task performance, demonstrating the potential of
Bairong Zhu, Amar Salehi, Tingting Yu
wiley   +1 more source

Applications of large‐scale artificial intelligence models in bioinformatics

open access: yesQuantitative Biology
Large‐scale artificial intelligence (AI) models can mine potential patterns from massive amounts of data and provide more accurate analyses. This capability has enabled its gradual application in various areas of bioinformatics. However, few reviews have
Mingjing Li   +5 more
doaj   +1 more source

Jailbreaking Attack against Multimodal Large Language Model

open access: yesCoRR
This paper focuses on jailbreaking attacks against multi-modal large language models (MLLMs), seeking to elicit MLLMs to generate objectionable responses to harmful user queries. A maximum likelihood-based algorithm is proposed to find an \emph{image Jailbreaking Prompt} (imgJP), enabling jailbreaks against MLLMs across multiple unseen prompts and ...
Zhenxing Niu   +4 more
openaire   +2 more sources

Intelligent Maintenance Review for Robots: Multimodal Information, Deep Diagnosis and Embodied Artificial Intelligence

open access: yesAdvanced Robotics Research, EarlyView.
This review maps the methods to monitor robots’ health by fusing vibration, sound, control signals, vision, force, and oil information with artificial intelligence. It identifies deep learning, transfer learning, digital twins, and physics‐informed models as key methodological pathways enabling earlier diagnosis, safer human–robot collaboration, and ...
Yuting Qiao   +6 more
wiley   +1 more source

Home - About - Disclaimer - Privacy