Results 111 to 120 of about 15,745,009 (280)
Visual quality assessment is entering a new frontier as media evolve from static images to temporally dynamic videos and 3D content. These visual signals are typically captured by sensing devices such as cameras and depth sensors, whose acquisition ...
Qihang Ge +4 more
doaj +1 more source
MoME: Mixture of Multimodal Experts for Generalist Multimodal Large Language Models
Multimodal large language models (MLLMs) have demonstrated impressive capabilities across various vision-language tasks. However, a generalist MLLM typically underperforms compared with a specialist MLLM on most VL tasks, which can be attributed to task interference.
Leyang Shen +4 more
openaire +3 more sources
Eliciting Metaknowledge in Large Language Models
Advances in Natural Language Processing led to the introduction of Large Language Models (LLMs), that have been found endowed of enriched capabilities and improved performance results when increased in size.
Misael Mongiovì +3 more
core
The Future of Research in Cognitive Robotics: Foundation Models or Developmental Cognitive Models?
Research in cognitive robotics founded on principles of developmental psychology and enactive cognitive science would yield what we seek in autonomous robots: the ability to perceive its environment, learn from experience, anticipate the outcome of events, act to pursue goals, and adapt to changing circumstances without resorting to training with ...
David Vernon
wiley +1 more source
Continual Learning for Multimodal Data Fusion of a Soft Gripper
Models trained on a single data modality often struggle to generalize when exposed to a different modality. This work introduces a continual learning algorithm capable of incrementally learning different data modalities by leveraging both class‐incremental and domain‐incremental learning scenarios in an artificial environment where labeled data is ...
Nilay Kushawaha, Egidio Falotico
wiley +1 more source
A Survey on Multimodal Large Language Models for Autonomous Driving
With the emergence of Large Language Models (LLMs) and Vision Foundation Models (VFMs), multimodal AI systems benefiting from large models have the potential to equally perceive the real world, make decisions, and control tools as humans.
Ma, Yunsheng +20 more
core
Auditory–Tactile Congruence for Synthesis of Adaptive Pain Expressions in RoboPatients
In this work, we explore auditory–tactile congruence for synthesizing adaptive vocal pain expressions in robopatients. Using a robopatient platform that integrates vocal pain sounds with palpation forces, we conducted 7680 trials across 20 participants.
Saitarun Nadipineni +4 more
wiley +1 more source
Exploring Multimodal Prompt for Visualization Authoring with Large Language Models
Here is the repository for paper "Exploring Multimodal Prompt for Visualization Authoring with Large Language ...
Wen Zhen
core +1 more source
Multimodal Romanian language resources and tools: challenges and perspectives
In recent years, research into multimodal resources, models, and tools has seen significant advancements, primarily due to the development of deep artificial neural network architectures. Even though large foundation models exist for different multimodal
Maria Mitrofan +2 more
doaj +1 more source
Large language models in the management of chronic ocular diseases: a scoping review
Large language models, a cutting-edge technology in artificial intelligence, are reshaping the new paradigm of chronic ocular diseases management. In this study, we comprehensively examined the current status and trends in the application of large ...
Jiatong Zhang +6 more
doaj +1 more source

