Results 61 to 70 of about 61,033 (257)
Meta-tracking for video scene understanding [PDF]
This paper presents a novel method to extract dominant motion patterns (MPs) and the main entry/exit areas from a surveillance video. The method first computes motion histograms for each pixel and then converts it into orientation distribution functions (ODFs).
Jodoin, Pierre-Marc +2 more
openaire +2 more sources
Unification of Road Scene Segmentation Strategies Using Multistream Data and Latent Space Attention
Road scene understanding, as a field of research, has attracted increasing attention in recent years. The development of road scene understanding capabilities that are applicable to real-world road scenarios has seen numerous complications.
August J. Naudé, Herman C. Myburgh
doaj +1 more source
We mechanically program liquid crystal elastomer coatings as a stimuli‐responsive and digitally addressable platform for refreshable tactile displays; we validate the perceptual performance of our dynamic device matches conventional static tactile media.
Tom Bruining +9 more
wiley +1 more source
Attention Mechanism-Based Cognition-Level Scene Understanding
Given a question–image input, a visual commonsense reasoning (VCR) model predicts an answer with a corresponding rationale, which requires inference abilities based on real-world knowledge.
Xuejiao Tang, Wenbin Zhang
doaj +1 more source
Grounding Large Language Models for Robot Task Planning Using Closed‐Loop State Feedback
BrainBody‐Large Language Model (LLM) introduces a hierarchical, feedback‐driven planning framework where two LLMs coordinate high‐level reasoning and low‐level control for robotic tasks. By grounding decisions in real‐time state feedback, it reduces hallucinations and improves task reliability.
Vineet Bhat +4 more
wiley +1 more source
Visual teach‐and‐repeat (VTR) navigation allows robots to learn and follow routes without building a full metric map. We show that navigation accuracy for VTR can be improved by integrating a topological map with error‐drift correction based on stereo vision.
Fuhai Ling, Ze Huang, Tony J. Prescott
wiley +1 more source
Permanent magnet putty (PMP) integrates high‐coercivity NdFeB particles with a dynamic polyborosiloxane–Ecoflex matrix, achieving rapid self‐healing (90% mechanical recovery in 10 s) and magnetic recovery within 20 min. With twice the sensitivity of commercial putties, PMP enables precise 5–30 N force detection and discrimination between pressing and ...
Ruotong Zhao +5 more
wiley +1 more source
Referring Remote Sensing Image Segmentation method based on Scene-Aware Guided Network model
Referring Remote Sensing Image Segmentation (RRSIS) aims to achieve accurate segmentation of objects in remote sensing images under the guidance of natural language expression.
Kai Tan +6 more
doaj +1 more source
Multimodal Scene Editing Algorithm Integrating CLIP and 3D Gaussian [PDF]
To address the issues of excessive reliance on annotated data and high computational complexity in 3D scene editing algorithms, in this study a multimodal scene editing method named CLIP2Gaussian was proposed, which integrated CLIP with 3D Gaussian ...
CAO Yangjie +4 more
doaj +1 more source
Multimodal Human–Robot Interaction Using Human Pose Estimation and Local Large Language Models
A multimodal human–robot interaction framework integrates human pose estimation (HPE) and a large language model (LLM) for gesture‐ and voice‐based robot control. Speech‐to‐text (STT) enables voice command interpretation, while a safety‐aware arbitration mechanism prioritizes gesture input for rapid intervention.
Nasiru Aboki +2 more
wiley +1 more source

