Results 61 to 70 of about 61,033 (257)

Meta-tracking for video scene understanding [PDF]

open access: yes2013 10th IEEE International Conference on Advanced Video and Signal Based Surveillance, 2013
This paper presents a novel method to extract dominant motion patterns (MPs) and the main entry/exit areas from a surveillance video. The method first computes motion histograms for each pixel and then converts it into orientation distribution functions (ODFs).
Jodoin, Pierre-Marc   +2 more
openaire   +2 more sources

Unification of Road Scene Segmentation Strategies Using Multistream Data and Latent Space Attention

open access: yesSensors, 2023
Road scene understanding, as a field of research, has attracted increasing attention in recent years. The development of road scene understanding capabilities that are applicable to real-world road scenarios has seen numerous complications.
August J. Naudé, Herman C. Myburgh
doaj   +1 more source

Liquid Crystal Elastomer‐Based Haptic Pixel Arrays at Your Fingertips for Advanced Human–Machine Interfaces

open access: yesAdvanced Materials, EarlyView.
We mechanically program liquid crystal elastomer coatings as a stimuli‐responsive and digitally addressable platform for refreshable tactile displays; we validate the perceptual performance of our dynamic device matches conventional static tactile media.
Tom Bruining   +9 more
wiley   +1 more source

Attention Mechanism-Based Cognition-Level Scene Understanding

open access: yesInformation
Given a question–image input, a visual commonsense reasoning (VCR) model predicts an answer with a corresponding rationale, which requires inference abilities based on real-world knowledge.
Xuejiao Tang, Wenbin Zhang
doaj   +1 more source

Grounding Large Language Models for Robot Task Planning Using Closed‐Loop State Feedback

open access: yesAdvanced Robotics Research, EarlyView.
BrainBody‐Large Language Model (LLM) introduces a hierarchical, feedback‐driven planning framework where two LLMs coordinate high‐level reasoning and low‐level control for robotic tasks. By grounding decisions in real‐time state feedback, it reduces hallucinations and improves task reliability.
Vineet Bhat   +4 more
wiley   +1 more source

Improving the Robustness of Visual Teach‐and‐Repeat Navigation Using Drift Error Correction and Event‐Based Vision for Low‐Light Environments

open access: yesAdvanced Robotics Research, EarlyView.
Visual teach‐and‐repeat (VTR) navigation allows robots to learn and follow routes without building a full metric map. We show that navigation accuracy for VTR can be improved by integrating a topological map with error‐drift correction based on stereo vision.
Fuhai Ling, Ze Huang, Tony J. Prescott
wiley   +1 more source

A Self‐Healing Permanent Magnet Putty for Soft Robot Skins With Force Sensing and Functional Recovery

open access: yesAdvanced Robotics Research, EarlyView.
Permanent magnet putty (PMP) integrates high‐coercivity NdFeB particles with a dynamic polyborosiloxane–Ecoflex matrix, achieving rapid self‐healing (90% mechanical recovery in 10 s) and magnetic recovery within 20 min. With twice the sensitivity of commercial putties, PMP enables precise 5–30 N force detection and discrimination between pressing and ...
Ruotong Zhao   +5 more
wiley   +1 more source

Referring Remote Sensing Image Segmentation method based on Scene-Aware Guided Network model

open access: yesInternational Journal of Applied Earth Observations and Geoinformation
Referring Remote Sensing Image Segmentation (RRSIS) aims to achieve accurate segmentation of objects in remote sensing images under the guidance of natural language expression.
Kai Tan   +6 more
doaj   +1 more source

Multimodal Scene Editing Algorithm Integrating CLIP and 3D Gaussian [PDF]

open access: yesZhengzhou Daxue xuebao. Gongxue ban
To address the issues of excessive reliance on annotated data and high computational complexity in 3D scene editing algorithms, in this study a multimodal scene editing method named CLIP2Gaussian was proposed, which integrated CLIP with 3D Gaussian ...
CAO Yangjie   +4 more
doaj   +1 more source

Multimodal Human–Robot Interaction Using Human Pose Estimation and Local Large Language Models

open access: yesAdvanced Robotics Research, EarlyView.
A multimodal human–robot interaction framework integrates human pose estimation (HPE) and a large language model (LLM) for gesture‐ and voice‐based robot control. Speech‐to‐text (STT) enables voice command interpretation, while a safety‐aware arbitration mechanism prioritizes gesture input for rapid intervention.
Nasiru Aboki   +2 more
wiley   +1 more source

Home - About - Disclaimer - Privacy