Results 51 to 60 of about 55,056 (264)
Multimodal Human–Robot Interaction Using Human Pose Estimation and Local Large Language Models
A multimodal human–robot interaction framework integrates human pose estimation (HPE) and a large language model (LLM) for gesture‐ and voice‐based robot control. Speech‐to‐text (STT) enables voice command interpretation, while a safety‐aware arbitration mechanism prioritizes gesture input for rapid intervention.
Nasiru Aboki +2 more
wiley +1 more source
Power Requirements Evaluation of Embedded Devices for Real-Time Video Line Detection
In this paper, the comparison of the power requirements during real-time processing of video sequences in embedded systems was investigated. During the experimental tests, four modules were tested: Raspberry Pi 4B, NVIDIA Jetson Nano, NVIDIA Jetson ...
Jakub Suder +2 more
doaj +1 more source
LLM‐Integrated Human–Robot Interaction System for Microrobots
This paper proposes an LLM‐based control framework for guiding microrobots using human natural language. This framework can convert the natural human speech into safe and executable command sets for reliable navigation in complex environments. The experimental results show high accuracy and robustness in task performance, demonstrating the potential of
Bairong Zhu, Amar Salehi, Tingting Yu
wiley +1 more source
A Practical Probabilistic Benchmark for AI Weather Models
Since the weather is chaotic, it is necessary to forecast an ensemble of future states. Recently, multiple AI weather models have emerged claiming breakthroughs in deterministic skill. Unfortunately, it is hard to fairly compare ensembles of AI forecasts
Noah D. Brenowitz +8 more
doaj +1 more source
Objective. The paper defines the relevance of the task of increasing the efficiency of software, which in this case is understood as reducing the operating time of the designed software in the process of solving computationally complex problems.
A. Yu. Bezruchenko, V. A. Egunov
doaj +1 more source
We introduce Nemotron-Parse-1.1, a lightweight document parsing and OCR model that advances the capabilities of its predecessor, Nemoretriever-Parse-1.0. Nemotron-Parse-1.1 delivers improved capabilities across general OCR, markdown formatting, structured table parsing, and text extraction from pictures, charts, and diagrams.
Kateryna Chumachenko +32 more
openaire +2 more sources
Data‐Driven Bulldozer Blade Control for Autonomous Terrain Leveling
A simulation‐driven framework for autonomous bulldozer leveling is presented, combining high‐fidelity terramechanics simulation with a neural‐network‐based reduced‐order model. Gradient‐based optimization enables efficient, low‐level blade control that balances leveling quality and operation time.
Harry Zhang +5 more
wiley +1 more source
This work presents the MicroRoboScope, a highly integrated, compact, and portable microrobotic experimentation platform combining electromagnetic and acoustic actuation with real‐time visual feedback into a single, end‐to‐end device. The system enables closed‐loop control and tracking algorithm experimentation within an accessible and unified hardware ...
Max Sokolich +4 more
wiley +1 more source
Learning‐Based Soft Robotic Grasping: Recent Progress and Remaining Challenges
This review analyzes learning‐based soft robotic grasping from a pipeline‐oriented perspective, encompassing soft gripper design, multimodal sensing, and learning‐based planning and control. It surveys key neural network architectures and benchmark datasets and identifies critical challenges such as sim‐to‐real transfer, generalization, and continual ...
Arnab Majumder +3 more
wiley +1 more source
Hardware acceleration of ray tracing is an active research field, but only with the release of Nvidia Turing architecture GPUs it became widely available. Nvidia RTX is a proprietary hardware ray tracing acceleration technology available in Vulkan and DirectX APIs as well as through Nvidia OptiX.
Alexey Voloboy +3 more
openaire +1 more source

