Results 51 to 60 of about 1,468,270 (262)
Grounding Large Language Models for Robot Task Planning Using Closed‐Loop State Feedback
BrainBody‐Large Language Model (LLM) introduces a hierarchical, feedback‐driven planning framework where two LLMs coordinate high‐level reasoning and low‐level control for robotic tasks. By grounding decisions in real‐time state feedback, it reduces hallucinations and improves task reliability.
Vineet Bhat +4 more
wiley +1 more source
Action knowledge for video captioning with graph neural networks
Many existing video captioning methods capture action information in the video by exploiting features extracted from an action recognition model. However, directly using the action features without object-specific representation may not well capture the ...
Willy Fitra Hendria +4 more
doaj +1 more source
Intelligent Sky Guardians (InSkyGuard) is introduced as a four‐drone swarm that autonomously detects, tracks, and safely captures rogue drones using a coordinated net system. Computer vision and leader–follower control architecture enable synchronized enclosure, while integrated failsafes enhance system reliability. Validated through closed‐environment
Joshua Hastings +6 more
wiley +1 more source
Adaptive Curriculum Learning for Video Captioning
A portion of the data in video captioning datasets are noisy and unsuitable for models to learn at early stages, e.g., there could be a generic 4-word-long caption lacking distinctive details of video content and a 19-word-long description with rare ...
Shanhao Li, Bang Yang, Yuexian Zou
doaj +1 more source
Building correspondences across different modalities, such as video and language, has recently become critical in many visual recognition applications, such as video captioning. Inspired by machine translation, recent models tackle this task using an encoder-decoder strategy.
Silvio Olivastri +2 more
openaire +3 more sources
Poly(vinyl alcohol)‐modified liquid‐metal particles form conductive networks through spray‐induced self‐fusion without additional post‐treatment. Spray impact initiates oxide‐shell rupture and electrical contact, while subsequent Marangoni flow and capillary forces promote particle redistribution, neck growth, and network consolidation. This sequential
Keun‐Young Yook +13 more
wiley +1 more source
Enhanced Hybrid Framework and Comparative Analysis of Deep Learning Architectures for Video Captioning [PDF]
With the rapid development of multimedia content on digital platforms, there is more and more need for intelligent systems to understand and describe videos in natural language.
Bhusare Pranali Prabhakar +1 more
doaj +1 more source
SoccerNet-Caption: Dense Video Captioning for Soccer Broadcasts Commentaries
Soccer is more than just a game - it is a passion that transcends borders and unites people worldwide. From the roar of the crowds to the excitement of the commentators, every moment of a soccer match is a thrill. Yet, with so many games happening simultaneously, fans cannot watch them all live.
Mkhallati, Hassan +4 more
openaire +3 more sources
Smart Exploration of Perovskite Photovoltaics: From AI Driven Discovery to Autonomous Laboratories
In this review, we summarize the fundamentals of AI in automated materials science, and review AI applications in perovskite solar cells. Then, we sum up recent progress in AI‐guided manufacturing optimization, and highlight AI‐driven high‐throughput and autonomous laboratories.
Wenning Chen +4 more
wiley +1 more source
Video captioning serves as a tool which enables computers to understand visual content and convert video footage into written language. The current CNN-RNN video captioning system encounters difficulties when it tries to show both term timelines and ...
Reshma DSouza, Snigdha Sen
doaj +1 more source

