Results 51 to 60 of about 1,468,270 (262)

Grounding Large Language Models for Robot Task Planning Using Closed‐Loop State Feedback

open access: yesAdvanced Robotics Research, EarlyView.
BrainBody‐Large Language Model (LLM) introduces a hierarchical, feedback‐driven planning framework where two LLMs coordinate high‐level reasoning and low‐level control for robotic tasks. By grounding decisions in real‐time state feedback, it reduces hallucinations and improves task reliability.
Vineet Bhat   +4 more
wiley   +1 more source

Action knowledge for video captioning with graph neural networks

open access: yesJournal of King Saud University: Computer and Information Sciences, 2023
Many existing video captioning methods capture action information in the video by exploiting features extracted from an action recognition model. However, directly using the action features without object-specific representation may not well capture the ...
Willy Fitra Hendria   +4 more
doaj   +1 more source

Intelligent Sky Guardians (InSkyGuard): An Aerial Robotic Swarm for Autonomous Detection and Entrapment of Rogue Multirotors

open access: yesAdvanced Robotics Research, EarlyView.
Intelligent Sky Guardians (InSkyGuard) is introduced as a four‐drone swarm that autonomously detects, tracks, and safely captures rogue drones using a coordinated net system. Computer vision and leader–follower control architecture enable synchronized enclosure, while integrated failsafes enhance system reliability. Validated through closed‐environment
Joshua Hastings   +6 more
wiley   +1 more source

Adaptive Curriculum Learning for Video Captioning

open access: yesIEEE Access, 2022
A portion of the data in video captioning datasets are noisy and unsuitable for models to learn at early stages, e.g., there could be a generic 4-word-long caption lacking distinctive details of video content and a 19-word-long description with rare ...
Shanhao Li, Bang Yang, Yuexian Zou
doaj   +1 more source

End-to-End Video Captioning

open access: yes2019 IEEE/CVF International Conference on Computer Vision Workshop (ICCVW), 2019
Building correspondences across different modalities, such as video and language, has recently become critical in many visual recognition applications, such as video captioning. Inspired by machine translation, recent models tackle this task using an encoder-decoder strategy.
Silvio Olivastri   +2 more
openaire   +3 more sources

Synergistic Mechano‐Fluidic Self‐Fusion of Poly(Vinyl Alcohol)‐Modified Liquid‐Metal Particles for Post‐Treatment‐Free Soft Electronics

open access: yesAdvanced Science, EarlyView.
Poly(vinyl alcohol)‐modified liquid‐metal particles form conductive networks through spray‐induced self‐fusion without additional post‐treatment. Spray impact initiates oxide‐shell rupture and electrical contact, while subsequent Marangoni flow and capillary forces promote particle redistribution, neck growth, and network consolidation. This sequential
Keun‐Young Yook   +13 more
wiley   +1 more source

Enhanced Hybrid Framework and Comparative Analysis of Deep Learning Architectures for Video Captioning [PDF]

open access: yesEPJ Web of Conferences
With the rapid development of multimedia content on digital platforms, there is more and more need for intelligent systems to understand and describe videos in natural language.
Bhusare Pranali Prabhakar   +1 more
doaj   +1 more source

SoccerNet-Caption: Dense Video Captioning for Soccer Broadcasts Commentaries

open access: yes2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), 2023
Soccer is more than just a game - it is a passion that transcends borders and unites people worldwide. From the roar of the crowds to the excitement of the commentators, every moment of a soccer match is a thrill. Yet, with so many games happening simultaneously, fans cannot watch them all live.
Mkhallati, Hassan   +4 more
openaire   +3 more sources

Smart Exploration of Perovskite Photovoltaics: From AI Driven Discovery to Autonomous Laboratories

open access: yesAdvanced Energy Materials, EarlyView.
In this review, we summarize the fundamentals of AI in automated materials science, and review AI applications in perovskite solar cells. Then, we sum up recent progress in AI‐guided manufacturing optimization, and highlight AI‐driven high‐throughput and autonomous laboratories.
Wenning Chen   +4 more
wiley   +1 more source

A SwinBERT framework with temporal windowed cross attention for video captioning for visual language intelligence

open access: yesDiscover Applied Sciences
Video captioning serves as a tool which enables computers to understand visual content and convert video footage into written language. The current CNN-RNN video captioning system encounters difficulties when it tries to show both term timelines and ...
Reshma DSouza, Snigdha Sen
doaj   +1 more source

Home - About - Disclaimer - Privacy