Results 1 to 10 of about 3,798,789 (248)
Video captioning based on vision transformer and reinforcement learning [PDF]
Global encoding of visual features in video captioning is important for improving the description accuracy. In this paper, we propose a video captioning method that combines Vision Transformer (ViT) and reinforcement learning.
Hong Zhao +3 more
doaj +2 more sources
DeepRide: Dashcam Video Description Dataset for Autonomous Vehicle Location-Aware Trip Description
Video description is one of the most challenging task in the combined domain of computer vision and natural language processing. Captions for various open and constrained domain videos have been generated in the recent past but descriptions for driving ...
Ghazala Rafiq +4 more
doaj +1 more source
Video Description: Datasets & Evaluation Metrics
Rapid expansion and the novel phenomenon of deep learning have manifested a variety of proposals and concerns in the area of video description, particularly in the recent past.
Muhammad Rafiq +2 more
doaj +1 more source
The article explores art of cinematography as an objectified cultural reality, in spatial and temporal structures of video description. Genesis of art of photography has changed the habits of human perception and thinking process – from photographic ...
Nadiia Korabliova, Hanna Chmil
doaj +1 more source
Video Description Model Based on Temporal-Spatial and Channel Multi-Attention Mechanisms
Video description plays an important role in the field of intelligent imaging technology. Attention perception mechanisms are extensively applied in video description models based on deep learning.
Jie Xu +4 more
doaj +1 more source
A thorough analysis and comprehension of the entire cue set in visual data are indispensable for an ideal video description model. As outlined in recent algorithm proposals, video descriptions have primarily been generated by learning RGB and optical ...
Ghazala Rafiq +2 more
doaj +1 more source
ST-VLAD: Video Face Recognition Based on Aggregated Local Spatial-Temporal Descriptors
How to integrate the temporal and spatial continuity information, when designing the video texture description operator, is crucial to realize video face recognition and facilitate video analysis and understanding, however, it has still yet to be ...
Yu Wang, Yong-Ping Huang, Xuan-Jing Shen
doaj +1 more source
Automated video game parameter tuning with XVGDL+ [PDF]
Usually, human participation is required in order to provide feedback during the game tuning or balancing process. Moreover, this is commonly an iterative process in which play-testing is required as well as human interaction for gathering all important ...
Jorge Ruiz Quiñones +1 more
doaj +3 more sources
XML-Based Video Game Description Language
This paper presents the XML-based Video Game Description Language (XVGDL), a new language for specifying Video games which is based on the Extensible Markup Language (XML).
Jorge R. Quinones +1 more
doaj +1 more source
Natural Language Description of Video Streams Using Task-Specific Feature Encoding
In recent years, deep learning approaches have gained great attention due to their superior performance and the availability of high speed computing resources.
Aniqa Dilawari +5 more
doaj +1 more source

