Results 31 to 40 of about 1,468,270 (262)
Video Content Caption Generation Based on ViT and Semantic Guidance [PDF]
This paper proposes a video captioning method based on Vision Transformer(ViT) and semantic guidance to alleviate the problems of poor readability and low accuracy of caption text generated by exsisting video content captioning models.First,the visual ...
ZHAO Hong, CHEN Zhiwen, GUO Lan, AN Dong
doaj +1 more source
Reconstruction Network for Video Captioning [PDF]
Accepted by CVPR ...
Bairui Wang +3 more
openaire +2 more sources
Describing a video automatically with natural language is a challenging task in the area of computer vision. In most cases, the on-site situation of great events is reported in news, but the situation of the off-site spectators in the entrance and exit is neglected which also arouses people's interest.
Liqi Yan, Mingjian Zhu, Changbin Yu
openaire +2 more sources
While describing visual data is a trivial task for humans, it is an intricate task for a computer. This is even more challenging if the visual data is a video. Comprehending a video and describing it is called Video Captioning.
Khushboo Khurana, Umesh Deshpande
doaj +1 more source
With the advancement of the technological field, day by day, people from around the world are having easier access to internet abled devices, and as a result, video data is growing rapidly. The increase of portable devices such as various action cameras,
Shakil Ahmed +10 more
doaj +1 more source
Dense-Captioning Events in Videos [PDF]
Most natural videos contain numerous events. For example, in a video of a "man playing a piano", the video might also contain "another man dancing" or "a crowd clapping". We introduce the task of dense-captioning events, which involves both detecting and describing events in a video.
Ranjay Krishna +4 more
openaire +3 more sources
As an increasingly popular format of input, the affordances of audio-visual materials have been widely studied. Past research has provided evidence that audio-visual input combined with different captioning strategies could benefit learners in terms of ...
Long He
doaj +1 more source
Parallel Dense Video Caption Generation with Multi-Modal Features
The task of dense video captioning is to generate detailed natural-language descriptions for an original video, which requires deep analysis and mining of semantic captions to identify events in the video. Existing methods typically follow a localisation-
Xuefei Huang +3 more
doaj +1 more source
Abstract The Internet of Things is emerging as a crucial technology in aiding humans and making their lives easier. Among the human population, a large percentage of people suffer from disabilities resulting in challenges in everyday life particularly people with visual disabilities.
Hania Tarik +8 more
wiley +1 more source
The video-based commonsense captioning task aims to add multiple commonsense descriptions to video captions to understand video content better. This paper aims to consider the importance of cross-modal mapping.
Haitao Xiong +4 more
doaj +1 more source

