Results 31 to 40 of about 1,468,270 (262)

Video Content Caption Generation Based on ViT and Semantic Guidance [PDF]

open access: yesJisuanji gongcheng, 2023
This paper proposes a video captioning method based on Vision Transformer(ViT) and semantic guidance to alleviate the problems of poor readability and low accuracy of caption text generated by exsisting video content captioning models.First,the visual ...
ZHAO Hong, CHEN Zhiwen, GUO Lan, AN Dong
doaj   +1 more source

Reconstruction Network for Video Captioning [PDF]

open access: yes2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2018
Accepted by CVPR ...
Bairui Wang   +3 more
openaire   +2 more sources

Crowd Video Captioning

open access: yesCoRR, 2019
Describing a video automatically with natural language is a challenging task in the area of computer vision. In most cases, the on-site situation of great events is reported in news, but the situation of the off-site spectators in the entrance and exit is neglected which also arouses people's interest.
Liqi Yan, Mingjian Zhu, Changbin Yu
openaire   +2 more sources

Video Question-Answering Techniques, Benchmark Datasets and Evaluation Metrics Leveraging Video Captioning: A Comprehensive Survey

open access: yesIEEE Access, 2021
While describing visual data is a trivial task for humans, it is an intricate task for a computer. This is even more challenging if the visual data is a video. Comprehending a video and describing it is called Video Captioning.
Khushboo Khurana, Umesh Deshpande
doaj   +1 more source

Att-BiL-SL: Attention-Based Bi-LSTM and Sequential LSTM for Describing Video in the Textual Formation

open access: yesApplied Sciences, 2021
With the advancement of the technological field, day by day, people from around the world are having easier access to internet abled devices, and as a result, video data is growing rapidly. The increase of portable devices such as various action cameras,
Shakil Ahmed   +10 more
doaj   +1 more source

Dense-Captioning Events in Videos [PDF]

open access: yes2017 IEEE International Conference on Computer Vision (ICCV), 2017
Most natural videos contain numerous events. For example, in a video of a "man playing a piano", the video might also contain "another man dancing" or "a crowd clapping". We introduce the task of dense-captioning events, which involves both detecting and describing events in a video.
Ranjay Krishna   +4 more
openaire   +3 more sources

How Do Different Keyword Captioning Strategies Impact Students’ Performance in Oral and Written Production Tasks? A Pilot Study

open access: yesThe EUROCALL Review, 2019
As an increasingly popular format of input, the affordances of audio-visual materials have been widely studied. Past research has provided evidence that audio-visual input combined with different captioning strategies could benefit learners in terms of ...
Long He
doaj   +1 more source

Parallel Dense Video Caption Generation with Multi-Modal Features

open access: yesMathematics, 2023
The task of dense video captioning is to generate detailed natural-language descriptions for an original video, which requires deep analysis and mining of semantic captions to identify events in the video. Existing methods typically follow a localisation-
Xuefei Huang   +3 more
doaj   +1 more source

Empowering and conquering infirmity of visually impaired using AI‐technology equipped with object detection and real‐time voice feedback system in healthcare application

open access: yesCAAI Transactions on Intelligence Technology, EarlyView., 2023
Abstract The Internet of Things is emerging as a crucial technology in aiding humans and making their lives easier. Among the human population, a large percentage of people suffer from disabilities resulting in challenges in everyday life particularly people with visual disabilities.
Hania Tarik   +8 more
wiley   +1 more source

Class-dependent and cross-modal memory network considering sentimental features for video-based captioning

open access: yesFrontiers in Psychology, 2023
The video-based commonsense captioning task aims to add multiple commonsense descriptions to video captions to understand video content better. This paper aims to consider the importance of cross-modal mapping.
Haitao Xiong   +4 more
doaj   +1 more source

Home - About - Disclaimer - Privacy