Results 21 to 30 of about 1,468,270 (262)
CLIP4Caption: CLIP for Video Caption [PDF]
Video captioning is a challenging task since it requires generating sentences describing various diverse and complex videos. Existing video captioning models lack adequate visual representation due to the neglect of the existence of gaps between videos and texts. To bridge this gap, in this paper, we propose a CLIP4Caption framework that improves video
Mingkang Tang +5 more
openaire +2 more sources
Empirical autopsy of deep video captioning encoder-decoder architecture
Contemporary deep learning based video captioning methods adopt encoder-decoder framework. In encoder, visual features are extracted with 2D/3D Convolutional Neural Networks (CNNs) and a transformed version of those features is passed to the decoder. The
Nayyer Aafaq +3 more
doaj +1 more source
Multimodal feature fusion based on object relation for video captioning
Video captioning aims at automatically generating a natural language caption to describe the content of a video. However, most of the existing methods in the video captioning task ignore the relationship between objects in the video and the correlation ...
Zhiwen Yan +3 more
doaj +1 more source
nasib-ullah/video-captioning-models-in-Pytorch: Video Captioning Models in Pytorch
A PyTorch implementation of state of the art video captioning models from 2015-2019 on MSVD and MSRVTT ...
Nasib Ullah
core +1 more source
Beyond caption to narrative: Video captioning with multiple sentences [PDF]
Recent advances in image captioning task have led to increasing interests in video captioning task. However, most works on video captioning are focused on generating single input of aggregated features, which hardly deviates from image captioning process and does not fully take advantage of dynamic contents present in videos.
Andrew Shin +2 more
openaire +2 more sources
Thinking Hallucination for Video Captioning
18 ...
Nasib Ullah, Partha Pratim Mohanta
openaire +3 more sources
A novel Multi-Layer Attention Framework for visual description prediction using bidirectional LSTM
The massive influx of text, images, and videos to the internet has recently increased the challenge of computer vision-based tasks in big data. Integrating visual data with natural language to generate video explanations has been a challenge for decades.
Dinesh Naik, C. D. Jaidhar
doaj +1 more source
Methodologies that utilize Deep Learning offer great potential for applications that automatically attempt to generate captions or descriptions about images and video frames.
Soheyla Amirian +3 more
doaj +1 more source
Video Captioning with Tube Features [PDF]
Visual feature plays an important role in the video captioning task. Considering that the video content is mainly composed of the activities of salient objects, it has restricted the caption quality of current approaches which just focus on global frame features while paying less attention to the salient objects.
Bin Zhao 0001 +2 more
openaire +2 more sources
Streamlined Dense Video Captioning [PDF]
Dense video captioning is an extremely challenging task since accurate and coherent description of events in a video requires holistic understanding of video contents as well as contextual reasoning of individual events. Most existing approaches handle this problem by first detecting event proposals from a video and then captioning on a subset of the ...
Jonghwan Mun +4 more
openaire +3 more sources

