Results 21 to 30 of about 1,468,270 (262)

CLIP4Caption: CLIP for Video Caption [PDF]

open access: yesProceedings of the 29th ACM International Conference on Multimedia, 2021
Video captioning is a challenging task since it requires generating sentences describing various diverse and complex videos. Existing video captioning models lack adequate visual representation due to the neglect of the existence of gaps between videos and texts. To bridge this gap, in this paper, we propose a CLIP4Caption framework that improves video
Mingkang Tang   +5 more
openaire   +2 more sources

Empirical autopsy of deep video captioning encoder-decoder architecture

open access: yesArray, 2021
Contemporary deep learning based video captioning methods adopt encoder-decoder framework. In encoder, visual features are extracted with 2D/3D Convolutional Neural Networks (CNNs) and a transformed version of those features is passed to the decoder. The
Nayyer Aafaq   +3 more
doaj   +1 more source

Multimodal feature fusion based on object relation for video captioning

open access: yesCAAI Transactions on Intelligence Technology, 2023
Video captioning aims at automatically generating a natural language caption to describe the content of a video. However, most of the existing methods in the video captioning task ignore the relationship between objects in the video and the correlation ...
Zhiwen Yan   +3 more
doaj   +1 more source

nasib-ullah/video-captioning-models-in-Pytorch: Video Captioning Models in Pytorch

open access: yes, 2023
A PyTorch implementation of state of the art video captioning models from 2015-2019 on MSVD and MSRVTT ...
Nasib Ullah
core   +1 more source

Beyond caption to narrative: Video captioning with multiple sentences [PDF]

open access: yes2016 IEEE International Conference on Image Processing (ICIP), 2016
Recent advances in image captioning task have led to increasing interests in video captioning task. However, most works on video captioning are focused on generating single input of aggregated features, which hardly deviates from image captioning process and does not fully take advantage of dynamic contents present in videos.
Andrew Shin   +2 more
openaire   +2 more sources

Thinking Hallucination for Video Captioning

open access: yes, 2023
18 ...
Nasib Ullah, Partha Pratim Mohanta
openaire   +3 more sources

A novel Multi-Layer Attention Framework for visual description prediction using bidirectional LSTM

open access: yesJournal of Big Data, 2022
The massive influx of text, images, and videos to the internet has recently increased the challenge of computer vision-based tasks in big data. Integrating visual data with natural language to generate video explanations has been a challenge for decades.
Dinesh Naik, C. D. Jaidhar
doaj   +1 more source

Automatic Image and Video Caption Generation With Deep Learning: A Concise Review and Algorithmic Overlap

open access: yesIEEE Access, 2020
Methodologies that utilize Deep Learning offer great potential for applications that automatically attempt to generate captions or descriptions about images and video frames.
Soheyla Amirian   +3 more
doaj   +1 more source

Video Captioning with Tube Features [PDF]

open access: yesProceedings of the Twenty-Seventh International Joint Conference on Artificial Intelligence, 2018
Visual feature plays an important role in the video captioning task. Considering that the video content is mainly composed of the activities of salient objects, it has restricted the caption quality of current approaches which just focus on global frame features while paying less attention to the salient objects.
Bin Zhao 0001   +2 more
openaire   +2 more sources

Streamlined Dense Video Captioning [PDF]

open access: yes2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019
Dense video captioning is an extremely challenging task since accurate and coherent description of events in a video requires holistic understanding of video contents as well as contextual reasoning of individual events. Most existing approaches handle this problem by first detecting event proposals from a video and then captioning on a subset of the ...
Jonghwan Mun   +4 more
openaire   +3 more sources

Home - About - Disclaimer - Privacy