Results 11 to 20 of about 400 (142)
A latent topic‐aware network for dense video captioning
Multiple events in a long untrimmed video possess the characteristics of similarity and continuity. These characteristics can be considered as a kind of topic semantic information, which probably behaves as same sports, similar scenes, same objects etc ...
Tao Xu +3 more
doaj +2 more sources
Multi-modal Dense Video Captioning [PDF]
Dense video captioning is a task of localizing interesting events from an untrimmed video and producing textual description (captions) for each localized event. Most of the previous works in dense video captioning are solely based on visual information and completely ignore the audio track.
Vladimir Iashin, Esa Rahtu
openaire +4 more sources
Streamlined Dense Video Captioning [PDF]
Dense video captioning is an extremely challenging task since accurate and coherent description of events in a video requires holistic understanding of video contents as well as contextual reasoning of individual events. Most existing approaches handle this problem by first detecting event proposals from a video and then captioning on a subset of the ...
Jonghwan Mun +4 more
openaire +3 more sources
Dense-Captioning Events in Videos [PDF]
Most natural videos contain numerous events. For example, in a video of a "man playing a piano", the video might also contain "another man dancing" or "a crowd clapping". We introduce the task of dense-captioning events, which involves both detecting and describing events in a video.
Ranjay Krishna +4 more
openaire +3 more sources
SoccerNet-Caption: Dense Video Captioning for Soccer Broadcasts Commentaries
Soccer is more than just a game - it is a passion that transcends borders and unites people worldwide. From the roar of the crowds to the excitement of the commentators, every moment of a soccer match is a thrill. Yet, with so many games happening simultaneously, fans cannot watch them all live.
Mkhallati, Hassan +4 more
openaire +3 more sources
Weakly Supervised Dense Video Captioning [PDF]
This paper focuses on a novel and challenging vision task, dense video captioning, which aims to automatically describe a video clip with multiple informative and diverse caption sentences. The proposed method is trained without explicit annotation of fine-grained sentence to video region-sequence correspondence, but is only based on weak video-level ...
Zhiqiang Shen +6 more
openaire +4 more sources
Semantic-Aware Pretraining for Dense Video Captioning
This report describes the details of our approach for the event dense-captioning task in ActivityNet Challenge 2021. We present a semantic-aware pretraining method for dense video captioning, which empowers the learned features to recognize high-level semantic concepts.
Teng Wang 0007 +5 more
openaire +3 more sources
Post-Attention Modulator for Dense Video Captioning
Peer ...
Wang, Tzu-Jui Julius +3 more
openaire +4 more sources
End-to-End Dense Video Captioning with Parallel Decoding [PDF]
Accepted by ICCV ...
Teng Wang 0007 +5 more
openaire +4 more sources
DVC‐Net: A deep neural network model for dense video captioning
Dense video captioning (DVC) detects multiple events in an input video and generates natural language sentences to describe each event. Previous studies predominantly used convolutional neural networks to extract visual features from videos but failed to
Sujin Lee, Incheol Kim
doaj +1 more source

