Results 11 to 20 of about 400 (142)

A latent topic‐aware network for dense video captioning

open access: yesIET Computer Vision, 2023
Multiple events in a long untrimmed video possess the characteristics of similarity and continuity. These characteristics can be considered as a kind of topic semantic information, which probably behaves as same sports, similar scenes, same objects etc ...
Tao Xu   +3 more
doaj   +2 more sources

Multi-modal Dense Video Captioning [PDF]

open access: yes2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), 2020
Dense video captioning is a task of localizing interesting events from an untrimmed video and producing textual description (captions) for each localized event. Most of the previous works in dense video captioning are solely based on visual information and completely ignore the audio track.
Vladimir Iashin, Esa Rahtu
openaire   +4 more sources

Streamlined Dense Video Captioning [PDF]

open access: yes2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019
Dense video captioning is an extremely challenging task since accurate and coherent description of events in a video requires holistic understanding of video contents as well as contextual reasoning of individual events. Most existing approaches handle this problem by first detecting event proposals from a video and then captioning on a subset of the ...
Jonghwan Mun   +4 more
openaire   +3 more sources

Dense-Captioning Events in Videos [PDF]

open access: yes2017 IEEE International Conference on Computer Vision (ICCV), 2017
Most natural videos contain numerous events. For example, in a video of a "man playing a piano", the video might also contain "another man dancing" or "a crowd clapping". We introduce the task of dense-captioning events, which involves both detecting and describing events in a video.
Ranjay Krishna   +4 more
openaire   +3 more sources

SoccerNet-Caption: Dense Video Captioning for Soccer Broadcasts Commentaries

open access: yes2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), 2023
Soccer is more than just a game - it is a passion that transcends borders and unites people worldwide. From the roar of the crowds to the excitement of the commentators, every moment of a soccer match is a thrill. Yet, with so many games happening simultaneously, fans cannot watch them all live.
Mkhallati, Hassan   +4 more
openaire   +3 more sources

Weakly Supervised Dense Video Captioning [PDF]

open access: yes2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017
This paper focuses on a novel and challenging vision task, dense video captioning, which aims to automatically describe a video clip with multiple informative and diverse caption sentences. The proposed method is trained without explicit annotation of fine-grained sentence to video region-sequence correspondence, but is only based on weak video-level ...
Zhiqiang Shen   +6 more
openaire   +4 more sources

Semantic-Aware Pretraining for Dense Video Captioning

open access: yesCoRR, 2022
This report describes the details of our approach for the event dense-captioning task in ActivityNet Challenge 2021. We present a semantic-aware pretraining method for dense video captioning, which empowers the learned features to recognize high-level semantic concepts.
Teng Wang 0007   +5 more
openaire   +3 more sources

Post-Attention Modulator for Dense Video Captioning

open access: yes2022 26th International Conference on Pattern Recognition (ICPR), 2022
Peer ...
Wang, Tzu-Jui Julius   +3 more
openaire   +4 more sources

End-to-End Dense Video Captioning with Parallel Decoding [PDF]

open access: yes2021 IEEE/CVF International Conference on Computer Vision (ICCV), 2021
Accepted by ICCV ...
Teng Wang 0007   +5 more
openaire   +4 more sources

DVC‐Net: A deep neural network model for dense video captioning

open access: yesIET Computer Vision, 2021
Dense video captioning (DVC) detects multiple events in an input video and generates natural language sentences to describe each event. Previous studies predominantly used convolutional neural networks to extract visual features from videos but failed to
Sujin Lee, Incheol Kim
doaj   +1 more source

Home - About - Disclaimer - Privacy