Results 21 to 30 of about 400 (142)
Multimodal Pretraining for Dense Video Captioning
AACL-IJCNLP ...
Gabriel Huang +4 more
openaire +2 more sources
An Efficient Framework for Dense Video Captioning
Dense video captioning is an extremely challenging task since an accurate and faithful description of events in a video requires a holistic knowledge of the video contents as well as contextual reasoning of individual events. Most existing approaches handle this problem by first proposing event boundaries from a video and then captioning on a subset of
Maitreya Suin, A. N. Rajagopalan 0001
openaire +2 more sources
Dense Procedure Captioning in Narrated Instructional Videos [PDF]
Understanding narrated instructional videos is important for both research and real-world web applications. Motivated by video dense captioning, we propose a model to generate procedure captions from narrated instructional videos which are a sequence of step-wise clips with description.
Botian Shi +6 more
openaire +1 more source
Leveraging auxiliary image descriptions for dense video captioning
Abstract Collecting textual descriptions is an especially costly task for dense video captioning, since each event in the video needs to be annotated separately and a long descriptive paragraph needs to be provided. In this paper, we investigate a way to mitigate this heavy burden and propose to leverage captions of visually similar images as ...
Emre Boran +5 more
openaire +3 more sources
End-to-End Dense Video Captioning with Masked Transformer [PDF]
Dense video captioning aims to generate text descriptions for all events in an untrimmed video. This involves both detecting and describing events. Therefore, all previous methods on dense video captioning tackle this problem by building two models, i.e. an event proposal and a captioning model, for these two sub-problems. The models are either trained
Luowei Zhou +4 more
openaire +2 more sources
Weakly Supervised Dense Event Captioning in Videos
Dense event captioning aims to detect and describe all events of interest contained in a video. Despite the advanced development in this area, existing methods tackle this task by making use of dense temporal annotations, which is dramatically source-consuming.
Xuguang Duan +5 more
openaire +4 more sources
Jointly Localizing and Describing Events for Dense Video Captioning [PDF]
Automatically describing a video with natural language is regarded as a fundamental challenge in computer vision. The problem nevertheless is not trivial especially when a video contains multiple events to be worthy of mention, which often happens in real videos. A valid question is how to temporally localize and then describe events, which is known as
Yehao Li +4 more
openaire +4 more sources
End-to-end Dense Video Captioning as Sequence Generation
Dense video captioning aims to identify the events of interest in an input video, and generate descriptive captions for each event. Previous approaches usually follow a two-stage generative process, which first proposes a segment for each event, then renders a caption for each identified segment.
Wanrong Zhu +4 more
openaire +3 more sources
Dense video captioning involves identifying, localizing, and describing multiple events within a video. Capturing temporal and contextual dependencies between events is essential for generating coherent and accurate captions.
Dvijesh Bhatt, Priyank Thakkar
doaj +1 more source
Streaming Dense Video Captioning
An ideal model for dense video captioning -- predicting captions localized temporally in a video -- should be able to handle long input videos, predict rich, detailed textual descriptions, and be able to produce outputs before processing the entire video. Current state-of-the-art models, however, process a fixed number of downsampled frames, and make a
Xingyi Zhou +7 more
openaire +4 more sources

