Results 131 to 140 of about 1,668,820 (165)
Some of the next articles are maybe not open access.

TrafficVLM: A Controllable Visual Language Model for Traffic Video Captioning

2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW)
Traffic video description and analysis have received much attention recently due to the growing demand for efficient and reliable urban surveillance systems.
Quang Minh Dinh   +3 more
semanticscholar   +1 more source

Learning Comprehensive Visual Grounding for Video Captioning

IEEE transactions on circuits and systems for video technology (Print)
The grounding accuracy of existing video captioners is still behind the expectation. The majority of existing methods perform grounded video captioning on sparse entity annotations.
Wenhui Jiang   +5 more
semanticscholar   +1 more source

Emotional Video Captioning With Vision-Based Emotion Interpretation Network

IEEE Transactions on Image Processing
Effectively summarizing and re-expressing video content by natural languages in a more human-like fashion is one of the key topics in the field of multimedia content understanding.
Peipei Song   +4 more
semanticscholar   +1 more source

Retrieval-Augmented Egocentric Video Captioning

Computer Vision and Pattern Recognition
Understanding human actions from videos offirst-person view poses significant challenges. Most prior approaches explore representation learning on egocentric videos only, while overlooking the potential benefit of exploiting existing large-scale third ...
Jilan Xu   +6 more
semanticscholar   +1 more source

DIBS: Enhancing Dense Video Captioning with Unlabeled Videos via Pseudo Boundary Enrichment and Online Refinement

Computer Vision and Pattern Recognition
We present Dive Into the BoundarieS (DIBS), a novel pretraining framework for dense video captioning (DVC), that elaborates on improving the quality of the generated event captions and their associated pseudo event bound-aries from unlabeled videos.
Hao Wu, Hua-Bin Liu, Yu Qiao, Xiao Sun
semanticscholar   +1 more source

Memory-Based Augmentation Network for Video Captioning

IEEE transactions on multimedia
Video captioning focuses on generating natural language descriptions according to the video content. Existing works mainly explore this multimodal learning with the paired source video and corresponding sentence, which have achieved competitive ...
Shuaiqi Jing   +5 more
semanticscholar   +1 more source

Learnability Matters: Active Learning for Video Captioning

Neural Information Processing Systems
This work focuses on the active learning in video captioning. In particular, we propose to address the learnability problem in active learning, which has been brought up by collective outliers in video captioning and neglected in the literature. To start
Yiqian Zhang   +5 more
semanticscholar   +1 more source

PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

arXiv.org
Vision-language pre-training has significantly elevated performance across a wide range of image-language applications. Yet, the pre-training process for video-related tasks demands exceptionally large computational and data resources, which hinders the ...
Lin Xu   +5 more
semanticscholar   +1 more source

Enhancing Traffic Safety with Parallel Dense Video Captioning for End-to-End Event Analysis

2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW)
This paper introduces our solution for Track 2 in AI City Challenge 2024. The task aims to solve traffic safety description and analysis with the dataset of Woven Traffic Safety (WTS), a real-world Pedestrian-Centric Traffic Video Dataset for Fine ...
Maged Shoman   +3 more
semanticscholar   +1 more source

AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

International Conference on Learning Representations
Video detailed captioning is a key task which aims to generate comprehensive and coherent textual descriptions of video content, benefiting both video understanding and generation.
Wen-Hao Chai   +8 more
semanticscholar   +1 more source

Home - About - Disclaimer - Privacy