Results 121 to 130 of about 1,668,820 (165)
Some of the next articles are maybe not open access.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking

arXiv.org
While recent advances in reinforcement learning have significantly enhanced reasoning capabilities in large language models (LLMs), these techniques remain underexplored in multi-modal LLMs for video captioning.
Desen Meng   +10 more
semanticscholar   +1 more source

VidChain: Chain-of-Tasks with Metric-based Direct Preference Optimization for Dense Video Captioning

AAAI Conference on Artificial Intelligence
Despite the advancements of Video Large Language Models (VideoLLMs) in various tasks, they struggle with fine-grained temporal understanding, such as Dense Video Captioning (DVC).
Ji Soo Lee   +4 more
semanticscholar   +1 more source

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning

ACM Multimedia
Intent-oriented controlled video captioning aims to generate targeted descriptions for specific targets in a video based on customized user intent.
Tianheng Qiu   +8 more
semanticscholar   +1 more source

Player-Centric Multimodal Prompt Generation for Large Language Model Based Identity-Aware Basketball Video Captioning

IEEE International Conference on Computer Vision
Existing sports video captioning methods often focus on the action yet overlook player identities, limiting their applicability. Although some methods integrate extra information to generate identity-aware descriptions, the player identities are ...
Zeyu Xi   +7 more
semanticscholar   +1 more source

Action-Driven Semantic Representation and Aggregation for Video Captioning

IEEE transactions on circuits and systems for video technology (Print)
Video captioning, a challenging task that entails generating natural language descriptions of visual content, often fails to effectively grasp the essence of action semantics.
Tingting Han   +4 more
semanticscholar   +1 more source

OwlCap: Harmonizing Motion-Detail for Video Captioning via HMD-270K and Caption Set Equivalence Reward

AAAI Conference on Artificial Intelligence
Video captioning aims to generate comprehensive and coherent descriptions of the video content, contributing to the advancement of both video understanding and generation.
Chunlin Zhong   +7 more
semanticscholar   +1 more source

Change3D: Revisiting Change Detection and Captioning from A Video Modeling Perspective

Computer Vision and Pattern Recognition
In this paper, we present Change3D, a framework that reconceptualizes the change detection and captioning tasks through video modeling. Recent methods have achieved remarkable success by regarding each pair of bi-temporal images as separate frames.
Duo-Wang Zhu   +4 more
semanticscholar   +1 more source

Do You Remember? Dense Video Captioning with Cross-Modal Memory Retrieval

Computer Vision and Pattern Recognition
There has been significant attention to the research on dense video captioning, which aims to automatically localize and caption all events within untrimmed video.
Minkuk Kim   +4 more
semanticscholar   +1 more source

Frame-by-Frame Multi-Object Tracking-Guided Video Captioning

IEEE transactions on circuits and systems for video technology (Print)
Video captioning through deep learning presents a multifaceted challenge that encompasses the extraction of complex spatio-temporal visual features and the synthesis of meaningful natural language descriptions.
Huilan Luo, Xia Cai, L. Shark
semanticscholar   +1 more source

Dual-path Collaborative Generation Network for Emotional Video Captioning

ACM Multimedia
Emotional Video Captioning (EVC) is an emerging task that aims to describe factual content with the intrinsic emotions expressed in videos. The essential of the EVC task is to effectively perceive subtle and ambiguous visual emotional cues during the ...
C. Ye   +4 more
semanticscholar   +1 more source

Home - About - Disclaimer - Privacy