Results 121 to 130 of about 1,668,820 (165)
Some of the next articles are maybe not open access.
VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking
arXiv.orgWhile recent advances in reinforcement learning have significantly enhanced reasoning capabilities in large language models (LLMs), these techniques remain underexplored in multi-modal LLMs for video captioning.
Desen Meng +10 more
semanticscholar +1 more source
VidChain: Chain-of-Tasks with Metric-based Direct Preference Optimization for Dense Video Captioning
AAAI Conference on Artificial IntelligenceDespite the advancements of Video Large Language Models (VideoLLMs) in various tasks, they struggle with fine-grained temporal understanding, such as Dense Video Captioning (DVC).
Ji Soo Lee +4 more
semanticscholar +1 more source
IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning
ACM MultimediaIntent-oriented controlled video captioning aims to generate targeted descriptions for specific targets in a video based on customized user intent.
Tianheng Qiu +8 more
semanticscholar +1 more source
IEEE International Conference on Computer Vision
Existing sports video captioning methods often focus on the action yet overlook player identities, limiting their applicability. Although some methods integrate extra information to generate identity-aware descriptions, the player identities are ...
Zeyu Xi +7 more
semanticscholar +1 more source
Existing sports video captioning methods often focus on the action yet overlook player identities, limiting their applicability. Although some methods integrate extra information to generate identity-aware descriptions, the player identities are ...
Zeyu Xi +7 more
semanticscholar +1 more source
Action-Driven Semantic Representation and Aggregation for Video Captioning
IEEE transactions on circuits and systems for video technology (Print)Video captioning, a challenging task that entails generating natural language descriptions of visual content, often fails to effectively grasp the essence of action semantics.
Tingting Han +4 more
semanticscholar +1 more source
AAAI Conference on Artificial Intelligence
Video captioning aims to generate comprehensive and coherent descriptions of the video content, contributing to the advancement of both video understanding and generation.
Chunlin Zhong +7 more
semanticscholar +1 more source
Video captioning aims to generate comprehensive and coherent descriptions of the video content, contributing to the advancement of both video understanding and generation.
Chunlin Zhong +7 more
semanticscholar +1 more source
Change3D: Revisiting Change Detection and Captioning from A Video Modeling Perspective
Computer Vision and Pattern RecognitionIn this paper, we present Change3D, a framework that reconceptualizes the change detection and captioning tasks through video modeling. Recent methods have achieved remarkable success by regarding each pair of bi-temporal images as separate frames.
Duo-Wang Zhu +4 more
semanticscholar +1 more source
Do You Remember? Dense Video Captioning with Cross-Modal Memory Retrieval
Computer Vision and Pattern RecognitionThere has been significant attention to the research on dense video captioning, which aims to automatically localize and caption all events within untrimmed video.
Minkuk Kim +4 more
semanticscholar +1 more source
Frame-by-Frame Multi-Object Tracking-Guided Video Captioning
IEEE transactions on circuits and systems for video technology (Print)Video captioning through deep learning presents a multifaceted challenge that encompasses the extraction of complex spatio-temporal visual features and the synthesis of meaningful natural language descriptions.
Huilan Luo, Xia Cai, L. Shark
semanticscholar +1 more source
Dual-path Collaborative Generation Network for Emotional Video Captioning
ACM MultimediaEmotional Video Captioning (EVC) is an emerging task that aims to describe factual content with the intrinsic emotions expressed in videos. The essential of the EVC task is to effectively perceive subtle and ambiguous visual emotional cues during the ...
C. Ye +4 more
semanticscholar +1 more source

