Results 111 to 120 of about 1,668,820 (165)

IcoCap: Improving Video Captioning by Compounding Images [PDF]

open access: yesIEEE Transactions on Multimedia
Video captioning is a more challenging task compared to image captioning, primarily due to differences in content density. Video data contains redundant visual content, making it difficult for captioners to generalize diverse content and avoid being ...
Xiaohan Wang, Linchao Zhu, Yi Yang
exaly   +3 more sources

Streaming Dense Video Captioning [PDF]

open access: yesComputer Vision and Pattern Recognition
An ideal model for dense video captioning - predicting captions localized temporally in a video - should be able to handle long input videos, predict rich, detailed textual descriptions, and be able to produce outputs before processing the entire video ...
Xingyi Zhou   +7 more
semanticscholar   +3 more sources

Video Captioning by Adversarial LSTM

open access: yesIEEE Transactions on Image Processing, 2018
In this paper, we propose a novel approach to video captioning based on adversarial learning and long short-term memory (LSTM). With this solution concept, we aim at compensating for the deficiencies of LSTM-based video captioning methods that generally ...
Yanli Ji, Yi Bin, Heng Tao Shen
exaly   +3 more sources
Some of the next articles are maybe not open access.

Related searches:

Analysis of Image Captioning Techniques in English and Arabic: A Review

2025 3rd International Conference on Business Analytics for Technology and Security (ICBATS)
Image translation is needed because the amount of visual material on the internet has grown quickly. The computer keeps track of the picture's parts and features and then writes a complete description of it. There are several ways to use this model. This
Shams A. Ahmed, A. T. Abdulameer
semanticscholar   +1 more source

Arabic Text Detection From Natural Images Using Transformers

International Symposium on Signal, Image, Video and Communications
Text detection in the wild is an important stage for recognizing and interpreting text present in images or videos captured in uncontrolled environments in the wild.
Ayad Abdessamad   +2 more
semanticscholar   +1 more source

Dense Video Captioning Using Graph-Based Sentence Summarization

IEEE transactions on multimedia
Recently, dense video captioning has made attractive progress in detecting and captioning all events in a long untrimmed video. Despite promising results were achieved, most existing methods do not sufficiently explore the scene evolution within an event
Zhiwang Zhang   +3 more
semanticscholar   +1 more source

Emotion-Oriented Cross-Modal Prompting and Alignment for Human-Centric Emotional Video Captioning

IEEE transactions on multimedia
Human-centric Emotional Video Captioning (H-EVC) aims to generate fine-grained, emotion-related sentences for human-based videos, enhancing the understanding of human emotions and facilitating human-computer emotional interaction. However, existing video
Yu Wang   +6 more
semanticscholar   +1 more source

Multi-round Mutual Emotion-Cause Pair Extraction for Emotion-Attributed Video Captioning

ACM Multimedia
Emotional Video Captioning (EVC) is an emerging task that aims to describe factual content with the intrinsic emotions expressed in videos. Existing EVC methods perceive global emotional cues through visual features at first, and then combine them with ...
C. Ye   +5 more
semanticscholar   +1 more source

Event-Equalized Dense Video Captioning

Computer Vision and Pattern Recognition
Dense video captioning aims to localize and caption all events in arbitrary untrimmed videos. Although previous methods have achieved appealing results, they still face the issue of temporal bias, i.e, models tend to focus more on events with certain ...
Kangyi Wu   +7 more
semanticscholar   +1 more source

VidCapBench: A Comprehensive Benchmark of Video Captioning for Controllable Text-to-Video Generation

Annual Meeting of the Association for Computational Linguistics
The training of controllable text-to-video (T2V) models relies heavily on the alignment between videos and captions, yet little existing research connects video caption evaluation with T2V generation assessment. This paper introduces VidCapBench, a video
Xinlong Chen   +9 more
semanticscholar   +1 more source

Home - About - Disclaimer - Privacy