Results 221 to 230 of about 1,468,270 (262)
Lightweight visual accessibility LLaVA architecture. [PDF]
Han Z, Liu X, Hao J.
europepmc +1 more source
Video Captioning by Adversarial LSTM
In this paper, we propose a novel approach to video captioning based on adversarial learning and long short-term memory (LSTM). With this solution concept, we aim at compensating for the deficiencies of LSTM-based video captioning methods that generally ...
Yanli Ji, Yi Bin, Heng Tao Shen
exaly +7 more sources
Some of the next articles are maybe not open access.
Related searches:
Related searches:
Dense Video Captioning for Incomplete Videos
2021Incomplete video or partially-missing video situations are rarely considered in video captioning research. Previous approaches are mainly trained and evaluated on complete video clip datasets where all the events involved are thoroughly observed. In this work, we formulate the issue of video content description for partially-missing videos.
Xuan Dang +3 more
openaire +1 more source
Video Captioning of Future Frames
2021 IEEE Winter Conference on Applications of Computer Vision (WACV), 2021Being able to anticipate and describe what may happen in the future is a fundamental ability for humans. Given a short clip of a scene about "a person is sitting behind a piano", humans can describe what will happen afterward, i.e. "the person is playing the piano".
Mehrdad Hosseinzadeh, Yang Wang 0003
openaire +1 more source
2019 49th Annual IEEE/IFIP International Conference on Dependable Systems and Networks Workshops (DSN-W), 2019
In recent years, developments in the field of computer vision have allowed deep learning-based techniques to surpass human-level performance. However, these advances have also culminated in the advent of adversarial machine learning techniques, capable of launching targeted image captioning attacks that easily fool deep learning models.
Suman Kalyan Adari +2 more
openaire +1 more source
In recent years, developments in the field of computer vision have allowed deep learning-based techniques to surpass human-level performance. However, these advances have also culminated in the advent of adversarial machine learning techniques, capable of launching targeted image captioning attacks that easily fool deep learning models.
Suman Kalyan Adari +2 more
openaire +1 more source
Multi-Perspective Video Captioning
Proceedings of the 29th ACM International Conference on Multimedia, 2021This work targets at the problems of comprehensive video captioning and the generation of multiple descriptions from different perspectives, termed asMulti-Perspective Video Captioning. We build and release a dataset named VidOR-MPVC, the first dataset for multi-perspective video captioning, where each video is annotated with multiple descriptions from
Yi Bin +4 more
openaire +2 more sources
IcoCap: Improving Video Captioning by Compounding Images [PDF]
Video captioning is a more challenging task compared to image captioning, primarily due to differences in content density. Video data contains redundant visual content, making it difficult for captioners to generalize diverse content and avoid being ...
Xiaohan Wang, Linchao Zhu, Yi Yang
exaly +2 more sources
Rethink video retrieval representation for video captioning
Pattern RecognitionQingming Huang +2 more
exaly +2 more sources

