Results 231 to 240 of about 1,468,270 (262)
Some of the next articles are maybe not open access.
State-aware Video Procedural Captioning
Proceedings of the 29th ACM International Conference on Multimedia, 2021Video procedural captioning (VPC), which generates procedural text from instructional videos, is an essential task for scene understanding and real-world applications. The main challenge of VPC is to describe how to manipulate materials accurately. This paper focuses on this challenge by designing a new VPC task, generating a procedural text from the ...
Taichi Nishimura +4 more
openaire +3 more sources
Sequence in sequence for video captioning
Pattern Recognition Letters, 2020Abstract For video captioning, the words in the caption are closely related to an overall understanding of the video. Thus, a suitable representation for the video is rather important for the description. For more precise words in the task of video captioning, we aim to encode the video feature for current word at each time-stamp of the generation ...
Huiyun Wang, Chongyang Gao, Yahong Han
openaire +2 more sources
Video Captioning with Listwise Supervision
Proceedings of the AAAI Conference on Artificial Intelligence, 2017Automatically describing video content with natural language is a fundamental challenging that has received increasing attention. However, existing techniques restrict the model learning on the pairs of each video and its own sentences, and thus fail to capture more holistically semantic relationships among all sentences. In this paper,
Yuan Liu, Xue Li, Zhongchao Shi
openaire +1 more source
Video caption duration extraction
2008 19th International Conference on Pattern Recognition, 2008Caption detection in the video is an active research topic in recent years. In the conventional methods, one of most difficult problems is to effectively and quickly extract the durations of the different-size captions in the complex background. To solve this problem, a novel and effective method is presented to locate and track the captions in the ...
Hongliang Bai +5 more
openaire +1 more source
A Review of Deep Learning for Video Captioning
42 pages, 10 ...
Abbas Khosravi +2 more
exaly +7 more sources
Video Captioning with Semantic Guiding
2018 IEEE Fourth International Conference on Multimedia Big Data (BigMM), 2018Video captioning is to generate descriptions of videos. Most existing approaches adopt the encoder-decoder architecture, which usually use different kinds of visual features, such as temporal features and motion features, but they neglect the abundant semantic information in the video. To address this issue, we propose a framework that jointly explores
Jin Yuan +4 more
openaire +1 more source
Multirate Multimodal Video Captioning
Proceedings of the 25th ACM international conference on Multimedia, 2017Automatically describing videos with natural language is a crucial challenge of video understanding. Compared to images, videos have specific spatial-temporal structure and various modality information. In this paper, we propose a Multirate Multimodal Approach for video captioning.
Ziwei Yang 0001 +4 more
openaire +1 more source
Multiple Videos Captioning Model for Video Storytelling
2019 IEEE International Conference on Big Data and Smart Computing (BigComp), 2019In this paper, We propose a novel video captioning model that utilizes context information of correlated clips. Unlike the ordinary “one clip - one caption” algorithms, we concatenate multiple neighboring clips as a chunk and train the network in “one chunk - multiple caption” manner.
Seung-Ho Han 0001 +2 more
openaire +1 more source
Creating Accessible Videos: Captions and Transcripts
Communications of the Association for Information Systems, 2021The rapid shift to online teaching due to the coronavirus disease of 2019 (COVID-19) exponentially increased the extent to which faculty use videoconferencing/virtual classroom tools such as Zoom, Google Meet, and Microsoft Teams It also exposed the challenge of ensuring that all students could access all video content Faculty may need to implement ...
openaire +1 more source
2014
Video contains two types of texts. The first type pertains to caption texts which are edited texts or graphics texts artificially superimposed into video and are relevant to the content of the video. The second type belongs to scene texts, which are naturally existing texts, usually embedded in objects in the video. This chapter focuses on the state-of-
Tong Lu +3 more
openaire +1 more source
Video contains two types of texts. The first type pertains to caption texts which are edited texts or graphics texts artificially superimposed into video and are relevant to the content of the video. The second type belongs to scene texts, which are naturally existing texts, usually embedded in objects in the video. This chapter focuses on the state-of-
Tong Lu +3 more
openaire +1 more source

