Results 11 to 20 of about 13,237 (113)
Long-text caption generation for surgical image with a concept retrieval augmented large multimodal model [PDF]
Jiquan Liu +6 more
doaj +2 more sources
Video Content Caption Generation Based on ViT and Semantic Guidance [PDF]
This paper proposes a video captioning method based on Vision Transformer(ViT) and semantic guidance to alleviate the problems of poor readability and low accuracy of caption text generated by exsisting video content captioning models.First,the visual ...
ZHAO Hong, CHEN Zhiwen, GUO Lan, AN Dong
doaj +1 more source
Image Captioning Using Motion-CNN with Object Detection
Automatic image captioning has many important applications, such as the depiction of visual contents for visually impaired people or the indexing of images on the internet.
Kiyohiko Iwamura +4 more
doaj +1 more source
Multifaceted Feature Coding Image Caption Generation Algorithm Based on Transformer [PDF]
Object features extracted by object detection algorithms play an increasingly critical role in the generation of image captions.However, only using the features of object detection as the input of an image caption task can lead to the loss of other ...
HENG Hongjun, FAN Yuchen, WANG Jialiang
doaj +1 more source
Image Caption Generation via Unified Retrieval and Generation-Based Method
Image captioning is a multi-modal transduction task, translating the source image into the target language. Numerous dominant approaches primarily employed the generation-based or the retrieval-based method.
Shanshan Zhao +4 more
doaj +1 more source
Image Caption Generation Using Contextual Information Fusion With Bi-LSTM-s
The image caption generation algorithm necessitates the expression of image content using accurate natural language. Given the existing encoder-decoder algorithm structure, the decoder solely generates words one by one in a front-to-back order and is ...
Huawei Zhang +3 more
doaj +1 more source
A Rapid Review of Image Captioning
Image captioning is an automatic process for generating text based on the content observed in an image. We do review, create framework, and build application model. We review image captioning into 4 categories based on input model, process model, output
Adriyendi Adriyendi
doaj +1 more source
Vision Transformer and Language Model Based Radiology Report Generation
Recent advancements in transformers exploited computer vision problems which results in state-of-the-art models. Transformer-based models in various sequence prediction tasks such as language translation, sentiment classification, and caption generation ...
Mashood Mohammad Mohsan +5 more
doaj +1 more source
Self-Learning for Few-Shot Remote Sensing Image Captioning
Large-scale caption-labeled remote sensing image samples are expensive to acquire, and the training samples available in practical application scenarios are generally limited.
Haonan Zhou +3 more
doaj +1 more source
Learn and Tell: Learning Priors for Image Caption Generation
In this work, we propose a novel priors-based attention neural network (PANN) for image captioning, which aims at incorporating two kinds of priors, i.e., the probabilities being mentioned for local region proposals (PBM priors) and part-of-speech clues ...
Pei Liu, Dezhong Peng, Ming Zhang
doaj +1 more source

