Results 61 to 70 of about 14,313,923 (129)
AEHRC CSIRO at ImageCLEFmed caption 2021
We describe our participation in the ImageCLEFmed Caption task of 2021. The task required participants to automatically compose coherent captions for a set of medical images.
Dowling, J, Nicolson, A, Koopman, B
core
Enhanced Image Captioning Using Bahdanau Attention Mechanism and Heuristic Beam Search Algorithm
Captioning images is a challenging task at the intersection of Computer Vision (CV) and Natural Language Processing (NLP), that involves generating descriptive text to depict the content of an image.
S. Abinaya +2 more
doaj +1 more source
Mamba-caption: Long-range sequence modelling for efficient and accurate image captioning
Image captioning has been a problem in vision–language research for a long time. Long-range dependencies and efficiency are challenges for the standard models, such as recurrent neural networks (RNNs) and Transformers.
Tariq Shahzad +5 more
doaj +1 more source
Tweet-Image-Caption conditioned diffusion model for multimodal Aspect-Based sentiment analysis
Multimodal Aspect-Based Sentiment Analysis (MABSA) aims to simultaneously extract aspects and predict their sentiment polarities from paired textual and visual content.
Haomei Jia +4 more
doaj +1 more source
Image captioning is a fascinating and fast-evolving research project that integrates two domains: Natural Language Processing and Computer Vision. Creating appropriate captions is a difficult task due to the many activities portrayed in the backdrop ...
P. V. Kavitha, V. Karpagam
doaj +1 more source
Grounded Video Caption Generation
We propose a new task, dataset and model for grounded video caption generation. This task unifies captioning and object grounding in video, where the objects in the caption are grounded in the video via temporally consistent bounding boxes.
Schmid, Cordelia +2 more
core
Diverse and Creative Image Caption Generation
© 2025 Dalin WangNatural language generation (NLG) is a cornerstone of artificial intelligence, aiming to replicate the human capacity to communicate complex ideas through symbolic language.
Wang, Dalin
core +1 more source
ISAR-Mamba: A Dual-Stream Gated Mamba with Hierarchical Spatial Summaries for ISAR Image Captioning. [PDF]
He Y +8 more
europepmc +1 more source
A Multimodal AI Framework for Medical Education: Integrating Adaptive Image Retrieval, Fast Synthesis, and LLM-Based Clinical Auditing. [PDF]
Díaz-Benito M +4 more
europepmc +1 more source
Contextual image caption creation using object positional embedding and generative models. [PDF]
Danyal M, Roman M, Shahid A, Yahya M.
europepmc +1 more source

