Attention-Guided Hierarchical Parsing for Fine-Grained Person-Centric Image Captioning
Although significant progress in the task of producing fine-grained captions for portrait images has been made by the current models for generating detailed descriptions in captions, they still face challenges in attention allocation and in capturing the
Zhengcheng Gu, Jing Jin
doaj +1 more source
AI caption generation model for digital pathology of adenocarcinoma in endoscopic histopathology using multi-instance attention mechanisms. [PDF]
Lee Y, Bai K, Kim YJ, Kim J, Kim KG.
europepmc +1 more source
ISAR-Mamba: A Dual-Stream Gated Mamba with Hierarchical Spatial Summaries for ISAR Image Captioning. [PDF]
He Y +8 more
europepmc +1 more source
Contextual image caption creation using object positional embedding and generative models. [PDF]
Danyal M, Roman M, Shahid A, Yahya M.
europepmc +1 more source
Correction: Long non-coding RNAs in glial cells: key drivers of neuroinflammation in cognitive disorders. [PDF]
Malatino C +6 more
europepmc +1 more source
A Multimodal AI Framework for Medical Education: Integrating Adaptive Image Retrieval, Fast Synthesis, and LLM-Based Clinical Auditing. [PDF]
Díaz-Benito M +4 more
europepmc +1 more source
Generative Distribution Prediction: A Unified Approach to Multimodal Learning. [PDF]
Tian X, Shen X.
europepmc +1 more source
TC3-VLM: A Vision-Language Model for Tactical Combat Casualty Care. [PDF]
Kim J +5 more
europepmc +1 more source
A framework for efficient scientific diagram captioning using mixture-of-experts and low-rank adaptation. [PDF]
Kamboj D, Harit G.
europepmc +1 more source
Rare-Disease Diagnosis on the ZebraMap Multimodal Case Report Dataset: A Hybrid Pipeline with Grounded Explainability. [PDF]
Islam MS, Jamal A, Alkhathlan A.
europepmc +1 more source

