Results 51 to 60 of about 14,313,923 (129)
CornCare: A Knowledge-Graph-Enhanced Multimodal Diagnostic Reporting System for Corn Diseases
Accurate and actionable crop disease diagnosis requires not only visual recognition of disease symptoms but also the ability to generate grounded reports that integrate symptom interpretation with agronomic knowledge.
Yang Liu, Yushan Xie, Xue Wu, Qi Wang
doaj +1 more source
Evolution of Neural Image Caption Generation: From RNN towards Transformers
Image caption generation is one of the core problems in artificial intelligence that leverages enhancements in computer vision and natural language processing.
Olimov Farrukh
core +1 more source
Generating captions automatically for images has been a challenging task, requiring the integration of image processing and natural language processing techniques.
Smita Bharne, Pawan Bhaladhare
doaj +1 more source
A study on diffusion probabilistic models for image generation [PDF]
Diffusion probabilistic models have emerged as powerful tools for image generation and synthesis tasks. This research delves into the intricate relationship between the hyperparameters of these models and the underlying hardware, aiming to provide ...
Munoz, Roman
core
Making Images Speak: Human-Inspired Image Description Generation
Despite significant advances in deep learning-based image captioning, many state-of-the-art models still struggle to balance visual grounding (i.e., accurate object and scene descriptions) with linguistic coherence (i.e., grammatical fluency and ...
Chifaa Sebbane +2 more
doaj +1 more source
A Systematic Approach for News Caption Generation
– Captions are essential components associated with images to make search engines to respond easily with user queries. Making appropriate captions for images is a difficult task.
Final Year M. Tech Cse, Philo Sumi
core
Localization and recognition of the scoreboard in sports video based on SIFT point matching [PDF]
In broadcast sports video, the scoreboard is attached at a fixed location in the video and generally the scoreboard always exists in all video frames in order to help viewers to understand the match’s progression quickly.
Guo, Jinlin +4 more
core
Using user-generated content (UGC) is of utmost importance for e-commerce platforms to extract valuable commercial information. In this paper, we propose an explainable multimodal learning approach named the visual–semantic embedding model with a self ...
Chengwen Sun, Feng Liu
doaj +1 more source
Multiple Conditions-Guided Diffusion Model for Remote Sensing Image Generation
To address the issues of large scenes and detail attributions for generating remote sensing images (RSIs), this study proposes a multiple conditions-guided diffusion (MCGD) model.
Ran Zhang +3 more
doaj +1 more source
An Overview of Image Caption Generation Methods. [PDF]
Wang H, Zhang Y, Yu X.
europepmc +1 more source

