Results 41 to 50 of about 2,289 (259)
Multi‐task learning for captioning images with novel words
Recent captioning models are limited in their ability to describe concepts unseen in paired image–sentence pairs. This study presents a framework of multi‐task learning for describing novel words not present in existing image‐captioning datasets.
He Zheng +4 more
doaj +1 more source
Image-Captioning Model Compression
Image captioning is a very important task, which is on the edge between natural language processing (NLP) and computer vision (CV). The current quality of the captioning models allows them to be used for practical tasks, but they require both large ...
Viktar Atliha, Dmitrij Šešok
doaj +1 more source
Retrieval-augmented Image Captioning
Inspired by retrieval-augmented language generation and pretrained Vision and Language (V&L) encoders, we present a new approach to image captioning that generates sentences given the input image and a set of captions retrieved from a datastore, as opposed to the image alone.
Ramos, Rita +2 more
openaire +3 more sources
Captioning Transformer with Stacked Attention Modules
Image captioning is a challenging task. Meanwhile, it is important for the machine to understand the meaning of an image better. In recent years, the image captioning usually use the long-short-term-memory (LSTM) as the decoder to generate the sentence ...
Xinxin Zhu +4 more
doaj +1 more source
Building machine‐readable vocabularies for materials science is slow, expert‐driven work. This study benchmarks 13 large language models on two of its first steps: finding candidate terms in engineering articles and deciding where they belong in a class hierarchy.
Thomas Bjarsch +3 more
wiley +1 more source
Where to put the image in an image caption generator [PDF]
AbstractWhen a recurrent neural network (RNN) language model is used for caption generation, the image information can be fed to the neural network either by directly incorporating it in the RNN – conditioning the language model by ‘injecting’ image features – or in a layer following the RNN – conditioning the language model by ‘merging’ image features.
Marc Tanti +2 more
openaire +4 more sources
Can Audio Captions Be Evaluated With Image Caption Metrics?
ICASSP ...
Zelin Zhou +5 more
openaire +2 more sources
Reproduction of stacking fault energy calculations from literature with a semi‐automated large language model‐assisted extraction procedure: extraction of simulation protocol, atomistic structures, computational parameters, and reported results, ontology alignment, knowledge graph construction and, finally, recomputation forvalidation.
Sepideh Baghaee Ravari +5 more
wiley +1 more source
Image Captioning Method Based on Transformer Visual Features Fusion [PDF]
Existing image captioning methods only use regional visual features to generate description statements and ignore the importance of grid visual features. Moreover, as these methods are two-stage approaches, image captioning quality is affected.
Xuebing BAI, Jin CHE, Jinman WU, Yumin CHEN
doaj +1 more source
Enhanced Image Captioning with Color Recognition Using Deep Learning Methods
Automatically describing the content of an image is an interesting and challenging task in artificial intelligence. In this paper, an enhanced image captioning model—including object detection, color analysis, and image captioning—is proposed to ...
Yeong-Hwa Chang +3 more
doaj +1 more source

