Results 11 to 20 of about 24,887 (267)
A Review of Transformer-Based Approaches for Image Captioning
Visual understanding is a research area that bridges the gap between computer vision and natural language processing. Image captioning is a visual understanding task in which natural language descriptions of images are automatically generated using ...
Oscar Ondeng, Heywood Ouma, Peter Akuon
doaj +1 more source
code: https://github.com/OpenNLPLab/Vicinity-Vision ...
Weixuan Sun +9 more
openaire +3 more sources
STHarDNet: Swin Transformer with HarDNet for MRI Segmentation
In magnetic resonance imaging (MRI) segmentation, conventional approaches utilize U-Net models with encoder–decoder structures, segmentation models using vision transformers, or models that combine a vision transformer with an encoder–decoder model ...
Yeonghyeon Gu +2 more
doaj +1 more source
Building Extraction With Vision Transformer [PDF]
Submitted to ...
Libo Wang +3 more
openaire +2 more sources
Human vision possesses a special type of visual processing systems called peripheral vision. Partitioning the entire visual field into multiple contour regions based on the distance to the center of our gaze, the peripheral vision provides us the ability to perceive various visual features at different regions.
Juhong Min +3 more
openaire +3 more sources
Vision Transformer with Progressive Sampling [PDF]
Accepted to ICCV ...
Xiaoyu Yue +6 more
openaire +3 more sources
Understanding actions in videos remains a significant challenge in computer vision, which has been the subject of several pieces of research in the last decades.
Oumaima Moutik +6 more
doaj +1 more source
Transformer-based ripeness segmentation for tomatoes
With the recent development of computer vision technology, various computer vision techniques have been applied to agriculture. Recently, the Transformer network has been introduced to image recognition, which allows a different approach to extracting ...
Risa Shinoda +3 more
doaj +1 more source
Reversible Vision Transformers
We present Reversible Vision Transformers, a memory efficient architecture design for visual recognition. By decoupling the GPU memory requirement from the depth of the model, Reversible Vision Transformers enable scaling up architectures with efficient memory usage.
Karttikeya Mangalam +6 more
openaire +2 more sources
Transformer architectures for computer vision: A comprehensive review and future research directions [PDF]
Long-range dependencies and contextual relationships in videos were captured by using Convolutional Neural Networks (CNNs) in past. Recently the use of Transformers is started for capturing the long-range dependencies and contextual relationships in ...
Ugile Tukaram, Uke Nilesh
doaj +1 more source

