Results 11 to 20 of about 1,457,631 (249)
CSiT: A Multiscale Vision Transformer for Hyperspectral Image Classification
The hyperspectral image (HSI) has nearly continuous spectral information; thus, the target of interest can be accurately identified by the subtle details of spectral properties.
Wenxuan He +4 more
doaj +1 more source
Semantic segmentation with deep learning networks has become an important approach to the extraction of objects from very high-resolution remote sensing images.
Jia Song, A-Xing Zhu, Yunqiang Zhu
doaj +1 more source
Prior works have proposed several strategies to reduce the computational cost of self-attention mechanism. Many of these works consider decomposing the self-attention procedure into regional and local feature extraction procedures that each incurs a much smaller computational complexity.
Ting Yao 0003 +5 more
openaire +3 more sources
Attention-based neural networks such as the Vision Transformer (ViT) have recently attained state-of-the-art results on many computer vision benchmarks. Scale is a primary ingredient in attaining excellent results, therefore, understanding a model's scaling properties is a key to designing future generations effectively.
Xiaohua Zhai +3 more
openaire +4 more sources
A Review of Transformer-Based Approaches for Image Captioning
Visual understanding is a research area that bridges the gap between computer vision and natural language processing. Image captioning is a visual understanding task in which natural language descriptions of images are automatically generated using ...
Oscar Ondeng, Heywood Ouma, Peter Akuon
doaj +1 more source
STHarDNet: Swin Transformer with HarDNet for MRI Segmentation
In magnetic resonance imaging (MRI) segmentation, conventional approaches utilize U-Net models with encoder–decoder structures, segmentation models using vision transformers, or models that combine a vision transformer with an encoder–decoder model ...
Yeonghyeon Gu +2 more
doaj +1 more source
LAVT: Language-Aware Vision Transformer for referring image segmentation [PDF]
Referring image segmentation is a fundamental vision-language task that aims to segment out an object referred to by a natural language expression from an image.
Zhao, H +12 more
core +1 more source
Vision Transformer with Progressive Sampling [PDF]
Accepted to ICCV ...
Xiaoyu Yue +6 more
openaire +4 more sources
Diverse features discovery transformer for pedestrian attribute recognition [PDF]
Recently, Swin Transformer has been widely explored as a general backbone for computer vision, which helps to improve the performance of vision tasks due to the ability to establish associations for long-range dependencies of different spatial locations.
Hussain, Amir +5 more
core +1 more source
Building Extraction With Vision Transformer [PDF]
Submitted to ...
Libo Wang +3 more
openaire +2 more sources

