Results 11 to 20 of about 25,505 (263)
Transformers in Vision: A Survey [PDF]
Astounding results from Transformer models on natural language tasks have intrigued the vision community to study their application to computer vision problems. Among their salient benefits, Transformers enable modeling long dependencies between input sequence elements and support parallel processing of sequence as compared to recurrent networks, e.g.,
Salman H. Khan 0001 +5 more
openaire +4 more sources
Human vision possesses a special type of visual processing systems called peripheral vision. Partitioning the entire visual field into multiple contour regions based on the distance to the center of our gaze, the peripheral vision provides us the ability to perceive various visual features at different regions.
Juhong Min +3 more
openaire +3 more sources
CSiT: A Multiscale Vision Transformer for Hyperspectral Image Classification
The hyperspectral image (HSI) has nearly continuous spectral information; thus, the target of interest can be accurately identified by the subtle details of spectral properties.
Wenxuan He +4 more
doaj +1 more source
Semantic segmentation with deep learning networks has become an important approach to the extraction of objects from very high-resolution remote sensing images.
Jia Song, A-Xing Zhu, Yunqiang Zhu
doaj +1 more source
Prior works have proposed several strategies to reduce the computational cost of self-attention mechanism. Many of these works consider decomposing the self-attention procedure into regional and local feature extraction procedures that each incurs a much smaller computational complexity.
Ting Yao 0003 +5 more
openaire +3 more sources
Attention-based neural networks such as the Vision Transformer (ViT) have recently attained state-of-the-art results on many computer vision benchmarks. Scale is a primary ingredient in attaining excellent results, therefore, understanding a model's scaling properties is a key to designing future generations effectively.
Xiaohua Zhai +3 more
openaire +4 more sources
A Review of Transformer-Based Approaches for Image Captioning
Visual understanding is a research area that bridges the gap between computer vision and natural language processing. Image captioning is a visual understanding task in which natural language descriptions of images are automatically generated using ...
Oscar Ondeng, Heywood Ouma, Peter Akuon
doaj +1 more source
STHarDNet: Swin Transformer with HarDNet for MRI Segmentation
In magnetic resonance imaging (MRI) segmentation, conventional approaches utilize U-Net models with encoder–decoder structures, segmentation models using vision transformers, or models that combine a vision transformer with an encoder–decoder model ...
Yeonghyeon Gu +2 more
doaj +1 more source
Vision Transformer with Progressive Sampling [PDF]
Accepted to ICCV ...
Xiaoyu Yue +6 more
openaire +4 more sources
Building Extraction With Vision Transformer [PDF]
Submitted to ...
Libo Wang +3 more
openaire +2 more sources

