Results 1 to 10 of about 1,457,631 (249)
A Survey on Vision Transformer [PDF]
Transformer, first applied to the field of natural language processing, is a type of deep neural network mainly based on the self-attention mechanism. Thanks to its strong representation capabilities, researchers are looking at ways to apply transformer to computer vision tasks.
Xinghao Chen, Yunhe Wang, Jianyuan Guo
exaly +3 more sources
ViTT: Vision Transformer Tracker
This paper presents a new model for multi-object tracking (MOT) with a transformer. MOT is a spatiotemporal correlation task among interest objects and one of the crucial technologies of multi-unmanned aerial vehicles (Multi-UAV).
Xiaoning Zhu +4 more
doaj +2 more sources
Review of Transformer in Computer Vision [PDF]
Transformer is an attention-based encoder-decoder architecture.Due to its long-range sequence modeling and parallel computing capability,Transformer have made a significant breakthrough in natural language processing and is gradually expanding to ...
CHEN Luoxuan, LIN Chengchuang, ZHENG Zhaoliang, MO Zefeng, HUANG Xinyi, ZHAO Gansen
doaj +2 more sources
code: https://github.com/OpenNLPLab/Vicinity-Vision ...
Weixuan Sun +9 more
openaire +6 more sources
Vision Transformer in Industrial Visual Inspection
Artificial intelligence as an approach to visual inspection in industrial applications has been considered for decades. Recent successes, driven by advances in deep learning, present a potential paradigm shift and have the potential to facilitate an ...
Nils Hütten +2 more
doaj +3 more sources
Multiscale Vision Transformers [PDF]
Technical ...
Haoqi Fan 0001 +6 more
openaire +3 more sources
Efficient and accurate rice identification based on high spatial and temporal resolution remote sensing imagery is essential for achieving precision agriculture and ensuring food security.
Huiyao Xu, Jia Song, Yunqiang Zhu
doaj +1 more source
Reversible Vision Transformers
We present Reversible Vision Transformers, a memory efficient architecture design for visual recognition. By decoupling the GPU memory requirement from the depth of the model, Reversible Vision Transformers enable scaling up architectures with efficient memory usage.
Karttikeya Mangalam +6 more
openaire +3 more sources
Transformers in Vision: A Survey [PDF]
Astounding results from Transformer models on natural language tasks have intrigued the vision community to study their application to computer vision problems. Among their salient benefits, Transformers enable modeling long dependencies between input sequence elements and support parallel processing of sequence as compared to recurrent networks, e.g.,
Salman H. Khan 0001 +5 more
openaire +4 more sources
Human vision possesses a special type of visual processing systems called peripheral vision. Partitioning the entire visual field into multiple contour regions based on the distance to the center of our gaze, the peripheral vision provides us the ability to perceive various visual features at different regions.
Juhong Min +3 more
openaire +3 more sources

