Results 1 to 10 of about 1,457,631 (249)

A Survey on Vision Transformer [PDF]

open access: yesIEEE Transactions on Pattern Analysis and Machine Intelligence, 2023
Transformer, first applied to the field of natural language processing, is a type of deep neural network mainly based on the self-attention mechanism. Thanks to its strong representation capabilities, researchers are looking at ways to apply transformer to computer vision tasks.
Xinghao Chen, Yunhe Wang, Jianyuan Guo
exaly   +3 more sources

ViTT: Vision Transformer Tracker

open access: yesSensors, 2021
This paper presents a new model for multi-object tracking (MOT) with a transformer. MOT is a spatiotemporal correlation task among interest objects and one of the crucial technologies of multi-unmanned aerial vehicles (Multi-UAV).
Xiaoning Zhu   +4 more
doaj   +2 more sources

Review of Transformer in Computer Vision [PDF]

open access: yesJisuanji kexue, 2023
Transformer is an attention-based encoder-decoder architecture.Due to its long-range sequence modeling and parallel computing capability,Transformer have made a significant breakthrough in natural language processing and is gradually expanding to ...
CHEN Luoxuan, LIN Chengchuang, ZHENG Zhaoliang, MO Zefeng, HUANG Xinyi, ZHAO Gansen
doaj   +2 more sources

Vicinity Vision Transformer

open access: yesIEEE Transactions on Pattern Analysis and Machine Intelligence, 2023
code: https://github.com/OpenNLPLab/Vicinity-Vision ...
Weixuan Sun   +9 more
openaire   +6 more sources

Vision Transformer in Industrial Visual Inspection

open access: yesApplied Sciences, 2022
Artificial intelligence as an approach to visual inspection in industrial applications has been considered for decades. Recent successes, driven by advances in deep learning, present a potential paradigm shift and have the potential to facilitate an ...
Nils Hütten   +2 more
doaj   +3 more sources

Multiscale Vision Transformers [PDF]

open access: yes2021 IEEE/CVF International Conference on Computer Vision (ICCV), 2021
Technical ...
Haoqi Fan 0001   +6 more
openaire   +3 more sources

Evaluation and Comparison of Semantic Segmentation Networks for Rice Identification Based on Sentinel-2 Imagery

open access: yesRemote Sensing, 2023
Efficient and accurate rice identification based on high spatial and temporal resolution remote sensing imagery is essential for achieving precision agriculture and ensuring food security.
Huiyao Xu, Jia Song, Yunqiang Zhu
doaj   +1 more source

Reversible Vision Transformers

open access: yes2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022
We present Reversible Vision Transformers, a memory efficient architecture design for visual recognition. By decoupling the GPU memory requirement from the depth of the model, Reversible Vision Transformers enable scaling up architectures with efficient memory usage.
Karttikeya Mangalam   +6 more
openaire   +3 more sources

Transformers in Vision: A Survey [PDF]

open access: yesACM Computing Surveys, 2022
Astounding results from Transformer models on natural language tasks have intrigued the vision community to study their application to computer vision problems. Among their salient benefits, Transformers enable modeling long dependencies between input sequence elements and support parallel processing of sequence as compared to recurrent networks, e.g.,
Salman H. Khan 0001   +5 more
openaire   +4 more sources

Peripheral Vision Transformer

open access: yesAdvances in Neural Information Processing Systems 35, 2022
Human vision possesses a special type of visual processing systems called peripheral vision. Partitioning the entire visual field into multiple contour regions based on the distance to the center of our gaze, the peripheral vision provides us the ability to perceive various visual features at different regions.
Juhong Min   +3 more
openaire   +3 more sources

Home - About - Disclaimer - Privacy