Results 11 to 20 of about 1,457,471 (199)
ViTT: Vision Transformer Tracker
This paper presents a new model for multi-object tracking (MOT) with a transformer. MOT is a spatiotemporal correlation task among interest objects and one of the crucial technologies of multi-unmanned aerial vehicles (Multi-UAV).
Xiaoning Zhu +4 more
doaj +2 more sources
Review of Transformer in Computer Vision [PDF]
Transformer is an attention-based encoder-decoder architecture.Due to its long-range sequence modeling and parallel computing capability,Transformer have made a significant breakthrough in natural language processing and is gradually expanding to ...
CHEN Luoxuan, LIN Chengchuang, ZHENG Zhaoliang, MO Zefeng, HUANG Xinyi, ZHAO Gansen
doaj +2 more sources
Vision Transformer in Industrial Visual Inspection
Artificial intelligence as an approach to visual inspection in industrial applications has been considered for decades. Recent successes, driven by advances in deep learning, present a potential paradigm shift and have the potential to facilitate an ...
Nils Hütten +2 more
doaj +3 more sources
Distinguishing Malicious Drones Using Vision Transformer
Drones are commonly used in numerous applications, such as surveillance, navigation, spraying pesticides in autonomous agricultural systems, various military services, etc., due to their variable sizes and workloads.
Sonain Jamil +2 more
doaj +3 more sources
Privacy-Preserving Semantic Segmentation Using Vision Transformer
In this paper, we propose a privacy-preserving semantic segmentation method that uses encrypted images and models with the vision transformer (ViT), called the segmentation transformer (SETR).
Hitoshi Kiya +3 more
doaj +3 more sources
Gait-ViT: Gait Recognition with Vision Transformer
Identifying an individual based on their physical/behavioral characteristics is known as biometric recognition. Gait is one of the most reliable biometrics due to its advantages, such as being perceivable at a long distance and difficult to replicate ...
Jashila Nair Mogan +3 more
doaj +3 more sources
Efficient and accurate rice identification based on high spatial and temporal resolution remote sensing imagery is essential for achieving precision agriculture and ensuring food security.
Huiyao Xu, Jia Song, Yunqiang Zhu
doaj +1 more source
CSiT: A Multiscale Vision Transformer for Hyperspectral Image Classification
The hyperspectral image (HSI) has nearly continuous spectral information; thus, the target of interest can be accurately identified by the subtle details of spectral properties.
Wenxuan He +4 more
doaj +1 more source
Semantic segmentation with deep learning networks has become an important approach to the extraction of objects from very high-resolution remote sensing images.
Jia Song, A-Xing Zhu, Yunqiang Zhu
doaj +1 more source
A Review of Transformer-Based Approaches for Image Captioning
Visual understanding is a research area that bridges the gap between computer vision and natural language processing. Image captioning is a visual understanding task in which natural language descriptions of images are automatically generated using ...
Oscar Ondeng, Heywood Ouma, Peter Akuon
doaj +1 more source

