Results 21 to 30 of about 25,505 (263)

Convolutional Neural Networks or Vision Transformers: Who Will Win the Race for Action Recognitions in Visual Data?

open access: yesSensors, 2023
Understanding actions in videos remains a significant challenge in computer vision, which has been the subject of several pieces of research in the last decades.
Oumaima Moutik   +6 more
doaj   +1 more source

Super Vision Transformer

open access: yesInternational Journal of Computer Vision, 2023
We attempt to reduce the computational costs in vision transformers (ViTs), which increase quadratically in the token number. We present a novel training paradigm that trains only one ViT model at a time, but is capable of providing improved image recognition performance with various computational costs. Here, the trained ViT model, termed super vision
Mingbao Lin   +6 more
openaire   +3 more sources

Transformer-based ripeness segmentation for tomatoes

open access: yesSmart Agricultural Technology, 2023
With the recent development of computer vision technology, various computer vision techniques have been applied to agriculture. Recently, the Transformer network has been introduced to image recognition, which allows a different approach to extracting ...
Risa Shinoda   +3 more
doaj   +1 more source

Transformer architectures for computer vision: A comprehensive review and future research directions [PDF]

open access: yesEPJ Web of Conferences
Long-range dependencies and contextual relationships in videos were captured by using Convolutional Neural Networks (CNNs) in past. Recently the use of Transformers is started for capturing the long-range dependencies and contextual relationships in ...
Ugile Tukaram, Uke Nilesh
doaj   +1 more source

Vicinity Vision Transformer

open access: yesIEEE Transactions on Pattern Analysis and Machine Intelligence, 2023
code: https://github.com/OpenNLPLab/Vicinity-Vision ...
Weixuan Sun   +9 more
openaire   +5 more sources

Vision Transformers Are Robust Learners

open access: yesProceedings of the AAAI Conference on Artificial Intelligence, 2022
Transformers, composed of multiple self-attention layers, hold strong promises toward a generic learning primitive applicable to different data modalities, including the recent breakthroughs in computer vision achieving state-of-the-art (SOTA) standard accuracy. What remains largely unexplored is their robustness evaluation and attribution.
Sayak Paul, Pin-Yu Chen
openaire   +3 more sources

Art authentication with vision transformers

open access: yesNeural Computing and Applications, 2023
AbstractIn recent years, transformers, initially developed for language, have been successfully applied to visual tasks. Vision transformers have been shown to push the state of the art in a wide range of tasks, including image classification, object detection, and semantic segmentation.
Schaerf, Ludovica   +2 more
openaire   +4 more sources

Supervised deep learning with vision transformer predicts delirium using limited lead EEG

open access: yesScientific Reports, 2023
As many as 80% of critically ill patients develop delirium increasing the need for institutionalization and higher morbidity and mortality. Clinicians detect less than 40% of delirium when using a validated screening tool.
Malissa A. Mulkey   +4 more
doaj   +1 more source

Optimal Topology of Vision Transformer for Real-Time Video Action Recognition in an End-To-End Cloud Solution

open access: yesMachine Learning and Knowledge Extraction, 2023
This study introduces an optimal topology of vision transformers for real-time video action recognition in a cloud-based solution. Although model performance is a key criterion for real-time video analysis use cases, inference latency plays a more ...
Saman Sarraf, Milton Kabia
doaj   +1 more source

QuadTree Attention for Vision Transformers

open access: yesCoRR, 2022
Transformers have been successful in many vision tasks, thanks to their capability of capturing long-range dependency. However, their quadratic computational complexity poses a major obstacle for applying them to vision tasks requiring dense predictions, such as object detection, feature matching, stereo, etc.
Tang, Shitao   +3 more
openaire   +4 more sources

Home - About - Disclaimer - Privacy