Results 11 to 20 of about 21,351 (259)
Similarity and Consistency by Self-distillation Method [PDF]
Due to high data pre-processing costs and missing local features detection in self-distillation methods for models compression,a similarity and consistency by self-distillation(SCD) method is proposed to improve model classification accuracy.Firstly ...
WAN Xu, MAO Yingchi, WANG Zibo, LIU Yi, PING Ping
doaj +1 more source
Recurrent Knowledge Distillation [PDF]
Knowledge distillation compacts deep networks by letting a small student network learn from a large teacher network. The accuracy of knowledge distillation recently benefited from adding residual layers. We propose to reduce the size of the student network even further by recasting multiple residual layers in the teacher network into a single recurrent
Pintea, S. (author) +2 more
openaire +3 more sources
Decoupled Knowledge Distillation
Accepted by CVPR2022, fix ...
Borui Zhao +4 more
openaire +2 more sources
Annealing Knowledge Distillation [PDF]
Significant memory and computational requirements of large deep neural networks restrict their application on edge devices. Knowledge distillation (KD) is a prominent model compression technique for deep neural networks in which the knowledge of a trained large teacher model is transferred to a smaller student model.
Aref Jafari +3 more
openaire +2 more sources
A Virtual Knowledge Distillation via Conditional GAN
Knowledge distillation aims at transferring the knowledge from a pre-trained complex model, called teacher, to a relatively smaller and faster one, called student. Unlike previous works that transfer the teacher’s softened distributions or feature
Sihwan Kim
doaj +1 more source
On the Efficacy of Knowledge Distillation [PDF]
13 pages, including ...
Jang Hyun Cho, Bharath Hariharan
openaire +2 more sources
Feature fusion-based collaborative learning for knowledge distillation
Deep neural networks have achieved a great success in a variety of applications, such as self-driving cars and intelligent robotics. Meanwhile, knowledge distillation has received increasing attention as an effective model compression technique for ...
Yiting Li +4 more
doaj +1 more source
Triplet Knowledge Distillation
In Knowledge Distillation, the teacher is generally much larger than the student, making the solution of the teacher likely to be difficult for the student to learn. To ease the mimicking difficulty, we introduce a triplet knowledge distillation mechanism named TriKD. Besides teacher and student, TriKD employs a third role called anchor model.
Xijun Wang 0002 +5 more
openaire +2 more sources
Knowledge distillation in deep learning and its applications [PDF]
Deep learning based models are relatively large, and it is hard to deploy such models on resource-limited devices such as mobile phones and embedded devices.
Abdolmaged Alkhulaifi +2 more
doaj +2 more sources
Reverse Self-Distillation Overcoming the Self-Distillation Barrier
Deep neural networks generally cannot gather more helpful information with limited data in image classification, resulting in poor performance. Self-distillation, as a novel knowledge distillation technique, integrates the roles of teacher and student ...
Shuiping Ni +4 more
doaj +1 more source

