Quantization and Deployment of Deep Neural Networks on Microcontrollers
Embedding Artificial Intelligence onto low-power devices is a challenging task that has been partly overcome with recent advances in machine learning and hardware design.
Pierre-Emmanuel Novac +4 more
doaj +1 more source
Optimal Brain Compression: A Framework for Accurate Post-Training Quantization and Pruning [PDF]
We consider the problem of model compression for deep neural networks (DNNs) in the challenging one-shot/post-training setting, in which we are given an accurate trained model, and must compress it without any retraining, based only on a small amount of ...
Elias Frantar, Dan Alistarh
semanticscholar +1 more source
Training and Inference of Optical Neural Networks with Noise and Low-Bits Control
Optical neural networks (ONNs) are getting more and more attention due to their advantages such as high-speed and low power consumption. However, in a non-ideal environment, the noise and low-bits control may heavily lead to a decrease in the accuracy of
Danni Zhang +9 more
doaj +1 more source
OliVe: Accelerating Large Language Models via Hardware-friendly Outlier-Victim Pair Quantization [PDF]
Transformer-based large language models (LLMs) have achieved great success with the growing model size. LLMs' size grows by 240× every two years, which outpaces the hardware progress and makes model inference increasingly costly.
Cong Guo +8 more
semanticscholar +1 more source
PTQD: Accurate Post-Training Quantization for Diffusion Models [PDF]
Diffusion models have recently dominated image synthesis tasks. However, the iterative denoising process is expensive in computations at inference time, making diffusion models less practical for low-latency and scalable real-world applications.
Yefei He +5 more
semanticscholar +1 more source
Post-Training Quantization on Diffusion Models [PDF]
Denoising diffusion (score-based) generative models have recently achieved significant accomplishments in generating realistic and diverse data. Unfortunately, the generation process of current denoising diffusion models is notoriously slow due to the ...
Yuzhang Shang +4 more
semanticscholar +1 more source
Geometric Quantization and Foliation Reduction [PDF]
A standard question in the study of geometric quantization is whether symplectic reduction interacts nicely with the quantized theory, and in particular whether “quantization commutes with reduction.” Guillemin and Sternberg first proposed ...
Skerritt, Paul Michael
core +1 more source
LoftQ: LoRA-Fine-Tuning-Aware Quantization for Large Language Models [PDF]
Quantization is an indispensable technique for serving Large Language Models (LLMs) and has recently found its way into LoRA fine-tuning. In this work we focus on the scenario where quantization and LoRA fine-tuning are applied together on a pre-trained ...
Yi-Xiao Li +6 more
semanticscholar +1 more source
Quantized Feature Distillation for Network Quantization
Neural network quantization aims to accelerate and trim full-precision neural network models by using low bit approximations. Methods adopting the quantization aware training (QAT) paradigm have recently seen a rapid growth, but are often conceptually complicated.
Ke Zhu, Yin-Yin He, Jianxin Wu 0001
openaire +3 more sources
QDrop: Randomly Dropping Quantization for Extremely Low-bit Post-Training Quantization [PDF]
Recently, post-training quantization (PTQ) has driven much attention to produce efficient neural networks without long-time retraining. Despite its low cost, current PTQ works tend to fail under the extremely low-bit setting.
Xiuying Wei +4 more
semanticscholar +1 more source

