Results 21 to 30 of about 298,825 (343)

Quantization and Deployment of Deep Neural Networks on Microcontrollers

open access: yesSensors, 2021
Embedding Artificial Intelligence onto low-power devices is a challenging task that has been partly overcome with recent advances in machine learning and hardware design.
Pierre-Emmanuel Novac   +4 more
doaj   +1 more source

Optimal Brain Compression: A Framework for Accurate Post-Training Quantization and Pruning [PDF]

open access: yesNeural Information Processing Systems, 2022
We consider the problem of model compression for deep neural networks (DNNs) in the challenging one-shot/post-training setting, in which we are given an accurate trained model, and must compress it without any retraining, based only on a small amount of ...
Elias Frantar, Dan Alistarh
semanticscholar   +1 more source

Training and Inference of Optical Neural Networks with Noise and Low-Bits Control

open access: yesApplied Sciences, 2021
Optical neural networks (ONNs) are getting more and more attention due to their advantages such as high-speed and low power consumption. However, in a non-ideal environment, the noise and low-bits control may heavily lead to a decrease in the accuracy of
Danni Zhang   +9 more
doaj   +1 more source

OliVe: Accelerating Large Language Models via Hardware-friendly Outlier-Victim Pair Quantization [PDF]

open access: yesInternational Symposium on Computer Architecture, 2023
Transformer-based large language models (LLMs) have achieved great success with the growing model size. LLMs' size grows by 240× every two years, which outpaces the hardware progress and makes model inference increasingly costly.
Cong Guo   +8 more
semanticscholar   +1 more source

PTQD: Accurate Post-Training Quantization for Diffusion Models [PDF]

open access: yesNeural Information Processing Systems, 2023
Diffusion models have recently dominated image synthesis tasks. However, the iterative denoising process is expensive in computations at inference time, making diffusion models less practical for low-latency and scalable real-world applications.
Yefei He   +5 more
semanticscholar   +1 more source

Post-Training Quantization on Diffusion Models [PDF]

open access: yesComputer Vision and Pattern Recognition, 2022
Denoising diffusion (score-based) generative models have recently achieved significant accomplishments in generating realistic and diverse data. Unfortunately, the generation process of current denoising diffusion models is notoriously slow due to the ...
Yuzhang Shang   +4 more
semanticscholar   +1 more source

Geometric Quantization and Foliation Reduction [PDF]

open access: yes, 2013
A standard question in the study of geometric quantization is whether symplectic reduction interacts nicely with the quantized theory, and in particular whether “quantization commutes with reduction.” Guillemin and Sternberg first proposed ...
Skerritt, Paul Michael
core   +1 more source

LoftQ: LoRA-Fine-Tuning-Aware Quantization for Large Language Models [PDF]

open access: yesInternational Conference on Learning Representations, 2023
Quantization is an indispensable technique for serving Large Language Models (LLMs) and has recently found its way into LoRA fine-tuning. In this work we focus on the scenario where quantization and LoRA fine-tuning are applied together on a pre-trained ...
Yi-Xiao Li   +6 more
semanticscholar   +1 more source

Quantized Feature Distillation for Network Quantization

open access: yesProceedings of the AAAI Conference on Artificial Intelligence, 2023
Neural network quantization aims to accelerate and trim full-precision neural network models by using low bit approximations. Methods adopting the quantization aware training (QAT) paradigm have recently seen a rapid growth, but are often conceptually complicated.
Ke Zhu, Yin-Yin He, Jianxin Wu 0001
openaire   +3 more sources

QDrop: Randomly Dropping Quantization for Extremely Low-bit Post-Training Quantization [PDF]

open access: yesInternational Conference on Learning Representations, 2022
Recently, post-training quantization (PTQ) has driven much attention to produce efficient neural networks without long-time retraining. Despite its low cost, current PTQ works tend to fail under the extremely low-bit setting.
Xiuying Wei   +4 more
semanticscholar   +1 more source

Home - About - Disclaimer - Privacy