Results 1 to 10 of about 298,825 (343)

Least squares quantization in PCM

open access: yesIEEE Transactions on Information Theory, 1982
S. Lloyd
exaly   +2 more sources

SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models [PDF]

open access: yesInternational Conference on Machine Learning, 2022
Large language models (LLMs) show excellent performance but are compute- and memory-intensive. Quantization can reduce memory and accelerate inference. However, existing methods cannot maintain accuracy and hardware efficiency at the same time.
Guang-Xuan Xiao   +5 more
semanticscholar   +1 more source

Finite Scalar Quantization: VQ-VAE Made Simple [PDF]

open access: yesInternational Conference on Learning Representations, 2023
We propose to replace vector quantization (VQ) in the latent representation of VQ-VAEs with a simple scheme termed finite scalar quantization (FSQ), where we project the VAE representation down to a few dimensions (typically less than 10). Each dimension
Fabian Mentzer   +3 more
semanticscholar   +1 more source

Quantized Graph Neural Networks for Image Classification

open access: yesMathematics, 2023
Researchers have resorted to model quantization to compress and accelerate graph neural networks (GNNs). Nevertheless, several challenges remain: (1) quantization functions overlook outliers in the distribution, leading to increased quantization errors; (
Xinbiao Xu   +3 more
doaj   +1 more source

A Survey of Quantization Methods for Efficient Neural Network Inference [PDF]

open access: yesLow-Power Computer Vision, 2021
As soon as abstract mathematical computations were adapted to computation on digital computers, the problem of efficient representation, manipulation, and communication of the numerical values in those computations arose.
A. Gholami   +5 more
semanticscholar   +1 more source

OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models [PDF]

open access: yesInternational Conference on Learning Representations, 2023
Large language models (LLMs) have revolutionized natural language processing tasks. However, their practical deployment is hindered by their immense memory and computation requirements.
Wenqi Shao   +9 more
semanticscholar   +1 more source

Geographic distribution and impacts of climate change on the suitable habitats of Rhamnus utilis Decne in China

open access: yesBMC Plant Biology, 2023
Background Rhamnus utilis Decne (Rhamnaceae) is an ecologically and economically important tree species. The growing market demands and recent anthropogenic impacts to R. utilis forests has negatively impacted its populations severely. However, little is
Song Guiquan   +7 more
doaj   +1 more source

Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference [PDF]

open access: yes2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2017
The rising popularity of intelligent mobile devices and the daunting computational cost of deep learning-based models call for efficient and accurate on-device inference schemes.
Benoit Jacob   +7 more
semanticscholar   +1 more source

QuIP: 2-Bit Quantization of Large Language Models With Guarantees [PDF]

open access: yesNeural Information Processing Systems, 2023
This work studies post-training parameter quantization in large language models (LLMs). We introduce quantization with incoherence processing (QuIP), a new method based on the insight that quantization benefits from incoherent weight and Hessian matrices,
Jerry Chee   +3 more
semanticscholar   +1 more source

SqueezeLLM: Dense-and-Sparse Quantization [PDF]

open access: yesInternational Conference on Machine Learning, 2023
Generative Large Language Models (LLMs) have demonstrated remarkable results for a wide range of tasks. However, deploying these models for inference has been a significant challenge due to their unprecedented resource requirements.
Sehoon Kim   +7 more
semanticscholar   +1 more source

Home - About - Disclaimer - Privacy