Results 291 to 300 of about 298,825 (343)
Some of the next articles are maybe not open access.

SpinQuant: LLM quantization with learned rotations

International Conference on Learning Representations
Post-training quantization (PTQ) techniques applied to weights, activations, and the KV cache greatly reduce memory usage, latency, and power consumption of Large Language Models (LLMs), but may lead to large quantization errors when outliers are present.
Zechun Liu   +8 more
semanticscholar   +1 more source

QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

International Conference on Machine Learning
Post-training quantization (PTQ) reduces the memory footprint of LLMs by quantizing their weights to low-precision. In this work, we introduce QuIP#, a weight-only PTQ method that achieves state-of-the-art results in extreme compression regimes (≤ 4 bits
Albert Tseng   +4 more
semanticscholar   +1 more source

Extreme Compression of Large Language Models via Additive Quantization

International Conference on Machine Learning
The emergence of accurate open large language models (LLMs) has led to a race towards performant quantization techniques which can enable their execution on end-user devices.
Vage Egiazarian   +5 more
semanticscholar   +1 more source

QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving

Conference on Machine Learning and Systems
Quantization can accelerate large language model (LLM) inference. Going beyond INT8 quantization, the research community is actively exploring even lower precision, such as INT4.
Yu-Jun Lin   +6 more
semanticscholar   +1 more source

No Lagrangian? No quantization!

Journal of Mathematical Physics, 1991
This work starts with classical equations of motion and sets very general quantization conditions (commutation relations). It is proved that these conditions imply that the equations of motion are equivalent to the Euler–Lagrange equations of a Lagrangian L. The result is a generalization of work by Feynman, recently reported by Dyson [Am. J. Phys. 58,
Hojman, Sergio A., Shepley, L. C.
openaire   +2 more sources

Quantization of chaos

Physical Review Letters, 1991
Summary: We study a rule for quantizing chaos based on the dynamical zeta function defined by a Euler product over the classical periodic orbits as suggested by Gutzwiller's semiclassical trace formula. A test of our approximate quantization formula is carried out for the planar hyperbold billiard, which shows that at least the first 150 quantum energy
Sieber, M., Steiner, F.
openaire   +3 more sources

EfficientQAT: Efficient Quantization-Aware Training for Large Language Models

Annual Meeting of the Association for Computational Linguistics
Large language models (LLMs) are crucial in modern natural language processing and artificial intelligence. However, they face challenges in managing their significant memory requirements.
Mengzhao Chen   +7 more
semanticscholar   +1 more source

Optimal Input and Quantization Interval for Quantized Feedback System With Variable Quantizer

IEEE Transactions on Industrial Electronics, 2017
Networked control systems (NCSs) have received much attention in the field of robot teleoperation, telesurgery, and other applications. In an NCS, it is important to compress data transmitted over a limited communication network while preserving the data required for control.
Tadanao Zanma   +2 more
openaire   +2 more sources

FlatQuant: Flatness Matters for LLM Quantization

International Conference on Machine Learning
Recently, quantization has been widely used for the compression and acceleration of large language models (LLMs). Due to the outliers in LLMs, it is crucial to flatten weights and activations to minimize quantization error with equally spaced ...
Yuxuan Sun   +12 more
semanticscholar   +1 more source

ViDiT-Q: Efficient and Accurate Quantization of Diffusion Transformers for Image and Video Generation

International Conference on Learning Representations
Diffusion transformers have demonstrated remarkable performance in visual generation tasks, such as generating realistic images or videos based on textual instructions.
Tianchen Zhao   +12 more
semanticscholar   +1 more source

Home - About - Disclaimer - Privacy