Results 291 to 300 of about 298,825 (343)
Some of the next articles are maybe not open access.
SpinQuant: LLM quantization with learned rotations
International Conference on Learning RepresentationsPost-training quantization (PTQ) techniques applied to weights, activations, and the KV cache greatly reduce memory usage, latency, and power consumption of Large Language Models (LLMs), but may lead to large quantization errors when outliers are present.
Zechun Liu +8 more
semanticscholar +1 more source
QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
International Conference on Machine LearningPost-training quantization (PTQ) reduces the memory footprint of LLMs by quantizing their weights to low-precision. In this work, we introduce QuIP#, a weight-only PTQ method that achieves state-of-the-art results in extreme compression regimes (≤ 4 bits
Albert Tseng +4 more
semanticscholar +1 more source
Extreme Compression of Large Language Models via Additive Quantization
International Conference on Machine LearningThe emergence of accurate open large language models (LLMs) has led to a race towards performant quantization techniques which can enable their execution on end-user devices.
Vage Egiazarian +5 more
semanticscholar +1 more source
QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving
Conference on Machine Learning and SystemsQuantization can accelerate large language model (LLM) inference. Going beyond INT8 quantization, the research community is actively exploring even lower precision, such as INT4.
Yu-Jun Lin +6 more
semanticscholar +1 more source
No Lagrangian? No quantization!
Journal of Mathematical Physics, 1991This work starts with classical equations of motion and sets very general quantization conditions (commutation relations). It is proved that these conditions imply that the equations of motion are equivalent to the Euler–Lagrange equations of a Lagrangian L. The result is a generalization of work by Feynman, recently reported by Dyson [Am. J. Phys. 58,
Hojman, Sergio A., Shepley, L. C.
openaire +2 more sources
Physical Review Letters, 1991
Summary: We study a rule for quantizing chaos based on the dynamical zeta function defined by a Euler product over the classical periodic orbits as suggested by Gutzwiller's semiclassical trace formula. A test of our approximate quantization formula is carried out for the planar hyperbold billiard, which shows that at least the first 150 quantum energy
Sieber, M., Steiner, F.
openaire +3 more sources
Summary: We study a rule for quantizing chaos based on the dynamical zeta function defined by a Euler product over the classical periodic orbits as suggested by Gutzwiller's semiclassical trace formula. A test of our approximate quantization formula is carried out for the planar hyperbold billiard, which shows that at least the first 150 quantum energy
Sieber, M., Steiner, F.
openaire +3 more sources
EfficientQAT: Efficient Quantization-Aware Training for Large Language Models
Annual Meeting of the Association for Computational LinguisticsLarge language models (LLMs) are crucial in modern natural language processing and artificial intelligence. However, they face challenges in managing their significant memory requirements.
Mengzhao Chen +7 more
semanticscholar +1 more source
Optimal Input and Quantization Interval for Quantized Feedback System With Variable Quantizer
IEEE Transactions on Industrial Electronics, 2017Networked control systems (NCSs) have received much attention in the field of robot teleoperation, telesurgery, and other applications. In an NCS, it is important to compress data transmitted over a limited communication network while preserving the data required for control.
Tadanao Zanma +2 more
openaire +2 more sources
FlatQuant: Flatness Matters for LLM Quantization
International Conference on Machine LearningRecently, quantization has been widely used for the compression and acceleration of large language models (LLMs). Due to the outliers in LLMs, it is crucial to flatten weights and activations to minimize quantization error with equally spaced ...
Yuxuan Sun +12 more
semanticscholar +1 more source
International Conference on Learning Representations
Diffusion transformers have demonstrated remarkable performance in visual generation tasks, such as generating realistic images or videos based on textual instructions.
Tianchen Zhao +12 more
semanticscholar +1 more source
Diffusion transformers have demonstrated remarkable performance in visual generation tasks, such as generating realistic images or videos based on textual instructions.
Tianchen Zhao +12 more
semanticscholar +1 more source

