weight and Hessian matrices, i.e., from the weights being even in magnitude and the directions in which it is important to round them accurately being unaligned with the coordinate axes. QuIP consists of two steps: (1) an adaptive rounding procedure minimizing a quadratic proxy objective; (2) efficient pre- and post-processing that ensures weight and Hessian incoherence via multiplication by random orthogonal matrices. We complement QuIP with the first theoretical analysis for an LLM-scale quantization algorithm, and show that our theory also applies to an existing method, OPTQ. Empirically, we find that our incoherence preprocessing improves several existing quantization algorithms and yields the first LLM quantization methods that produce viable results using only two bits per weight. Our code can be found at https://github.com/Cornell-RelaxML/QuIP.
Paper
Similar papers
Peer review
Summary
This work proposes a unified framework for weight quantization with error feedback together with preprocessing and postprocessing transformation that makes the model more quantization-friendly. Authors derive theoretical bounds on the quantization error and investigate the failure cases of OPTQ quantization. The introduced LDLQ method with incoherence processing is evaluated on quantization of LLM for 2,3 and 4 bit quantization.
Strengths
* Paper proposes a unified view on quantization with the objective of minimization of layer-wise MSE loss that involves well-known OPTQ as a particular case. * Authors derive explicit average and worst-case bounds on quantization error. * The introduced incoherence processing is well-motivated and supported by quantitative analysis. * The demonstrated LDLQ failure case example despite being very different from typical cases occurring in practice is still interesting and shows potential limitations of the optimal quantization methods. * Different variants of LDLQ achieve strong performance on quantization of models from OPT family at various bit widths considered. The most impressive result is that the model attains reasonable perplexity and zero-shot accuracy for 2 bit quantization, which is known to be very challenging and all the competitive methods experience massive performance drop. * Overall, the work is well-structured and accompanied with thorough theoretical analysis and empirical study.
Weaknesses
* Method is evaluated only on a single model LLM family. To be sure that the method achieves strong performance on LLM quantization in general one should consider at least one more family of LLM. * Minor. $U^{\prime}$ in formula (4) is not upper unit triangular but rather strictly upper triangular (I guess $U^{\prime} + I$ is supposed to be upper unit triangular). * Minor. This will result in each of its eigenvalues being a random unit vector. I guess the authors meant eigenvectors instead of eigenvalues.
Questions
How well does LDLQ+QuIP perform on other families of LLM? Would be interesting to evaluate QuIP on LLaMA given the popularity and impressive performance of this model family. Since it is not fully open-sourced one could consider the recent Falcon family (Falcon-7B and Falcon-40B) instead.
Rating
7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.
Confidence
4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.
Soundness
3 good
Presentation
3 good
Contribution
3 good
Limitations
Not applicable
Summary
The work presented in this paper introduces quantization with incoherence processing, which enables better quantization with fewer bits per parameter. The authors provide a theoretical analysis for adaptive rounding methods and present experimental results demonstrating performance of 2-bit quantization..To do this, the paper introduces the LDLQ adaptive rounding method and demonstrates its optimality compared to other rounding methods which specify linear feedback U for hessians, and rounding to integers (see Theorem 1). The authors define incoherence and demonstrate its effectiveness in achieving a theoretically superior asymptotic bound for LDLQ in terms of the spectrum of H. Additionally, the authors propose efficient pre and post incoherence processing techniques to transform W and H matrices, eliminating the need for nxn matrix multiplications.
Strengths
The paper encompasses numerous interesting new concepts, theorems, and their corresponding proofs. The extensive experimentation serves to prove the authors' claims effectively. The writing is solid, making the paper relatively accessible despite the number of theorems and proofs. Proposed incoherence seems to improve baseline methods, for example OPTQ (table 6,7 appendix). The utilization of the incoherence technique has the potential to contribute significantly to achieve usable 2-bit quantization. They provided code, which is always a plus.
Weaknesses
1) One significant drawback of the paper is that despite the aforementioned enhancements, the performance of 30-bit 2-bit quantization, although much better than 2 bit OPTQ, still falls short compared to 13b model 4-bit quantization with OPTQ(and even 4 bit RTN), making it practically unusable. This limitation diminishes the practical usability surrounding the work on 2 bit quantization. 2) The concept of incoherence remains unclear until page 4 of the paper, which is problematic considering it is one of the main focal points. I believe the introduction should include a clear definition and intuitive explanation of what incoherence entails. Furthermore, I found it challenging to establish a connection between the intuition provided in line 22 and the definition presented in line 134. 3) It is possible I missed, but it appears that there is no mention or measurement of inference speed presented in the paper.
Questions
1)I may have overlooked some details, but it appears that the paper only show how making H incoherent reduce loss upper bounds, nothing about W incoherence. What is benefit for making also W incoherent? 2) In line 169 you write that "to make symmetric matrix incoherent is to conjugate it by uniform random orthogonal matrix". Can you please add citation or explanation why this is true? 3) Again I may have overlooked some details, in 172- 176 you said that the procedure described there makes the matrix H and W incoherent with high probability/ Why is that, and where is the proof? Suggestions: 1. In line 64, it seems that "OBC" might be a typographical error, and "OBQ" could be the correct term. 2. It would be helpful to include a citation for LDL decomposition. 3.In line 99, it would enhance clarity to present a equation that demonstrates the result obtained when equation (4) is applied in equation (3). 4.In line 217, it appears that "OTPQ" might be a typographical error, and you mean "OPTQ".
Rating
5: Borderline accept: Technically solid paper where reasons to accept outweigh reasons to reject, e.g., limited evaluation. Please use sparingly.
Confidence
3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.
Soundness
3 good
Presentation
3 good
Contribution
3 good
Limitations
Yes
Summary
- The authors coin a family of adaptive rounding methods for layer-wise quantization, establish that the state-of-the-art method GPTQ is optimal within this family, and prove quality guarantees. - The paper further introduces incoherence preprocessing, together with a Kronecker-factor based inference scheme, which leads to significant accuracy improvements for very low bitwidth compression.
Strengths
- The paper introduces a non-obvious and useful optimization to GPTQ which fully eliminates the matrix inversion while obtaining equivalent results. - The authors present what seem like the first theoretical guarantees for an adaptive layerwise rounding algorithm. - The proposed incoherence preprocessing leads to greatly improved results for 2-bit compression. While the idea of multiplying with an orthogonal matrix to produce more uniform and thus easier to compress data is not new, I have not seen it applied in the context of LLM quantization before. - The paper also studies the impact of some additional heuristics, like greedy post-processing passes on top of GPTQ. While those are mostly small tricks, it is good to have those implemented and evaluated. - Code is provided for reproducabilty, including also a script to verify the equivalence of their GPTQ optimization.
Weaknesses
- The method is only evaluated on OPT, which is not considered a great LLM anymore by today's standards; good results on e.g. LLaMa would be significantly stronger. - There do not seem to be any runtime numbers for inference via the proposed efficient Kronecker-factored scheme. I think that the tripling of required FLOPs (2 additional o(n^2) matmuls if I understood this part correctly) may be challenging to implement in practice without significant overheads (both for low- and large-batch inference), which would limit the practicality of incoherence preprocessing. - The algorithm family seems to be somewhat designed around how GPTQ works and the corresponding optimality is thus not exactly super surprising.
Questions
- Given that the largest models are usually more robust to quantization, I am wondering how QuiP performs on the 66B and 175B variants, is there any reason why those results are missing? - How does QuiP perform with groups, which are generally very effective for standard GPTQ / RTN? - 221: GPTQ only requires one inverse + one Cholesky decomposition - 282: It would be good to reference that the official GPTQ repository also proposed a similar trick in the context of LLaMa models several months ago. - The incoherence post-processing code currently returns a matrix that is not quantized anymore, which is a bit confusing. Related to that, I would also suggest to provide pseudocode for the efficient Kronecker-factor inference scheme.
Rating
7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.
Confidence
4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.
Soundness
3 good
Presentation
3 good
Contribution
3 good
Limitations
The paper briefly discusses limitations and broader impact in the Appendix.
Summary
This paper introduces QuIP, an algorithm for weight quantization in large language models, with theoretical guarantees. The experimental results demonstrate that QuIP achieves nearly lossless performance when using 3-bit quantization for models larger than 3B, and it shows good performance with 2-bit quantization for models larger than 7B.
Strengths
First of all, QuIP is the first post-training quantization algorithm that compresses weights to 2 bits with reasonable loss of performance, significantly pushing the boundary of LLM on-device deployment and making larger LLMs available to ordinary users. Second, the paper proposes LDLQ rounding method, which is both worst-case and average-case optimal in a family of adaptive rounding methods, i.e. iterative quantization through each column of weights. The family is carefully selected so that OPTQ (previously known as GPTQ) falls within this family. Third, solid proof is shown that LDLQ is optimal in its family assuming rounding to integers, while a rigid study, with counterexamples and more careful analysis is conducted in more realistic cases in Section 5.2.
Weaknesses
The experiments are mostly conducted on OPT, which might not be the most standard model at the time of reviewing. Models like Llama, Falcon and MPT might be more popular LLMs. However, the reviewer understand the paper was submitted months ago and there might be limited time to conduct all the experiments. Also, the main experiments are conducted on decoder-only auto-regressive language models, while there are other models/domains of interest, for example, FastChat-T5 which is encoder-decoder architecture, Vision transformers, etc.
Questions
As mentioned in Section 5.1., OPTQ is equivalent to LDLQ in the class of adaptive rounding methods with linear feedback (proof in Supplement E), and thus is a special case of LDLQ in general. Is there any further case study why OPTQ performs worse in general? Is there any analysis on quantization outliers under QuIP/LDLQ and how would QuIP handle those cases?
Rating
7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.
Confidence
3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.
Soundness
3 good
Presentation
3 good
Contribution
4 excellent
Limitations
We believe that considering activation quantization could lead to further improvements in the future. However, the implementation of QuIP as a weight quantization method has already overcome major obstacles and enabled the deployment of LLMs. Currently, there is a lack of realistic experiments to assess the performance, in terms of compute speed. Nevertheless, implementing and optimizing 2/3-bit performance on CUDA would likely be a highly challenging task, and it could be worthwhile to explore this further in a separate research paper.
Response to rebuttal
After reading the responses, I decided to keep the original score. LLaMA-2-chat is a good choice for the method validation, but the results provided in response involve only 2-bit compression for a single model on two tasks from `lm-eval-harness`. Individual tasks are known to be quite noisy and limited, therefore, in order to make any definite conclusions about the LLM performance it would be desirable to consider a larger number of tasks (at least 5-6). In addition, I would recommend comparing the performance against the fp16 baseline, since these metrics, without having a reference point (one can indeed check the number of LLaMA-2 paper), are not illustrative. Anyway, given that 2 bits is a challenging target, a significant drop in performance is expected and having reasonable performance at this point is still a good result.
Updated results on LLaMa-2
We conducted a suite of further experiments quantizing Llama-2: 7b and 13b parameter models, both pretrained and chat finetuned. Overall, we’ve verified that our method QuIP works well on this additional model, and can provide a step function improvement in quantization at 2 bits compared to OPTQ. We adapted code from the GPTQ-for-LLaMa repo, which we use to evaluate WikiText and C4 perplexity. We use the lm-evaluation-harness to evaluate downstream zeroshot accuracy on BoolQ, PiQA, WinoGrande, ARC easy, and ARC challenge. | Llama-2 Model | Quant Method | Wbits | Wiki | C4 | BOOLQ | PIQA | WinoGrande | ARC-e | ARC-c | | ------------- | ------------ | ----- | -------- | -------- | ------ | ------ | ---------- | ------ | ------ | | 13b | Full Precision | 16 | 4.884 | 6.727 | 80.52% | 80.52% | 72.22% | 77.48% | 49.15% | | | QuIP | 4 | 5.011 | 6.887 | 80.90% | 80.10% | 72.60% | 77.30% | 49.20% | | | | 3 | 5.340 | 7.343 | 77.40% | 79.10% | 71.10% | 76.30% | 49.00% | | | | 2 | 10.094 | 13.131 | 63.90% | 69.60% | 57.90% | 55.80% | 31.50% | | | OPTQ | 4 | 5.203 | 7.060 | 78.80% | 80.40% | 70.00% | 76.10% | 48.80% | | | | 3 | 6.666 | 8.910 | 74.50% | 76.70% | 69.30% | 69.50% | 42.49% | | | | 2 | 3086.167 | 406.934 | 40.20% | 48.80% | 48.70% | 27.00% | 27.90% | | 13b-chat | Full Precision | 16 | 6.108 | 8.489 | 81.71% | 79.11% | 71.27% | 73.74% | 50.26% | | | QuIP | 4 | 6.275 | 8.733 | 80.80% | 78.40% | 71.20% | 72.50% | 48.60% | | | | 3 | 6.713 | 9.367 | 78.60% | 77.80% | 70.50% | 71.30% | 47.10% | | | | 2 | 16.046 | 20.034 | 58.30% | 67.50% | 56.90% | 53.00% | 33.10% | | | OPTQ | 4 | 6.454 | 8.962 | 79.60% | 78.00% | 70.50% | 72.80% | 49.40% | | | | 3 | 8.393 | 11.399 | 70.80% | 74.40% | 64.70% | 66.70% | 42.20% | | | | 2 | 3136.833 | 1138.701 | 41.90% | 47.90% | 47.50% | 25.50% | 29.60% | | Llama-2 Model | Quant Method | Wbits | Wiki | C4 | BOOLQ | PIQA | WinoGrande | ARC-e | ARC-c | | ------------- | ------------ | ----- | ------ | -------- | ------ | ------ | ---------- | ------ | ------ | | 7b | Full Precision | 16 | 5.472 | 7.263 | 77.77% | 79.11% | 69.06% | 74.54% | 46.25% | | | QuIP | 4 | 5.940 | 8.010 | 75.87% | 77.26% | 67.88% | 71.00% | 42.58% | | | | 3 | 6.499 | 8.738 | 74.92% | 76.28% | 67.01% | 69.15% | 41.47% | | | | 2 | 27.125 | 31.333 | 52.72% | 60.66% | 51.70% | 39.14% | 26.19% | | | OPTQ | 4 | 6.067 | 7.845 | 76.09% | 78.51% | 67.96% | 72.52% | 44.03% | | | | 3 | 9.505 | 11.956 | 68.84% | 74.32% | 63.38% | 62.16% | 37.37% | | | | 2 | NaN | 1794.547 | 41.31% | 48.31% | 48.46% | 26.09% | 27.56% | | 7b-chat | Full Precision | 16 | 7.077 | 9.528 | 80.67% | 76.66% | 66.22% | 69.65% | 44.37% | | | QuIP | 4 | 7.431 | 10.147 | 80.24% | 76.71% | 66.46% | 67.26% | 42.49% | | | | 3 | 8.090 | 11.052 | 72.48% | 76.33% | 65.35% | 67.76% | 40.19% | | | | 2 | 66.586 | 61.662 | 50.18% | 57.67% | 50.43% | 35.27% | 27.90% | | | OPTQ | 4 | 7.791 | 10.965 | 79.02% | 75.68% | 66.54% | 69.36% | 41.13% | | | | 3 | 11.847 | 17.736 | 64.25% | 70.67% | 62.19% | 53.62% | 33.62% | | | | 2 | NaN | NaN | 45.54% | 50.00% | 49.88% | 27.57% | 29.10% |
Response
Thanks for the updates. In my own evaluations, GPTQ performs considerably better for 3,4 bit quantization than the numbers reported in your table (4 bit GPTQ has lower perplexity than the numbers reported). However, for 2 bits I can confirm that the application of GPTQ to Llama breaks the model completely, whereas QuIP quantized models are still usable. Nevertheless, considering again the overall contribution of the work, I decided to keep the original score.
Updated results on LLaMa-2
We conducted a suite of further experiments quantizing Llama-2: 7b and 13b parameter models, both pretrained and chat finetuned. Overall, we’ve verified that our method QuIP works well on this additional model, and can provide a step function improvement in quantization at 2 bits compared to OPTQ. We adapted code from the GPTQ-for-LLaMa repo, which we use to evaluate WikiText and C4 perplexity. We use lm-evaluation-harness to evaluate downstream zeroshot accuracy on BoolQ, PiQA, WinoGrande, ARC easy, and ARC challenge. | Llama-2 Model | Quant Method | Wbits | Wiki | C4 | BOOLQ | PIQA | WinoGrande | ARC-e | ARC-c | | ------------- | ------------ | ----- | -------- | -------- | ------ | ------ | ---------- | ------ | ------ | | 13b | Full Precision | 16 | 4.884 | 6.727 | 80.52% | 80.52% | 72.22% | 77.48% | 49.15% | | | QuIP | 4 | 5.011 | 6.887 | 80.90% | 80.10% | 72.60% | 77.30% | 49.20% | | | | 3 | 5.340 | 7.343 | 77.40% | 79.10% | 71.10% | 76.30% | 49.00% | | | | 2 | 10.094 | 13.131 | 63.90% | 69.60% | 57.90% | 55.80% | 31.50% | | | OPTQ | 4 | 5.203 | 7.060 | 78.80% | 80.40% | 70.00% | 76.10% | 48.80% | | | | 3 | 6.666 | 8.910 | 74.50% | 76.70% | 69.30% | 69.50% | 42.49% | | | | 2 | 3086.167 | 406.934 | 40.20% | 48.80% | 48.70% | 27.00% | 27.90% | | 13b-chat | Full Precision | 16 | 6.108 | 8.489 | 81.71% | 79.11% | 71.27% | 73.74% | 50.26% | | | QuIP | 4 | 6.275 | 8.733 | 80.80% | 78.40% | 71.20% | 72.50% | 48.60% | | | | 3 | 6.713 | 9.367 | 78.60% | 77.80% | 70.50% | 71.30% | 47.10% | | | | 2 | 16.046 | 20.034 | 58.30% | 67.50% | 56.90% | 53.00% | 33.10% | | | OPTQ | 4 | 6.454 | 8.962 | 79.60% | 78.00% | 70.50% | 72.80% | 49.40% | | | | 3 | 8.393 | 11.399 | 70.80% | 74.40% | 64.70% | 66.70% | 42.20% | | | | 2 | 3136.833 | 1138.701 | 41.90% | 47.90% | 47.50% | 25.50% | 29.60% | | Llama-2 Model | Quant Method | Wbits | Wiki | C4 | BOOLQ | PIQA | WinoGrande | ARC-e | ARC-c | | ------------- | ------------ | ----- | ------ | -------- | ------ | ------ | ---------- | ------ | ------ | | 7b | Full Precision | 16 | 5.472 | 7.263 | 77.77% | 79.11% | 69.06% | 74.54% | 46.25% | | | QuIP | 4 | 5.940 | 8.010 | 75.87% | 77.26% | 67.88% | 71.00% | 42.58% | | | | 3 | 6.499 | 8.738 | 74.92% | 76.28% | 67.01% | 69.15% | 41.47% | | | | 2 | 27.125 | 31.333 | 52.72% | 60.66% | 51.70% | 39.14% | 26.19% | | | OPTQ | 4 | 6.067 | 7.845 | 76.09% | 78.51% | 67.96% | 72.52% | 44.03% | | | | 3 | 9.505 | 11.956 | 68.84% | 74.32% | 63.38% | 62.16% | 37.37% | | | | 2 | NaN | 1794.547 | 41.31% | 48.31% | 48.46% | 26.09% | 27.56% | | 7b-chat | Full Precision | 16 | 7.077 | 9.528 | 80.67% | 76.66% | 66.22% | 69.65% | 44.37% | | | QuIP | 4 | 7.431 | 10.147 | 80.24% | 76.71% | 66.46% | 67.26% | 42.49% | | | | 3 | 8.090 | 11.052 | 72.48% | 76.33% | 65.35% | 67.76% | 40.19% | | | | 2 | 66.586 | 61.662 | 50.18% | 57.67% | 50.43% | 35.27% | 27.90% | | | OPTQ | 4 | 7.791 | 10.965 | 79.02% | 75.68% | 66.54% | 69.36% | 41.13% | | | | 3 | 11.847 | 17.736 | 64.25% | 70.67% | 62.19% | 53.62% | 33.62% | | | | 2 | NaN | NaN | 45.54% | 50.00% | 49.88% | 27.57% | 29.10% |
Post-Rebuttal Comments
Thank you for the detailed reply! I am not entirely convinced by your new experimental results, in particular the kernel numbers. 377s for 512 tokens implies ~736ms per token, GPTQ reports 589ms on the same GPU for OPT-175B, a 2x larger model, at FP16. An ideal implementation of a 2-bit kernel for the 2x smaller LLaMa-70B should be close to 16x faster than this (as this is the difference in memory that needs to be loaded, which dominates inference for generation). Hence, your low overheads currently seem to be measured relative to a rather uncompetitive baseline. As for the two Cholesky decompositions, I believe `torch.linalg.cholesky(H)` and `torch.cholesky_inverse(H)` combine to a symmetric matrix inverse that should be faster and more stable than a general inversion. Nevertheless, I will maintain my score based on the merits of the work discussed in my initial review.
Clarification on timing experiments
Thanks for your response. We were not precise enough in our original rebuttal; our timing results were to generate a batchsize of 8, with sequences length 512 each. Therefore our baseline model had an inference throughput of 92ms/token, not 736ms/token. While this is faster than GPTQ's 589ms/token for OPT-175B at fp16 precision, we've identified several factors that make a direct comparison difficult. (1) Our code measures the whole huggingface token generation loop including encoding and decoding, while GPTQ's code just measures the model evaluation. (2) We used 1 GPU, while GPTQ's timing results use 8 GPUs. (3) We used a 2-bit compressed version of Llama-2-70b-chat, while they are using OPT-175B. (4) We used a batch size of 8 while the GPTQ's results used a batch size of 1.
Sorry for the delay in the response.Thank you for taking time to address my questions. I have read the feedback from other reviewers and the authors' rebuttal. Regarding Weakness 1: Thank you for pointing out results for 3-bit quantization. I understand the fact that 2-bit quantization is challenging, and making it work somewhat decently is achievement by itself. However, from a practical standpoint, despite the fact that QuIP outperforms competitive methods at 2-bit, it is still impractical since for a given memory budget it is better to take smaller model with higher bit-width. Taking into account, the paper focus on 2-bit quantization, in my opinion this is still major drawback. I think one way to address this, is to make paper more transparent, for example, by including tables with fixed memory budget to compare different options. Inference speed: Thank you for measuring the inference speed. I read your discussion with Av5n. I would recommend to include to the paper more precise measurements and fair comparison with OPTQ. Apart from these points, I am satisfied with the authors' response. After revisiting the paper, reading other authors' reviews, and the authors' responses, I would like to increase my score from 4 to 5.
Decision
Accept (spotlight)