Abstract
Smaller activation outliers do not necessarily imply better low-bit quantization: their alignment with the quantizer matters. We introduce PrismQuant, a quantizer-aware rotation framework that aligns the leading activation eigenspace with the constant group subspace of asymmetric grouped INT4. The affine offsets represent the energy in this subspace without widening the range within the group. We formulate rotation design as a Ky Fan trace maximization and derive a closed-form solution that is provably optimal for this alignment objective. Compact Householder transformations and their compact-WY representation enable gradient-free construction and efficient application at both foldable and online sites. A predictive range law further connects unaligned activation energy and group size to quantization-relevant variation. Experiments on Llama, Qwen, and Mistral span dense models up to 70B parameters and a 30B mixture-of-experts model. Under W4A4KV4, PrismQuant sets the state of the art on Llama-3.2-3B among the compared methods in both perplexity and accuracy. On Llama-3.1-70B, it attains 3.85 perplexity and 72.46% average zero-shot accuracy, only 0.22 percentage points below full precision. In the deployment study on Llama-3.1-8B, our optimized implementation achieves 1.51x prefill and 1.22x CUDA Graph decode speedups over matched FP16 baselines, with 56.34% lower decode peak memory and only 2.35% additional Graph decode latency over Hadamard. Code is available at https://github.com/ForeverBlue816/PrismQuant.
Community
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- SubRot: Signed Gradient Subspace Calibration for VLM Rotation Quantization (2026)
- SPHQuant: Efficient extreme low bit weight quantization for Vision-Language Models (2026)
- ThinQuant: Scalable Rotation Learning for Weight and Activation Quantization of LLMs (2026)
- Component Type, Not Reconstruction Error, Predicts Attention Quantization Sensitivity (2026)
- When Local Variance Optimality Is Not Enough: RoPE-Aligned Q/K Rotations for Dynamic 4-Bit Quantisation (2026)
- Low-Dimensional High-Leverage Subspace Optimization: Beyond Full-Parameter Coupled Training for Neural Network Quantization (2026)
- SandwichQuant: Which Parameters Matter Before and After Quantization? (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
🔷 What if better 4-bit quantization is not about making outliers smaller, but putting them in the right place? Meet PrismQuant: we rotate dominant activation directions into group-constant components that the quantizer’s affine offsets can absorb, leaving less variation for the 4-bit codes to represent. The rotation is constructed in closed form, is provably optimal for this alignment objective, and requires no gradient-based training. On Llama-3.1-70B with W4A4KV4, PrismQuant achieves 72.46% average zero-shot accuracy, just 0.22 percentage points below full precision. We’ve released code, calibrated rotation factors, and research checkpoints for eight Llama and Qwen models, spanning 0.6B to 70B, including Qwen3-30B-A3B MoE. Start with the 0.6B model in a few lines of Python, explore the paper, and let us know what you think: are we optimizing the size of outliers when we should be optimizing their alignment?
💻 Code: https://github.com/ForeverBlue816/PrismQuant
🤗 Models: https://huggingface.co/collections/ForeverBlue/prsimquant
🤖 Want a quick AI-generated walkthrough of the paper?
AlphaXiv has a structured explanation of PrismQuant here:
https://www.alphaxiv.org/abs/2609.32429
Get this paper in your agent:
hf papers read 2609.32429 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 5
ForeverBlue/PrismQuant-Llama-3.1-8B
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
