PrismQuant-Qwen3-30B-A3B-Base

Paper · Code · Loading guide

Official checkpoints for PrismQuant: Optimal Null-Space Rotations for Grouped Quantizers. PrismQuant aligns dominant activation directions with the constant directions of asymmetric quantization groups.

Models

Model key Base model Quantization
qwen3_30b_a3b_base Qwen/Qwen3-30B-A3B-Base W4A4KV4

The MoE checkpoint includes all 48 layers and all 128 experts per layer, with calibrated rotations for each expert. The router remains unquantized.

Quick start

git clone --depth 1 https://github.com/ForeverBlue816/PrismQuant.git
cd PrismQuant
pip install -e .
from prismquant import load_model

loaded = load_model("qwen3_30b_a3b_base", device_map="balanced")
print(loaded.generate("The key idea behind quantization is", max_new_tokens=64))

The loader selects the default PrismQuant checkpoint, downloads its required files at a pinned revision, and obtains the matching base model and tokenizer. These are base models for text completion.

Checkpoint format

The GPTQ INT4 values are stored as dequantized floating-point tensors. The reference runtime applies activation and KV quantization while retaining floating-point storage. Use the PrismQuant loader; this repository is not a standalone Transformers checkpoint or a packed INT4 serving model.

This fp32 reference requires roughly 122 GB for model parameters, plus runtime overhead. Use sufficient aggregate GPU memory and, if needed, max_memory={...} with balanced placement.

config.json declares the PrismQuant artifact format and default models. checkpoints/ contains decoder weights; rotations/ contains the factors required by the loader.

License

Qwen-derived weights retain Apache-2.0.

Downloads last month
396
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ForeverBlue/PrismQuant-Qwen3-30B-A3B-Base

Finetuned
(71)
this model

Collection including ForeverBlue/PrismQuant-Qwen3-30B-A3B-Base

Paper for ForeverBlue/PrismQuant-Qwen3-30B-A3B-Base