Vishva007/Qwen3.5-9B-W4A16-AutoRound-AWQ
This is a W4A16 (4-bit weight, 16-bit activation) AWQ-format quantized version of Qwen/Qwen3.5-9B, produced using AutoRound — Intel's sign gradient descent based quantization method designed for production-grade accuracy retention.
Quantization Details
| Parameter |
Value |
| Method |
AutoRound (W4A16, AWQ format) |
| Group Size |
128 |
| Symmetric |
Yes |
| Iterations |
800 |
| Calibration Samples |
512 |
| Sequence Length |
2048 |
| Torch Compile |
Enabled |
Key Notes
- AWQ format — Exported in the standard AWQ format, optimized for efficient inference with activation-aware weight clipping.
- High accuracy configuration — 800 iterations with 512 calibration samples targets production-grade quality with minimal degradation from the base model.
- W4A16 — Weights are quantized to 4-bit integers; activations remain in FP16 for inference stability.
- ~50% memory reduction compared to the FP16 base model, enabling deployment on consumer and mid-range GPUs.
Usage
This model is compatible with transformers, AutoAWQ, vLLM, and SGLang — any backend supporting AWQ-format weights works out of the box. For full model details, architecture, and capabilities, refer to the base model page.
🚀 Deploy on RunPod
One-click launch environments pre-configured with PyTorch, CUDA, and dependencies for fine-tuning or quantization.
🎁 Need GPU compute? Sign up via RunPod and get $5–$500 in free credits when you add your first $10.
PyTorch 2.14
| Template |
CUDA Version |
Docker Image |
Template ID |
Deploy |
| PyTorch 2.14 (CUDA 12.6) |
12.6 |
vishva123/cuda-12.6-pytorch-2.14-runpod |
d7lxsa4w9m |
 |
| PyTorch 2.14 (CUDA 13.0) |
13.0 |
vishva123/cuda-13.0-pytorch-2.14-runpod |
yk0y6j6rpg |
 |
| PyTorch 2.14 (CUDA 13.2) |
13.2 |
vishva123/cuda-13.2-pytorch-2.14-runpod |
gsp4gwx0nw |
 |
PyTorch 2.13
| Template |
CUDA Version |
Docker Image |
Template ID |
Deploy |
| PyTorch 2.13 (CUDA 12.6) |
12.6 |
vishva123/cuda-12.6-pytorch-2.13-runpod |
gmlupxnxfk |
 |
| PyTorch 2.13 (CUDA 13.0) |
13.0 |
vishva123/cuda-13.0-pytorch-2.13-runpod |
y3j8xvk4f4 |
 |
| PyTorch 2.13 (CUDA 13.2) |
13.2 |
vishva123/cuda-13.2-pytorch-2.13-runpod |
vigpissn5w |
 |
PyTorch 2.12
| Template |
CUDA Version |
Docker Image |
Template ID |
Deploy |
| PyTorch 2.12 (CUDA 12.6) |
12.6 |
vishva123/cuda-12.6-pytorch-2.12-runpod |
ctmz86zmf0 |
 |
| PyTorch 2.12 (CUDA 13.0) |
13.0 |
vishva123/cuda-13.0-pytorch-2.12-runpod |
qjko5yiwzi |
 |
| PyTorch 2.12 (CUDA 13.2) |
13.2 |
vishva123/cuda-13.2-pytorch-2.12-runpod |
ifg6xmye0f |
 |