BitNest-Llama-3-8B

BitNest nested W4/W8 weight package for Meta-Llama-3-8B, from the paper BitNest: Bit-Nested Speculative Decoding for Memory-Efficient LLM Inference Acceleration. Code: https://github.com/YCC-DAVID/BitNest.

A single 8-bit weight tensor serves two precisions: the target reads all 8 bits (W8A8, near-lossless) and the draft reads only the high 4-bit nibble (W4A8). The high nibble is a GPTQ-W4 model fitted first; the low nibble is the closed-form residual back to W8. The two nibbles are stored as separate planes, so a draft step reads half the bytes and the draft costs no extra memory. Speculative decoding verifies every draft token with the target, so the output is exactly the W8A8 target's greedy output.

Base model meta-llama/Meta-Llama-3-8B
Rotation SpinQuant R1/R2 (learned, W4A8 recipe) fused offline, R4 online Hadamard on down_proj inputs
Weights GPTQ-W4 high plane + residual low plane, group size 128 (embedding 8-bit RTN, lm_head bf16)
Activations symmetric int8, per token (o_proj: per 128-group)
KV cache nested at runtime: draft reads KV4, target reads KV8 (--kv_planes)
Default draft length γ = 3

Files

  • planes.safetensors: per decoder Linear <name>.hi / <name>.lo (uint8 [N, K/2], unsigned nibbles u = q + 8; byte j holds column 2j in its low nibble and 2j+1 in its high nibble), <name>.scale (fp16 [N, K/128]), <name>.bias (if any). draft W4 = (hi - 8) * scale, target W8 = (16 * (hi - 8) + (lo - 8)) * scale / 16.
  • aux.safetensors: rotated embed_tokens, final_norm, R4 Hadamard matrices, rotated lm_head (bf16, final norm fused).
  • meta.json: layout, formulas, activation / KV quantization and RoPE details. config.json, tokenizer: from the base model.
  • rotation/R.bin: the learned SpinQuant rotations (R1, per-layer R2). With it the package can be rebuilt from the base model without rotation training: ROT=rotation/R.bin bash scripts/build_package.sh llama3.

Usage

git clone https://github.com/YCC-DAVID/BitNest && cd BitNest && pip install -r requirements.txt
huggingface-cli download YccHugAi/BitNest-Llama-3-8B --local-dir bitnest_llama3
python tools/make_prompts.py --tokenizer bitnest_llama3 --out prompts/llama3 --domains sharegpt,wiki2
# fp16 AR vs W8A8 AR vs BitNest speculative decoding (2K-token prompts, 256 new tokens)
python bitnest/generate.py --pkg bitnest_llama3 --prompts prompts/llama3/sharegpt.pt \
    --v2_attn --kv_planes --kv_r3 --gamma 3 --gen 256 --n 10 --out results/llama3_sharegpt.json
# PPL of target / draft through the engine
python bitnest/generate.py --pkg bitnest_llama3 --prompts prompts/llama3/wiki2.pt --v2_attn --kv_planes --kv_r3 \
    --skip_fp16 --accept_only --n 1 --gen 16 --ppl_check prompts/llama3/wiki2.pt --out results/llama3_ppl.json
# pure PyTorch reference loader (CPU or GPU, no Triton)
python bitnest/reference_model.py --pkg bitnest_llama3 --mode target --a8

License

These weights are a derivative of meta-llama/Meta-Llama-3-8B and are distributed under the base model's license (Meta Llama 3 Community License). Built with Meta Llama 3.

Citation

@article{yang2026bitnest,
  title   = {BitNest: Bit-Nested Speculative Decoding for Memory-Efficient LLM Inference Acceleration},
  author  = {Yang, Chence and Cheng, Ningxi and Akbari, Arash and Tan, Qitao and Zhu, Qingchan and Zhang, Ci and Yang, Changdi and Wang, Yanzhi and Niu, Wei and Wang, Jinhui and Lu, Jin and Yuan, Geng},
  journal = {arXiv preprint arXiv:2610.02800},
  year    = {2026}
}
Downloads last month
2
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for YccHugAi/BitNest-Llama-3-8B

Finetuned
(613)
this model

Paper for YccHugAi/BitNest-Llama-3-8B