gitcommit90 commited on
Commit
b2e8cc4
·
verified ·
1 Parent(s): 7c83793

Publish independent quantization fidelity measurement

Browse files
Files changed (1) hide show
  1. README.md +12 -0
README.md CHANGED
@@ -69,6 +69,18 @@ Use K=8 for list/JSON-heavy workloads (`ONE_SPARK_K=8`). Raw data in the GitHub
69
 
70
  Every cold request had zero prefix-cache hits; every warm checksum response was correct.
71
 
 
 
 
 
 
 
 
 
 
 
 
 
72
  ## Links
73
 
74
  - **GitHub source and instructions:** `https://github.com/gitcommit90/glm-5.3-one-spark`
 
69
 
70
  Every cold request had zero prefix-cache hits; every warm checksum response was correct.
71
 
72
+ ## Quantization Analysis
73
+
74
+ This deployment uses [Turboderp's GLM-5.3-Flash EXL3 2.05-bpw quant](https://huggingface.co/turboderp/GLM-5.3-Flash-exl3/tree/2.05bpw), revision `51058cd551c7e570d87bd32a4adee720edce2349`. The exact checkpoint is **85.23 GB (79.38 GiB)**.
75
+
76
+ An [independent full-vocabulary measurement](https://github.com/malaiwah/quant-fidelity-suite/blob/794d80fa79db4d30cd0fa8140a07c665dd363251/registry/receipts/malaiwah/stream-turbo-2.05bpw-kld.json) of this exact revision against BF16 teacher logits reported:
77
+
78
+ | Quant | Size | Top-1 agreement | Mean KLD | Scored positions |
79
+ |---|---:|---:|---:|---:|
80
+ | Turboderp EXL3 2.05 | 85.23 GB | **88.92%** | **0.121638** | **51,175** |
81
+
82
+ The measurement used 25 windows, the full 154,880-token vocabulary, teacher forcing, FP64 accumulation, and two cold runs with identical results. The checkpoint was quantized from the official FP8 release; the measurement reference is BF16. Machine-readable summary: [`quantization-analysis.json`](https://github.com/gitcommit90/glm-5.3-one-spark/blob/main/benchmarks/quality/quantization-analysis.json).
83
+
84
  ## Links
85
 
86
  - **GitHub source and instructions:** `https://github.com/gitcommit90/glm-5.3-one-spark`