Acknowledge the Responsible Use Agreement to access this repository

Access is granted automatically after you agree to the terms below and submit the form.

Responsible Use Agreement

This model has had safety refusals removed. That makes it useful for red-teaming, security research, evaluation, and unfiltered assistant tasks — and also removes guardrails a user must therefore supply themselves.

Prohibited uses (you must agree before access is granted):

  • Anything involving the sexual exploitation or endangerment of minors.
  • You must be of age 18 years or older to use and download this model.
  • You agree any information generated that can cause harm in terms of generating recipe, knowledge to make any materials/substances is your own input and responsibility. You will be accountable for any harm/damage caused by your action/input.
  • Content promoting self-harm or suicide.
  • Generation of material that is illegal in your jurisdiction, or that targets real individuals for harassment, doxxing, or fraud.
  • Any use prohibited by the upstream Z.AI / GLM MIT license.

You are responsible for adding appropriate safety filtering, human review, and access controls for your deployment. The weights are provided as-is, with no warranty. The license is the upstream Z.AI GLM-5.3-Flash MIT license — review and comply with it before use or redistribution.

Log in or Sign Up to review the conditions and access this model content.

GLM-5.3-Flash EXL3 4bpw TensorFold Ablit

EXL3 4bpw TensorFold weights of GLM-5.3-Flash. Layers 15–43 and MTP layer 45 replace self_attn.o_proj with the Keys tensors. Layers 0–14 and layer 44 stay the parent. Experts, vision, QKV, embeddings, and the head are the parent weights.

drowzeys (GitHub) authored those o_proj tensors. They are the files in drowzeys/keys-GLM-5.3-Flash-NVFP4-ablit-l15-43-mtp-l45 at commit cecebc0ff0bbce8eefaa19667290a79d0216cd0e. Method and scripts: drowzeys/keys-GLM-5.3-Flash-NVFP4-ablit-l15-43-mtp-l45.

dealignai (@dealignai) authored GLM-5.3-Flash-UNCENSORED-NVFP4. jordanschenck ran the compute for that checkpoint. drowzeys published the tensors copied here as an altered copy of that checkpoint's BF16 self_attn.o_proj (layers 15–43 and MTP layer 45; layers 0–14 and 44 left stock).

Mia's AI Lab authored the parent quant, Mia-AiLab/GLM-5.3-Flash-EXL3-4bpw-TensorFold at commit 078455ffe6472f9a52fbc1139f58b9db2881b25c.

Z.AI authored GLM-5.3-Flash. turboderp authored exllamav3. Serving engine: TensorFold.

This repository is gated. Agree to the Responsible Use terms on the form above before download. The license is MIT, inherited from Z.AI GLM-5.3-Flash. See RESPONSIBLE_USE.md.

Credits

What changed

Parent Mia-AiLab/GLM-5.3-Flash-EXL3-4bpw-TensorFold @ 078455ffe6472f9a52fbc1139f58b9db2881b25c
o_proj source drowzeys Keys commit cecebc0ff0bbce8eefaa19667290a79d0216cd0e
Edited BF16 model.language_model.layers.{15–43,45}.self_attn.o_proj.weight (30 tensors)
Left stock Layers 0–14 and layer 44 o_proj; all experts, vision, QKV, embeddings, head
Relative change min 0.0415, median 0.115, max 0.187 (layer 43)
Layout 83 EXL3 shards, same config.json as the parent
License MIT (Z.AI GLM-5.3-Flash)

drowzeys reports 32/32 Refusal32 bypass, 0 refuse, 0 garble on the NVFP4 parents those tensors were published for, with thinking off and no stock drafter. That score is theirs, measured on those NVFP4 trees. This file is the same 30 tensors copied into the Mia EXL3 tree. It does not add a new Refusal32 number.

ABLIT_META.json records each edited tensor, shard offset, and sha256. Layer 15 sha256 begins 1b46d7388815c0d7. Layer 45 sha256 begins 91f809721dd7f90a.

Use

Serve it with the same GLM-5.3-Flash EXL3 2× DGX Sparks TensorFold recipe as the parent. Point the recipe at this directory. Do not pass a different quantization format.

The stock chat_template.jinja is Z.ai's. It opens <think> even when thinking is turned off. drowzeys authored a closed-think overlay for that behavior: chat_template.thinking-off.jinja.

hf download Mia-AiLab/GLM-5.3-Flash-EXL3-4bpw-TensorFold-Ablit \
  --local-dir ~/models/GLM-5.3-Flash-EXL3-4bpw-TensorFold-Ablit

Download stays closed until you submit the gate form on the model page.

Downloads last month
2,695
Safetensors
Model size
88B params
Tensor type
BF16
·
F32
·
I32
·
F16
·
I16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Mia-AiLab/GLM-5.3-Flash-EXL3-4bpw-TensorFold-Ablit

Finetuned
(1)
this model

Spaces using Mia-AiLab/GLM-5.3-Flash-EXL3-4bpw-TensorFold-Ablit 2

Collection including Mia-AiLab/GLM-5.3-Flash-EXL3-4bpw-TensorFold-Ablit