decider-4b v2 — Q4_K_M GGUF (revision-pinned)

decider-4b at Hub tag v2 — exactly the revision that holds JevBench v1.4.2 rank 1 (Score 64.1 · Intel 49 · Calib 75 · Speed 93 · Cost 61), self-quantized with full provenance. Community GGUFs rarely pin which revision they quantized — this one does.

Which revision, and why it matters

The upstream repo root is v2.1; v2 lives under the v2 tag. Per the upstream card, v2 wins on hard decisions and is better calibrated there (0.676 vs 0.649), while v2.1 trades some of that for sampled play. For logit-read typed decisions (SemIf-style: state + lettered options → probabilities in one forward pass, no text generation) you want v2.

Source Mapika/decider-4b @ tag v2 (bf16, 8.4 GB)
Quant Q4_K_M via llama.cpp convert_hf_to_gguf.py (9575389) + llama-quantize
Size / SHA256 2708804544 bytes / f7e2e510ef51d212ea9b7fb8bf27906b5f516d7939ca847428fb91f6a8acfa79
Base Qwen3.5-4B-Base (4.2B params, Apache-2.0)
Quantized by mindchain (JEV stack), Kaggle CPU, pipeline in provenance.json

Usage — System-One decisions from logits (no generation)

llama-server -m decider-4b.v2-Q4_K_M.gguf -ngl 99 -c 2048 --port 8080

Send the SemIf minimal prompt (state + criteria + lettered options) and read the option-letter probabilities straight from the logits — one forward pass, no decoding loop. Apply the per-answer-type temperature from the upstream card for calibrated confidences. Works with any llama.cpp runtime (server, Termux/Android, ChatterUI).

Honest benchmark note (protocol separation)

JevBench v1.4.2 ranks decider-4b v2 #1 (64.1). On the independent JEV-stack gold set, a decider-4b Q4 measured 0.5975 accuracy vs 0.7635 for a fine-tuned encoder-decider (laya) — different protocols crown different winners. Measure on YOUR domain before shipping.

  • Benchmark: JevBench v1.4.2
  • JEV decision-model stack: calibration-first System-One models, gold-set gated deployments
Downloads last month
599
GGUF
Model size
4B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for mindchain/decider-4b-v2-GGUF

Quantized
(6)
this model