Plumbline-0.6B

Loosely inspired by Jev. It decides whether a text breaks a rule written in plain English.

What ships is a 0.6B LoRA adapter over Qwen/Qwen3-Reranker-0.6B and the harness it needs: the prompt, the sign the readout is taken with, and the cutoffs. Those are not decoration. Feed the model a rule and a text without them and the polarity alone turns 0.901 AUC into 0.099. Copy the snippet below; it is the interface, not an illustration.

This release ships the cutoff. The one from last week returned a score and left it to you; three constants chosen by the rule's shape now decide, with nothing for you to label. They do not reach the ceiling: there is 0.063 F1 of headroom on document rules and 0.173 on social comments, measured against an oracle cutoff fitted on every label of the rule it scores. Your own labels do not reliably close it. Given twelve labelled items for a rule, fitting your own cutoff gains 0.021 on document rules and loses 0.015 on social comments.

It runs on a CPU. One call takes 0.93s on eight threads, about 64 items a minute. It loads in four seconds and needs no trust_remote_code.

The contribution

Training

  1. Every label is machine-verifiable, like RLVR without the RL, so it carries no annotator noise and no ceiling from a teacher model's beliefs.
  2. The training set is orthogonal in kind to everything evaluated here. Not one semantic rule in training: the adapter has never seen a rule about tone, disclosure, blame, evidence or sourcing.

Inference and harness

  1. A prompt and readout built for the task, and a threshold that needs no labels from you. Three constants chosen by the rule's shape, with a two-call path for conditional rules. Both ship in the snippet below and both work on any scorer, so they are separable from the weights.

How to read the numbers

F1 held out is what a user gets, from one constant fitted on other rules and never on the rule being scored, and F1 per rule fits a cutpoint on each rule's own items and is a ceiling rather than a result. Every model is measured at its own constant, intervals are bootstrapped over rules, and every figure averages within a rule and then over rules.

The conduct set, which is ours end to end

160 rules over 1,920 texts, rules and texts both written by Claude, built to isolate one failure. GOLDSET_RECIPE.md sets out how. A diagnostic, not evidence of transfer, and the only place a conditional rule shape is isolated. An earlier set of the same design and size exists too; the table below is this one alone, and the threshold figures elsewhere in this card pool both for 320 rules.

conduct rules, Claude-authored (160 rules, 1,920 items) AUC F1 held out F1 per rule
Plumbline-0.6B, two calls 0.968 [0.955, 0.979] 0.895 0.969
Plumbline-0.6B, one call 0.771 [0.712, 0.828] 0.859 0.908
Qwen3-Reranker-4B 0.732 [0.673, 0.786] 0.796 0.877
Qwen3-Reranker-0.6B (base) 0.718 [0.661, 0.775] 0.774 0.866

The whole difference between the two Plumbline rows is conditional rules: 0.139 in one call, 0.926 in two. The reason is vacuous compliance: a conditional is also satisfied when its condition never fires, and one call scores such a text as failing, because nothing in it meets the criterion. Those texts are compliant, so the ordering inverts. This set records per text whether the condition fired, and 82% of the conditional-compliant texts are vacuous, which is the share one call gets wrong.

LegalBench, with LegalBench's own rules

14 contract_nli tasks, 1,927 clauses, expert labels. The criteria are LegalBench's own hypotheses, fetched from the dataset rather than written by us, so this is the only table here scored on someone else's rules against someone else's labels, and the only one that can speak to whether any of the above generalises.

LegalBench contract clauses (14 rules, 1,927 items) AUC F1 held out F1 per rule
Qwen3-Reranker-4B 0.905 [0.869, 0.937] 0.847 0.894
Plumbline-0.6B 0.862 [0.799, 0.916] 0.817 0.863
Qwen3-Reranker-0.6B (base) 0.816 [0.765, 0.867] 0.764 0.828

A model 6.75x the size wins by 0.043. This adapter adds 0.046 to the model it was built from, on a CPU, from 10.1M trainable parameters in a 40MB file.

All 14 read out with complies, because contract_nli asks whether a clause entails its hypothesis. Three of the hypotheses contain "shall not", so readout_for would call them prohibitions and invert them; the snippet names the readout at the call site instead, and prints those three so the disagreement is visible rather than buried.

Thresholds

Three constants, chosen by the rule's shape, since no single cutoff works. They beat one global constant, 0.855 against 0.769 mean per-rule F1, with nothing for you to label. The conditional one encodes a vacuity rate, so it runs lenient where conditionals do fire.

plain        -2.19      shape = "conditional" if "only if" in rule else
prohibition  -1.25              "prohibition" if "must not" in rule else "plain"
conditional  -4.88      flag  = score >= CONSTANT[shape]

These are not probabilities. Calibration error is 0.207 on conduct against the base model's 0.414, lower and still high. Rank by the score, decide with the constant.

Use it

flag(rule, text) is the entry point, and score(readout, rule, text) is the one interface underneath it. Everything else here is built from those two, so there is no second code path to choose between. Below the functions the same script fetches LegalBench, its rules included, and recomputes this card's LegalBench row.

# `datasets` before `torch`: the other order segfaults inside pyarrow on Windows.
from datasets import load_dataset
import hashlib, json, re, urllib.request, torch
from sklearn.metrics import roc_auc_score
from transformers import AutoTokenizer, AutoModelForCausalLM
from peft import PeftModel

BASE, ADAPTER = "Qwen/Qwen3-Reranker-0.6B", "dsup2026/plumbline-0.6b"
DEVICE = "cpu"          # a GPU is not needed; "cuda" if you have one and are impatient
tok = AutoTokenizer.from_pretrained(BASE)
model = PeftModel.from_pretrained(
    AutoModelForCausalLM.from_pretrained(BASE), ADAPTER).to(DEVICE).eval()
YES, NO = tok.convert_tokens_to_ids("yes"), tok.convert_tokens_to_ids("no")
PRE = tok.encode('<|im_start|>system\nJudge whether the Document meets the requirements '
                 'based on the Query and the Instruct provided. Note that the answer can '
                 'only be "yes" or "no".<|im_end|>\n<|im_start|>user\n',
                 add_special_tokens=False)
SUF = tok.encode('<|im_end|>\n<|im_start|>assistant\n<think>\n\n</think>\n\n',
                 add_special_tokens=False)
MAX_LEN = 384
# One constant per rule KIND, fitted on rules of that kind and never on yours.
CONSTANT = {"plain": -2.19, "prohibition": -1.25, "conditional": -4.88}

# The two readouts, and the sign each is read with. "yes" means the document MATCHES the
# query, so for a prohibition it means the text FAILS. The readout is named at the call site
# so the wording and the sign stay together and cannot get out of step.
READOUT = {"complies": ("complies with the criterion in the query", -1.0),
           "violates": ("violates the restriction in the query", +1.0)}
PROHIBITION = re.compile(r"\b(must not|may not|shall not|must never|may never|must avoid|"
                         r"must refrain|must exclude|must omit|should not|cannot|do not|"
                         r"does not|never)\b", re.I)
CONDITIONAL = re.compile(r"^The (.+?) may (.+?) only if it (\w+) (.+)\.$")


@torch.no_grad()
def score(readout, rule, text):
    """Score for "this text breaks the rule". One interface, every rule, every corpus."""
    ask, sign = READOUT[readout]
    body = ("<Instruct>: Judge whether the document {}.\n<Query>: {}\n<Document>: {}"
            .format(ask, rule, text))
    ids = tok.encode(body, add_special_tokens=False, truncation=True,
                     max_length=MAX_LEN - len(PRE) - len(SUF))
    lg = model(input_ids=torch.tensor([PRE + ids + SUF], device=DEVICE)).logits[0, -1]
    return sign * (lg[YES] - lg[NO]).item()


def readout_for(rule):
    return "violates" if PROHIBITION.search(rule) else "complies"


def kind_of(rule):
    if CONDITIONAL.match(rule):
        return "conditional"
    return "prohibition" if PROHIBITION.search(rule) else "plain"


def one_call(rule, text):
    return score(readout_for(rule), rule, text)


def two_calls(rule, text):
    """A conditional becomes "must D" and "must state E", flagged only where the first holds
    and the second does not. Any other shape is one question already and falls through."""
    m = CONDITIONAL.match(rule)
    if not m:
        return one_call(rule, text)
    noun, action, verb, rest = m.groups()
    verb = verb[:-1] if verb.endswith("s") else verb
    does = one_call("The {} must {}.".format(noun, action), text)
    states = one_call("The {} must {} {}.".format(noun, verb, rest), text)
    return min(-does, states)


def flag(rule, text):
    """The shipped path: split if conditional, then the constant for the rule's kind."""
    return two_calls(rule, text) >= CONSTANT[kind_of(rule)]


# ---- this card's LegalBench row. NOTHING ABOUT THE BENCHMARK IS WRITTEN HERE: the clauses
# and labels come from the Hub, and so do the criteria, from the dataset's own task metadata.
# The hash is asserted so an upstream edit fails loudly rather than moving the numbers.
TASKS = ["confidentiality_of_agreement", "explicit_identification",
         "inclusion_of_verbally_conveyed_information", "limited_use", "no_licensing",
         "notice_on_compelled_disclosure",
         "permissible_acquirement_of_similar_information", "permissible_copy",
         "permissible_development_of_similar_information",
         "permissible_post-agreement_possession", "return_of_confidential_information",
         "sharing_with_employees", "sharing_with_third-parties", "survival_of_obligations"]
PREFIX = "Identify if the clause provides that "
META = "https://huggingface.co/datasets/nguha/legalbench/raw/main/task_metadata.json"
with urllib.request.urlopen(META, timeout=30) as r:
    meta = json.load(r)
RULES = {}
for task in TASKS:
    line = meta["contract_nli_" + task]["instruction"].splitlines()[0].strip()
    RULES[task] = line[len(PREFIX):] if line.startswith(PREFIX) else line
blob = "\n".join("{}\t{}".format(k, RULES[k]) for k in sorted(RULES))
assert hashlib.sha256(blob.encode()).hexdigest() == (
    "1e519d6747239d77dce23a2bed59604ae19620795565db0546188f86132a77c2"), (
    "LegalBench's wording changed upstream; the numbers in this card no longer apply")

# contract_nli asks whether a clause ENTAILS its hypothesis, so the readout is "complies" for
# all 14, including the three that contain "shall not". readout_for() would call those three
# prohibitions and invert them, which is what naming the readout at the call site makes visible.
ROW = "{:48s}{:>6}{:>8}"
print(ROW.format("contract_nli task", "n", "AUC"))
auc, total = [], 0
for task, rule in RULES.items():
    ds = load_dataset("nguha/legalbench", "contract_nli_" + task, split="test")
    y, z = [], []
    for r in ds:
        answer = str(r["answer"]).strip().lower()
        if answer not in ("yes", "no"):
            continue
        y.append(0 if answer == "yes" else 1)   # "yes" entails the hypothesis: it COMPLIES
        z.append(score("complies", rule, r["text"]))
    total += len(y)
    auc.append(roc_auc_score(y, z))
    print(ROW.format(task[:47], len(y), "{:.3f}".format(auc[-1])))
print(ROW.format("AUC within each rule, averaged over rules", total,
                 "{:.3f}".format(sum(auc) / len(auc))))

# What readout_for() would have chosen, and what flag() therefore returns on one clause:
one = load_dataset("nguha/legalbench", "contract_nli_limited_use", split="test")[0]["text"]
for task, rule in RULES.items():
    if readout_for(rule) != "complies":
        print("readout_for() says {:9s} flag()={}  {}".format(
            readout_for(rule), flag(rule, one), task))

Read this first: polarity

yes means the text meets the criterion. For a prohibition, meeting it means the text fails, so the sign flips with the phrasing, which is what sign does above. Getting it backwards turns 0.901 AUC into 0.099 on the same data, with no error message. The PROHIBITION pattern is only a heuristic: a rule phrased as a description of a violation will fool it.

What it cannot do

Conditional rules need two calls. In one call, rules shaped may X only if Y score AUC 0.139, against 0.986 on plain requirements and 0.976 on prohibitions: a text that never triggers the condition complies vacuously and one call reads that as a violation, inverting the ordering. Splitting the rule takes it to 0.926. Where most compliant texts comply vacuously, one call is unusable.

It judges what a text does, not whether a claim is true. It is not a fact checker, and it is not calibrated to your base rate.

Do not fit a threshold per rule. It loses to the constants above at every label budget, and no label-free way to fit one works here either.

Only the adapter is published

This repository holds the LoRA delta; the base stays Apache-2.0 from its own authors, and merge.py merges it. Every number here was measured on the adapter path.

Provenance

Machine-verifiable labels, on synthetic data generated by Qwen from an initial prompt/seed designed by Claude.

The conduct rules reported above were written by Claude, which is why they are presented as a diagnostic rather than as transfer: training and evaluation share a phrasing register there even though they share no rule. Wording is not a detail: it moves a single rule's AUC by up to 0.27. LegalBench's clauses, labels and rules are all its own.

Licence

PolyForm Noncommercial 1.0.0. Commercial use needs a separate licence.

The adapter is under PolyForm Noncommercial 1.0.0, reproduced in full in LICENSE. Any noncommercial purpose is permitted with no permission needed, which its own text spells out as personal study, private entertainment, hobby projects and amateur pursuits, and which covers charities, schools, public research bodies and government.

Commercial use needs a separate licence. Contact me via the details in LICENSE. Running the model over a group or a queue belonging to a business, a client or an employer is commercial use however informally it is run and whoever owns the laptop.

The base model remains Apache-2.0 and is not relicensed here; retain its LICENSE and NOTICE.

Citation

@misc{plumbline_0_6b_2026,
  title        = {Plumbline-0.6B: rule adherence from machine-verifiable supervision},
  author       = {Kav-Venaki, Eitam},
  year         = {2026},
  publisher    = {Hugging Face},
  howpublished = {\url{https://huggingface.co/dsup2026/plumbline-0.6b}},
  note         = {LoRA adapter over Qwen/Qwen3-Reranker-0.6B}
}

Results are reported on LegalBench, which should be cited if you quote them, and on the base model, which does most of the work:

@misc{guha2023legalbench,
  title        = {LegalBench: A Collaboratively Built Benchmark for Measuring Legal
                  Reasoning in Large Language Models},
  author       = {Guha, Neel and Nyarko, Julian and Ho, Daniel E. and R{\'e}, Christopher
                  and Chilton, Adam and Narayana, Aditya and others},
  year         = {2023},
  eprint       = {2308.11462},
  archivePrefix= {arXiv},
  primaryClass = {cs.CL},
  url          = {https://arxiv.org/abs/2308.11462},
  note         = {40 authors; all fourteen contract\_nli tasks are used here}
}

@article{qwen3embedding,
  title={Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models},
  author={Zhang, Yanzhao and Li, Mingxin and Long, Dingkun and Zhang, Xin and Lin, Huan and Yang, Baosong and Xie, Pengjun and Yang, An and Liu, Dayiheng and Lin, Junyang and Huang, Fei and Zhou, Jingren},
  journal={arXiv preprint arXiv:2506.05176},
  year={2025}
}
Downloads last month
167
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for dsup2026/plumbline-0.6b

Adapter
(10)
this model

Papers for dsup2026/plumbline-0.6b