clef-vision-0.8b

A 0.92B-parameter image decision model, distilled from Cloudflare's Clef-flash (Qwen3.5-9B) into a Qwen3.5-0.8B backbone with Clef's joint schema head. Same contract and API as Clef / Jev / SystemOne: a state (text or JSON), optional images, and typed questions (noul true/false, choice, ordered score) go in; one probability for every allowed option of every question comes out of a single forward pass. Nothing is generated.

This is the PyTorch release (any CUDA / CPU device, e.g. Jetson). The Apple-silicon build is FluidInference/clef-vision-0.8b-coreml.

Files

The layout matches Cloudflare's Clef releases, so Clef's own code loads it unchanged.

File Purpose
model.safetensors, config.json, generation_config.json Qwen3.5-0.8B backbone with the distillation LoRA merged in, including the vision encoder
joint_head.safetensors, joint_head_config.json Joint schema head (hidden 1024, width 1024, 2 routing + 2 layers, 16 heads; 69M params)
joint_schema_model.py Cloudflare's Clef module, unchanged (Apache-2.0): record encoding, batching, load_release_model, systemone
tokenizer.json, tokenizer_config.json, chat_template.jinja, processor_config.json Tokenizer and image processor
LICENSE Apache-2.0

Usage

Needs torch, transformers ≥ 5.10 (Qwen3.5) and pillow. On CUDA, install causal-conv1d and flash-linear-attention for the fast Gated DeltaNet kernels; without them transformers falls back to slower reference code.

import sys

import torch
from huggingface_hub import snapshot_download
from PIL import Image

path = snapshot_download("FluidInference/clef-vision-0.8b")
sys.path.insert(0, path)
from joint_schema_model import collate_records, encode_record, load_release_model

model, processor = load_release_model(path, device="cuda")

record = {
    "state": {"task": "Check the camera frame."},
    "images": [Image.open("frame.jpg").convert("RGB")],
    # The student was trained at this image size; use the same on inference.
    "media_kwargs": {"size": {"shortest_edge": 16384, "longest_edge": 200704}},
    "questions": {
        "person": {"type": "noul", "instructions": "Is there a person in the image?"},
        "scene": {
            "type": "choice",
            "instructions": "Where was this taken?",
            "criteria": {"indoor": "Inside a building.", "outdoor": "Outside."},
        },
    },
}

encoded = encode_record(processor.tokenizer, record, processor=processor)
batch = collate_records([encoded], processor.tokenizer.pad_token_id, torch.device("cuda"))
with torch.inference_mode():
    logits = model(batch)[0]

for question, question_logits in zip(encoded.questions, logits):
    probabilities = question_logits.float().softmax(-1).tolist()
    print(question.question_id, dict(zip(question.option_ids, probabilities)))

systemone(model, processor, request) from the same module takes a Jev/SystemOne POST /v1/systemone request body and returns the response body; see the Clef-flash card for the full input format.

Verified by loading this exact folder with load_release_model (CPU, fp32): token ids identical to the reference encoder on 4 image records, 8/8 argmax agreement and max probability difference 4.3e-3 against the shipped student. On Apple silicon, use the Core ML build; loading this checkpoint onto MPS through device_map crashed in our tests.

Quality

Held-out test, 1,500 records / 4,094 questions over Oxford-IIIT Pets, Food-101 and COCO val images (classification, reference matching, tile grids):

Teacher agreement Gold accuracy KL to teacher
this model (0.92B) 92.5% 92.2% 0.095
teacher Clef-flash 9B (MLX 8-bit) — 94.1% —

Per task (gold): COCO object presence 98%, Food-101 dish 97%, pets breed 95%, reference-match grids 96–100%, tile grids 88–93%. The training images cover pets, food and COCO objects; expect lower accuracy on other domains (e.g. surveillance or industrial camera feeds) until you check it on your own data.

Training

14,246 records (train 5,648 / dev 2,167 / test 6,431, hash-split by image) labelled by Clef-flash running locally (MLX 8-bit backbone, exact port of the head). Student: LoRA r64 on the language model (merged into this checkpoint) + a fresh joint schema head; loss = KL(T=2) to the teacher + 0.3 × label-smoothed CE on dataset gold; 530 steps × 16 records on a 24 GB Apple M5 Pro (6.2 h). Vision tower frozen.

Licenses

Apache-2.0 (this model, Qwen3.5-0.8B, Clef and joint_schema_model.py © Cloudflare). Training images: Oxford-IIIT Pet (CC BY-SA 4.0), Food-101 (research use), COCO 2017 val (CC BY 4.0).

Downloads last month
31
Safetensors
Model size
0.9B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for FluidInference/clef-vision-0.8b

Finetuned
(493)
this model