Instructions to use FluidInference/clef-vision-0.8b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use FluidInference/clef-vision-0.8b with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="FluidInference/clef-vision-0.8b") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("FluidInference/clef-vision-0.8b") model = AutoModelForMultimodalLM.from_pretrained("FluidInference/clef-vision-0.8b", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use FluidInference/clef-vision-0.8b with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "FluidInference/clef-vision-0.8b" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "FluidInference/clef-vision-0.8b", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/FluidInference/clef-vision-0.8b
- SGLang
How to use FluidInference/clef-vision-0.8b with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "FluidInference/clef-vision-0.8b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "FluidInference/clef-vision-0.8b", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "FluidInference/clef-vision-0.8b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "FluidInference/clef-vision-0.8b", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use FluidInference/clef-vision-0.8b with Docker Model Runner:
docker model run hf.co/FluidInference/clef-vision-0.8b
clef-vision-0.8b
A 0.92B-parameter image decision model, distilled from Cloudflare's
Clef-flash (Qwen3.5-9B) into a
Qwen3.5-0.8B backbone with Clef's joint schema head.
Same contract and API as Clef / Jev / SystemOne: a state (text or JSON), optional images, and typed questions
(noul true/false, choice, ordered score) go in; one probability for every allowed option of every
question comes out of a single forward pass. Nothing is generated.
This is the PyTorch release (any CUDA / CPU device, e.g. Jetson). The Apple-silicon build is FluidInference/clef-vision-0.8b-coreml.
Files
The layout matches Cloudflare's Clef releases, so Clef's own code loads it unchanged.
| File | Purpose |
|---|---|
model.safetensors, config.json, generation_config.json |
Qwen3.5-0.8B backbone with the distillation LoRA merged in, including the vision encoder |
joint_head.safetensors, joint_head_config.json |
Joint schema head (hidden 1024, width 1024, 2 routing + 2 layers, 16 heads; 69M params) |
joint_schema_model.py |
Cloudflare's Clef module, unchanged (Apache-2.0): record encoding, batching, load_release_model, systemone |
tokenizer.json, tokenizer_config.json, chat_template.jinja, processor_config.json |
Tokenizer and image processor |
LICENSE |
Apache-2.0 |
Usage
Needs torch, transformers ≥ 5.10 (Qwen3.5) and pillow. On CUDA, install causal-conv1d and
flash-linear-attention for the fast Gated DeltaNet kernels; without them transformers falls back to slower
reference code.
import sys
import torch
from huggingface_hub import snapshot_download
from PIL import Image
path = snapshot_download("FluidInference/clef-vision-0.8b")
sys.path.insert(0, path)
from joint_schema_model import collate_records, encode_record, load_release_model
model, processor = load_release_model(path, device="cuda")
record = {
"state": {"task": "Check the camera frame."},
"images": [Image.open("frame.jpg").convert("RGB")],
# The student was trained at this image size; use the same on inference.
"media_kwargs": {"size": {"shortest_edge": 16384, "longest_edge": 200704}},
"questions": {
"person": {"type": "noul", "instructions": "Is there a person in the image?"},
"scene": {
"type": "choice",
"instructions": "Where was this taken?",
"criteria": {"indoor": "Inside a building.", "outdoor": "Outside."},
},
},
}
encoded = encode_record(processor.tokenizer, record, processor=processor)
batch = collate_records([encoded], processor.tokenizer.pad_token_id, torch.device("cuda"))
with torch.inference_mode():
logits = model(batch)[0]
for question, question_logits in zip(encoded.questions, logits):
probabilities = question_logits.float().softmax(-1).tolist()
print(question.question_id, dict(zip(question.option_ids, probabilities)))
systemone(model, processor, request) from the same module takes a Jev/SystemOne POST /v1/systemone request
body and returns the response body; see the Clef-flash card for
the full input format.
Verified by loading this exact folder with load_release_model (CPU, fp32): token ids identical to the reference
encoder on 4 image records, 8/8 argmax agreement and max probability difference 4.3e-3 against the shipped
student. On Apple silicon, use the Core ML build; loading this checkpoint onto MPS through device_map crashed
in our tests.
Quality
Held-out test, 1,500 records / 4,094 questions over Oxford-IIIT Pets, Food-101 and COCO val images (classification, reference matching, tile grids):
| Teacher agreement | Gold accuracy | KL to teacher | |
|---|---|---|---|
| this model (0.92B) | 92.5% | 92.2% | 0.095 |
| teacher Clef-flash 9B (MLX 8-bit) | — | 94.1% | — |
Per task (gold): COCO object presence 98%, Food-101 dish 97%, pets breed 95%, reference-match grids 96–100%, tile grids 88–93%. The training images cover pets, food and COCO objects; expect lower accuracy on other domains (e.g. surveillance or industrial camera feeds) until you check it on your own data.
Training
14,246 records (train 5,648 / dev 2,167 / test 6,431, hash-split by image) labelled by Clef-flash running locally (MLX 8-bit backbone, exact port of the head). Student: LoRA r64 on the language model (merged into this checkpoint) + a fresh joint schema head; loss = KL(T=2) to the teacher + 0.3 × label-smoothed CE on dataset gold; 530 steps × 16 records on a 24 GB Apple M5 Pro (6.2 h). Vision tower frozen.
Licenses
Apache-2.0 (this model, Qwen3.5-0.8B, Clef and joint_schema_model.py © Cloudflare). Training images:
Oxford-IIIT Pet (CC BY-SA 4.0), Food-101 (research use), COCO 2017 val (CC BY 4.0).
- Downloads last month
- 31