Upload README.md with huggingface_hub
Browse files
README.md
ADDED
|
@@ -0,0 +1,258 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
language:
|
| 3 |
+
- en
|
| 4 |
+
- multilingual
|
| 5 |
+
license: apache-2.0
|
| 6 |
+
tags:
|
| 7 |
+
- rust
|
| 8 |
+
- cpu-inference
|
| 9 |
+
- quantized
|
| 10 |
+
- q4
|
| 11 |
+
- image-classification
|
| 12 |
+
- zero-shot-classification
|
| 13 |
+
- image-embedding
|
| 14 |
+
- siglip
|
| 15 |
+
- vision-transformer
|
| 16 |
+
- pure-rust
|
| 17 |
+
- no-python
|
| 18 |
+
- no-cuda
|
| 19 |
+
- contrastive-learning
|
| 20 |
+
base_model: google/siglip2-base-patch16-224
|
| 21 |
+
library_name: qora
|
| 22 |
+
pipeline_tag: zero-shot-image-classification
|
| 23 |
+
model-index:
|
| 24 |
+
- name: QORA-Vision-Image
|
| 25 |
+
results:
|
| 26 |
+
- task:
|
| 27 |
+
type: zero-shot-image-classification
|
| 28 |
+
dataset:
|
| 29 |
+
name: ImageNet-1K
|
| 30 |
+
type: imagenet-1k
|
| 31 |
+
metrics:
|
| 32 |
+
- name: Zero-shot Accuracy
|
| 33 |
+
type: accuracy
|
| 34 |
+
value: 69.8
|
| 35 |
+
---
|
| 36 |
+
|
| 37 |
+
# QORA-Vision (Image) - Native Rust Image Encoder
|
| 38 |
+
|
| 39 |
+
Pure Rust image understanding engine based on SigLIP 2. Zero-shot image classification, image embeddings, and image-text similarity. No Python runtime, no CUDA, no external dependencies.
|
| 40 |
+
|
| 41 |
+
## Overview
|
| 42 |
+
|
| 43 |
+
| Property | Value |
|
| 44 |
+
|----------|-------|
|
| 45 |
+
| **Engine** | QORA-Vision (Pure Rust) |
|
| 46 |
+
| **Base Model** | SigLIP 2 Base (google/siglip2-base-patch16-224) |
|
| 47 |
+
| **Vision Params** | ~93M |
|
| 48 |
+
| **Text Params** | ~283M (256K vocab) |
|
| 49 |
+
| **Quantization** | Q4 (4-bit symmetric, group_size=32) |
|
| 50 |
+
| **Vision Model Size** | 58 MB (Q4 binary) |
|
| 51 |
+
| **Executable** | 4.4 MB |
|
| 52 |
+
| **Input** | 224x224 RGB images (PNG/JPEG) |
|
| 53 |
+
| **Output** | 768-dim embeddings + zero-shot classification scores |
|
| 54 |
+
| **Platform** | Windows x86_64 (CPU-only) |
|
| 55 |
+
|
| 56 |
+
## Architecture
|
| 57 |
+
|
| 58 |
+
### Vision Encoder (12-layer ViT-Base)
|
| 59 |
+
|
| 60 |
+
| Component | Details |
|
| 61 |
+
|-----------|---------|
|
| 62 |
+
| **Layers** | 12 transformer layers |
|
| 63 |
+
| **Hidden Size** | 768 |
|
| 64 |
+
| **Attention Heads** | 12 (head_dim=64) |
|
| 65 |
+
| **MLP (Intermediate)** | 3,072 (GELU-Tanh activation) |
|
| 66 |
+
| **Patch Size** | 16x16 (non-overlapping) |
|
| 67 |
+
| **Sequence Length** | 196 patches (14x14 grid) |
|
| 68 |
+
| **Normalization** | LayerNorm with bias (eps=1e-6) |
|
| 69 |
+
| **Attention** | Bidirectional (no causal mask) |
|
| 70 |
+
| **Position Encoding** | Learned position embeddings |
|
| 71 |
+
| **Pooling** | MAP (Multi-head Attention Pooling) |
|
| 72 |
+
|
| 73 |
+
### Text Encoder (12-layer ViT-Base)
|
| 74 |
+
|
| 75 |
+
| Component | Details |
|
| 76 |
+
|-----------|---------|
|
| 77 |
+
| **Layers** | 12 transformer layers |
|
| 78 |
+
| **Hidden Size** | 768 |
|
| 79 |
+
| **Vocabulary** | 256,000 tokens |
|
| 80 |
+
| **Max Position** | 64 tokens |
|
| 81 |
+
| **Pooling** | Last token + linear head |
|
| 82 |
+
|
| 83 |
+
### Contrastive Scoring
|
| 84 |
+
|
| 85 |
+
```
|
| 86 |
+
score = sigmoid(cosine_sim(image_embed, text_embed) * exp(logit_scale) + logit_bias)
|
| 87 |
+
```
|
| 88 |
+
|
| 89 |
+
## Pipeline
|
| 90 |
+
|
| 91 |
+
```
|
| 92 |
+
Image (224x224) β Patch Embedding (196 patches)
|
| 93 |
+
β Add Position Embeddings
|
| 94 |
+
β 12x ViT Transformer Layers (bidirectional)
|
| 95 |
+
β Post-LayerNorm
|
| 96 |
+
β MAP Pooling (cross-attention with learned probe)
|
| 97 |
+
β L2 Normalize
|
| 98 |
+
β 768-dim Image Embedding
|
| 99 |
+
|
| 100 |
+
Text β Tokenize β Token + Position Embedding
|
| 101 |
+
β 12x ViT Transformer Layers
|
| 102 |
+
β Final LayerNorm (last token)
|
| 103 |
+
β Linear Head
|
| 104 |
+
β L2 Normalize
|
| 105 |
+
β 768-dim Text Embedding
|
| 106 |
+
|
| 107 |
+
Score = sigmoid(cosine_sim * exp(scale) + bias)
|
| 108 |
+
```
|
| 109 |
+
|
| 110 |
+
## Files
|
| 111 |
+
|
| 112 |
+
```
|
| 113 |
+
siglip-model/
|
| 114 |
+
qora-vision.exe - 4.4 MB Inference engine
|
| 115 |
+
model.qora-vision - 58 MB Vision encoder (Q4)
|
| 116 |
+
tokenizer.json - 33 MB Text tokenizer (256K vocab)
|
| 117 |
+
config.json - 611 B QORA-branded config
|
| 118 |
+
README.md - This file
|
| 119 |
+
```
|
| 120 |
+
|
| 121 |
+
## Usage
|
| 122 |
+
|
| 123 |
+
```bash
|
| 124 |
+
# Image embedding
|
| 125 |
+
qora-vision.exe siglip --image photo.jpg --model-path ./siglip-model/
|
| 126 |
+
|
| 127 |
+
# Zero-shot classification
|
| 128 |
+
qora-vision.exe siglip --image photo.jpg --labels "cat,dog,bird,car" --model-path ../SigLIP2/
|
| 129 |
+
|
| 130 |
+
# Image-text similarity
|
| 131 |
+
qora-vision.exe siglip --image photo.jpg --text "a photo of a sunset" --model-path ../SigLIP2/
|
| 132 |
+
|
| 133 |
+
# Load from binary (vision encoder only)
|
| 134 |
+
qora-vision.exe siglip --load model.qora-vision --image photo.jpg
|
| 135 |
+
```
|
| 136 |
+
|
| 137 |
+
### CLI Arguments
|
| 138 |
+
|
| 139 |
+
| Flag | Default | Description |
|
| 140 |
+
|------|---------|-------------|
|
| 141 |
+
| `--model-path <path>` | `.` | Path to model directory (safetensors) |
|
| 142 |
+
| `--image <path>` | - | Input image (PNG/JPEG) |
|
| 143 |
+
| `--labels <list>` | - | Comma-separated labels for zero-shot |
|
| 144 |
+
| `--text <string>` | - | Text for similarity scoring |
|
| 145 |
+
| `--load <path>` | - | Load vision binary (.qora-vision) |
|
| 146 |
+
| `--save <path>` | - | Save vision binary |
|
| 147 |
+
| `--f16` | off | Use F16 weights instead of Q4 |
|
| 148 |
+
|
| 149 |
+
## Published Benchmarks
|
| 150 |
+
|
| 151 |
+
### SigLIP 2 Base (224px) - Published Scores
|
| 152 |
+
|
| 153 |
+
| Benchmark | Score |
|
| 154 |
+
|-----------|-------|
|
| 155 |
+
| **ImageNet-1K Zero-shot** | ~69.8% |
|
| 156 |
+
| **Multilingual support** | Yes (trained on WebLI) |
|
| 157 |
+
|
| 158 |
+
SigLIP 2 improves over the original SigLIP with enhanced semantic understanding, localization, and dense features. The sigmoid loss enables better calibrated scores compared to CLIP's softmax-based approach.
|
| 159 |
+
|
| 160 |
+
### Model Comparison
|
| 161 |
+
|
| 162 |
+
| Model | Params | Image Size | Architecture | Zero-shot ImageNet |
|
| 163 |
+
|-------|--------|------------|-------------|-------------------|
|
| 164 |
+
| **QORA-Vision (SigLIP 2 Base)** | 93M | 224 | ViT-B/16 | ~69.8% |
|
| 165 |
+
| CLIP ViT-B/16 | 86M | 224 | ViT-B/16 | 68.3% |
|
| 166 |
+
| SigLIP Base (v1) | 86M | 224 | ViT-B/16 | 66.2% |
|
| 167 |
+
| OpenCLIP ViT-B/16 | 86M | 224 | ViT-B/16 | 67.0% |
|
| 168 |
+
|
| 169 |
+
## Test Results
|
| 170 |
+
|
| 171 |
+
All tests run with Q4 quantization on CPU.
|
| 172 |
+
|
| 173 |
+
### Test 1: Red Image Classification
|
| 174 |
+
|
| 175 |
+
**Input:** Solid red 224x224 image
|
| 176 |
+
**Labels:** red, blue, green, yellow
|
| 177 |
+
|
| 178 |
+
| Label | Score |
|
| 179 |
+
|-------|-------|
|
| 180 |
+
| **red** | **0.0022** |
|
| 181 |
+
| blue | 0.0000 |
|
| 182 |
+
| green | 0.0000 |
|
| 183 |
+
| yellow | 0.0000 |
|
| 184 |
+
|
| 185 |
+
| Metric | Value |
|
| 186 |
+
|--------|-------|
|
| 187 |
+
| Result | PASS (correctly identified "red") |
|
| 188 |
+
| Vision Forward | 42.0s |
|
| 189 |
+
| Embedding Dim | 768, L2 norm = 1.0000 |
|
| 190 |
+
|
| 191 |
+
### Test 2: Blue Image Classification
|
| 192 |
+
|
| 193 |
+
**Input:** Solid blue 224x224 image
|
| 194 |
+
**Labels:** red, blue, green, yellow
|
| 195 |
+
|
| 196 |
+
| Label | Score |
|
| 197 |
+
|-------|-------|
|
| 198 |
+
| red | 0.0000 |
|
| 199 |
+
| **blue** | **0.0014** |
|
| 200 |
+
| green | 0.0000 |
|
| 201 |
+
| yellow | 0.0000 |
|
| 202 |
+
|
| 203 |
+
| Metric | Value |
|
| 204 |
+
|--------|-------|
|
| 205 |
+
| Result | PASS (correctly identified "blue") |
|
| 206 |
+
| Vision Forward | 31.5s |
|
| 207 |
+
|
| 208 |
+
### Test 3: Green Image with Natural Language Labels
|
| 209 |
+
|
| 210 |
+
**Input:** Solid green 224x224 image
|
| 211 |
+
**Labels:** "a photo of a cat", "a photo of a dog", "a solid green image", "a landscape"
|
| 212 |
+
|
| 213 |
+
| Label | Score |
|
| 214 |
+
|-------|-------|
|
| 215 |
+
| a photo of a cat | 0.0000 |
|
| 216 |
+
| a photo of a dog | 0.0000 |
|
| 217 |
+
| **a solid green image** | **0.0176** |
|
| 218 |
+
| a landscape | 0.0000 |
|
| 219 |
+
|
| 220 |
+
| Metric | Value |
|
| 221 |
+
|--------|-------|
|
| 222 |
+
| Result | PASS (correctly identified natural language description) |
|
| 223 |
+
| Vision Forward | 39.2s |
|
| 224 |
+
| Note | Highest score by far, demonstrating text understanding |
|
| 225 |
+
|
| 226 |
+
### Test Summary
|
| 227 |
+
|
| 228 |
+
| Test | Input | Best Label | Correct? | Score |
|
| 229 |
+
|------|-------|------------|----------|-------|
|
| 230 |
+
| Color (red) | Solid red | "red" | PASS | 0.0022 |
|
| 231 |
+
| Color (blue) | Solid blue | "blue" | PASS | 0.0014 |
|
| 232 |
+
| NL Description | Solid green | "a solid green image" | PASS | 0.0176 |
|
| 233 |
+
| **Overall** | | | **3/3 (100%)** | |
|
| 234 |
+
|
| 235 |
+
## Performance
|
| 236 |
+
|
| 237 |
+
| Metric | Value |
|
| 238 |
+
|--------|-------|
|
| 239 |
+
| **Model Load** | ~25-30s (from safetensors) |
|
| 240 |
+
| **Vision Forward** | ~31-42s (196 tokens, 12 layers) |
|
| 241 |
+
| **Text Forward** | ~25s per label |
|
| 242 |
+
| **Total (4 labels)** | ~120-150s |
|
| 243 |
+
| **Memory (Vision Q4)** | 58 MB |
|
| 244 |
+
| **Memory (Text Q4)** | 151 MB |
|
| 245 |
+
| **Binary Save** | 41ms (58 MB) |
|
| 246 |
+
|
| 247 |
+
## QORA Model Family
|
| 248 |
+
|
| 249 |
+
| Engine | Model | Params | Size (Q4) | Purpose |
|
| 250 |
+
|--------|-------|--------|-----------|---------|
|
| 251 |
+
| **QORA** | SmolLM3-3B | 3.07B | 1.68 GB | Text generation, reasoning, chat |
|
| 252 |
+
| **QORA-TTS** | Qwen3-TTS | 1.84B | 1.5 GB | Text-to-speech synthesis |
|
| 253 |
+
| **QORA-Vision (Image)** | SigLIP 2 Base | 93M | 58 MB | Image embeddings, zero-shot classification |
|
| 254 |
+
| **QORA-Vision (Video)** | ViViT Base | 89M | 60 MB | Video action classification |
|
| 255 |
+
|
| 256 |
+
---
|
| 257 |
+
|
| 258 |
+
*Built with QORA - Pure Rust AI Inference*
|