drdraq commited on
Commit
ebf9f7c
Β·
verified Β·
1 Parent(s): 1c7b424

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +258 -0
README.md ADDED
@@ -0,0 +1,258 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ language:
3
+ - en
4
+ - multilingual
5
+ license: apache-2.0
6
+ tags:
7
+ - rust
8
+ - cpu-inference
9
+ - quantized
10
+ - q4
11
+ - image-classification
12
+ - zero-shot-classification
13
+ - image-embedding
14
+ - siglip
15
+ - vision-transformer
16
+ - pure-rust
17
+ - no-python
18
+ - no-cuda
19
+ - contrastive-learning
20
+ base_model: google/siglip2-base-patch16-224
21
+ library_name: qora
22
+ pipeline_tag: zero-shot-image-classification
23
+ model-index:
24
+ - name: QORA-Vision-Image
25
+ results:
26
+ - task:
27
+ type: zero-shot-image-classification
28
+ dataset:
29
+ name: ImageNet-1K
30
+ type: imagenet-1k
31
+ metrics:
32
+ - name: Zero-shot Accuracy
33
+ type: accuracy
34
+ value: 69.8
35
+ ---
36
+
37
+ # QORA-Vision (Image) - Native Rust Image Encoder
38
+
39
+ Pure Rust image understanding engine based on SigLIP 2. Zero-shot image classification, image embeddings, and image-text similarity. No Python runtime, no CUDA, no external dependencies.
40
+
41
+ ## Overview
42
+
43
+ | Property | Value |
44
+ |----------|-------|
45
+ | **Engine** | QORA-Vision (Pure Rust) |
46
+ | **Base Model** | SigLIP 2 Base (google/siglip2-base-patch16-224) |
47
+ | **Vision Params** | ~93M |
48
+ | **Text Params** | ~283M (256K vocab) |
49
+ | **Quantization** | Q4 (4-bit symmetric, group_size=32) |
50
+ | **Vision Model Size** | 58 MB (Q4 binary) |
51
+ | **Executable** | 4.4 MB |
52
+ | **Input** | 224x224 RGB images (PNG/JPEG) |
53
+ | **Output** | 768-dim embeddings + zero-shot classification scores |
54
+ | **Platform** | Windows x86_64 (CPU-only) |
55
+
56
+ ## Architecture
57
+
58
+ ### Vision Encoder (12-layer ViT-Base)
59
+
60
+ | Component | Details |
61
+ |-----------|---------|
62
+ | **Layers** | 12 transformer layers |
63
+ | **Hidden Size** | 768 |
64
+ | **Attention Heads** | 12 (head_dim=64) |
65
+ | **MLP (Intermediate)** | 3,072 (GELU-Tanh activation) |
66
+ | **Patch Size** | 16x16 (non-overlapping) |
67
+ | **Sequence Length** | 196 patches (14x14 grid) |
68
+ | **Normalization** | LayerNorm with bias (eps=1e-6) |
69
+ | **Attention** | Bidirectional (no causal mask) |
70
+ | **Position Encoding** | Learned position embeddings |
71
+ | **Pooling** | MAP (Multi-head Attention Pooling) |
72
+
73
+ ### Text Encoder (12-layer ViT-Base)
74
+
75
+ | Component | Details |
76
+ |-----------|---------|
77
+ | **Layers** | 12 transformer layers |
78
+ | **Hidden Size** | 768 |
79
+ | **Vocabulary** | 256,000 tokens |
80
+ | **Max Position** | 64 tokens |
81
+ | **Pooling** | Last token + linear head |
82
+
83
+ ### Contrastive Scoring
84
+
85
+ ```
86
+ score = sigmoid(cosine_sim(image_embed, text_embed) * exp(logit_scale) + logit_bias)
87
+ ```
88
+
89
+ ## Pipeline
90
+
91
+ ```
92
+ Image (224x224) β†’ Patch Embedding (196 patches)
93
+ β†’ Add Position Embeddings
94
+ β†’ 12x ViT Transformer Layers (bidirectional)
95
+ β†’ Post-LayerNorm
96
+ β†’ MAP Pooling (cross-attention with learned probe)
97
+ β†’ L2 Normalize
98
+ β†’ 768-dim Image Embedding
99
+
100
+ Text β†’ Tokenize β†’ Token + Position Embedding
101
+ β†’ 12x ViT Transformer Layers
102
+ β†’ Final LayerNorm (last token)
103
+ β†’ Linear Head
104
+ β†’ L2 Normalize
105
+ β†’ 768-dim Text Embedding
106
+
107
+ Score = sigmoid(cosine_sim * exp(scale) + bias)
108
+ ```
109
+
110
+ ## Files
111
+
112
+ ```
113
+ siglip-model/
114
+ qora-vision.exe - 4.4 MB Inference engine
115
+ model.qora-vision - 58 MB Vision encoder (Q4)
116
+ tokenizer.json - 33 MB Text tokenizer (256K vocab)
117
+ config.json - 611 B QORA-branded config
118
+ README.md - This file
119
+ ```
120
+
121
+ ## Usage
122
+
123
+ ```bash
124
+ # Image embedding
125
+ qora-vision.exe siglip --image photo.jpg --model-path ./siglip-model/
126
+
127
+ # Zero-shot classification
128
+ qora-vision.exe siglip --image photo.jpg --labels "cat,dog,bird,car" --model-path ../SigLIP2/
129
+
130
+ # Image-text similarity
131
+ qora-vision.exe siglip --image photo.jpg --text "a photo of a sunset" --model-path ../SigLIP2/
132
+
133
+ # Load from binary (vision encoder only)
134
+ qora-vision.exe siglip --load model.qora-vision --image photo.jpg
135
+ ```
136
+
137
+ ### CLI Arguments
138
+
139
+ | Flag | Default | Description |
140
+ |------|---------|-------------|
141
+ | `--model-path <path>` | `.` | Path to model directory (safetensors) |
142
+ | `--image <path>` | - | Input image (PNG/JPEG) |
143
+ | `--labels <list>` | - | Comma-separated labels for zero-shot |
144
+ | `--text <string>` | - | Text for similarity scoring |
145
+ | `--load <path>` | - | Load vision binary (.qora-vision) |
146
+ | `--save <path>` | - | Save vision binary |
147
+ | `--f16` | off | Use F16 weights instead of Q4 |
148
+
149
+ ## Published Benchmarks
150
+
151
+ ### SigLIP 2 Base (224px) - Published Scores
152
+
153
+ | Benchmark | Score |
154
+ |-----------|-------|
155
+ | **ImageNet-1K Zero-shot** | ~69.8% |
156
+ | **Multilingual support** | Yes (trained on WebLI) |
157
+
158
+ SigLIP 2 improves over the original SigLIP with enhanced semantic understanding, localization, and dense features. The sigmoid loss enables better calibrated scores compared to CLIP's softmax-based approach.
159
+
160
+ ### Model Comparison
161
+
162
+ | Model | Params | Image Size | Architecture | Zero-shot ImageNet |
163
+ |-------|--------|------------|-------------|-------------------|
164
+ | **QORA-Vision (SigLIP 2 Base)** | 93M | 224 | ViT-B/16 | ~69.8% |
165
+ | CLIP ViT-B/16 | 86M | 224 | ViT-B/16 | 68.3% |
166
+ | SigLIP Base (v1) | 86M | 224 | ViT-B/16 | 66.2% |
167
+ | OpenCLIP ViT-B/16 | 86M | 224 | ViT-B/16 | 67.0% |
168
+
169
+ ## Test Results
170
+
171
+ All tests run with Q4 quantization on CPU.
172
+
173
+ ### Test 1: Red Image Classification
174
+
175
+ **Input:** Solid red 224x224 image
176
+ **Labels:** red, blue, green, yellow
177
+
178
+ | Label | Score |
179
+ |-------|-------|
180
+ | **red** | **0.0022** |
181
+ | blue | 0.0000 |
182
+ | green | 0.0000 |
183
+ | yellow | 0.0000 |
184
+
185
+ | Metric | Value |
186
+ |--------|-------|
187
+ | Result | PASS (correctly identified "red") |
188
+ | Vision Forward | 42.0s |
189
+ | Embedding Dim | 768, L2 norm = 1.0000 |
190
+
191
+ ### Test 2: Blue Image Classification
192
+
193
+ **Input:** Solid blue 224x224 image
194
+ **Labels:** red, blue, green, yellow
195
+
196
+ | Label | Score |
197
+ |-------|-------|
198
+ | red | 0.0000 |
199
+ | **blue** | **0.0014** |
200
+ | green | 0.0000 |
201
+ | yellow | 0.0000 |
202
+
203
+ | Metric | Value |
204
+ |--------|-------|
205
+ | Result | PASS (correctly identified "blue") |
206
+ | Vision Forward | 31.5s |
207
+
208
+ ### Test 3: Green Image with Natural Language Labels
209
+
210
+ **Input:** Solid green 224x224 image
211
+ **Labels:** "a photo of a cat", "a photo of a dog", "a solid green image", "a landscape"
212
+
213
+ | Label | Score |
214
+ |-------|-------|
215
+ | a photo of a cat | 0.0000 |
216
+ | a photo of a dog | 0.0000 |
217
+ | **a solid green image** | **0.0176** |
218
+ | a landscape | 0.0000 |
219
+
220
+ | Metric | Value |
221
+ |--------|-------|
222
+ | Result | PASS (correctly identified natural language description) |
223
+ | Vision Forward | 39.2s |
224
+ | Note | Highest score by far, demonstrating text understanding |
225
+
226
+ ### Test Summary
227
+
228
+ | Test | Input | Best Label | Correct? | Score |
229
+ |------|-------|------------|----------|-------|
230
+ | Color (red) | Solid red | "red" | PASS | 0.0022 |
231
+ | Color (blue) | Solid blue | "blue" | PASS | 0.0014 |
232
+ | NL Description | Solid green | "a solid green image" | PASS | 0.0176 |
233
+ | **Overall** | | | **3/3 (100%)** | |
234
+
235
+ ## Performance
236
+
237
+ | Metric | Value |
238
+ |--------|-------|
239
+ | **Model Load** | ~25-30s (from safetensors) |
240
+ | **Vision Forward** | ~31-42s (196 tokens, 12 layers) |
241
+ | **Text Forward** | ~25s per label |
242
+ | **Total (4 labels)** | ~120-150s |
243
+ | **Memory (Vision Q4)** | 58 MB |
244
+ | **Memory (Text Q4)** | 151 MB |
245
+ | **Binary Save** | 41ms (58 MB) |
246
+
247
+ ## QORA Model Family
248
+
249
+ | Engine | Model | Params | Size (Q4) | Purpose |
250
+ |--------|-------|--------|-----------|---------|
251
+ | **QORA** | SmolLM3-3B | 3.07B | 1.68 GB | Text generation, reasoning, chat |
252
+ | **QORA-TTS** | Qwen3-TTS | 1.84B | 1.5 GB | Text-to-speech synthesis |
253
+ | **QORA-Vision (Image)** | SigLIP 2 Base | 93M | 58 MB | Image embeddings, zero-shot classification |
254
+ | **QORA-Vision (Video)** | ViViT Base | 89M | 60 MB | Video action classification |
255
+
256
+ ---
257
+
258
+ *Built with QORA - Pure Rust AI Inference*