Indirakumar01 commited on
Commit
9175b66
Β·
verified Β·
1 Parent(s): 1f18e2a

Add model card

Browse files
Files changed (1) hide show
  1. README.md +188 -0
README.md ADDED
@@ -0,0 +1,188 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ base_model: Qwen/Qwen2.5-Coder-1.5B-Instruct
4
+ tags:
5
+ - text-to-sql
6
+ - nl2sql
7
+ - sql
8
+ - qlora
9
+ - gguf
10
+ - llama-cpp
11
+ - edge
12
+ - on-prem
13
+ language:
14
+ - en
15
+ library_name: gguf
16
+ pipeline_tag: text-generation
17
+ datasets:
18
+ - xlangai/spider
19
+ ---
20
+
21
+ # TinySQL-1.5B
22
+
23
+ **A private, on-prem Natural-Language-to-SQL model that runs on a laptop CPU.**
24
+
25
+ TinySQL is [Qwen2.5-Coder-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-Coder-1.5B-Instruct)
26
+ fine-tuned with QLoRA for text-to-SQL and quantized to 4-bit GGUF (Q4_K_M) so it
27
+ runs offline via `llama.cpp` β€” no GPU, no cloud, no per-query cost. It converts an
28
+ English question + a database schema into a **validated, read-only SQL SELECT**.
29
+
30
+ The point is **not** to beat frontier models on raw accuracy. It is to be the
31
+ *compliant, $0/query, offline* option for regulated data (finance, health, legal,
32
+ government) where the schema + data **cannot** leave the premises.
33
+
34
+ ---
35
+
36
+ ## Measured results
37
+
38
+ All numbers below are **measured**, not estimated. Evaluation is **execution
39
+ accuracy** on the full **Spider dev split (1,034 examples)**: run the predicted SQL
40
+ and the gold SQL against the real SQLite database and compare returned rows
41
+ (order-insensitive), over a read-only connection.
42
+
43
+ ### Fine-tuning lift (apples-to-apples)
44
+
45
+ Identical base model, identical Q4_K_M quantization, identical prompt template and
46
+ SELECT-only guardrail β€” **only the QLoRA adapter differs**.
47
+
48
+ | Model | Spider dev exec. acc | Valid (SELECT-only + executes) | Malformed outputs |
49
+ |-------|----------------------|--------------------------------|-------------------|
50
+ | Base Qwen2.5-Coder-1.5B | 50.87% | 76.9% | 54 |
51
+ | **TinySQL-1.5B** | **62.86%** | **87.99%** | **0** |
52
+
53
+ - **+11.99 points** execution accuracy from fine-tuning.
54
+ - **54 β†’ 0** malformed outputs: the base model emitted non-SQL chatter and
55
+ degenerate repetition loops; the fine-tune produces clean, parseable SELECTs.
56
+
57
+ ### Performance (laptop CPU, Intel Core Ultra 5 235U)
58
+
59
+ | Metric | Value |
60
+ |--------|-------|
61
+ | Mean latency | 0.57 s / query |
62
+ | p95 latency | 0.75 s |
63
+ | Peak RAM | ~1.73 GB |
64
+ | Cost | $0 / query (self-hosted) |
65
+
66
+ > Larger models and cloud APIs achieve higher accuracy, but require GPU/cloud and
67
+ > send your schema + data off-premises. TinySQL trades peak accuracy for privacy,
68
+ > $0 cost, and offline operation β€” the axes that matter for regulated data.
69
+
70
+ ---
71
+
72
+ ## Intended use
73
+
74
+ - Private/on-prem "text-to-SQL copilot" for non-technical users to query a database
75
+ in plain English.
76
+ - Embedded/offline analytics (edge devices, desktop apps) with no network.
77
+ - A cheap first-pass layer that handles routine queries locally, escalating only
78
+ hard ones to a larger model.
79
+
80
+ **Read-only by design.** Generated SQL is validated to be a single SELECT before it
81
+ is shown or executed. It cannot INSERT/UPDATE/DELETE/DROP.
82
+
83
+ ## Out of scope / limitations
84
+
85
+ - **Not for autonomous critical decisions** β€” ~63% accuracy means a human should
86
+ verify before acting on results.
87
+ - **No writes** β€” SELECT-only.
88
+ - **Schema size** β€” trained/served at 2048-token context; very large schemas are
89
+ pruned and accuracy drops.
90
+ - **SQLite dialect** β€” targets SQLite SQL.
91
+
92
+ ---
93
+
94
+ ## How to use
95
+
96
+ ### With `llama-cpp-python`
97
+
98
+ ```python
99
+ from llama_cpp import Llama
100
+
101
+ llm = Llama(model_path="tinysql-1.5b-q4_k_m.gguf", n_ctx=2048, verbose=False)
102
+
103
+ INSTRUCTION = ("You are a SQL expert. Given the database schema, write a single "
104
+ "SQLite SELECT query that answers the question. Return ONLY the SQL.")
105
+
106
+ schema = """CREATE TABLE orders (
107
+ id INTEGER PRIMARY KEY,
108
+ status TEXT,
109
+ customer_id INTEGER
110
+ );"""
111
+ question = "How many orders are completed?"
112
+
113
+ prompt = (f"### Instruction:\n{INSTRUCTION}\n\n"
114
+ f"### Schema:\n{schema}\n\n"
115
+ f"### Question:\n{question}\n\n"
116
+ f"### SQL:\n")
117
+
118
+ out = llm(prompt, max_tokens=256, temperature=0.0, stop=["###"])
119
+ print(out["choices"][0]["text"].strip())
120
+ # -> SELECT count(*) FROM orders WHERE status = 'completed';
121
+ ```
122
+
123
+ ### With `llama.cpp` CLI
124
+
125
+ ```bash
126
+ llama-cli -m tinysql-1.5b-q4_k_m.gguf -p "### Instruction:..." -n 256 --temp 0
127
+ ```
128
+
129
+ **Always enforce SELECT-only + run against a read-only DB connection** before
130
+ executing generated SQL. Do not run model output with write permissions.
131
+
132
+ ---
133
+
134
+ ## Prompt format
135
+
136
+ The model was trained with (and expects) this exact template:
137
+
138
+ ```
139
+ ### Instruction:
140
+ You are a SQL expert. Given the database schema, write a single SQLite
141
+ SELECT query that answers the question. Return ONLY the SQL.
142
+
143
+ ### Schema:
144
+ {CREATE TABLE statements}
145
+
146
+ ### Question:
147
+ {natural-language question}
148
+
149
+ ### Evidence: # optional external-knowledge hint (BIRD-style)
150
+ {hint}
151
+
152
+ ### SQL:
153
+ ```
154
+
155
+ ---
156
+
157
+ ## Training
158
+
159
+ - **Base:** Qwen/Qwen2.5-Coder-1.5B-Instruct
160
+ - **Method:** QLoRA (4-bit), LoRA r=16, Ξ±=16
161
+ - **Data:** Spider + BIRD, ~15.4k instruction examples, SELECT-only, FK-aware schema
162
+ pruning to fit 2048 tokens
163
+ - **Shipped checkpoint:** 500 steps (~0.26 epoch). A training-length ablation found
164
+ accuracy peaks early: 500 steps = 62.86%, 1000 = 61.22%, 1800 = 40.81% (overfit).
165
+ **Less was more.**
166
+ - **Export:** merged to 16-bit β†’ GGUF β†’ quantized Q4_K_M for CPU/edge.
167
+
168
+ ---
169
+
170
+ ## Datasets & attribution
171
+
172
+ - **Spider** (Yu et al., 2018) β€” CC BY-SA 4.0
173
+ - **BIRD** (Li et al., 2023) β€” CC BY-SA 4.0
174
+
175
+ Base model **Qwen2.5-Coder-1.5B-Instruct** is Apache-2.0. This fine-tune is released
176
+ under **Apache-2.0**; please also honor the CC BY-SA 4.0 attribution for Spider/BIRD.
177
+
178
+ ## Citation
179
+
180
+ ```bibtex
181
+ @misc{tinysql2026,
182
+ title = {TinySQL: Private On-Prem NL-to-SQL on a Laptop CPU},
183
+ author = {Indirakumar},
184
+ year = {2026},
185
+ howpublished = {\url{https://huggingface.co/Indirakumar01/tinysql-1.5b}},
186
+ note = {Fine-tuned Qwen2.5-Coder-1.5B, GGUF Q4_K_M}
187
+ }
188
+ ```