File size: 5,115 Bytes
00d9092
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
# SecureCoder β€” handoff state (2026-10-04)

Everything needed to pick this up later. All code lives in this repo and in
`C:\Users\Taimwe\.cline\data\workspaces\chat\securecoder-pipeline`.

## The bug we found and fixed

**v1 and v2 both emit literal `\n` (backslash-n) instead of real newlines**, so their
Python output does not parse. Measured with the same harness on the same prompts:

| Model | literal `\n` | valid Python |
|---|---:|---:|
| base `unsloth/Qwen3-Coder-30B-A3B-Instruct` | β€” | **93.3%** |
| `securecoder-30b-pro` (v1) | 14/15 | ~6.7% |
| `securecoder-30b-pro-v2` | 15/15 | **0.0%** |

**Root cause: parts of the training data mix store message text JSON-escaped.**
`_unescape_if_needed()` in `train_securecoder.py` now repairs them. Confirmed by
`--validate-only` against 40 rows/source:

| Source | repaired/40 |
|---|---:|
| `Trendyol/Cybersecurity-Instruction-Tuning` | 40 (100%) |
| `AlicanKiraz0/Cybersecurity-Dataset-Fenrir-v2.1` | 40 (100%) |
| `Humanlearning/CyberSecurity_OWASP-sft-dataset` | 40 (100%) |
| `ukcli/Cybersecurity-Dataset-Heimdall-v1.1` | 40 (100%) |
| `MrClipperz134/CTF-Instruct` | 23 (58%) |
| hermes-fc, xlam, Magicoder, threat-intel, ctf-solver, FineTome | 0 |

Every escaped source is a cybersecurity dataset. That is why tool-calling still scored
100% (clean sources dominate it) while code scored 0% (escaped sources dominate it).

## Files in this repo

| File | What changed |
|---|---|
| `train_securecoder.py` | **fixed** β€” `_unescape_if_needed()`, `repaired=` counter, BOM stripped so uv PEP 723 deps install |
| `eval_securecoder.py` | **fixed** β€” honours `--max-new-tokens` (was hardcoded 384), excludes truncation from the syntax score, falls back across dataset splits |
| `pipeline.py` | supervisor: `RESULT` seeded up front (fixes the `UnboundLocalError` that killed every supervised run), retries, `ADAPTER_REPO`/`GGUF_REPO` env overrides |
| `debug_v2_code.py` | dumps full replies β€” found the escaped newlines |
| `bisect_adapter.py` | proved base+adapter (no merge) is also affected β‡’ the adapter, not the merge |

## Two bugs worth remembering

1. **UTF-8 BOM before `# /// script`** β€” uv never saw the PEP 723 metadata block, so
   `torch` was never installed and every job died in 3–6s with `No module named 'torch'`.
   This is almost certainly what the old `patch-deps` jobs were fighting.
2. **`eval_securecoder.py` scored truncation as invalid syntax.** With `max_new_tokens=384`

   correct code was cut mid-statement and counted as a failure. Now truncated replies are

   tracked separately (`truncated`, `n_truncated`, `truncated_rate`).



## Repo state



- **Deleted** (intentionally, both shipped broken code): `securecoder-30b-pro-merged`,

  `securecoder-30b-pro-GGUF`, `securecoder-30b-pro-v2-merged`, `securecoder-30b-pro-v2-GGUF`.

- **Kept**: both adapters (`securecoder-30b-pro`, `securecoder-30b-pro-v2`) β€” re-mergeable.

- **Eval reports**: `Taimwe/securecoder-eval-v2`, `Taimwe/securecoder-eval-base`.



## How to resume



```powershell

# 1. retrain with the fixed data  (a100-large, ~2.4h, ~$6)

hf jobs run ghcr.io/astral-sh/uv:python3.12-bookworm `

  uv run --with torch --with transformers --with datasets --with huggingface_hub `

         --with bitsandbytes --with accelerate --with unsloth --with trl --with peft `

  https://huggingface.co/Taimwe/securecoder-scripts/resolve/main/train_securecoder.py `

  --base-model unsloth/Qwen3-Coder-30B-A3B-Instruct `

  --output-repo Taimwe/securecoder-30b-pro-v3 `

  --mix-scale 0.25 --max-steps 480 --max-seq-length 2048 --learning-rate 1e-4 `

  --save-steps 200 --eval-samples 100 --report-to none `

  --flavor a100-large --timeout 10800 --secrets HF_TOKEN --detach



# 2. then merge + quantise + eval

```



**Sizing matters:** the full mix is 83,781 rows; at seq 4096 that is 983 steps at
43.5 s/it β‰ˆ **11.8h / $29.50**. With `--mix-scale 0.25 --max-seq-length 2048` it is
480 steps at 17.7 s/it β‰ˆ **2.4h / $6**. Do not omit those two flags.

Also set `--timeout` generously but below the credit budget you intend to spend.

## Still to do (publishing, free)

- Fill the `_pending_` eval table on the model cards with the eval-14 numbers.
- The corrected harness is ready: `python publish_cards.py` or three `hf upload` calls.

## Money spent this session (~$13 of 20)

| Item | ~$ |
|---|---:|
| two evals (v2 + base, A/B) | 2.30 |
| adapter-vs-merge bisect | 0.11 |
| two `--validate-only` diagnostics | 0.06 |
| cancelled run A (bad scoping: full mix @4096) | 7.45 |
| cancelled run B (this one, stopped in error) | 2.50 |

The `$7.45` was avoidable β€” v2's original job command showed `--mix-scale 0.25` and
`--max-seq-length 2048`, and I should have read it before launching instead of guessing.

## Unrelated scaffold

`C:\Users\Taimwe\.cline\data\workspaces\chat\flux2-lora` β€” LoRA fine-tune + FastAPI
serving for `FLUX.2-klein-4B` (Apache-2.0). Not runnable here: HF large-file downloads
transfer at **0 MB/min** from this machine, and there is no GPU.