securecoder-scripts / HANDOFF.md
Taimwe's picture
Handoff: fixed supervisor + full findings/state so work can resume
00d9092 verified
|
Raw History Blame Contribute Delete
5.12 kB

SecureCoder β€” handoff state (2026-10-04)

Everything needed to pick this up later. All code lives in this repo and in C:\Users\Taimwe\.cline\data\workspaces\chat\securecoder-pipeline.

The bug we found and fixed

v1 and v2 both emit literal \n (backslash-n) instead of real newlines, so their Python output does not parse. Measured with the same harness on the same prompts:

Model literal \n valid Python
base unsloth/Qwen3-Coder-30B-A3B-Instruct β€” 93.3%
securecoder-30b-pro (v1) 14/15 ~6.7%
securecoder-30b-pro-v2 15/15 0.0%

Root cause: parts of the training data mix store message text JSON-escaped. _unescape_if_needed() in train_securecoder.py now repairs them. Confirmed by --validate-only against 40 rows/source:

Source repaired/40
Trendyol/Cybersecurity-Instruction-Tuning 40 (100%)
AlicanKiraz0/Cybersecurity-Dataset-Fenrir-v2.1 40 (100%)
Humanlearning/CyberSecurity_OWASP-sft-dataset 40 (100%)
ukcli/Cybersecurity-Dataset-Heimdall-v1.1 40 (100%)
MrClipperz134/CTF-Instruct 23 (58%)
hermes-fc, xlam, Magicoder, threat-intel, ctf-solver, FineTome 0

Every escaped source is a cybersecurity dataset. That is why tool-calling still scored 100% (clean sources dominate it) while code scored 0% (escaped sources dominate it).

Files in this repo

File What changed
train_securecoder.py fixed β€” _unescape_if_needed(), repaired= counter, BOM stripped so uv PEP 723 deps install
eval_securecoder.py fixed β€” honours --max-new-tokens (was hardcoded 384), excludes truncation from the syntax score, falls back across dataset splits
pipeline.py supervisor: RESULT seeded up front (fixes the UnboundLocalError that killed every supervised run), retries, ADAPTER_REPO/GGUF_REPO env overrides
debug_v2_code.py dumps full replies β€” found the escaped newlines
bisect_adapter.py proved base+adapter (no merge) is also affected β‡’ the adapter, not the merge

Two bugs worth remembering

  1. UTF-8 BOM before # /// script β€” uv never saw the PEP 723 metadata block, so torch was never installed and every job died in 3–6s with No module named 'torch'. This is almost certainly what the old patch-deps jobs were fighting.
  2. eval_securecoder.py scored truncation as invalid syntax. With max_new_tokens=384 correct code was cut mid-statement and counted as a failure. Now truncated replies are tracked separately (truncated, n_truncated, truncated_rate).

Repo state

  • Deleted (intentionally, both shipped broken code): securecoder-30b-pro-merged, securecoder-30b-pro-GGUF, securecoder-30b-pro-v2-merged, securecoder-30b-pro-v2-GGUF.
  • Kept: both adapters (securecoder-30b-pro, securecoder-30b-pro-v2) β€” re-mergeable.
  • Eval reports: Taimwe/securecoder-eval-v2, Taimwe/securecoder-eval-base.

How to resume

# 1. retrain with the fixed data  (a100-large, ~2.4h, ~$6)
hf jobs run ghcr.io/astral-sh/uv:python3.12-bookworm `
  uv run --with torch --with transformers --with datasets --with huggingface_hub `
         --with bitsandbytes --with accelerate --with unsloth --with trl --with peft `
  https://huggingface.co/Taimwe/securecoder-scripts/resolve/main/train_securecoder.py `
  --base-model unsloth/Qwen3-Coder-30B-A3B-Instruct `
  --output-repo Taimwe/securecoder-30b-pro-v3 `
  --mix-scale 0.25 --max-steps 480 --max-seq-length 2048 --learning-rate 1e-4 `
  --save-steps 200 --eval-samples 100 --report-to none `
  --flavor a100-large --timeout 10800 --secrets HF_TOKEN --detach

# 2. then merge + quantise + eval

Sizing matters: the full mix is 83,781 rows; at seq 4096 that is 983 steps at 43.5 s/it β‰ˆ 11.8h / $29.50. With --mix-scale 0.25 --max-seq-length 2048 it is 480 steps at 17.7 s/it β‰ˆ 2.4h / $6. Do not omit those two flags.

Also set --timeout generously but below the credit budget you intend to spend.

Still to do (publishing, free)

  • Fill the _pending_ eval table on the model cards with the eval-14 numbers.
  • The corrected harness is ready: python publish_cards.py or three hf upload calls.

Money spent this session (~$13 of 20)

Item ~$
two evals (v2 + base, A/B) 2.30
adapter-vs-merge bisect 0.11
two --validate-only diagnostics 0.06
cancelled run A (bad scoping: full mix @4096) 7.45
cancelled run B (this one, stopped in error) 2.50

The $7.45 was avoidable β€” v2's original job command showed --mix-scale 0.25 and --max-seq-length 2048, and I should have read it before launching instead of guessing.

Unrelated scaffold

C:\Users\Taimwe\.cline\data\workspaces\chat\flux2-lora β€” LoRA fine-tune + FastAPI serving for FLUX.2-klein-4B (Apache-2.0). Not runnable here: HF large-file downloads transfer at 0 MB/min from this machine, and there is no GPU.