Gemma-2-9B-it Natural Language Autoencoder — layer 20

Describes a single residual-stream activation of google/gemma-2-9b-it at layer 20 (d_model = 3584) in English, and reconstructs the vector from that description. Two adapters: a verbalizer (LoRA; the vector is injected at layer 1 and the model writes <explanation>...</explanation>) and a reconstructor (21-block backbone + Linear(d, d) head) that maps the text back to the vector.

Scored as FVE — fraction of variance explained, 1 − MSE / MSE(mean), on a document-disjoint held-out split.

Two RL stages

Held-out FVE across both RL stages

steps domain held-out FVE
rl_vllm/ — stage 1 0 → 400 FineFineWeb (web text) 45.2% → 61.6%
rl_wildchat/ — stage 2 400 → 800 WildChat (chat) 51.3% → 61.8%

Stage 2 continues from stage 1's step-400 adapter. The gap between the rows is the point: those same stage-1 weights score 61.6% on web text but 51.3% on chat (−10.3 pp). Stage 2 recovers all of it. Each curve is scored on its own domain's held-out split, so the two are separate measurements, not one line.

Training data

Activations harvested from gemma-2-9b-it at layer 20, 10 positions per document.

used drawn from
SFT (av_sft/, ar_sft/) 245,344 pairs (one epoch) FineFineWeb
Stage 1 RL 51,200 activations 449,828-row pool (~11%), 100k FFW documents
Stage 2 RL 51,200 activations 150,000-row pool (~34%), 20k WildChat conversations

Each RL stage: 400 GRPO steps × 128 prompts × 8 rollouts = 409,600 generations. Neither completed an epoch over its pool. WildChat rows are allenai/WildChat-1M train, English only, non-toxic, ≥400 chars after the chat template.

Both stages: on-policy GRPO, reward −MSE(reconstruction, v), KL β=0.01 (k3), lr 1e-4, critic lr 8e-5, max_new_tokens 256, temperature 1.0. Stage 2 adds --salvage-truncated (chat explanations often hit the token cap before closing their tag; salvage scores the partial text instead of discarding the rollout).

Use

av_sft/, ar_sft/   SFT adapters          merged/    SFT adapters merged into full models
rl_vllm/           stage 1, iter_000025 … iter_000400
rl_wildchat/       stage 2, iter_000425 … iter_000800
  critic_latest/   reconstructor + value_head.safetensors
  nla_meta.yaml    extraction contract (layer, marker token, prompt templates)

rl_wildchat/iter_000800 for chat activations, rl_vllm/iter_000400 for web text. You need this repo plus google/gemma-2-9b-it — the actor is built from the raw base with the adapter attached. Each RL checkpoint also carries a reference/ copy of the frozen SFT adapter used for the KL term.

Gotchas: the injection marker is ㈀ (token id 255423); corpus text is stored with BOS stripped so tokenizing adds exactly one; held-out splits are keyed by a hash of "{corpus path}:{split}:{row index}", so relocating a corpus changes the split.

Notes

Activation parquets are not published — the WildChat ones hold verbatim excerpts of real user–chatbot conversations. Weights only; regenerate activations from allenai/WildChat-1M if you need them.

Modified Gemma weights (LoRA adapters + a value head), governed by the Gemma Terms of Use §3.1, and additionally by the WildChat-1M terms. A research artefact: explanations are model-generated and unverified.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Yooniel/gemma2-9b-it-nla-L20

Adapter
(494)
this model

Dataset used to train Yooniel/gemma2-9b-it-nla-L20