good writer

#1
by Naphula - opened

Xenon has a unique way of writing with noticeably less slop than other Gemma finetunes, well done.

1 Epoch on highly distilled data. 1 epoch is the gold standard for high-signal distillation; it forces the model to internalize the routing and tool-formatting patterns without memorizing the exact traces.

This confirms what @Gryphe and others have reported

Turns out you really shouldn't use this technique with multiple epochs. I delved deep into the data and found that a second epoch does all sorts of nasty stuff to MoE models, so V2 is a single epoch of an otherwise unchanged technique.

Owner
โ€ข
edited Sep 2

Preciate it!

this is my first fine tune like ever

kinda sucks at coding (probably because of my old OPAL quant that relied on a static map)- but im glad the full precision weights are useful!

@el4 can you please provide a clean repository with just the merged weights in the root folder? This will allow quants to be created successfully.

@el4 since the lora is small i have uploaded these for you much quicker than a merge would take. See the link here https://huggingface.co/26B-Suite/Xenon-26B-A4B-safetensors

@redaihf
I can also upload merged safetensors for Calplus/GemmaWiki-Gemma-4-26b-a4b and nbeerbower/Gemma4-Gutenberg-26B-A4B if you would like so that @mradermacher and/or others can quantize them.

In fact the merge I started to upload OrionOmegaFiction would take too long at only 4mb/s upload so it might get deleted if necessary for space. HF uploader usually goes much faster for finetunes than merges, minutes instead of hours. It uploaded the entire 50GB safetensors for Xenon in under 30 seconds, while just one 5GB shard for OrionOmegaFiction took 20 minutes.

Owner

@el4 since the lora is small i have uploaded these for you much quicker than a merge would take. See the link here https://huggingface.co/26B-Suite/Xenon-26B-A4B-safetensors

@redaihf
I can also upload merged safetensors for Calplus/GemmaWiki-Gemma-4-26b-a4b and nbeerbower/Gemma4-Gutenberg-26B-A4B if you would like so that @mradermacher and/or others can quantize them.

In fact the merge I started to upload OrionOmegaFiction would take too long at only 4mb/s upload so it might get deleted if necessary for space. HF uploader usually goes much faster for finetunes than merges, minutes instead of hours. It uploaded the entire 50GB safetensors for Xenon in under 30 seconds, while just one 5GB shard for OrionOmegaFiction took 20 minutes.

Youre the goat ๐Ÿ™

Sign up or log in to comment