Instructions to use el4/Xenon-26B-A4B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use el4/Xenon-26B-A4B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="el4/Xenon-26B-A4B") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("el4/Xenon-26B-A4B") model = AutoModelForMultimodalLM.from_pretrained("el4/Xenon-26B-A4B", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use el4/Xenon-26B-A4B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "el4/Xenon-26B-A4B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "el4/Xenon-26B-A4B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/el4/Xenon-26B-A4B
- SGLang
How to use el4/Xenon-26B-A4B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "el4/Xenon-26B-A4B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "el4/Xenon-26B-A4B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "el4/Xenon-26B-A4B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "el4/Xenon-26B-A4B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use el4/Xenon-26B-A4B with Docker Model Runner:
docker model run hf.co/el4/Xenon-26B-A4B
good writer
Xenon has a unique way of writing with noticeably less slop than other Gemma finetunes, well done.
1 Epoch on highly distilled data. 1 epoch is the gold standard for high-signal distillation; it forces the model to internalize the routing and tool-formatting patterns without memorizing the exact traces.
This confirms what @Gryphe and others have reported
Turns out you really shouldn't use this technique with multiple epochs. I delved deep into the data and found that a second epoch does all sorts of nasty stuff to MoE models, so V2 is a single epoch of an otherwise unchanged technique.
Preciate it!
this is my first fine tune like ever
kinda sucks at coding (probably because of my old OPAL quant that relied on a static map)- but im glad the full precision weights are useful!
@el4 can you please provide a clean repository with just the merged weights in the root folder? This will allow quants to be created successfully.
@el4 since the lora is small i have uploaded these for you much quicker than a merge would take. See the link here https://huggingface.co/26B-Suite/Xenon-26B-A4B-safetensors
@redaihf
I can also upload merged safetensors for Calplus/GemmaWiki-Gemma-4-26b-a4b and nbeerbower/Gemma4-Gutenberg-26B-A4B if you would like so that @mradermacher and/or others can quantize them.
In fact the merge I started to upload OrionOmegaFiction would take too long at only 4mb/s upload so it might get deleted if necessary for space. HF uploader usually goes much faster for finetunes than merges, minutes instead of hours. It uploaded the entire 50GB safetensors for Xenon in under 30 seconds, while just one 5GB shard for OrionOmegaFiction took 20 minutes.
@el4 since the lora is small i have uploaded these for you much quicker than a merge would take. See the link here https://huggingface.co/26B-Suite/Xenon-26B-A4B-safetensors
@redaihf
I can also upload merged safetensors forCalplus/GemmaWiki-Gemma-4-26b-a4bandnbeerbower/Gemma4-Gutenberg-26B-A4Bif you would like so that @mradermacher and/or others can quantize them.In fact the merge I started to upload
OrionOmegaFictionwould take too long at only 4mb/s upload so it might get deleted if necessary for space. HF uploader usually goes much faster for finetunes than merges, minutes instead of hours. It uploaded the entire 50GB safetensors for Xenon in under 30 seconds, while just one 5GB shard for OrionOmegaFiction took 20 minutes.
Youre the goat ๐