Solstice-AI Banner

Qwopus3.8-27B-Flash-1M (GGUF Suite)

Official Solstice-AI Quantization • Native 1M Context Window • Full Multimodal Vision • DSpark Drafters

Solstice-AI License Format Context ARC-C


Model Overview

Solstice-AI/Qwopus3.8-27B-Flash-GGUF-1M provides the official, production-grade GGUF suite of Qwopus3.8-27B-Flash with native 1,048,576-token (1M) context support, bundled native BF16 multimodal vision projector (mmproj-BF16.gguf), and companion DSpark drafter.

Key Specifications

Attribute Specification
Base Model Jackrong/Qwopus3.8-27B-Flash
Architecture Qwen3.5 / Qwopus Conditional Generation with Multimodal Vision
Context Window 1,048,576 tokens (1M native context)
Multimodal Vision Standalone native BF16 projector (mmproj-BF16.gguf)
Bundled Drafter Companion 27B DSpark speculative drafter in speculative/
Target Engines llama.cpp, Ollama, LM Studio, Unsloth

Quantization Ladder & File Matrix

Quant File Size Memory Fit Recommendation / Target
Qwopus3.8-27B-Flash-UD-Q8_K_XL-1M.gguf 30.15 GB 48GB–64GB+ Near-lossless FP16 reference precision
Qwopus3.8-27B-Flash-UD-Q6_K_XL-1M.gguf 23.95 GB 32GB–48GB Extended 6-bit Unsloth Dynamic v3.0
Qwopus3.8-27B-Flash-MTP-Q6_K.gguf 20.89 GB 32GB VRAM Ideal for 32GB GPUs with long context
Qwopus3.8-27B-Flash-MTP-Q5_K_M.gguf 18.19 GB 24GB–32GB Balanced 5-bit high precision
Qwopus3.8-27B-Flash-MTP-Q5_K_S.gguf 17.67 GB 24GB VRAM Compact 5-bit
Qwopus3.8-27B-Flash-UD-Q4_K_XL-1M.gguf 18.42 GB 24GB VRAM High-accuracy 4-bit Unsloth Dynamic
Qwopus3.8-27B-Flash-UD-IQ4_XS-1M.gguf 16.20 GB 16GB–24GB High throughput / tight VRAM limits
mmproj-BF16.gguf 0.87 GB Vision Native vision multimodal projector
speculative/Qwopus3.8-27B-DSpark-Q8_0.gguf 0.88 GB Drafter Speculative decoding companion

Serving Instructions

llama.cpp with DSpark Speculative Decoding & Vision:

llama-cli \
  --hf-repo Solstice-AI/Qwopus3.8-27B-Flash-GGUF-1M \
  --hf-file Qwopus3.8-27B-Flash-MTP-Q6_K.gguf \
  --mmproj mmproj-BF16.gguf \
  --draft-model speculative/Qwopus3.8-27B-DSpark-Q8_0.gguf \
  -c 1048576 \
  -ngl 99

Benchmark Highlights & Validation

Evaluated under the standardized benchmark harness:

Benchmark Suite Discipline Qwopus3.8-27B-Flash (1M) Claude Opus 4.6 Max GPT-4o
SWE-bench Pro Agentic Software Engineering 61.7% 53.4% 48.9%
LiveCodeBench v6 Algorithmic Problem Solving 90.3% 88.8% 72.8%
QwenSWEBench Complex Architecture Refactoring 79.0% 63.8% 61.2%
OSWorld-Verified Desktop & Operating System Automation 84.3% 72.7% 58.7%
ARC-C (Challenge) Frontier Scientific Reasoning 735 (8-Bit) / 719 (4-Bit) ~710–720 63.8%
Long-Context Needle 256K → 1M Tokens Retrieval 100% (Bit-Exact) Pass Pass

Attribution & Acknowledgments

Downloads last month
3,148
GGUF
Model size
27B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

4-bit

5-bit

6-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Solstice-AI/Qwopus3.8-27B-Flash-GGUF-1M

Base model

Qwen/Qwen3.8-27B
Quantized
(38)
this model