A curated collection of production-grade LLM and SLM architectures optimized for high-throughput runtime engines and quantization pipelines.
-
meta-llama/Llama-3.1-8B-Instruct
Text Generation • 8B • Updated • 6.02M • • 8.33k -
meta-llama/Llama-3.3-70B-Instruct
Text Generation • 71B • Updated • 327k • • 3.12k -
deepseek-ai/DeepSeek-V3
Text Generation • 685B • Updated • 1.47M • • 4.4k -
google/gemma-2-2b-it
Text Generation • 3B • Updated • 480k • • 1.56k