Open source models
Open source (and open-weight) models publish downloadable parameters so you can self-host, inspect configuration, fine-tune and deploy inside private infrastructure β subject to the licence.
Models in this categoryβ
| Model | Publisher | Highlights | Guide |
|---|---|---|---|
| Kimi K3 | Moonshot AI | ~2.8T MoE, ~104B active, 1M context, KDA, custom licence | Kimi K3 |
| Qwen 3.5 | Alibaba / Qwen | Hybrid Gated DeltaNet + MoE, 397B-A17B, Apache 2.0 | Qwen 3.5 |
| Gemma 4 | Google DeepMind | Dense/MoE ladder, local multimodal, Apache 2.0 | Gemma 4 |
| Llama 4 | Meta | Maverick / Scout MoE, early fusion, Community Licence | Llama 4 |
| DeepSeek | DeepSeek | MLA, DeepSeekMoE, GRPO, V3/R1/V4 long-context | DeepSeek |
| Mistral | Mistral AI | Small 4 unified MoE, Medium 3.5 dense, Apache 2.0 | Mistral |
| GLM-5 | Z.ai | Sparse attention, agentic engineering, Apache 2.0 | GLM-5 |
| Nemotron 3 | NVIDIA | Hybrid Mamba + LatentMoE, open recipes, NVFP4 | Nemotron 3 |
| Command A+ | Cohere | 218B/25B MoE, RAG/agents, sovereign deploy, Apache 2.0 | Command A+ |
| gpt-oss | OpenAI | 120B/20B reasoning MoE, MXFP4, Apache 2.0 | gpt-oss |
How to read an open-weight guideβ
- What kind of model β open-weight vs open-data vs open-code; licence implications.
- Configuration β total vs activated parameters, layers, experts, context.
- Architecture β attention, MoE, vision, residuals and stability tricks.
- Training β pre-train β long context β SFT β RL β distillation β quantisation.
- Deployment β hosted API vs true self-host; GPU shapes; serving engines; ops checklist.
Discussion
Commentsβ
Share feedback or questions about this page. No account required.
Loading commentsβ¦