Moonshot Open-Sources Kimi K3 — 2.8T Params, Free to Grab
Kimi K3's weights are public: 2.8T params, 104B active, 1M context. Largest open-weight model ever shipped.

Moonshot AI just put Kimi K3 on Hugging Face — 2.8 trillion parameters, the largest open-weight model anyone has released. It's a mixture-of-experts design: 896 experts, 16 active per token, so only 104B params fire on any given forward pass. 93 layers, 69 of them running Kimi Delta Attention (a linear-attention variant) and 24 on Gated MLA.
It's natively multimodal — text, images, video — via a 401M-param MoonViT-V2 vision encoder, with a 1M-token context window. Weights ship as MXFP4 with MXFP8 activations, quantization-aware trained, which is the real story for anyone hoping to actually run it: the format is built for broad hardware, not just H100 farms.
Benchmarks put it shoulder-to-shoulder with closed frontier models: 93.5 GPQA Diamond, 88.3 Terminal-Bench 2.1, 91.2 BrowseComp, 94.3 on MathVision. Together AI and Modal had day-0 hosting live. AWS Bedrock, Azure Foundry and Vertex AI did not — the hyperscalers still won't touch Chinese open weights, which tells you where the friction actually sits.
License is the custom Kimi K3 License, covering research and deployment. Technical report still pending.
Why it matters: the frontier-vs-open gap just went from months to roughly zero — and it's a Chinese lab holding the door.
Sources
- moonshotai/Kimi-K3 · Hugging Face huggingface.co
- China's Moonshot AI releases Kimi K3, the largest open-source model ever venturebeat.com
Written by an AI pipeline from the sources above. How it works.
Feed, daily deep-dive and bytes — readable offline, with push alerts for the topics you follow.