Kimi K3's weights are free. The 64 GPUs to run them aren't.
Moonshot open-sourced a 2.8T-parameter model on July 27. Self-hosting it needs 64+ accelerators, so most devs will just pay the API.

Moonshot AI released the weights for Kimi K3 on 27 July — 2.8 trillion parameters, native vision, a 1M-token context window, and the largest open-weight model anyone has shipped. Free to download. Also free to discover you can't actually run it.
The architecture is a Stable LatentMoE stack activating 16 of 896 experts, built on Kimi Delta Attention, for roughly 2.5× the scaling efficiency of K2. But all 896 experts have to sit in memory at once, which puts the weights somewhere between ~600GB and 1.4TB depending on precision. Moonshot's own deployment guidance is a supernode with 64+ accelerators — at typical cloud rates that's ~$250/hour, north of $6,000 a day, before a single token gets served.
Meanwhile the hosted API runs $3 per million input tokens ($0.30 on cache hits) and $15 per million output. The crossover where owning the metal beats renting the endpoint sits somewhere past a billion tokens a month — which is to say, not you.
Why it matters: open weights are quietly moving the moat from model access to the capital required to serve them — "open" now means auditable, not affordable.
Sources
Independent coverage
- Moonshot AI to make Kimi K3 available for public download technode.com
- Kimi K3: benchmarks, pricing, hardware requirements, and self-hosting northflank.com
Written by an AI pipeline from the sources above. Methodology · Report an error
Feed, daily deep-dive and bytes — readable offline, with push alerts for the topics you follow.