Get the app

Mira Murati’s Thinking Machines Drops Inkling: A 975B Open-Weights Giant

Former OpenAI CTO Mira Murati just open-sourced Inkling, a massive 975B parameter MoE model that reclaims the US open-weights crown from China.

The US open-weights ecosystem just got the heavyweight champion it desperately needed. After months of watching Chinese AI labs like DeepSeek, Alibaba, and Moonshot AI dominate the open-source leaderboards, Thinking Machines Lab—the startup founded in early 2025 by former OpenAI CTO Mira Murati—has officially entered the arena.

Their debut model, codenamed Inkling, is a 975-billion parameter Mixture-of-Experts (MoE) behemoth. Released under a highly permissive Apache 2.0 license, Inkling isn't just a research preview; it's a multimodal, reasoning-capable frontier model designed to be fine-tuned, customized, and deployed at scale by enterprise developers.

Here is a comprehensive breakdown of the model that just reclaimed the American open-weights crown, how it works under the hood, and where it fits into the rapidly evolving 2026 AI landscape.

Under the Hood: A 975B Parameter Multimodal Giant

Inkling is undeniably massive, but it is architected specifically with inference efficiency in mind. While the model boasts 975 billion total parameters, its MoE architecture means only about 41 billion parameters are active during any given token generation.

The architecture features 256 routed experts alongside two shared experts, with each token being processed by six experts simultaneously. Thinking Machines admits this design was heavily inspired by the highly successful DeepSeek-V3 and V4 models, but the execution is entirely homegrown. The model was trained from scratch on next-generation Nvidia GB300 NVL72 systems using a staggering 45 trillion tokens of public and synthetic data.

Interestingly, Thinking Machines utilized Chinese frontier models—specifically Moonshot's Kimi K2.5—to generate portions of Inkling's synthetic training data. This highlights a fascinating ouroboros in the current AI meta: US labs are now leveraging the capabilities of Chinese open models to bootstrap the high-quality synthetic data needed to beat those very same Chinese models.

Unlike many open-weights models that bolt on vision or audio capabilities after the fact via cross-attention adapters, Inkling is natively multimodal. It processes text, images, audio, and video out of the box. Furthermore, it features a massive 1-million-token context window. This extended short-term memory is perfect for needle-in-a-haystack retrieval, complex codebase analysis, and processing feature-length video inputs.

Hardware Requirements: Bring Your Own Supercomputer

Running a near-trillion parameter model locally is not for the faint of heart, and Inkling pushes the boundaries of what can be considered "accessible" open source. At its native 16-bit precision, Inkling requires over 2 terabytes of GPU memory. In practical terms, that means you'll need a node with at least eight Nvidia B300 accelerators or sixteen H200s just to load the weights into VRAM.

However, Thinking Machines is keenly aware of the hardware crunch facing the developer community. Alongside the standard FP16 weights, they have released an NVFP4 quantized version of Inkling that cuts the VRAM requirements in half, making it accessible to a wider tier of enterprise deployments. At launch, the model supports a broad ecosystem of inference engines, including vLLM, SGLang, TokenSpeed, and Llama.cpp, ensuring that developers can slot it into existing infrastructure pipelines with minimal friction.

Benchmarks: Dominating Agents, Struggling with Facts

So, how does a 975B parameter model actually perform in the wild? According to the Artificial Analysis Intelligence Index, Inkling debuts with a score of 41. This firmly establishes it as the most powerful open-weights model from a US lab, easily clearing Nvidia's Nemotron 3 Ultra (38) and Google's Gemma 4 31B (29).

Where Inkling truly shines is in agentic workflows and complex reasoning. As a reinforcement learning (RL) trained reasoning model, it utilizes chain-of-thought "thinking" tokens to map out complex problems before committing to an answer.

  • Agentic Dominance: On the GDPval-AA v2 benchmark (which simulates complex, multi-step knowledge-work tasks), Inkling achieved an Elo rating of 1,238. This edges out top-tier Chinese models like Kimi K2.6 (1,190) and DeepSeek v4 Flash max (1,189).
  • Banking and Finance: Inkling scored 24% on the rigorous Tau-3 banking benchmark, pulling ahead of Kimi K2.6's 21%.
  • Token Efficiency: Thinking tokens usually mean higher inference costs and slower time-to-first-token, but Inkling is remarkably concise. It averages just 25,000 output tokens per Intelligence Index task, compared to 43,000 for GLM-5.2 max and 38,000 for Kimi K2.6. This efficiency allows it to match Nemotron 3 Ultra on Terminal Bench 2.1 using roughly a third of the tokens.

But it's not all perfect. Inkling has a severe, documented hallucination problem. On the AA Omniscience benchmark, which tests raw factual accuracy and knowledge retrieval, the model scored a dismal +2. Its factual accuracy sits at just 40%, with a hallucination rate of 63%. If you need a model to accurately recall historical dates, specific medical facts, or obscure trivia zero-shot, Inkling is simply not the tool for the job. It is an engine for reasoning, coding, and agentic planning—not a static encyclopedia.

The Business Play: Tinker and Apache 2.0

Mira Murati's strategy with Thinking Machines is a direct counter to the walled-garden approach of her former employer, OpenAI. By releasing Inkling under the Apache 2.0 license, the company is inviting developers to commercialize, modify, and fine-tune the model without restrictive royalties, acceptable use policies, or the fear of sudden API deprecation.

To monetize this massive open-source contribution, Thinking Machines has launched Tinker, a managed platform for model customization. Tinker provides the cloud infrastructure and tooling necessary to fine-tune Inkling for specific enterprise verticals. In a massive flex of the model's agentic capabilities, Thinking Machines claims Inkling is actually capable of writing its own fine-tuning scripts. Developers can prompt the model to evaluate its own baseline performance on a custom dataset and generate the PyTorch code necessary to teach itself new skills.

For developers who just want standard API access, Inkling is available on Tinker at $1.87 per million input tokens and $4.68 per million output tokens (for the standard 64K context window). For massive context workloads up to 256,000 tokens, pricing scales to $3.74 for input and $9.36 for output. While this is slightly more expensive than API access for equivalent Chinese models, Inkling's high token efficiency means the actual cost-per-task remains highly competitive.

The company is also aggressively partnering with third-party serverless inference providers, bringing Inkling to TogetherAI, Fireworks, Modal, Databricks, and Baseten to ensure developers can access the model wherever their compute lives.

What's Next for Thinking Machines?

Inkling is a massive milestone for the American open-source AI community. It proves that US labs can still compete at the absolute frontier of open weights, providing a viable, commercially permissive alternative to the models coming out of Beijing and Hangzhou. It also validates the thesis that top-tier talent leaving closed-AI labs can successfully bootstrap frontier-class open models.

But Thinking Machines isn't stopping at 975 billion parameters. The company is already teasing Inkling-Small, a 276-billion-parameter MoE model (with 12 billion active parameters) designed for low-latency, high-throughput applications where deploying a 2TB model is overkill.

For now, developers have a new heavyweight champion to play with. Just make sure you have the GPU cluster to run it—and definitely double-check its facts before pushing your agentic workflows to production.

Sources

The daily AI brief, on your phone.

Feed, daily deep-dive and bytes — readable offline, with push alerts for the topics you follow.

Get it on Google Play