Inkling AI: Mira Murati’s 975B Open-Weight Giant Changes the MoE Game
Thinking Machines Lab just dropped a massive multimodal MoE with a 1-million token context and a controllable "thinking effort" dial—and it's completely open-source.
The AI community has spent the last 48 hours dissecting the real-world implications of Inkling, the highly anticipated debut model from Thinking Machines Lab. When Mira Murati stepped down as OpenAI’s Chief Technology Officer in late 2024, the industry speculated heavily on her next move. Having had a front-row seat to the development of the world's most powerful closed, proprietary AI systems, her pivot to an open-weights philosophy is a massive statement.
Thinking Machines Lab has taken a radically different approach from her former employer. They have built a staggering 975-billion parameter multimodal model, released it fully under the permissive Apache 2.0 license, and openly admitted it isn't the smartest model on the leaderboard.
Instead of chasing benchmark supremacy at all costs, Inkling is engineered for developers who need deep customization, manageable inference costs, and native multimodality without the vendor lock-in of a closed API. Here is a deep dive into the architecture, the multimodal capabilities, and the "thinking dial" that makes Inkling one of the most pragmatic AI releases of 2026.
Under the Hood: 975B Parameters, but Only 41B Active
At first glance, a 975-billion parameter model sounds impossible to self-host without a massive, dedicated data center. However, Inkling is built on a highly optimized sparse Mixture-of-Experts (MoE) architecture that drastically reduces its computational footprint during inference. While the model boasts nearly a trillion total parameters, it only activates 41 billion parameters per inference step.
This extreme sparsity is achieved through a massive and intricate routing network:
- The architecture features 256 routed experts and 2 shared experts within each MoE feed-forward layer.
- For every single token processed, the model activates just 6 routed experts alongside the 2 shared experts.
- The network spans 66 decoder-only transformer layers, utilizing a sophisticated mix of local sliding-window attention and global attention layers to handle massive contexts efficiently.
This design allows Inkling to retain the vast knowledge capacity and reasoning capabilities of a frontier-class model while keeping inference costs remarkably low. You simply don't need to wake up 975 billion parameters to answer a basic routing query or extract a date from a document.
Furthermore, the model supports a native 1-million-token context window in its weights (though most third-party hosted APIs currently cap it at 256K for memory management). This makes it uniquely capable of ingesting entire enterprise codebases, massive legal archives, or hours of transcribed audio in a single, unbroken prompt.
Encoder-Free Native Multimodality
Most open-weight models available today treat vision and audio as afterthoughts. The standard industry practice has been to bolt separate vision encoders (like CLIP) onto a pre-trained text model. While this works for basic image captioning, it severely limits the model's ability to perform deep, cross-modal reasoning.
Inkling was built to be multimodal from the ground up. It utilizes an encoder-free architecture that ingests different modalities directly into the shared token representation space:
- Images are processed directly as 40×40 pixel patches, bypassing the need for a separate visual translation layer.
- Audio is fed in natively as dMel spectrograms, supporting 16 kHz WAV files up to roughly 20 minutes in length.
Because image, audio, and text understanding are baked into the exact same neural architecture, Inkling excels at tasks that require synthesizing multiple data types. It doesn't just transcribe an audio file; it can reason about the emotional tone of a voice recording while simultaneously analyzing a UI screenshot and a block of code to debug a frontend application.
The Killer Feature: Controllable "Thinking Effort"
Perhaps the most innovative production feature in Inkling is its controllable thinking effort.
In the current AI landscape, inference is generally treated as a fixed computational cost. Whether you ask an LLM to output the capital of France or write a complex Python script, the model expends a similar amount of baseline computational energy per token. Thinking Machines Lab has completely upended this paradigm by exposing a "thinking dial" (ranging from 0.2 to 0.99) that allows developers to dictate the model's reasoning budget per request.
- Low Effort (0.2 - 0.4): Ideal for basic classification, data extraction, or simple customer support routing. The model responds almost instantly with minimal compute, drastically reducing API costs.
- Medium Effort (0.5 - 0.7): The sweet spot for standard conversational agents, RAG (Retrieval-Augmented Generation) synthesis, and standard copywriting.
- High Effort (0.8 - 0.99): Forces the model to allocate maximum compute for complex multi-step reasoning, advanced coding tasks, or autonomous agent planning.
In enterprise deployments, treating every prompt equally is an expensive mistake. The thinking dial gives engineers granular control over their token economics, allowing them to dynamically scale compute based on the complexity of the user's request. A customer service bot, for example, can use low effort to classify an incoming ticket, and then automatically crank the dial up to 0.9 if the ticket requires deep technical troubleshooting.
Benchmarks and Honest Positioning
In an industry plagued by cherry-picked benchmarks and exaggerated marketing claims, Thinking Machines Lab's launch strategy is remarkably candid. The company states outright that Inkling "is not the strongest model available today, closed or open."
Independent evaluations back this up. On the Artificial Analysis Intelligence Index, Inkling scores a 41/100. While this places it behind closed-source behemoths like Anthropic's Claude Opus 5 and OpenAI's GPT-5.6 series, it holds its own incredibly well in the open-weights arena.
When pitted against open-weights rivals like Nemotron 3 Ultra, Kimi K2.6, and DeepSeek V4 Pro, Inkling shows highly competitive performance, particularly in reasoning-heavy benchmarks (MMLU, GPQA) and visual mathematical reasoning (MathVista). Its MoE architecture routes inputs to specialized expert clusters, which demonstrably improves performance on structured logic tasks compared to dense models of similar active parameter counts.
The Business Play: Open Weights, Paid Fine-Tuning
Why would a heavily funded AI lab spend millions of dollars in compute to train a 975B parameter model, only to give it away for free? The strategy is a direct counter to the closed-API moats of OpenAI and Anthropic.
Thinking Machines Lab isn't trying to monetize the raw model weights. Instead, they are monetizing the enterprise customization layer through their Tinker fine-tuning platform. By releasing Inkling under the permissive Apache 2.0 license, they are commoditizing the foundation model layer to drive adoption of their paid post-training, alignment, and deployment tools.
For teams building specialized AI agents, the choice is no longer between a weak open-source model and an expensive, restrictive closed API. Inkling provides a massive, highly capable, and customizable foundation that enterprises can actually own and control. It represents a maturation of the open-source AI ecosystem—proving that you don't need to win the benchmark war to build the most useful model for real-world engineering.
Sources
Feed, daily deep-dive and bytes — readable offline, with push alerts for the topics you follow.