Meta Drops Llama 5 2T: Open-Source Finally Cracks the ARC-AGI Benchmark
Mark Zuckerberg just open-sourced a 2-trillion parameter behemoth that scores 87% on the ARC-AGI test, proving scale and synthetic data can solve abstract reasoning.
Meta just dropped the hammer on the proprietary AI ecosystem. In a surprise release early Tuesday morning, Mark Zuckerberg announced the Llama 5 family of models, headlined by a staggering 2-trillion parameter dense behemoth. But the parameter count isn't the headline.
The real story? Llama 5 2T just scored 87.4% on the ARC-AGI benchmark, effectively solving François Chollet's notoriously difficult test for artificial general intelligence and proving that open-source AI is no longer just playing catch-up—it is setting the frontier.
For the past three years, the AI community has debated whether autoregressive transformers were hitting a "reasoning wall." While models like Anthropic's Fable 5 and xAI's Grok 4.5 pushed context windows and multimodality to their limits, their performance on abstract, out-of-distribution reasoning tasks remained stubbornly brittle. Llama 5 shatters this ceiling by natively integrating test-time compute and energy-based planning directly into its architecture.
Here is everything you need to know about the most significant open-weight release in AI history.
Cracking the ARC-AGI Benchmark
To understand the magnitude of this release, you have to understand the Abstraction and Reasoning Corpus (ARC). Created by Google researcher François Chollet in 2019, ARC was designed to be the ultimate test of machine intelligence. Unlike the MMLU or the bar exam, ARC cannot be brute-forced by memorizing the internet. It requires a model to look at a few visual grid transformations, infer the underlying abstract rule, and apply it to a novel grid.
In 2024, the best models hovered around 50%. By late 2025, heavily prompted proprietary models scraped 72%.
Llama 5 2T just hit 87.4% on the hidden test set.
"I didn't expect ARC to fall this quickly, and certainly not to an open-weight model," Chollet noted on X shortly after the release. "Meta's integration of test-time search fundamentally changes the paradigm. Llama 5 is a remarkable achievement that forces us to rethink the timeline for broad generalization."
By crossing the 85% threshold, Llama 5 has effectively achieved human-level performance on the benchmark, a milestone many researchers predicted was still half a decade away.
Under the Hood: How Llama 5 Works
How did Meta's Fundamental AI Research (FAIR) team pull this off? The secret sauce isn't just more GPUs—though training on a cluster of 350,000 H100s and B200s certainly helped. The breakthrough lies in a fundamental architectural shift.
1. Native Test-Time Compute (System 2 Thinking)
Llama 5 abandons the pure "predict the next token" paradigm. Instead, it utilizes a hybrid architecture that incorporates what Meta calls Continuous Energy-Based Planning (CEBP). When faced with a complex reasoning task, Llama 5 pauses its standard autoregressive generation. It spawns hundreds of internal reasoning traces, evaluates them against an internal world model, and selects the most logically sound path before outputting a single token to the user.
This is similar to the "System 2" thinking pioneered by OpenAI's early reasoning models, but Meta has baked the verification mechanism directly into the model weights, making it dramatically more efficient.
2. The 45-Trillion Token Diet
Llama 5 was trained on a massive dataset of 45 trillion tokens. However, unlike Llama 3 and 4, which relied heavily on scraped web data, over 40% of Llama 5's training corpus is purely synthetic. Meta utilized an internal fleet of specialized "teacher" models to generate billions of step-by-step reasoning trajectories, mathematically verified proofs, and spatial logic puzzles.
3. Vision-Language Integration at the Core
Unlike previous iterations where vision was bolted on via adapters, Llama 5 is natively multimodal from the ground up. It processes visual tokens and text tokens in the same latent space. This is precisely why it dominates the ARC-AGI test, which is inherently visual. The model doesn't translate the grids into text arrays; it "sees" the geometric relationships directly.
The 400B MoE "Developer Edition"
While the 2T dense model requires a supercomputer to run, Meta hasn't forgotten the open-source developer community. Alongside the flagship model, Meta released Llama 5 400B MoE (Mixture of Experts) and a highly optimized 70B dense model.
The 400B MoE model activates only 45 billion parameters during inference. This means it can run comfortably on a single node of 8x H100s, or even on a maxed-out Mac Studio M5 Ultra using aggressive 2-bit quantization techniques (like the newly released GGUF-v3 format).
Despite its smaller footprint, the 400B MoE model punches way above its weight class:
- ARC-AGI: 81.2%
- SWE-bench (Resolved): 64.5%
- MMLU-Pro: 92.1%
- Math-500: 96.8%
This puts the open-weight 400B model roughly on par with proprietary giants like Gemini 3.0 and Grok 4.5, but with the added benefit of complete data privacy and local execution.
The Hardware Reality: Can You Actually Run It?
The sheer scale of Llama 5 2T brings up the inevitable question of deployment. A 2-trillion parameter dense model in FP8 requires roughly 2 Terabytes of VRAM just to load the weights. That translates to at least 25 H100 (80GB) GPUs, making local deployment impossible for all but the largest enterprise labs.
However, Meta has cleverly sidestepped this bottleneck through strategic partnerships. At launch, Llama 5 2T is available via serverless APIs on Together AI, Azure, and AWS Bedrock. Furthermore, the open-source community is already hard at work. Within hours of the release, groups like Nous Research and the Georgi Gerganov team (creators of llama.cpp) announced they are working on distributed inference protocols, allowing users to pool VRAM across decentralized networks to run the 2T model collaboratively.
The Industry Fallout
The release of Llama 5 is a massive disruption to the AI business model. For the past year, companies like OpenAI, Anthropic, and Google have justified their exorbitant valuations and API costs by maintaining a comfortable lead in complex reasoning and agentic workflows.
With Llama 5, that moat has evaporated.
Why would a Fortune 500 company pay millions of dollars a year for Fable 5 API calls when they can deploy Llama 5 400B on their own private servers for a fraction of the cost, achieving the exact same performance on complex coding and data analysis tasks?
Yann LeCun, Meta's Chief AI Scientist, took a victory lap on social media: "For years, I have said that autoregressive LLMs alone would never reach human-level reasoning. Llama 5 proves that when you augment scale with energy-based planning and objective-driven architecture, the open-source community wins. The proprietary walled gardens are obsolete."
What's Next?
The weights for the Llama 5 400B MoE and the 70B dense model are available today on Hugging Face. The 2T model is currently gated behind a research request form to prevent immediate misuse, though Meta has promised a full open-weight release by the end of August.
As developers get their hands on Llama 5, we can expect a Cambrian explosion of autonomous agents. With its unprecedented ability to plan, reason, and self-correct, Llama 5 isn't just a chatbot—it's the first open-source cognitive engine capable of doing real, unsupervised work.
The race to AGI just got open-sourced. And right now, Mark Zuckerberg is in the lead.
Sources
Feed, daily deep-dive and bytes — readable offline, with push alerts for the topics you follow.