Induction Labs Unveils Photon-1: A 106B 'Imagination Model' That Learns by Watching
By predicting future latent states instead of pixels, Photon-1 achieves 100x better compression and beats LLMs trained with 30x more compute.
Internet video is the ultimate repository of human knowledge, capturing millions of hours of people using computers, interacting in the physical world, and performing skilled work. But extracting that knowledge has traditionally forced AI researchers into a brutal trade-off: either rely on expensive, hard-to-scale action labels, or burn massive amounts of compute generating raw pixels.
Induction Labs has just introduced a third path. With the release of Photon-1, a sparse 106B-A5B Mixture-of-Experts (MoE) transformer, the company has debuted a new architecture it calls "imagination models." Instead of generating pixels or relying on labeled actions, Photon-1 learns by predicting future states in a highly compressed, learned representation space.
The results are staggering. On internal computer-use benchmarks, Photon-1 outperforms a production LLM trained with 30x more FLOPs, while being 3x cheaper to serve. It proves that models can learn to act purely through observation—watching 18 years of computer demonstration video—and then translating those observations into emergent capabilities like simulating physics, playing checkers, and steering other AI tools.
The Architecture: Next-Latent-Token Prediction
At the core of Photon-1 is a shift away from the standard autoregressive pixel generation that has bottlenecked video models. Standard video generation models are forced to dedicate the vast majority of their parameters to modeling high-frequency pixel noise—the exact texture of a cursor, the slight compression artifacts in a YouTube video, or the exact shade of a background. This is computationally wasteful if the goal is simply to understand what is happening.
Imagination models bypass this by predicting future frames autoregressively using a next-latent-token-prediction objective. During pretraining, Photon-1 never generates a single pixel. Everything is modeled in a highly compressed representation space. By predicting what the next semantic state of a computer screen will look like, the model builds an internal world model of cause and effect. It learns to act implicitly, despite seeing zero action labels in its training data.
The Secret Sauce: Finite Scalar Quantization (FSQ)
To make latent prediction computationally feasible at scale, Induction Labs engineered a vision encoder that achieves 100x better compression compared to existing OCR and multimodal-model representations.
The math behind this compression is elegant and solves a major problem in representation learning. Instead of using continuous latents or traditional Vector Quantized Variational Autoencoders (VQ-VAE)—which are notoriously prone to codebook collapse where the model ignores large portions of its vocabulary—Induction Labs utilizes Finite Scalar Quantization (FSQ).
- The Tokenization: The encoder compresses each video frame into exactly 960 discrete tokens.
- The Vector Space: Each token is an 8-dimensional vector.
- The Constraints: Each dimension can only take one of five fixed values: {-1, -1/2, 0, 1/2, 1}.
- The Vocabulary: This yields a vocabulary of 5⁸ (390,625) possible codes.
- The Payload: The resulting encoding totals just 2.2KB per frame ($log_2(5^8) \times 960$ bits).
Despite this extreme compression, the representation preserves the text, layout, and state changes required to fully understand a complex video interface.
Differential Encoding: Focusing on the Delta
Furthermore, Photon-1 utilizes a differential latent encoder. In a computer environment, 90% of the screen remains static between actions. A traditional model re-encodes the entire screen every frame, wasting compute on static information.
Photon-1's encoder processes video frames as pairs and encodes only the differences between them. This differential approach drastically reduces the entropy the model needs to predict, focusing its capacity entirely on the actual state changes—a mouse click opening a menu, text appearing in a terminal, or a window minimizing.
Training at Scale: 18 Years of Video
Photon-1 is a sparse MoE transformer with 106 billion total parameters, but it only activates 5 billion parameters per token (106B-A5B). This sparsity is key to its inference efficiency, making it 3x cheaper to serve than comparable dense models.
The pretraining dataset is a masterclass in data curation:
- The Source: Induction Labs started with an internal index of 2 billion publicly available internet videos.
- The Filter: Through extensive video and frame-level filtering, they aggressively pruned non-computer-use content, narrowing the dataset down to approximately 2 million high-quality computer screen recordings.
- Keyframe Detection: An internal keyframe detection model was deployed to strip out redundant frames, ensuring the model wasn't wasting compute on static screens.
- The Scale: The final pretraining corpus consisted of 575 million frames. Sampled at 1 frame per second, this equates to 18 years of continuous video, or 552 billion tokens.
Photon-1 was pretrained from scratch for a single epoch on this dataset.
Performance: Beating the 30x FLOP Goliath
The most shocking claim from Induction Labs is Photon-1's compute efficiency. On internal computer-use benchmarks, Photon-1 outperforms a production LLM trained with 30x more FLOPs.
Let that sink in. In an era where frontier models cost hundreds of millions of dollars to train, Induction Labs has achieved superior agentic performance using roughly 3% of the compute budget of a flagship LLM. This proves that for specific domains like computer use and physical simulation, predicting latent video states is a vastly more sample-efficient learning objective than predicting the next text token.
The Contrast with Text-Based Agents
In July 2026, the standard approach to building AI agents—seen in models like Anthropic's Claude Fable 5 or OpenAI's GPT-5.6 Sol—is to take a massive text-based LLM and give it access to computer tools via API calls. The model reads the screen (often via standard OCR or heavy vision-language encoders), reasons in text, and outputs a text command to move the mouse or click a button.
While effective, this approach is fundamentally misaligned with how humans operate. We don't translate our visual field into text before deciding to click a button. We operate in a continuous, visual-spatial world.
Photon-1 eliminates this translation layer. Because it is natively trained on video and operates entirely in a visual representation space, it doesn't need to convert a screen into text to understand it. It simply imagines the next visual state. This native visual understanding is precisely why it can outperform models trained with 30x more compute—it isn't wasting FLOPs translating between modalities.
Emergent Capabilities: From Screen Pixels to World Physics
Because Photon-1 was trained exclusively on computer-use videos, one might expect its capabilities to be strictly limited to UI navigation. However, the model demonstrates remarkable out-of-domain generalization, proving that computer interfaces can serve as a microcosm for broader world understanding.
Some of the most notable emergent capabilities include:
- Simulating Computers: Photon-1 acts as a robust computer world model. Given a single seed screenshot, it can accurately generate all subsequent states based on a proposed action.
- Modeling Physics: After finetuning on a dataset of billiard videos, Photon-1 successfully simulates complex physical interactions, accurately predicting the future positions of billiard balls.
- Playing Games: When finetuned on 20,000 tournament games of checkers, Photon-1 learned to play the game at a high level, outright beating traditional LLMs that were finetuned on the exact same dataset.
- Tool Steering: Following reinforcement learning, Photon-1 exhibits human-like behavior in how it interacts with other AI models. It can effectively "steer" ChatGPT, prompting the LLM to generate specific outputs needed to complete a broader task.
From Imagination to Action: The RL Phase
How does a model that only predicts the future actually take action? The answer lies in Reinforcement Learning (RL).
During pretraining, Photon-1 learns the dynamics of the world—if action X happens, state Y will follow. It builds a causal map of its environment. Once this world model is established, Induction Labs applies RL to teach the model to search for the actions that lead to a desired future state.
When given a task like "Open the Steam page for Counter-Strike 2," the model imagines multiple possible future trajectories. It evaluates which trajectory successfully navigates the UI to reach the goal, and then executes the action that leads to that specific future. The imagination is predicted entirely in the representation space, but Induction Labs notes it can be visualized using an auxiliary finetuned image generation model for debugging and interpretability.
The Road Ahead
The AI industry is currently obsessed with scaling laws, assuming that simply throwing more compute at next-token text prediction will eventually yield AGI. Induction Labs' Photon-1 challenges this orthodoxy.
By proving that we don't need to generate pixels to understand video, and we don't need action labels to learn how to act, imagination models offer a highly scalable, compute-efficient path forward. As we look beyond the current paradigm of static text and image generation, architectures like Photon-1—which learn by watching and act by imagining—are poised to become the foundation for the next generation of autonomous agents.
Sources
- Scaling Video Pretraining with Imagination Models — Induction Labs inductionlabs.com
- Anthropic's Opus 5 surprise - The Rundown AI therundown.ai
Feed, daily deep-dive and bytes — readable offline, with push alerts for the topics you follow.