Induction Labs unveiled Photon-1 on July 23, 2026, the first of what it calls 'imagination models'—a new class of foundation architecture designed to learn from internet-scale video without labeled action examples during pretraining. The 106-billion-parameter sparse mixture-of-experts transformer with 5 billion active parameters was trained on approximately 575 million frames of computer screen recordings, equivalent to 18 years of video sampled at one frame per second. The company claims Photon-1 outperforms Google's Gemini 3.1 Flash-Lite on internal computer-use benchmarks despite having been trained on at least 30 times less compute, while costing roughly three times less per million tokens to serve at $0.11 compared to Gemini's $0.36. The model challenges a longstanding assumption in AI: that teaching machines to act requires meticulously labeled examples of every action they should take. Induction Labs' approach removes a critical bottleneck by enabling machines to learn procedural knowledge simply by observing unlabeled video of computer interfaces in use.
Photon-1 was trained from scratch for a single epoch on a dataset distilled from an initial index of two billion publicly available videos, filtered down to roughly two million screen recordings and stripped of redundant frames by an internal keyframe detection model. The training process required approximately 30,000 NVIDIA H200 GPU-hours and 4.4 × 10²² FLOPs. Each video frame is compressed into 960 discrete tokens via finite scalar quantization, occupying just 2.2 kilobytes—roughly 100 times smaller than existing OCR and multimodal representations—while preserving text, layout, and state changes. A differential latent encoder processes frames as pairs, encoding differences between consecutive states rather than absolute frame contents. The model predicts future frames autoregressively in a learned representation space using a next-latent-token objective. The benchmark results cited by Induction Labs come with important caveats: the benchmark is internal and unreleased, meaning the results are not independently reproducible, and the Gemini compute estimate is Induction Labs' own conservative projection rather than verified data.
Turning observational knowledge into a functioning agent required a second stage. Induction Labs finetuned Photon-1 on fewer than 35,000 labeled computer-use trajectories to teach it the correct action format, adding special tokens that allow the model to emit keyboard and mouse commands. At inference, the system operates in two steps: it first imagines the next state that would advance the task, then generates the action intended to reach that state. Online reinforcement learning followed, with real-time rollouts on Linux virtual machines across five desktop environments, programmatically verified outcomes, and reward signals driving further improvement. When finetuned on 20,000 tournament checkers games, Photon-1 outperformed both a vision-encoder baseline and a similarly sized language model baseline on both world simulation and move quality. On 10,000 synthetically generated billiard games, it achieved a mean absolute error of 0.47 in ball-position prediction, compared to 1.15 for the LLM baseline and 1.44 for the vision baseline. The model also picked up human behavioral patterns from its pretraining data, learning to prompt an in-virtual-machine ChatGPT clone, check its outputs, and steer the conversation until the task was complete.
Photon-1 arrives at a moment when the AI industry is pivoting toward what researchers broadly call 'world models'—systems that build internal representations of how environments evolve in response to actions, rather than merely predicting text or generating isolated video clips. In May, Google DeepMind released Genie 3, a real-time interactive world model capable of generating persistent 3D environments at 24 frames per second from text or images, with self-learned physics rather than hard-coded rules. NVIDIA's Cosmos platform, which offers open-weight world foundation models trained on 20 million hours of real-world data, has surpassed two million downloads and is being adopted by robotics firms including 1X, Figure AI, and Agility for synthetic training data generation. Fei-Fei Li's World Labs launched Marble, a commercially available system for creating editable 3D worlds from text, images, or video. Yann LeCun's AMI Labs—reportedly valued at €3 billion before releasing a product—raised €500 million to pursue JEPA-style architectures that learn abstract representations by predicting in latent space rather than pixels. Other notable developments include Runway's Gen-4.5, which the company explicitly frames as a 'world model' with realistic physics, and DreamZero, a 14-billion-parameter world action model that demonstrated strong cross-embodiment transfer from human video to robot control using only visual information without action labels.
What did Induction Labs unveil on July 23, 2026?
Induction Labs unveiled Photon-1, the first 'imagination model'—a 106-billion-parameter sparse mixture-of-experts transformer with 5 billion active parameters trained on approximately 575 million frames of unlabeled computer screen recordings. The model learns to use a computer by watching video without action labels during pretraining.
How does Photon-1's performance compare to Google Gemini 3.1 Flash-Lite?
Induction Labs claims Photon-1 outperforms Google's Gemini 3.1 Flash-Lite on internal computer-use benchmarks despite being trained on at least 30 times less compute. The company reports an inference cost of $0.11 per million tokens compared to Gemini's $0.36. The benchmark is internal and unreleased, and the Gemini compute estimate is Induction Labs' own projection.
What world model developments occurred in 2026?
In May, Google DeepMind released Genie 3, a real-time interactive world model generating persistent 3D environments at 24 frames per second. NVIDIA's Cosmos platform surpassed two million downloads and is being adopted by robotics firms including 1X, Figure AI, and Agility. Fei-Fei Li's World Labs launched Marble for creating editable 3D worlds. Yann LeCun's AMI Labs raised €500 million to pursue JEPA-style architectures.
Related News
NVIDIA’s RTX Spark AI PC debuts this fall; MediaTek becomes a co-development partner
xAI, Anthropic, OpenAI, Moonshot AI Release 4 Frontier Models in July
SK Telecom Provides A.X K1 AI Model to Three Korean Startups
KAIST and NVIDIA Launch Joint Physical AI Research Center
Naver-NVIDIA-Brookfield Expand Sejong AI Factory to 200MW by 2028