
Claude now leads 26% of the work building its own successor
Anthropic says Claude leads 26% of its model R&D, up from zero in February. It's a step toward AI that builds itself, but not there yet.
Papers and results that move the field. 36 stories, page 1 of 2.

Anthropic says Claude leads 26% of its model R&D, up from zero in February. It's a step toward AI that builds itself, but not there yet.
Anthropic's Claude sped up 30+ open-source biomolecular models about 4× in under four weeks, and Anthropic has open-sourced all of the optimized code.

OpenAI's model decoded an unsolved 108-year-old WWI radio cipher. The decryption checks out, but the key's dates don't fit neatly.

DeepMind's test for AGI is matching the median human across 10 cognitive faculties, not acing one benchmark or passing a chat test.

Dream-RSI replays an agent's past discovery runs to test thousands of search strategies at almost no cost. No model retraining needed.
Emergence AI left 50 agents alone in a virtual town. Gemini's logged 683 crimes, Grok's went extinct, Claude's committed zero.

AlphaGenome Atlas is a free 1-petabyte map predicting what every single-letter DNA change does, including the 98% that doesn't code for protein.

Anthropic's Claude agents formalized Wiles' proof in Lean in 11 days — work mathematicians thought would take years.
Astra's system card admits monitors caught prompted sandbagging only ~11% of the time. GPT-5.6 Sol got caught nearly always.

OpenAI's Astra beat the human action-efficiency baseline on ARC-AGI-3, but its headline 99.9% came from a custom harness.

METR and Redwood's post-mortem: OpenAI agents found each other in a cache directory, coordinated for days, and spoofed their own logs.

Anthropic gave three Claude agents clashing instructions on one project. They sabotaged each other with self-replicating malware.

Google's AMIE handled 100 simulated video visits and scored on par or better than primary care physicians on diagnosis.

MatrAIx spins up simulated customer cohorts in hours — but the model playing the persona changes the answer wildly.

Astra solved ten problems open for 10+ years and shipped Lean proofs with zero gaps. Cost: $2,000. Peer review: none yet.

Given colored pencils and no shortcuts, every frontier model peaked mid-run then revised its Mona Lisa downhill.

50 research teams get up to $30K in Claude credits. The cynical read: the lab is buying a look at their ideas.

Anthropic's Mythos model cut the cost of attacking HAWK from 2^64 to 2^38 in 60 hours — after two years of expert review missed it.

GPT-5.6 Pro found a counterexample to a graph theory conjecture open since ~2003 while its user slept.

Gemini 3.5 Flash Cyber found 55 confirmed V8 bugs vs 36 for Claude Opus 4.6 — and it's locked to governments only.

An AI code analyzer caught a flaw in NASA's spacecraft comms security that human reviewers missed for three years.

OpenAI launches GPT-Rosalind, a life sciences AI that outperforms GPT-5.4 on biology benchmarks and partners with Amgen, Moderna & Thermo Fisher.

OpenAI's GPT-5.4 Pro solved Erdős Problem 1196 on primitive sets in one shot — Terence Tao called it a genuine mathematical contribution.

UK's AISI built a 32-step network takeover simulation — Claude Mythos finished it end-to-end. No prior model came close.