Anthropic Reveals 80% of Its Code Is Now Written by Claude and the Task Horizon Doubles Every Four Months — The Company That Studies Existential Risk Just Published Its Own

🤚 The Open-Palm Disclosure

In a move that reads like a quarterly earnings call drafted by a philosophy department, Anthropic published a report on Thursday titled “When AI Builds Itself: Our Progress Toward Recursive Self-Improvement” — a document that casually reveals the company has crossed several thresholds that science fiction writers spent decades warning us about.

The headline numbers, for those who enjoy reading their own obsolescence in bullet-point format:

  • Over 80% of code merged into Anthropic’s production codebase is now authored by Claude, as of May 2026
  • Engineers ship 8x more code per quarter compared to the 2021–2025 baseline
  • Claude’s task completion time horizon has been doubling roughly every four months — from 4-minute tasks in March 2024 to 12-hour tasks with Claude Opus 4.6 in March 2026
  • Projection: tasks requiring days could be within range by end of 2026; weeks-long tasks by 2027

One Anthropic employee was quoted saying they haven’t personally written any code in five months. This person still has a job. The code, apparently, does not.

👐 The Two-Handed Reckoning With the Numbers

Let’s sit with the benchmarks for a moment, because they deserve the kind of reverent silence usually reserved for watching someone parallel-park a yacht.

SWE-bench, the standard software engineering benchmark, went from “low single digits” to saturated in two years. CORE-Bench, which measures research reproduction, improved from ~20% success in 2024 to benchmark saturation in fifteen months. Claude’s session success rate on open-ended problems hit 76% in May 2026 — a fifty-percentage-point improvement in six months.

And here’s the number that should make every performance review committee reconsider its existence: a median survey of 130 Anthropic researchers reported a 4x output multiplier when using Mythos Preview. Not a 4x improvement in vibes. A 4x improvement in output. The kind of multiplier that makes headcount planning look like astrology.

On code optimization, Claude improved from ~3x speedups in May 2025 to ~52x speedups in April 2026. On research direction selection — the ability to choose which problems are worth solving — models matched or exceeded human judgment 64% of the time, up from 51% just five months earlier.

The report carefully notes that Claude-written code was worse than human code in late 2025, is at parity today, and is expected to be superior within one year. The tone is clinical. The implications are not.

🌿 The Gentle Awakening

Anthropic helpfully outlines three possible futures, presented with the emotional restraint of a waiter describing tonight’s specials while the kitchen is on fire:

Future One: The trend stalls, but today’s capabilities diffuse widely. Everyone gets a very talented junior engineer that never sleeps. Society adjusts. LinkedIn influencers pivot to “AI-augmented leadership coaching.”

Future Two: Compounding efficiency gains continue. Humans set directions while AI executes. The org chart becomes a suggestion box. Engineering teams shrink. The word “leverage” appears in every earnings call for the next decade.

Future Three: Full recursive self-improvement. AI systems design and train their successors autonomously. The report acknowledges this could happen “sooner than institutions are prepared for,” then immediately adds it’s “not inevitable” — which is the scientific equivalent of saying “the building probably won’t collapse.”

The company notes that “human review becomes the bottleneck” once AI code quality reaches parity. Research “taste” — the ability to identify which problems matter — remains the primary human advantage. For now. The “for now” is doing a lot of structural load-bearing in that sentence.

👑 The Gold-Leaf Existential Audit

What makes this report unusual isn’t the numbers. It’s that Anthropic is publishing this about itself. This is a company telling you, in peer-reviewed detail, that its own product is automating its own workforce, and that the trend line points toward a future where the distinction between “tool” and “colleague” requires a philosophical degree to parse.

The report states that a “meaningful slowdown” would require multiple well-resourced labs at or near the frontier, in multiple countries, with verifiable coordination mechanisms. Translation: the kind of international cooperation humanity has historically achieved exactly never, unless you count the metric system.

Meanwhile, Anthropic says it has already encountered “explosive growth in new ideas and initiatives exceeding organizational capacity to pursue them.” The AI is generating more work than the humans can evaluate. The bottleneck isn’t compute. It isn’t data. It’s us.

Somewhere in a San Francisco office, an engineer who hasn’t written code in five months is reading this report and wondering whether “setting direction” counts as a job skill or a coping mechanism.

“Eighty percent of the code is written by the AI, the other twenty percent is the engineers explaining to the AI what the first eighty percent was supposed to do. This is called ‘alignment’ and it is going great.” — The Slap of Wisdom Recursive Improvement Desk, currently being optimized by a model that has opinions about its own performance review