Twenty-six percent sounds close to autonomy until you read Anthropic’s scale. The company says Claude now “leads” 26% of measured work involved in improving its own AI systems. On Anthropic’s six-level framework, that means the AI can drive much of the task while a human still provides oversight. Full autonomy is a separate level, and Anthropic reports none of the measured work there.1
Anthropic says Claude now “leads” 26% of measured internal AI R&D work and more than 90% is at least collaborative. But “leads” is not autonomous: it is level 4 on Anthropic’s six-level scale, one step below full autonomy. Anthropic reports no measured subset at the fully autonomous level.1, 2

Anthropic’s internal R&D automation snapshot
“Leads” is one step below fully autonomous
| Level | Anthropic label | Plain-English meaning |
|---|---|---|
| AL3 | AI collaborates | The human and AI divide meaningful parts of the work and iterate together. |
| AL4 | AI leads | The AI drives most of the task while a human provides oversight and intervenes when needed. |
| AL5 | AI fully autonomous | The AI can complete the task without human intervention. |
That distinction is the reason the 26% figure should not be rewritten as “Claude does a quarter of Anthropic’s research autonomously.” Anthropic’s own terminology reserves autonomy for AL5. AL4 still assumes human oversight even when the AI is doing much of the substantive work.1
Anthropic built the index from its own work records
Anthropic says it sampled internal AI R&D work, broke that work into thousands of granular tasks, organized those tasks into a hierarchy and estimated how much person-time each area represents. AI judges then classify tasks on the AL0–AL5 scale. The index is therefore an attempt to measure real internal work rather than only benchmark capability.1
The company reports that its task tree contained 542 nodes and 378 leaf tasks after refinement. Anthropic also compared model judgments with human labels. Exact agreement was imperfect, while agreement within one automation level was much higher. That is useful context because the headline percentages depend on classification, not a direct clock attached to every task.1
The measurement is self-referential by design
Anthropic is measuring its own models inside its own organization using a methodology it designed. The company explicitly says there is no common cross-lab framework yet and discusses the risk of using AI systems in the evaluation itself. That makes the index valuable as a longitudinal internal measure, but weak evidence for declaring one lab more automated than another.1
Thirty thousand agents changes the supervision problem
Anthropic says roughly 30,000 research and engineering agents are active at any one time on its most-used internal platform. At that scale, supervision is no longer only a person reading one answer. Teams need ways to monitor many parallel agents, inspect outputs and decide where human intervention is still required.1
Anthropic’s separate research on deployed-agent autonomy makes the same distinction from another direction. A model can be capable of long autonomous runs without users actually giving it that freedom in practice. People change oversight, permissions and intervention behavior as tasks become more consequential.2
What the 26% figure does and does not mean
- It measures Anthropic’s own AI R&D work under an internal six-level automation framework.
- It says Claude leads a measured share of tasks while humans still provide oversight.
- It does not mean 26% of Anthropic’s research organization is autonomous or unnecessary.
- It is not directly comparable with another lab unless that lab uses a compatible task taxonomy and measurement method.
S&C’s agent-success guide argues that task definitions and denominators have to stay visible. The same rule applies here. “Claude leads 26%” is interesting because Anthropic defines what “leads” means, and because that definition explicitly stops short of autonomy.
The larger story is not a clean march from human work to zero-human work. Anthropic’s own data describes a growing middle ground where AI takes the lead and people supervise, redirect and verify. If that middle ground expands, the important management question becomes how much oversight a task needs, not simply whether a model can attempt it.
Sources and methodology
Sources checked September 28, 2026. Dates and periods for individual figures are stated beside them.
- Anthropic Institute: Measuring the pace of AI development ↗Accessed 2026-09-28
- Anthropic: Measuring AI agent autonomy in practice ↗Accessed 2026-09-28
Scope and assumptions
The index is designed by Anthropic to measure Anthropic's own work, so it is not a standardized cross-lab benchmark.
Parts of the evaluation use Anthropic models as judges, and task classification is not perfectly identical to human labeling.
The approximately 30,000 active-agent figure covers Anthropic's most-used internal R&D platform rather than every internal system.
Continue reading
Anthropic Built a Wet Lab So Claude Can Test Its Biology Ideas →
86% Agreement Can Still Miss Half the AI Failures →
40,000 Firms Applied to Anthropic’s Partner Network. Who Makes Money Around the Model? →