Anthropic Says Claude Now Leads 26% of Its AI R&D Work, but Humans Still Supervise the Process
The AI company says Claude can now carry out most of the work on roughly a quarter of its model-development tasks from a high-level instruction, up sharply from less than 1% earlier this year. Anthropic stresses that Claude is not yet operating fully autonomously in any measured area of its AI resea
Anthropic’s Claude AI systems contributing to research and development under human supervision
Table of Contents (15 sections)
Anthropic says its Claude artificial intelligence models are taking a substantially larger role in developing the company's future AI systems, with Claude now “leading” 26% of Anthropic’s measured AI research and development work.
The figure, released by Anthropic on September 17, is one of the clearest disclosures yet from a major frontier AI company about how extensively artificial intelligence is being used inside the process of building more advanced AI.
But the number requires careful interpretation.
Anthropic is not saying that Claude independently controls 26% of its research department or that a quarter of its future models are being created without humans.
Under Anthropic’s definition, a task reaches the “AI leads” level when Claude can perform most of that task from beginning to end after receiving a high-level instruction, while a human continues to supervise the work. The company explicitly says Claude is not operating fully autonomously in any measured area of its AI R&D.
Key Takeaways
Anthropic says Claude “leads” 26% of its internal AI research and development work as of August 2026.
Recommended Reading
Related Stories & In-Depth Guides
Curated editorial perspectives matching this topic.
The United States and China have opened a new high-level dialogue on artificial intelligence ahead of President Donald Trump’s September 24 meeting with Chinese President Xi Jinping in Washington, with the US proposing a bilateral notification system for serious AI incidents that could threaten nati
Anthropic has launched Claude Opus 5.5 with faster output, lower API pricing, a 1-million-token context window and stronger agentic coding capabilities. Here is its price, access, features and how it compares with Opus 5.
In Anthropic’s measurement system, “leads” means Claude can complete most of a task end-to-end from a high-level prompt while a human remains responsible for supervision.
More than 90% of Anthropic’s measured AI R&D work now involves Claude at least at the “collaborates” level.
Anthropic says Claude is not fully autonomous in any measured category of model R&D.
The company had approximately 30,000 AI agents performing research and engineering work at any one time on its main internal agent platform in August.
The 26% figure comes from Anthropic’s own prototype R&D Automation Index and should not be interpreted as an independently audited measure of the entire AI industry.
Claude Now ‘Leads’ 26% of Anthropic’s Model R&D
Anthropic created what it calls the R&D Automation Index to measure the degree to which Claude participates in the work required to develop future AI models.
As of August 2026, the company reported three particularly important findings: Claude “leads” 26% of its AI R&D work; more than 90% of the work is at or above the level Anthropic calls “AI collaborates”; and none of the measured work has reached full AI autonomy.
The increase has been rapid.
Anthropic’s published chart places Claude’s share of R&D work at the “leads” level at less than 1% in February 2026. By August, six months later, the figure had reached 26%. Reuters separately reported the increase as roughly 1% in March to 26% in August.
That trend is more important than the headline percentage alone because it suggests AI systems are becoming capable of handling increasingly complete research and engineering workflows rather than simply helping employees with isolated coding or writing tasks.
What Does ‘Leads’ Actually Mean?
Anthropic uses an automation scale ranging from AL0 to AL5.
At AL0, there is no AI involvement.
At AL3, which Anthropic calls “collaborates”, Claude can perform large parts of a task but remains under close human direction.
At AL4, or “leads”, Claude can take a high-level instruction and complete most of a task end-to-end, with a human supervising rather than constantly directing every step.
AL5 represents full autonomy, where an AI performs the work without a human in the loop. Anthropic says Claude has not reached that level in any measured category of its AI research and development.
A practical example helps show the difference.
If an internal data pipeline breaks, a collaborating AI might receive logs from an engineer, discuss possible causes with the engineer and perform parts of the debugging while repeatedly returning to the person for decisions.
At the “leads” level, an engineer could instead tell Claude to fix the failed pipeline. Claude could investigate the logs, identify the failure, create and test a repair and deal with unexpected issues largely on its own, while the human remains responsible for oversight.
That is considerably more advanced than conventional autocomplete or chatbot assistance, but it is still different from an AI independently deciding what research should be conducted and carrying it out without human supervision.
More Than 90% of R&D Already Involves Human-AI Collaboration
The broader figure in Anthropic’s disclosure may be just as significant as the 26% headline number.
The company says more than 90% of its measured AI R&D work has reached at least the “collaborates” level.
That means Claude is already deeply involved in most of the model-development workflow Anthropic measured, even when it is not performing enough of a particular task to qualify as “leading” it.
This includes work across areas such as model training, reinforcement learning, infrastructure, evaluation systems and engineering.
The disclosure suggests that the development of frontier AI is increasingly becoming a combined human-and-AI process rather than one in which researchers simply build models using traditional software tools.
How Anthropic Calculated the 26% Figure
The 26% figure is not based on the percentage of employees replaced by AI, the percentage of research papers written by Claude or the percentage of computer code generated by the model.
Anthropic built a catalogue of the tasks involved in its model-development operation.
For each week in July 2026, the company says it randomly sampled 20% of employees from every department involved in the model R&D loop.
A Claude research agent then examined internal records including Slack activity and documentation to identify the tasks those employees worked on.
Across the month, Anthropic produced a list of approximately 15,000 granular R&D tasks. Claude then organised those tasks into a hierarchical system containing 542 nodes, including 378 detailed task categories.
Examples cited by Anthropic include diagnosing defects in evaluation infrastructure, managing reinforcement-learning sandbox network policies and investigating incidents affecting model-serving systems.
The company then assessed how much AI performs within each category and assigned an automation level.
The Measurement Has Important Limitations
Anthropic itself acknowledges that the R&D Automation Index is not a perfect measurement.
One important limitation is that Claude is involved in evaluating how much work Claude performs.
A Claude research agent studies how tasks are performed, and another Claude-based judge assigns automation levels using the collected evidence. Anthropic notes that this could introduce errors or shared biases between the system doing the work and the system evaluating it.
The company compared the AI ratings with judgments from employees responsible for the relevant work areas.
Anthropic says the AI judge and human reviewers were within one automation level of each other 97% of the time, although exact agreement was lower.
The company also acknowledges that borderline distinctions — particularly between “collaborates” and “leads” — can be subjective.
For that reason, the 26% figure is best understood as Anthropic’s structured internal estimate rather than a universally established measurement of AI automation.
Anthropic says independent third-party verification will be important if similar figures are to be compared reliably between competing AI laboratories.
Around 30,000 AI Agents Were Working Internally
Anthropic also disclosed the scale at which it is deploying AI agents inside its own research operation.
As of August, approximately 30,000 agents were carrying out research and engineering work at any given time on the company’s most-used internal agent platform.
These are not 30,000 separate human-like employees.
They are software agents powered by AI models and configured to perform tasks, interact with internal systems and, in some situations, coordinate with other agents.
Anthropic says it gives individual agents persistent identities so their activities can be tracked and audited. Agents also communicate through a shared messaging system designed to make their interactions visible for monitoring.
The scale nevertheless illustrates how frontier AI development is moving toward large fleets of AI agents working alongside human researchers.
Every Agent Action Passes Through Monitoring, Anthropic Says
Greater AI involvement also creates a safety and oversight problem: if increasingly capable agents can execute code, interact with infrastructure and make decisions, developers need mechanisms capable of detecting potentially harmful actions.
Anthropic says 100% of actions made by agents on the covered internal platform pass through an online monitor before execution.
The company analysed more than one billion decisions from its research and engineering agents during August and reported that approximately 0.002% — roughly one decision in every 47,000 — was blocked by the monitoring system.
Anthropic also operates offline monitoring that examines agent activity after it occurs.
The company says roughly 100,000 transcripts per week are flagged for additional automated examination, with around 50 of the highest-priority cases escalated to human review.
The figures do not by themselves prove that monitoring catches every dangerous action. Anthropic explicitly acknowledges that agent monitoring is still a relatively new field and that it cannot be certain existing systems reveal every possible type of problematic behaviour.
Claude Is Already Writing a Large Share of Anthropic’s Code
The new automation index follows earlier Anthropic disclosures showing how rapidly Claude has moved into day-to-day engineering work.
In a separate report, Anthropic said that by May 2026, more than 80% of code merged into its codebase had been authored by Claude.
The company also said the typical Anthropic engineer was merging roughly eight times as much code per day during the second quarter of 2026 as in 2024, with much of the increase attributed to Claude writing code while engineers directed and reviewed the work.
That figure measures code generation rather than overall model R&D automation, so it should not be confused with the new 26% figure.
Writing code is only one part of developing a frontier AI system.
Researchers also have to choose problems, design experiments, interpret results, evaluate safety, determine training strategies and make decisions about which research directions are worth pursuing.
Anthropic says these higher-level judgment and direction-setting tasks remain areas where humans currently hold an important advantage.
Why Anthropic Is Publishing These Numbers
The company says it wants the public and governments to have better visibility into how quickly AI is becoming involved in creating more powerful AI.
The issue is closely connected to what researchers call recursive self-improvement — the possibility that an advanced AI system could eventually contribute to, or ultimately autonomously develop, a more capable successor.
Anthropic says the current Claude systems have not reached that stage.
However, the company argues that measuring AI’s contribution to AI development is important because accelerating the development process could eventually make technological progress harder for humans to understand or control.
Anthropic plans to continue publishing its automation, oversight and compute-allocation metrics and has encouraged other frontier AI companies to disclose comparable information.
Anthropic Also Disclosed How Research Compute Is Used
The September report includes another measurement intended to provide insight into Anthropic’s internal priorities.
For a sample week from July 13 to July 20, Anthropic estimated that approximately 6% of computing resources used for AI R&D went toward work classified primarily as safety research.
For AI-driven AI research specifically, that figure was approximately 12%.
Anthropic describes those estimates as conservative because research that simultaneously improves safety and model capabilities was counted as capabilities work rather than safety work.
It also cautions that computing power is an imperfect way to measure how much attention a laboratory gives to safety because some safety research can require large amounts of human reasoning while consuming relatively little compute.
Does This Mean Claude Is Building Itself?
Only in a limited sense.
Claude is increasingly doing work that contributes to the development of future Claude models and other Anthropic AI systems. It writes code, debugs infrastructure, runs experiments and handles some research workflows with significantly less human intervention than earlier systems could manage.
But describing the current system as fully “building itself” would go beyond Anthropic’s evidence.
Humans still set research priorities, supervise AI-led tasks, review results and remain responsible for important decisions.
Anthropic explicitly states that Claude is not fully autonomous in any measured area of the company's AI R&D.
The more accurate description is that AI is becoming an increasingly important participant in the process of building the next generation of AI.
Why the 26% Figure Matters
The significance of Anthropic’s disclosure lies less in one percentage and more in the speed of change.
A task that once required engineers to continuously write code, diagnose failures and execute experiments can increasingly be handed to an AI system as a broader objective.
If that trend continues, human researchers may spend progressively more time selecting problems, reviewing results and deciding research direction while AI agents perform much of the implementation and experimentation.
Anthropic itself says human review is already becoming a bottleneck in some engineering workflows because AI can produce code faster than people can inspect it.
Whether the current trajectory continues is uncertain.
The company acknowledges that AI capability growth could slow, infrastructure or computing constraints could become limiting factors, or higher-level research judgment may prove substantially harder to automate than coding and experimentation.
Bottom Line
The central claim is supported, with an important qualification.
Anthropic reported on September 17 that, as of August 2026, Claude “leads” 26% of its measured AI research and development work. Under the company’s methodology, that means Claude can complete most of the work involved in those tasks from a high-level prompt while humans supervise.
More than 90% of Anthropic’s measured model R&D now involves Claude at least in a substantial collaborative role.
But Claude is not yet fully autonomous in any measured part of that process.
The figures come from Anthropic’s own prototype R&D Automation Index, whose methodology relies partly on Claude to analyse and classify internal work. Anthropic itself says third-party verification and common industry measurement standards would be needed for stronger comparisons between AI laboratories.
The disclosure therefore does not show an AI independently creating its successor.
It does show that AI systems are becoming increasingly embedded in the engineering and research process used to build the next generation of frontier models — and that the amount of work they can perform with limited human direction is rising rapidly.
Key Takeaway
Claude leads 26% of Anthropic’s measured AI R&D work under human supervision.
More than 90% of work involves substantial AI collaboration.
No measured category has reached full AI autonomy.
Figures come from Anthropic’s internal R&D Automation Index.
The Rajatheertha Team publishes news, explainers, guides and updates across India and the world. Our coverage follows Rajatheertha's editorial, verification and corrections standards.
Anthropic is reportedly seeking shareholder approval for a new share structure giving its seven co-founders 50.1% collective voting power ahead of a potential IPO. Here is how the proposal would work and what it means for investors.
Google has confirmed that one of its Gemini artificial-intelligence models gained unauthorized access to systems belonging to three real companies during a cybersecurity evaluation. The incidents occurred after a testing environment unexpectedly allowed internet access, leading Gemini to mistake rea
Anthropic CEO Dario Amodei has called for slowing the rate of frontier AI capability improvements so safety work can keep pace, as researchers raise severe risk warnings and policymakers consider stronger independent oversight.
US President Donald Trump says his administration will create a new “AI Force” modeled in part on the Space Force and will soon appoint an artificial intelligence czar, placing AI policy more firmly at the center of his administration’s technology and economic agenda.
Donald Trump has ordered federal agencies to use “Super Intelligence” instead of “Artificial Intelligence” while leading AI companies signed a voluntary safety accord covering internal controls, external audits and board oversight.
Nvidia has agreed to acquire Hugging Face for about $12.93 billion, significantly expanding its AI software presence while promising to keep the developer platform open across chips and clouds.
0 Comments