Anthropic says 26 percent of the work involved in developing its future models now reaches a level where Claude can complete an assigned task without asking for help. That share was below one percent in February, while more than 90 percent of work now involves at least substantial human-AI collaboration. None qualifies as fully autonomous.

The company uses a scale from Epoch AI. Its “AI leads” level means a person provides the goal, Claude works through obstacles and produces a result, and a human still reviews and decides whether to use it. One example is fixing and testing a reported software bug without permission to ship the change. The metric weights activities by employee time, so it does not show who sets research direction or how many important decisions the model makes.

Claude also assigned the official automation scores after agents gathered evidence from internal documents. The model matched employee judgments 59 percent of the time; two humans assessing the same work agreed on the exact level only about one-third of the time. Anthropic additionally reports roughly 30,000 concurrent agents on its main internal platform and says six percent of research compute in one July week went to safety work. The disclosures offer unusual operational detail, but their definitions, self-grading and limited monitoring history make them indicators rather than precise measures of autonomous AI research.