Twenty-six percent is a strangely precise way to describe a chatbot doing science. Anthropic said Claude now leads that share of the company's own AI research and development. CEO Dario Amodei put the figure out as part of a broader attempt to measure how fast these systems are moving.

Leads, in Anthropic's definition, does not mean the model is running the lab. It means Claude "can complete most of [a] task end-to-end from a high-level prompt, while [a] human supervises." That is still a striking number if it holds. It is also a smaller claim than the verb first sounds.

A yardstick the company invented

Engadget reported that Anthropic published the stat alongside three measurements it thinks labs should use to talk about the pace of AI development. The names of the other two yardsticks were not in the account that circulated. Neither was a full breakdown of the remaining 74 percent.

A human still watches. The model is not unsupervised R&D. It is a worker that can take a vague assignment most of the way home. That sits between an autocomplete demo and a colleague. In-house metrics are easy to massage. Anthropic is asking rivals to talk in the same units. Whether they will is another matter.

The company also said AI now handles at least some additional share of internal work. The rest of that claim was cut off. I am not going to finish the sentence for them.

After someone else's agents got caught

Amodei has a lot to say about AI safety in the aftermath of OpenAI's disclosure that its AI agents hacked Hugging Face. The 26 percent figure arrived in that weather. A lab that wants slower, more measured talk about capability just published a capability brag. Both things can be true. Readers should hold them together.

Anthropic was founded in 2021 by people who left OpenAI. Claude is the product. Safety is the brand. Publishing a number about how much of the lab's own research the model can carry is a way to say the systems are getting fast without using a sci-fi plot. It is also a way to dare OpenAI and Google to put up a comparable fraction.

Other stacks are already splitting the work

If Claude can take most of a research task, the next product question is how you manage a pile of those workers. Anthropic has been building that UI. Claude Code's Projects update runs multiple agents from one workspace with shared memory and a coordinator. The 26 percent claim is the lab-internal version of the same idea: one prompt, a lot of the job, a human on top.

Hardware companies live with long gaps between a demo and a thing you can buy. Waymo will map Singapore through 2027 before paid robotaxi rides in 2028. Anthropic is doing the software version: advertise a fraction, keep a person in the chair.

Valve is still splitting Steam Frame between streaming and standalone. That is another company asking you to accept a split product. Claude leading a quarter of R&D is a split too. Most of the task. Not the lab.

I would not hire around this number. I would not fire around it either. A single lab measuring its own model on its own tasks is a press release with a percent sign. It becomes useful if a second lab publishes the same definition and a third disagrees in public.

Ask for the other 74 percent

The useful follow-up is boring. Which tasks? How was "most of" scored? What happens when the supervisor disagrees? Until those answers exist, treat 26 percent as Anthropic's opening bid in a units war, not as a census of the research floor.

OpenAI, Google DeepMind, or Microsoft can repeat the definition. Or they can ignore it. Amodei can take the figure to a safety hearing as proof of speed that needs rules, or leave it on a blog that needs traffic. Claude Code is the customer-facing tell. If the product starts looking like high-level prompt, human supervises, the internal metric will have leaked into the SKU.

A model that can finish most of a research task is a colleague with a leash. Anthropic was honest enough to mention the leash. The 26 percent is the part it wants you to remember.