The Wire
BusinessTechnologyArtificial IntelligencePolitics & Policy

Anthropic publishes metrics to track AI progress inside labs

Anthropic publishes metrics to track AI progress inside labs
Photo: cnbc.com

Anthropic published three measurements for tracking AI development inside frontier labs.

Why it matters: The metrics could make AI progress easier to compare across companies and easier for regulators to scrutinize. They also give customers and rivals a more concrete way to judge how quickly models are helping build better models.

  • Anthropic said on Sept. 17, 2026, that it is tracking AI-led R&D, oversight of AI agents and compute allocation.
  • The company said its prototype R&D Automation Index estimates how much of its research and development is performed by Claude.
  • AP reported Claude led 26% of Anthropic’s model R&D work, up from none in February 2026, and collaborated on more than 90% of R&D overall.
  • Anthropic said its most-used internal platform had about 30,000 agents doing research and engineering work as of August 2026.

Anthropic is trying to turn a fuzzy question into a measurable one: how quickly are frontier models helping build better models?

In a blog post published Sept. 17, 2026, the company said it is sharing three metrics for tracking AI development inside frontier labs - how much AI research and development is done by AI itself, how AI agents are overseen, and how compute is allocated.

Anthropic said it built a prototype R&D Automation Index to estimate how much of its research and development is performed by Claude. In reporting on the post, AP said Claude led 26% of Anthropic’s model R&D work, up from none in February 2026, and collaborated on more than 90% of R&D overall.

The company also described a framework for overseeing AI agents - software systems that can take actions on their own, such as helping with research or engineering tasks. Anthropic said it tracks coverage, review latency and escalation rate. As of August 2026, it said roughly 30,000 agents were doing research and engineering work on its most-used internal platform.

Anthropic said 100% of those agents’ actions passed through an online monitor before execution and 100% were reviewed by an offline monitor after the fact. The offline monitor flagged roughly 100,000 transcripts a week, with about 50 of the highest-priority flags escalated to human review.

The company also said that, over the examined week, about 6% of compute used for AI R&D and about 12% of compute used for AI-driven AI R&D was allocated toward safety work.

The broader implication goes beyond Anthropic’s safety framing. If labs start publishing common metrics, it could create a benchmark for investors, enterprise buyers and policymakers trying to compare how fast competing systems are advancing and how much human oversight remains in the loop. AP said Anthropic is urging other AI developers to publish similar metrics regularly using a public methodology so results can be compared over time and across labs.

Anthropic said the effort fits its warning that systems helping accelerate their own development could "make it more challenging for humans to understand or control these systems." The company also said, "we should do everything possible to minimize the gap between what frontier labs know and what the public knows," and called for "better measuring the development of AI, reporting on it publicly, and giving society an opportunity to decide how to use this information."

By the numbers

  • 26% - AP reported Claude led that share of Anthropic’s model R&D work, up from none in February 2026.
  • 30,000 - Approximate number of agents doing research and engineering work on Anthropic’s most-used internal platform as of August 2026.
  • 100,000 - Approximate number of transcripts the offline monitor flagged each week.

Yes, but: Anthropic’s figures come from its own internal methodology, so the metrics may be hard to compare unless other labs adopt similar reporting standards.

What's next: Anthropic said it wants other AI developers to publish similar metrics regularly using a public methodology.

Based on reporting from

  • CNBC

See how this story touches your network - open The Wire in Jane.

Open in Jane