Skip to main content
Back to Blog
AI Agents
Governance
Enterprise AI
Risk Management
Leadership
Implementation

Three Metrics Separate Agent Programs From Agent Theater.

Jason Oglesby

By Jason Oglesby · September 22, 2026

Anthropic published a scorecard for agent oversight this month and then filled in its own row.

Three numbers. Coverage, review latency, escalation rate.

Go try to fill in yours.

Disclosures. I pay Anthropic, and I pay OpenAI, xAI and Perplexity, so the company I am about to praise takes my money and so do three of its competitors. Ergon sells technology leadership work that includes exactly this kind of operating discipline, so a post about whether you can measure your agents is a post that sells what I do. And Anthropic benefits from being the lab that publishes its homework. The numbers below are self-reported and they say so themselves.

The Scorecard Nobody Asked For

On September 17 the Anthropic Institute published a piece by Marina Favaro and Phillie Wright proposing three measurements for what is actually happening inside an AI development organization. Three of them apply directly to anyone running agents in production.

Coverage. The share of an agent's actions that pass through a monitor, before execution or after it.

Review latency. The time between an action and its review, first by an automated monitor and then by a human.

Escalation rate. The share of agent activity that gets blocked, redirected, or flagged for further review.

That is the whole framework. It is not complicated, which is the point. It is three questions any operator should be able to answer about a system they are accountable for.

Then they published their own answers, which is the part most vendors skip.

What They Reported

Roughly 30,000 agents running at once on their most-used internal platform as of August. Online monitors covering 100 percent of agent actions before execution. Offline monitors covering 100 percent after.

Blocking rate of 0.002 percent, about one decision in 47,000, across more than a billion decisions analyzed in August. Humans review blocked actions within a week.

On the offline side, one to two transcripts per thousand get flagged. That works out to roughly 100,000 flagged weekly, with about 50 escalated to a human.

They also state the scope plainly: this covers their most-used internal platform, not everything they run. That caveat is worth more than the numbers. A vendor that tells you what its measurement does not cover is a vendor whose measurement you can use.

One figure from the same piece will get quoted all week by people selling you something. Anthropic reports that Claude leads 26 percent of their AI research and development work, with more than 90 percent at collaborates or above. Treat that as a fact about one company with unusual tooling, unusual talent, and an unusual incentive to push the number up. The measurement framework travels. The percentage does not.

Coverage Is the One That Hurts

Most companies running agents today cannot answer the first question.

Not because their monitoring is bad. Because nobody knows the denominator. You cannot report the share of agent actions that pass through a monitor if you do not know how many agent actions exist.

This is the same problem shadow IT was twenty years ago, arriving faster and with a bigger blast radius. Somebody in finance wired a workflow to an API key. Somebody in marketing has an agent posting to a system of record. Neither is in a diagram. Neither shows up in a count.

Coverage is not a monitoring question. It is an inventory question wearing a monitoring costume. If you want to know your number, start by finding out what is running.

The Funnel Is the Actual Lesson

Here is the part I would tape to a wall.

About 100,000 transcripts flagged in a week. About 50 reaching a person.

That ratio is not a failure. It is what oversight looks like at a scale where human review is arithmetically impossible. Fifty a week is roughly ten a day, which is a job somebody can actually do. A hundred thousand a week is not a job, it is a wish.

Every company adding agents is walking toward this ratio whether or not they have thought about it. The choice is not between machine review and human review. It is between designing the funnel deliberately and discovering later that your humans quietly stopped reading the queue.

And note what review latency does here. Blocked actions get human eyes within a week. Not within an hour. A week is a real answer from a company that measures itself, and it should recalibrate anyone who thinks their own untimed process is fine.

What I'd Do This Week

Count your agents before you measure them. Every automation with credentials that takes an action on your behalf. If the list takes more than an afternoon, that is your finding.

Pick a denominator and publish it internally. Even a bad coverage number that everyone agrees on beats an excellent one nobody can reproduce.

Time your review, do not describe it. Latency is a number. If your answer is a paragraph about process, you do not have the number.

Design the funnel before the volume arrives. Decide now what fraction reaches a person and what happens to the rest. That decision gets made either way. The only question is whether you make it.

The Part That Matters

A company with 30,000 agents published three numbers about how it watches them, said out loud what the numbers do not cover, and gave away the framework.

The framework costs nothing. Filling it in is the whole job.

Most agent programs are not going to fail on capability. They are going to fail on the fact that nobody could say how many there were.

Three Metrics Separate Agent Programs From Agent Theater | Ergon Insights