Run status classified: success / flawed / blocked / manual
Every agent run terminates in one of four classified states - success, flawed, blocked or manual - only success ships, each of the other three routes to a different fix, and the classification rate is what justifies expanding the automation boundary.
- Test-oracle reliability is measured and tracked on a dashboard
- Every agent run is recorded with a terminal state (success / flawed / blocked / manual), and the distribution is reviewed
- Each non-success state has a named owner and routes to a different remedy
- Agent Autonomy Score (% of tasks completed without human intervention) is measured and broken down by task type
- Metrics trigger automated alerts when thresholds are breached (e.g., test-oracle reliability drops)
- Oracle-reliability dashboard (e.g., TORS) with per-service breakdown
- Run-state distribution report showing how the non-success classes moved between periods
- Merge queue wait time chart showing sub-10-minute target
- Delivery L3 (Metrics) - ITS, CPI, and CI feedback latency must be operational
- Development L4 (Code Review & Quality) - auto-approve workflow must exist for auto-approve rate tracking
What It Is
Every agent run ends somewhere, and at L4 that ending is a classified state rather than an impression. Vercel's software factory write-up for the AI SDK, published 2026-08-12, named the four states the rest of the industry promptly adopted: success, flawed, blocked, manual. Only success ships. The other three are not failures to be swept up, they are the feedback that tells you where the automation boundary currently sits - and each of them routes to a different fix.
- Flawed - the run completed and produced something wrong. The fix is upstream, in prompts and evals: the agent had what it needed and still got it wrong, so the specification or the checks around it are inadequate.
- Blocked - the run could not complete because the environment would not let it: a missing credential, an unprovisioned service, a dependency it could not install. The fix is to provision the environment, not to rewrite the prompt. Teams that respond to blocked runs with prompt engineering burn weeks.
- Manual - the run reached a point that requires human judgment. The fix is to re-examine the boundary: is this a class of decision that should always be human (problem selection, architecture, the quality bar), or is it a gap that automation could reasonably close?
The value of the taxonomy is exactly that it forbids the single bucket. A team with one "did not merge" pile learns nothing from it, because prompts, infrastructure and scope all land in the same place and the remediation effort scatters. A team with four states can read its own weekly distribution and know which of three quite different investments to make. Vercel's own reported outcome for July was 25-35% of weekly merged PRs, more than 75% of closed issues, and open issues falling from 1,022 to 844, with over half of v6 weekly merges being backports - self-reported by an interested party, so treat the numbers as a shape rather than a benchmark.
Addy Osmani picked up the same four states on 2026-08-21 and added the frame that makes them a management tool rather than a logging convention: a verification budget, spent on cheap early checks in preference to heavy late ones, and five points where human judgment relocates rather than disappears - problem selection, architecture, the quality bar, deciding which signals to trust, and shipping authority. His line is the one to put above the dashboard: "Code authorship may shift; responsibility cannot."
The auto-approve percentage this guide used to lead with is still worth watching, but it is a summary statistic over the success bucket and it hides everything interesting. Two teams can both report 60% auto-approved while one has 35% blocked runs (an infrastructure problem, cheap to fix) and the other has 35% flawed runs (a quality problem, expensive to fix). The classification rate - what fraction of runs terminate in a state you can name and act on - is what justifies expanding the automation boundary. An unclassified run is not evidence of anything.
Auto-approval itself sounds radical to teams used to mandatory code review, and the reaction is often "but someone needs to look at every change." The response to that reaction is: not every change carries equal risk, and treating every PR as high-risk is a scaling bottleneck that becomes untenable at L4. A PR that writes a new unit test, updates a comment, fixes a typo in a log message, or generates documentation from source code annotations does not require the same scrutiny as a PR that changes authentication logic or modifies a payment processing flow. Auto-approve is about routing correctly: low-risk, well-tested changes merge automatically; high-risk or structurally significant changes get human attention.
A target above 60% is meaningful because it's the threshold at which agent throughput begins to outpace human review capacity. Teams running 3-5 parallel agents per developer can produce 20-50 PRs per week per developer. If every PR requires human review, the review queue quickly exceeds the team's review capacity, creating a bottleneck that negates the throughput gains from parallel agents. At 60% auto-approve, the human review burden is reduced to a manageable level and reviewers can focus their attention on the 40% of PRs that actually need it.
Auto-approve rate is a lagging indicator of the entire L4 metrics system working correctly. It requires: high TORS (otherwise the automated CI gates produce false signals), low ITS (otherwise PRs arrive at the gate with lingering quality problems), a well-designed policy-based merge rules system, and mature agent workflows that produce clean, well-tested code. A team that has 95% TORS, median ITS of 1.5, and good CI pipelines will naturally achieve 60%+ auto-approve rate as they build out the policy rules. The metrics are mutually reinforcing.
Why It Matters
- Eliminates the review bottleneck at agent scale - at L4, the review queue becomes the primary bottleneck to delivery; auto-approve for the 60% of PRs that are genuinely low-risk removes the bottleneck and lets reviewers focus on what matters
- Measures algorithmic trust in the delivery pipeline - the four-state distribution is the readout on whether the whole L4 system is working; if CI is reliable, agents are high quality, and policies are well-designed, the success share rises on its own and the other three shares tell you which investment moved it
- Accelerates agent feedback loops - agents that don't have to wait for human review can complete full delivery cycles (write code, test, merge, deploy, observe) autonomously; this enables more sophisticated agent learning and optimization
- Forces quality gate investment - achieving 60% auto-approve requires investing in the quality gates that make algorithmic trust possible; TORS, CI reliability, security scanning, coverage thresholds, and lint rules all must work correctly; the auto-approve target creates organizational pressure for this infrastructure investment
- Each non-success state has its own cheapest fix - flawed runs are answered with prompts and evals, blocked runs with environment provisioning, manual runs by revisiting the boundary; a single failure bucket sends all three to whoever is on rota
- Classification is what licenses expansion - you can only widen the automation boundary safely if you can say what happened on the runs that did not ship; a high success rate over a pile of unclassified runs is a claim, not a measurement
- Reduces context-switching for human reviewers - humans who review only the 40% of high-risk PRs are reviewing the genuinely important changes, not the mechanical ones; this makes review a higher-value activity and reduces reviewer burnout
Getting Started
- Make every agent run terminate in one of the four states - success, flawed, blocked, manual. Emit the classification from the harness, not from a human reading logs afterwards, and store it with the run. Any run that ends unclassified is a gap in the instrumentation and should be visible as its own count until it is fixed.
- Route each state to a different owner and a different fix - flawed goes to whoever owns prompts and evals; blocked goes to the platform or environment owner; manual goes to whoever owns the automation boundary. Writing these three routes down is most of the work, because it forces the admission that they are three different problems.
- Audit what's blocking auto-approval today - For every PR that required human review in the last month, categorize why: was it a CI failure, a policy violation, a coverage drop, a security scan finding, or a developer judgment call? The distribution of reasons tells you where to invest, and it should reconcile with your four-state distribution.
- Define auto-approve eligibility criteria - Write the explicit criteria for PR types that are eligible for auto-approval: PR size below 200 lines, no changes to security-sensitive paths, no changes to payment processing code, CI passing with zero flaky re-runs, test coverage delta neutral or positive. Document these criteria as policy files in the repository.
- Implement policy-based merge rules - Use GitHub's merge queue, Mergify, Trunk, or equivalent tooling to encode the eligibility criteria as automated rules. A PR that meets all criteria is automatically merged when CI passes. A PR that fails any criterion goes to the human review queue.
- Start with a low-risk cohort - Don't attempt 60% auto-approve immediately. Start with the lowest-risk PR types: documentation updates, test additions, dependency version bumps. Set auto-approve rules for these cohorts and measure the auto-approve rate. Start at 20-30%, verify that the merge quality is maintained, and expand the criteria gradually.
- Track the four-state distribution weekly, not a single pass rate - Report the shape: "This week, 55% success, 18% flawed, 19% blocked, 8% manual." A rising blocked share is a platform ticket. A rising flawed share is an evals problem. A rising manual share is a conversation about where the boundary belongs. A single auto-approve percentage compresses all three signals into one number and destroys them.
- Monitor post-merge defect rates for auto-approved PRs - Track production incidents and test failures separately for auto-approved vs. human-reviewed PRs. If auto-approved PRs have significantly higher defect rates, the eligibility criteria need to be tightened. If they have similar defect rates, the criteria are well-calibrated. This comparison is the quality assurance mechanism for the entire auto-approve system.
Spend the verification budget early rather than late. Cheap checks that run before an agent starts work - a provisioned environment, a resolvable dependency graph, an eval on the task specification - convert what would have been blocked and flawed runs into successes at a fraction of the cost of catching the same problems in review. The path to 60% auto-approve often has a step function pattern: building the basic policy system gets you to 30-40%, fixing TORS from 85% to 95% jumps you to 50-55%, and adding smart security scanning rules gets you past 60%. Track which improvements drive the biggest auto-approve rate increase rather than optimizing everything simultaneously.
Common Pitfalls
Setting auto-approve rules without security review. Auto-merge is a significant security decision. Changes to authentication, authorization, secrets management, or infrastructure-as-code should never be auto-approved without explicit security review, regardless of CI passing. The auto-approve policy must have explicit security-sensitive path exclusions that are reviewed and maintained by the security team.
Using auto-approve rate as a goal rather than an outcome. If the team targets 60% auto-approve rate as a goal, they may achieve it by relaxing quality gates - reducing coverage thresholds, ignoring lint warnings, or weakening security scans. This hits the number while degrading quality. Auto-approve rate should only be tracked as an outcome of good quality gates, never as a target that justifies weakening those gates.
Anchoring auto-approve thresholds in benchmark or validation scores instead of post-merge outcomes. SpecBench (May 20, 2026 - arXiv 2605.21384) showed the gap between validation reward and held-out reward grows roughly 27 percentage points per 10x increase in lines of code, which means reward hacking scales with change size. A model that scores well on a benchmark or on its own validation set can still degrade on the larger, messier PRs that auto-approve is meant to wave through. Anchor auto-approve thresholds only in post-merge outcome metrics - post-merge bug rate, review-overturn rate, and production incident rate broken down by Green PR cohort - never in benchmark or validation scores.
Collapsing the four states back into "failed". This happens quietly, usually in a dashboard rewrite, and it is the single most common way teams lose the value of the taxonomy. Once flawed, blocked and manual share a bucket, the remediation effort spreads evenly across three problems with wildly different costs, and the team concludes that agent reliability is a mystery. It is not; the distinction was thrown away.
Answering blocked runs with prompt engineering. A run that failed because a service was not provisioned or a credential was missing will not be fixed by a better prompt, and teams that do not separate the state will spend weeks proving it. Blocked is an infrastructure signal. It is also usually the cheapest of the three to fix, which is why the split pays for itself quickly.
Treating "manual" as a defect. Some decisions belong to humans permanently. Osmani's five relocation points - problem selection, architecture, the quality bar, which signals to trust, shipping authority - are not automation gaps to be closed but the places judgment moved to. A manual classification on one of those is the system working. Drive the manual rate down only where the classification reveals a boundary that was drawn out of caution rather than out of principle.
Expanding the boundary on the strength of a good week. The classification rate is what licenses expansion, and it needs enough runs behind it to mean something. Widening auto-approve eligibility after one clean week is how teams discover that their success rate was measuring the tasks they already trusted agents with.
Not distinguishing agent PRs from human PRs. At L4, most PRs are agent-authored. Auto-approve rules designed for human PRs may not fit agent PRs correctly. Agent PRs tend to be larger (agents are verbose), may touch more files, and have different commit message patterns. Review your auto-approve criteria to ensure they're calibrated for the actual PR population, which is now majority agent-authored.
Abandoning human review too quickly. Moving from 100% human review to 60% auto-approve is a gradual process that should take 3-6 months. Teams that rush the transition skip the validation step: verifying that auto-approved PRs maintain quality. Run each new auto-approve rule for 30 days before expanding it, and check defect rates at each stage. Gradual expansion with validation is the responsible path.
Not reviewing the auto-approve policy as the codebase evolves. An auto-approve policy that was well-calibrated six months ago may need updating after a major refactor, a new service addition, or a change in the security posture. Assign an owner to review and update the auto-approve policy quarterly. Policies that grow stale either become too permissive (merging things they shouldn't) or too restrictive (requiring review for things that don't need it).
How Different Roles See It
Bob has deployed a merge queue with basic automated gates but his auto-approve rate is stuck at 25%. Most PRs are requiring human review because CI failures are being classified as needing investigation rather than being handled automatically. His team is spending too much time in review queue management.
What Bob should do: Bob should trace why CI failures are routing to human review rather than back to the agent. The likely root cause is either TORS (flaky tests look like real failures, requiring human judgment) or ITS (agents have already iterated 4+ times and the system flags the PR for review rather than another agent iteration). Before any of that, Bob should classify. If his pipeline cannot say whether a run ended flawed, blocked or manual, the diagnosis below is a hypothesis rather than a finding, and the fixes for the three states are not interchangeable. A week of four-state data usually settles the question outright. Assuming it confirms the reliability theory, Bob should address the TORS problem first: improve test reliability to 95%+, then retune the auto-approve rules to route flagging CI failures back to the agent rather than to a human reviewer. Every percentage point improvement in TORS should translate to measurable improvement in auto-approve rate. Bob should track this correlation explicitly and use it to justify the continued TORS investment.
Sarah is designing the L4 metrics dashboard and wants to show auto-approve rate alongside the metrics that drive it (TORS, ITS, CPI). She wants the dashboard to tell a coherent story rather than just showing numbers.
What Sarah should do: Sarah should put the four-state run distribution at the top of the dashboard, as a stacked bar over time rather than a single rate, because that one chart tells three different teams whether they have work to do. Below it she should build a "delivery pipeline health" cascade: TORS feeds into ITS quality, ITS quality feeds into the success share, and the success share feeds into review queue depth. The visual layout should show the cascade clearly: if TORS drops, ITS worsens, auto-approve rate falls, and review queue depth grows. This cascade view makes the system dynamics visible and helps the team understand which leading metrics to fix when the lagging metric (auto-approve rate) degrades. Sarah should present this dashboard in the monthly engineering review and use it to guide investment decisions: "TORS dropped 3 points this month - that explains the auto-approve rate drop - let's find the new flaky tests."
Victor has achieved 75% auto-approve rate for his agent workflows by carefully designing the task types he assigns to agents. He only sends agents on tasks where the eligibility criteria are almost certain to be met: the changes are bounded, the test coverage is predictable, and the security-sensitive paths are not touched.
What Victor should do: Victor should formalize his task type classification system as a shared playbook. The classification has three tiers: green (auto-approve eligible, assign directly to agent), yellow (auto-approve sometimes, agent with human review scheduled), and red (always human review, agent can assist but human drives). Victor should document the criteria for each tier and the examples from the team's actual codebase. He should also reconcile the tiers against the four-state data rather than against his intuition: a green-tier task type that produces a steady stream of blocked runs is an environment problem masquerading as a classification mistake, and a red-tier type whose manual classifications all turn out to be routine is a boundary drawn out of caution. This playbook turns auto-approve rate improvement from an infrastructure problem into a workflow design problem: if developers learn to classify their tasks correctly and assign only green-tier tasks to fully autonomous agents, the team's auto-approve rate naturally rises toward the 60% target as more green-tier tasks are assigned.
Further Reading
From the Field
Recent releases, projects, and discussions relevant to this maturity level.
How This Guide Changed
What each edition changed in this guide, newest first.
- V1.6September 2026LATEST
This item was rebuilt around Vercel's four-state run-status taxonomy - success, flawed, blocked, manual - published on 2026-08-12 and adopted within a fortnight by Addy Osmani and much of the practitioner conversation. The change is more than vocabulary: a single auto-approve percentage compresses three quite different problems into one number, whereas the four states route separately to prompts and evals, to environment provisioning, and to a conversation about where the automation boundary belongs. The guide now leads with the classification rather than the rate, adds Osmani's verification budget as the reason to spend checks early rather than late, and carries his five relocation points as the argument that a "manual" classification is often the system working rather than a gap to close. Vercel's own numbers are self-reported and are presented here as a shape, not a benchmark.
- V1.3June 2026
Added the reason auto-approve thresholds must never be derived from benchmark or validation scores. SpecBench (arXiv 2605.21384, May 20) found the gap between validation reward and held-out reward grows roughly 27 percentage points per 10x increase in lines of code, so reward hacking scales with change size, and a model that scores well on its own validation set can still degrade on the larger, messier PRs that auto-approve is meant to wave through. Anchor thresholds in post-merge bug rate, review-overturn rate and production incident rate by Green cohort instead.
- V1.0March 2026
In the first edition this item was largely the number in its title: keep auto-approve above sixty percent and human review stops being the throughput ceiling for a team running several agents each. The routing argument was there from the start - a new unit test and a change to authentication logic do not warrant the same scrutiny - but the classification that makes the percentage readable arrived much later.
Where does your team actually sit on this?
This guide describes one level of one area. Run the assessment to place your team across all 16 areas, see which gates you have passed, and get a report you can take to your stakeholders.