PR throughput per dev

PRs merged per developer per week is a crude proxy, since PRs vary wildly in size, but it is the first output signal available without real instrumentation - useful as a diagnostic, ruinous as a target, and far better than suggestion acceptance rate.

L2 · DELEGATEDWhat this level takes
MUSTNot met, not at this level
  • A delivery-performance baseline (throughput, lead time, change failure rate, restore time - DORA, SPACE or an equivalent set) is on a dashboard the team can open
  • AI tool license count vs. active usage rate is measured
  • PR throughput per developer is tracked
SHOULDExpected in practice, not required
  • Cost per merged PR is measured per tool (acceptance rate is not used as a quality signal)
  • Metrics are reviewed in team retrospectives at least monthly
EVIDENCEHow you would check
  • Delivery-performance dashboard with current data
  • License utilization report (licenses purchased vs. active users)
  • PR throughput chart showing per-developer breakdown

What It Is

PR throughput per developer is the first meaningful per-developer productivity metric for AI-assisted development: how many pull requests does a developer merge per week, on average? At L2 (Delegated), this becomes the primary signal for tracking whether AI tools are changing developer output. It's not a perfect metric - PRs vary enormously in size and complexity - but it's the most directly measurable output signal available without sophisticated instrumentation.

Before AI tools, typical developer throughput in a well-functioning team is 3-6 PRs per week, depending on the codebase, review culture, and task complexity. With active AI tool usage at L2, teams consistently see this number move to 6-10 PRs per week for high-usage developers. At L3 and L4 with agent-heavy workflows, throughput can reach 15-30 PRs per week per developer. These numbers aren't universal - they depend heavily on PR size convention, codebase maturity, and CI speed - but the direction and magnitude of the shift is consistent across teams that have tracked it carefully.

PR throughput is the gateway metric to a richer productivity picture. By itself, it tells you output rate but not output quality. A developer running agents that produce 20 PRs per week of mediocre, heavily-edited code might have lower net output than a developer producing 8 carefully crafted PRs that merge cleanly. This is why PR throughput is always paired with PR cycle time (how long from PR open to merge?), review burden (how much effort do reviewers spend on each PR?), and defect rate (how often does the code break in production?). Together, these four metrics give a much more complete picture.

Two boundaries around this metric matter more than the metric itself. The first: it is a proxy, never a target. Everything in the pitfalls section about splitting PRs and merging half-finished work is what happens the moment throughput appears on a scorecard, and it happens fast.

The second boundary is about what you use instead when someone asks for a simpler number. The obvious candidate - suggestion acceptance rate, which every AI tool vendor reports by default - is a vanity metric, and now there is direct evidence. The DECODE study published 2026-07-27 instrumented 53.6K real in-IDE edits from more than 1,000 developers and found that 31% of accepted AI completions are deleted entirely, most of them edited within 15 minutes of acceptance. Acceptance is a click, not an outcome. An acceptance-rate dashboard reports roughly a third of its volume on code that no longer exists by lunchtime.

The reason PR throughput is the right starting metric at L2 is that it requires no new instrumentation. GitHub, GitLab, and Bitbucket all expose this data directly. A simple query or dashboard plugin will give you per-developer weekly PR counts going back as far as your git history. This makes it possible to compute a pre-AI baseline from historical data even if you didn't instrument anything before AI adoption began.

Why It Matters

  • Directly observable - PR throughput requires zero new tooling to track; it's the most accessible productivity signal available and can be computed retroactively from git history
  • Reveals the AI impact immediately - teams that add AI tools and then measure PR throughput almost always see an inflection point at the month when adoption ramped up; this is the clearest data point available for AI ROI conversations
  • Per-developer granularity - reporting average throughput across the team hides enormous variance; per-developer tracking reveals who has adopted effective AI workflows (high throughput) and who hasn't (unchanged throughput), enabling targeted intervention
  • Creates healthy benchmarks - once the team knows the distribution of PR throughput, the benchmark becomes the high-performer's output; developers can see that 10 PRs/week is achievable for their codebase and are motivated to find the AI workflows that get them there
  • It is the least-bad simple number - the alternatives on offer at L2 are worse: acceptance rate counts code that a third of the time is deleted within 15 minutes, lines of code rewards verbosity, and token spend rewards waste
  • Leading indicator for capacity planning - if developers are producing PRs faster, the bottleneck moves to review, to CI, or to planning; tracking PR throughput early reveals where the next constraint is before it becomes a crisis

Getting Started

  1. Pull historical PR throughput data - Query your version control API for the last 6-12 months of merged PRs, grouped by author and week. This gives you the pre-AI baseline. In GitHub, this is a simple GraphQL query against the repository PRs; LinearB, Jellyfish, and most engineering analytics platforms provide this out of the box.
  2. Normalize for PR size - Raw PR count is a noisy metric if PR size varies widely. Add a second dimension: PR size in lines changed. Track both "PRs per week" and "PR-weeks normalized for size" (e.g., PRs per week weighted by the inverse of lines changed). The size-normalized metric is more stable and more comparable across developers working on different areas of the codebase.
  3. Identify the pre-AI baseline - Find the date when each developer's AI tool usage began and compute their average throughput in the 8 weeks before. This is their personal baseline. Post-adoption throughput is measured against this baseline, not against team averages.
  4. Build a distribution view, not just averages - Don't report average PR throughput; report the distribution. Show the 25th, 50th, and 75th percentile. The distribution reveals the spread: some developers have dramatically higher throughput, others haven't moved. The gap between percentiles is where the adoption story lives.
  5. Track PR cycle time alongside throughput - A developer who pushes 20 PRs per week that each wait 3 days in review is not actually faster than a developer who pushes 8 PRs that merge in 4 hours. Pair throughput with cycle time (PR creation to merge) to get a more accurate productivity picture.
  6. Delete suggestion acceptance rate from every dashboard it appears on - it arrives free from the vendor and it is a vanity metric; DECODE's instrumented data has 31% of accepted completions deleted outright, mostly within 15 minutes. If you need a leading indicator of tool value, use surviving code (accepted completions still present in the file after 24 hours) rather than accepted completions.
  7. Publish the metric as a diagnostic and say so in writing - name the questions it is allowed to answer ("where did throughput not move, and what is different about those developers' setups?") and the ones it is not ("who is performing well?"). Do not set a throughput target for a quarter. A quarterly target on this number reliably produces smaller PRs rather than more delivery, and you will have taught the team to do it.
TIP

GitHub's GraphQL API lets you pull merged PR data with dates and author information with a single query. If you're not a data engineer, ask an AI agent to write the query for you - it's a straightforward task that produces the exact data you need. Prompt: "Write a GitHub GraphQL query to fetch all merged PRs for repository X in the last 6 months, including author login, merged date, and additions/deletions."

Common Pitfalls

Treating PR throughput as an individual performance metric. Displaying per-developer PR throughput on a leaderboard creates perverse incentives: developers will split large PRs into many small ones, merge half-finished work, and focus on maximizing PR count rather than shipping valuable features. Use throughput for team-level analysis and improvement identification, not individual performance evaluation.

Comparing developers on different work types. A developer maintaining a large legacy codebase with complex dependencies will have lower PR throughput than a developer working on a greenfield service, regardless of AI tool usage. Throughput comparisons are only valid within cohorts doing similar work. Cross-cohort comparisons are misleading and demotivating.

Not accounting for PR size conventions. Some teams have a culture of small, atomic PRs (5-50 lines each). Others merge large feature PRs (500-2000 lines). PR throughput numbers mean completely different things in these two cultures. If your team doesn't have a consistent PR size convention, establish one before using throughput as a benchmark.

Ignoring the review burden side. Agents that produce more PRs increase the review burden on human reviewers. If agent-authored PR throughput doubles but review cycle time also doubles (because reviewers are overwhelmed), net delivery speed hasn't improved. Track PR cycle time and review queue depth alongside throughput to see if the review process is keeping up.

Setting a target on it. This is the pitfall that swallows the others. Throughput responds instantly to PR-splitting, and PR-splitting looks identical to healthy work in the data. The only safe use of this metric is comparative and diagnostic - against a developer's own pre-adoption baseline, to find out where to go and ask questions.

Substituting acceptance rate because it is easier to get. Every AI coding tool reports acceptance rate on its admin dashboard, which is why it ends up in so many reports. DECODE's 53.6K instrumented in-IDE edits put a number on the problem: 31% of accepted completions are deleted entirely, most within 15 minutes. Whatever acceptance rate measures, it is not delivered work.

Reading agent-authored throughput as team throughput. LinearB's 2026 benchmarks, drawn from 8.1M pull requests, found elite organisations with AI assistance on 54% of PRs but agents actually authoring only 4.7% of them - and agentic PRs merging at 79% in elite organisations against 37% at the fair tier. The throughput that matters is what merges, and for agent-opened PRs that depends mostly on whether a named human owns the PR.

Using throughput without outcome correlation. A team that ships 30 PRs per week but sees an increase in production incidents has higher throughput with worse outcomes. Throughput is a means, not an end. Always correlate PR throughput with outcome metrics: production defect rate, customer-reported bugs, and deployment success rate. If throughput goes up but quality goes down, something in the agent workflow needs fixing.

How Different Roles See It

BobHEAD OF ENGINEERING

Bob's team has been using AI tools for six months. He wants to show the CTO that the tools are producing results, but his engineering managers are giving him conflicting reports about productivity impact. Some say the team is shipping faster, others say they're not sure.

What Bob should do: Bob should pull the PR throughput data himself. Query GitHub for the last 12 months of merged PRs, compute per-developer weekly averages, and find the inflection point at the month when the majority of the team adopted AI tools. If there's a visible increase in median throughput at that inflection point - even a 20-30% increase - that's the ROI data point Bob needs. He should present this analysis in the next engineering leadership review with the caveat that it's observational but directionally significant. Bob should also ask his engineering managers: "For the developers whose throughput hasn't increased, what's different about their setup or workflow?" The non-improvers are the adoption gap that needs attention.

SarahPRODUCTIVITY LEAD

Sarah has been asked to design a developer productivity scorecard. She's been given several options: PR throughput, lines of code, story points, commit frequency. She needs to pick the right metrics for a fair and useful scorecard.

What Sarah should do: Sarah should push back on the framing before picking metrics. A per-developer scorecard turns every metric on the list into a target, and each of these is gameable in a different direction: throughput rewards PR-splitting, lines of code rewards verbosity, commit frequency rewards noise. Lines of code and acceptance rate should be struck outright - the latter because DECODE found 31% of accepted completions deleted within 15 minutes. If a team-level view is what is actually wanted, Sarah should build one from PR throughput, PR cycle time, and a qualitative self-assessment from monthly developer surveys, reported as a distribution across the team rather than a score per person. Sarah should also add a quarterly calibration step where the team reviews whether the index is capturing the right behaviors. If developers are optimizing for the metric rather than for actual productivity, the index needs adjustment. The goal is a scorecard that helps developers see where they're working well and where they have room to improve, not a leaderboard.

VictorSTAFF ENGINEER - AI CHAMPION

Victor's personal PR throughput has tripled since he adopted parallel agent workflows. He produces 20-25 PRs per week, up from 7-8. He wants to help the rest of the team reach similar levels, but he knows that simply showing his numbers will seem unbelievable or intimidating.

What Victor should do: Victor should run a transparent "productivity audit" of his own workflow and publish it as a technical write-up. The write-up should break down where his PRs come from: what percentage are agent-authored, what percentage are human-written, how much time he spends on review vs. implementation, and what his CI pass rate looks like. The goal is to show that his throughput isn't magic - it's the result of specific, learnable workflow choices that any developer can adopt. Victor should also identify one other developer on the team who is willing to adopt his workflow for a 4-week experiment, document the before/after, and publish that as a case study. Two data points are twice as convincing as one.

How This Guide Changed

What each edition changed in this guide, newest first.

  1. V1.6September 2026LATEST

    The matrix item hardened into "proxy, never a target", and this guide had to change to match - the old getting-started step that asked teams to set a quarterly throughput goal was quietly teaching them to split pull requests, so it is gone and replaced with an instruction to write down which questions the metric is allowed to answer. The item's second new clause, ruling out suggestion acceptance rate, now rests on DECODE's instrumented study of 53.6K in-IDE edits, which found 31% of accepted completions deleted entirely and most edited within 15 minutes. Sarah's scorecard brief was reworked in the same direction, from weighting a per-developer index to arguing against having one.

  2. V1.5August 2026

    The ladder's second rung was renamed from Guided to Delegated, and the guide's opening was rewritten to put the caveat first: a crude measure, since pull requests vary wildly in size, but the only output signal a team has before it builds real instrumentation. The ranges, and the pairing with cycle time, review burden and defect rate, were left as they were.

  3. V1.3June 2026

    A maintenance month here. Both research references behind the throughput ranges had been reorganised by their publishers, and the guide was pointed at the material's current locations.

  4. V1.0March 2026

    In the first edition this was the guide that let a team say something quantitative about AI adoption without building any instrumentation. It shipped with the before-and-after ranges - three to six pull requests a week without AI tools, six to ten for heavy users - and with the insistence that throughput on its own means nothing until it is read alongside cycle time, review burden and defect rate.

Where does your team actually sit on this?

This guide describes one level of one area. Run the assessment to place your team across all 16 areas, see which gates you have passed, and get a report you can take to your stakeholders.

Start the assessment