UPDATED IN OCTOBER 2026

CPI (Cost-per-Iteration): target < $0.50

Cost-per-Iteration (CPI) measures what a single agent CI attempt costs across three components - model tokens, CI compute, and the compute the generated code goes on to burn in production - with cost per merged PR tracked as a trend.

L3 · SYSTEMATICWhat this level takes
MUSTNot met, not at this level
  • The team measures the cost of one agent iteration (cost per iteration, CPI). The cost includes model tokens and CI compute.
  • The team counts the iterations that one change needs to pass CI (iterations to success, ITS).
  • The team records the median time from a push to the CI result, for each reporting period.
SHOULDExpected in practice, not required
  • The team sets a limit for the cost per iteration and for the iteration count.
  • The team reports each measurement per team and per repository.
  • The team attributes cost to a delivered change, not only to a pull request.
EVIDENCEHow you would check
  • A cost record for each iteration. The record separates the model component from the CI component.
  • An iteration count for each change that entered CI, with the distribution across changes.
  • A chart of the push-to-result time. The chart shows the median and the 95th percentile.
DEPENDS ON
  • Delivery L2 (Metrics) - a delivery-performance baseline must be operational before the AI-specific metrics layer
  • Delivery L2 (CI/CD Pipeline) - CI must be fast enough for the iteration count to be meaningful

What It Is

Cost-per-Iteration (CPI) measures what a single agent CI attempt costs - in model API costs, CI compute, and related infrastructure. When an agent submits a commit and CI runs, that's one iteration. The cost of that iteration includes: the token cost of the agent's reasoning and code generation to produce the commit, plus the CI runner cost for executing the test suite. The target at L3 is below $0.50 per iteration.

The $0.50 target is not arbitrary. It's derived from the math of agent economics: at $0.50/iteration and an ITS (Iterations-to-Success) target of 1-3, the total agent cost per PR is $0.50-$1.50. A typical engineer-hour of review and direction costs $50-100 (fully loaded). An agent that costs $1.50 to produce a PR and needs 15 minutes of human review is delivering enormous leverage. But if CPI is $3-5 per iteration and ITS is 5-8, a single PR can cost $15-40 in raw compute - approaching the cost of the human time it was meant to save.

There is a third component, and 2026 is the year it acquired a number. A study of 3.52M code changes tracked from April 2025 to April 2026 in a single large enterprise brownfield C++ codebase with full production observability (arXiv 2608.06640, published 2026-08-06) found that AI-generated code carried a higher interface and coupling burden, more copy and allocation overhead, and explicit loops where an optimised standard API existed. Downstream, that showed up as a 5-8% increase in compute resource consumption, plus increased review effort. Targeted taxonomy-informed feedback to the model cut targeted static-analysis warnings by 11.1%, so this is tractable rather than fixed.

That 5-8% is the first production-scale number putting AI code quality on the cloud bill instead of the maintainability ledger, and it changes the shape of CPI. The token cost of an iteration is paid once. The runtime cost of the code that iteration produced is paid on every request, for as long as the code lives. A team measuring only tokens and CI minutes is measuring the cheap half.

CPI has three components that must be optimized separately. The AI token cost depends on model choice, context window size, and output length. Claude Sonnet is significantly cheaper per token than Opus; using the right model for the task type can cut AI costs by 5-10x without sacrificing quality for well-specified tasks. The CI compute cost depends on CI pipeline efficiency, test suite duration, and runner instance type. A test suite that takes 20 minutes to run on a large instance costs dramatically more per iteration than a 2-minute suite on a standard runner. The production compute cost depends on what the generated code actually does at runtime, and it is the only one of the three that keeps billing after the PR merges.

At L3, teams that track CPI for the first time frequently discover that 20% of their agent tasks are consuming 80% of their iteration costs. These high-cost tasks typically share two characteristics: high ITS (many failed iterations before success) and large context windows (agents consuming excessive tokens trying to understand complex requirements). Fixing these high-cost outliers - through better context management and task specification - produces dramatic improvements in overall CPI.

September 2026 settled a question that the price war had left open: CPI has to be measured per task, not inferred from the price per token. Three pieces of evidence point the same way. First, a cheaper model can cost more per task. Gemini 3.8 Flash went GA on 2026-09-02 at an introductory $0.75/$3.75 per MTok, yet its cost per task rose from $0.40 to $0.58 against 3.7 Flash, because it takes more reasoning steps and tool calls to finish the same job. The token price was not the problem; the number of tokens per task was. Second, cheaper tokens make sessions longer. Anthropic's own Claude Code usage data for March to September (Anthropic, 2026-09-24) shows Claude working 3.3x longer per prompt, making 40%+ more model calls per prompt, and sending 2.6x more context per request, with the input-to-output ratio going from 189:1 to 324:1. Third, the market-wide numbers show the same Jevons effect: the Ramp AI Index for August data found the effective token price down 41% since the March peak ($1.15 to $0.68 per 1M tokens), while median spend per employee at the top 1% of firms fell only 9.7%.

The bill bent down where teams changed policy, not where tokens got cheaper. Ramp saw frontier models' share of tokens fall from 53% to 45%, with companies setting default policies that restrict frontier use. Uber reports weekly agent requests up 9.4x and weekly active users up 7x from February to August with spend flat since April (Uber, 2026-08-29), using cheap models for subagents, a 400K-token context cap with auto-compaction, and MCP "code-mode" cutting tokens by 50-90% - after, as Thoughtworks noted, exhausting a 12-month AI budget in four months. The vendors have noticed too: they now market cost per task and fewer tokens rather than only price per token (Anthropic says Opus 5.5 is about 40% cheaper to run than Opus 5, Sonnet 5.5 claims up to 30% lower cost per task, Meta says Muse Spark 1.3 uses about 25% fewer tokens - all vendor-reported). Treat those claims as hypotheses to check against your own per-task numbers, which is exactly what CPI is for.

Why It Matters

  • Prevents agent cost from becoming a budget blocker - without CPI tracking, agent costs grow silently as the team adds more agents and more tasks; the first monthly cloud bill that shocks leadership can cause an overcorrection that cuts the entire agent program
  • Creates incentive for CI speed investment - CI pipeline cost is directly visible in CPI; teams that track CPI have a clear financial argument for CI infrastructure investment: "cutting CI from 20 minutes to 5 minutes saves $0.30/iteration and at 500 iterations/week, that's $6K/month"
  • Drives model right-sizing - not every task needs the most capable (and expensive) model; CPI tracking reveals that many tasks can use a cheaper model without quality loss, creating a data-driven argument for model tiering
  • Enables per-task-type cost optimization - some task types (test generation) have inherently low CPI; others (complex feature implementation) have inherently high CPI; knowing this allows teams to structure agent workflows to maximize value per dollar
  • The runtime bill outlives the iteration - tokens and CI minutes are one-time charges; the 5-8% production compute overhead measured across 3.52M changes is charged on every request the generated code serves, which is why a CPI that stops at the merge is an incomplete number
  • Refactoring is a CPI lever with a measured price - a controlled August experiment ran the same feature request against a codebase at 15 successive refactoring stages, with a fresh agent each time to eliminate learning bias, and watched input tokens fall from 159,564 to 27,360 - an 83% reduction, roughly $0.40 per change at Sonnet 5 pricing, compounding across every future modification to that code
  • Price per token no longer predicts cost per task - Gemini 3.8 Flash cost 45% more per task than its predecessor ($0.40 to $0.58) because it takes more steps; only a per-task measure catches a regression like that
  • Cheaper tokens lengthen sessions rather than shrink bills - 3.3x longer work per prompt and 2.6x more context per request in Claude Code (Anthropic), and a 41% fall in effective token price against a 9.7% fall in spend per employee (Ramp); routing, default policies and context caps are what moved the bill
  • Provides the unit economics for scaling - before expanding agent usage from 10 to 100 concurrent agents, you need to know what that costs; CPI * ITS * weekly PR volume gives you the monthly cost projection for any scale level

Getting Started

  1. Instrument token costs per agent session - Add token counting to your agent orchestration layer. Every time an agent completes a CI iteration (one commit + CI run), log: tokens in (context), tokens out (code generated), model used, and duration. These map directly to API costs using the model's published pricing.
  2. Instrument CI runner costs per iteration - Most CI platforms provide cost or usage data per run. GitHub Actions charges by minute; the per-iteration CI cost is (minutes per run * cost per minute). Log this alongside token costs for each iteration.
  3. Build a per-PR cost rollup - Sum the token cost and CI cost for all iterations of a PR to get total PR cost. Publish this as a dashboard: median PR cost, 90th percentile PR cost, total weekly agent cost, and weekly cost trend.
  4. Compute CPI as the mean iteration cost - CPI = (total week cost) / (total iterations in the week). Track this weekly and set alerts when CPI exceeds your target. A CPI above $1.00 in any week is a signal that something has gone wrong - a specific task type or agent configuration is producing expensive failures.
  5. Identify the high-CPI outliers - Run a weekly analysis: which 10% of PRs have the highest total cost? What do they have in common? Are they in a specific area of the codebase, on a specific task type, or using a specific agent configuration? These outliers are where CPI optimization effort pays off most.
  6. Add the production compute delta as a third CPI line - pick the services where agent-authored code lands most, and track compute consumption per unit of traffic before and after. You are looking for the 5-8% band the C++ study measured, or your own equivalent. If you cannot measure it directly, at least stop reporting CPI as if the number ends at merge.
  7. Feed the runtime findings back into the model, not just into review - the same study cut targeted static-analysis warnings by 11.1% by giving the model taxonomy-informed feedback about the specific inefficiencies it produced. This is cheaper than catching each one in review, and it is the only intervention on this list that reduces the runtime cost rather than accounting for it.
  8. Treat refactoring as a budgeted CPI investment - context cost scales with how hard the codebase is to understand, and refactoring is now measurable in tokens rather than in taste. Note the honest caveat on the August experiment: the cost of performing the refactor itself, an upper bound of roughly 5M tokens, was not precisely tracked, so the 83% figure is the benefit side of a ledger whose other side is estimated.
  9. Report cost per task, not price per token - define the task unit (a merged PR, a resolved ticket, a completed agent run) and report its full cost: all model calls, CI compute and, where you can, the production compute delta. When you evaluate a new model, compare cost per task on your own workload, never the list price. The Gemini 3.8 Flash case - $0.40 to $0.58 per task because it takes more steps - is the one to show anyone who wants to switch on price alone.
  10. Re-cost every workload, because your August routing config is wrong - top-end per-token prices fell 20-50% in September. Claude Opus 5.5 (2026-09-22) is $4/$20, down from $5/$25, with cache reads at $0.20. GPT-6 Sol (2026-09-22) is $2/$10 ($4/$15 above 272K context) and is the new Codex default; GPT-6 Luna is $0.10/$0.50, the new API price floor. Claude Sonnet 5.5 (2026-09-28) stays at $2/$10. Claude Fable 5.1 is $10/$50 with cache reads cut to $0.25, and GPT-6 Astra $10/$50. At the cheap end, DeepSeek V4.1 Flash is $0.15/$0.60 off-peak (double at weekday peak), Xiaomi MiMo-V2.6 Flash $0.14/$0.28, and Gemini 3.8 Flash $0.75/$3.75 until 31 December, then $1.50/$7.50 - put that date in the calendar now. Re-run your tiering experiment against current prices rather than remembered ones.
  11. Cap context and route side tasks down - the levers that bent real bills in September were structural: a context cap with auto-compaction (Uber uses 400K tokens), cheaper models for subagents and side tasks, and a default routing policy that reserves the frontier model for work that needs it. Tools now ship this: GitHub Copilot's Auto model selection gained Efficiency, Balance and Intelligence tiers on 2026-09-14, routing per prompt with a 10% discount on Auto.
  12. Experiment with model tiering - Run a two-week experiment: use a cheaper model (Sonnet vs. Opus, or Haiku for very simple tasks) for the task types where quality hasn't suffered. Measure ITS for both model tiers on the same task types. If ITS doesn't significantly increase with the cheaper model, you've found a cost reduction that doesn't sacrifice quality.
TIP

Model tiering still works, but the ratios have to be re-derived each time the price list moves, and it moved again in September. Sonnet 5.5 held at $2/$10 per MTok (the planned rise to $3/$15 was cancelled), which makes it a stable base to plan against. Watch the line items that are not the headline per-token price: GPT-6 Sol's long-context rate of $4/$15 above 272K tokens, Gemini 3.8 Flash's introductory price that doubles after 31 December, DeepSeek's peak-hour doubling, Claude Managed Agents' separate $0.08 per session-hour, and OpenAI's "Fast mode" charged at 2x standard. A tiering strategy that only compares input and output token prices will miss these.

Common Pitfalls

Tracking only token costs and ignoring CI compute. In many agent workflows, CI compute is the dominant cost, not token costs. A 30-minute test suite on a large runner can cost $2-5 per iteration - far more than the $0.05-0.20 in token costs for the agent's code generation. Teams that optimize only for token cost are optimizing the smaller part of the cost equation.

Setting a single CPI target across all task types. A CPI target of $0.50 is reasonable for test writing tasks (where the agent has a clear, bounded output) but may be unrealistically low for complex feature implementation (where more context and iteration is inherently needed). Set CPI targets by task type: aggressive targets for high-frequency, well-defined tasks, more lenient targets for high-complexity, exploratory tasks.

Optimizing CPI at the expense of ITS. If you reduce context window size to cut token costs and ITS goes from 2 to 6 as a result, you've saved on per-iteration cost but increased total PR cost. CPI and ITS must be optimized together. Total PR cost (CPI * ITS) is the number that matters, not either metric in isolation.

Not alerting on CPI spikes. CPI can spike suddenly if an agent configuration change, a library update, or a codebase change causes agents to consume dramatically more context. Without alerts, these spikes go unnoticed until the monthly cloud bill arrives. Set weekly CPI alerts: if this week's CPI is more than 50% above last week's, trigger an investigation.

Stopping the meter at merge. This is the pitfall the September edition exists to name. Every dashboard described above measures cost up to the moment a PR merges, and the 3.52M-change study says the generated code then goes on consuming 5-8% more compute for the rest of its life. If your CPI report has two components rather than three, it is a story about how cheap your agents are, not about what they cost.

Assuming a routing configuration stays correct. Prices moved in both directions in August, and in September top-end per-token prices fell another 20-50% while a new generation (GPT-6 Sol and Luna, Claude Opus 5.5 and Sonnet 5.5) replaced the defaults many configs were written against. A routing config that was optimal in August is now, at best, coincidentally correct. Put a recurring calendar entry on re-costing.

Assuming a cheaper token means a cheaper task. A model that needs more steps or more tool calls can cost more per task at the same or lower token price, as Gemini 3.8 Flash did. Every routing change should be validated on cost per task, measured on your own work.

Underestimating organizational overhead costs. CPI captures infrastructure costs but not human overhead costs: the time developers spend reviewing high-ITS PRs, the time spent investigating agent failures, the time spent updating context files. The true cost-per-PR includes human time. Track this separately as a monthly estimate rather than trying to instrument it precisely - but don't forget it when making decisions about agent complexity.

How Different Roles See It

BobHEAD OF ENGINEERING

Bob approved an expansion of the agent program last quarter and is now seeing unexpectedly high cloud bills. The AI API costs are three times what he projected. He doesn't know which agents are consuming the budget or why.

What Bob should do: Bob needs CPI instrumentation immediately. He should ask his platform engineer to add token logging to the agent orchestration layer and connect it to the billing data from the cloud provider. Within a week, he should have a report showing: which agent configurations are most expensive, what the per-PR cost distribution looks like, and which task types are producing the highest costs. The analysis will almost certainly reveal that a small number of high-ITS, high-context tasks are consuming a disproportionate share of the budget. Bob should put a temporary cap on context window size for new agent tasks (this is a single configuration change) and measure the impact over two weeks. The combination of CPI instrumentation and context window capping typically reduces AI API costs by 30-50% without significant quality impact. Bob should also brace for the second half of the bill: the C++ study's 5-8% production compute increase does not appear on the AI API line at all, and Uber exhausting its annual AI budget was the most-repeated enterprise cost story of the month for a reason. The follow-up is the useful part: Uber then held spend flat from April while agent requests grew 9.4x, using a context cap, cheap models for subagents and token-saving tool access. Bob should adopt the same three levers as default policy rather than leaving model choice to each developer, because Anthropic's data shows sessions get longer as tokens get cheaper, and a cheaper price list alone will not bring his bill back to plan.

SarahPRODUCTIVITY LEAD

Sarah is building the quarterly AI productivity report and wants to include unit economics alongside throughput metrics. She wants to show: "Here is the cost of producing each agent PR and here is the value it delivers."

What Sarah should do: Sarah should build a simple cost-vs-value model. Cost side: CPI * ITS = cost per PR. Value side: estimated time saved per PR (based on the type of task - a test-writing PR saves ~30 minutes of developer time, a bug fix PR saves ~90 minutes). The ratio of value to cost is the agent ROI per PR type. Sarah should add a third cost line to the model alongside tokens and CI: the runtime compute the merged code consumes, benchmarked against the 5-8% band the C++ study measured, even if her first version of it is an estimate rather than an instrumented number. A cost-versus-value model that omits the only cost component that recurs will systematically over-rate the cheapest-looking task types. Sarah should present this model with ranges and uncertainty estimates rather than false precision. The goal isn't an exact ROI number - it's a framework that the team can use to make decisions about which task types to prioritize for agent automation. High-value, low-cost tasks (test writing) should be automated first. High-cost, low-value tasks (simple boilerplate) should be the last priority.

VictorSTAFF ENGINEER - AI CHAMPION

Victor tracks CPI for all his agent workflows and has achieved sub-$0.30 CPI through a combination of model tiering (Haiku for simple tasks, Sonnet for complex tasks, Opus reserved for architectural reasoning), optimized context windows (only the most relevant files included, not the whole codebase), and a fast CI pipeline (2-minute test runs via incremental test selection).

What Victor should do: Victor should publish his model tiering configuration and context window strategy as a platform template. The specific configuration choices - which model for which task type, how to determine context window contents, how to structure the agent's working directory to avoid loading unnecessary files - are the optimizations that took Victor months to develop. Packaging them as a template that other developers can adopt with minimal modification is the highest-leverage contribution Victor can make. Victor should also set up a monthly CPI review where he helps other teams analyze their own CPI data and identify their highest-cost outliers. The combination of a good template and ongoing coaching is how the team moves from "Victor's CPI is 0.30" to "the team's median CPI is 0.45." Victor should also test the newest tier: decision models such as TypeSafe's Jev ($0.042 per 1M input tokens, output free; vendor-reported speed and cost claims not yet independently reproduced) return typed choices, scores and booleans rather than prose, and can take over routing, "compact now?" and "is this PR risky?" calls that currently burn frontier tokens. The three-tier template - a decision model decides, a cheap model executes, the frontier plans - is worth measuring on CPI before it is worth adopting.

How This Guide Changed

What each edition changed in this guide, newest first.

  1. V1.7October 2026LATEST

    The item now says CPI is measured per task, not per token, and the guide was rebuilt around why: Gemini 3.8 Flash cost more per task than its predecessor because it takes more steps, Anthropic's data showed Claude Code sessions running 3.3x longer as tokens got cheaper, and Ramp found a 41% fall in effective token price barely moving spend. The re-costing step was replaced with September's prices (GPT-6 Sol and Luna, Claude Opus 5.5 and Sonnet 5.5), and new steps cover per-task reporting and the context caps and routing defaults that actually bent Uber's bill.

  2. V1.6September 2026

    CPI gained a third component this month, and it is the one that does not stop billing. A study of 3.52M code changes in a single enterprise C++ codebase with full production observability put AI-generated code's higher coupling and copy overhead at a 5-8% increase in production compute consumption - the first production-scale number that moves AI code quality from the maintainability ledger onto the cloud bill. The guide now treats a two-component CPI as an incomplete measurement and adds the model-feedback loop that cut targeted static-analysis warnings by 11.1%. The price table was rebuilt too, since Sonnet 5's $2/$10 became permanent, GPT-5.6 Sol fell to $4/$20, Grok 4.6 arrived at $2/$6 and DeepSeek shipped a flagship with a price rise: any routing configuration written in July is now wrong.

  3. V1.5August 2026

    The July price war moved the cost floor. Claude Opus 5 launched at $5/$25 per million tokens - half of Fable 5 - with a low/medium/high effort toggle. GPT-5.6 arrived in three tiers: Sol $5/$30, Terra $2.50/$15, Luna $1/$6, with Sol about 54% more token-efficient on coding - and watch the billing gotcha, because the bare "gpt-5.6" alias defaults to Sol. Grok 4.5 at $2/$6 completes a typical coding task for $2.49 versus $11.80 on Fable 5, and GLM-5.2 lands at roughly a fifth of Opus cost. Model routing is now cost architecture, not tuning: "frontier plans, cheap executes" is the pattern behind Cursor's Router Intelligence/Balance/Cost modes, and mixmod reports a 75.5% cut in frontier-model tokens. The CFO vocabulary is shifting to match - OpenAI's scorecard for the AI age (July 18) proposes Useful Work, Cost per Successful Task, Dependability, and Return on Compute, close cousins of CPI * ITS. The stakes: Gartner (June 24) projects per-developer AI coding token costs will exceed the average developer salary by 2028 without governance, while DX benchmarks put healthy blended cost at $200-600/month/dev with 2.5-3.5x ROI. One caution from Simon Willison: public token leaderboards distort decisions - optimize your own CPI, not someone else's benchmark.

    In May 2026 cost-per-merged-PR crossed over from advanced telemetry into a CFO-facing line item. Microsoft cancelled Claude Code internally on cost (The Verge, May 14) and Uber's COO publicly questioned AI ROI after a team burned a full-year budget in four months (Fortune, May 26). The DORA ROI report (analyzed May 11) found roughly 39% first-year ROI but only about a 10% gain on complex legacy code - a J-curve with a reliability penalty - while Goldman Sachs (May 5) projected a 24x rise in token consumption. The lesson for CPI: cheaper per-token prices do not lower the bill, so per-session caps and the CPI * ITS total are now the numbers leadership will ask about by name.

  4. V1.4July 2026

    The market turn became explicit: spending pivoted from "tokenmaxxing" to efficiency, with the metric shifting from tokens consumed to value per token (CNBC, June 26). The dominant cost pattern is now token arbitrage: spend the premium model only on judgment and planning, then hand the bulk code-writing to a cheaper specialist model. This makes model tiering a CPI lever rather than a nice-to-have. Pricing itself is a live risk: Anthropic announced Agent SDK billing changes for June 15 and then paused them on June 16, so budget against re-pricing within the quarter and keep per-session caps in place.

  5. V1.3June 2026

    Recorded the month cost-per-merged-PR crossed from advanced telemetry into a CFO-facing line item. Microsoft cancelled Claude Code internally on cost (May 14) and Uber's COO publicly questioned AI ROI after a team burned a full-year budget in four months. The DORA ROI report found roughly 39% first-year ROI but only about 10% on complex legacy code, a J-curve with a reliability penalty, while Goldman Sachs projected a 24x rise in token consumption. The lesson for CPI: cheaper per-token prices do not lower the bill, so per-session caps and the CPI x ITS total became the numbers leadership asks about by name.

  6. V1.2May 2026

    Cost telemetry stopped being optional. ccusage (13.2k GitHub stars, ccusage.com) prints token spend per Claude Code session from local JSONL files - cache-aware, offline-capable, MCP-integrated. Claude-Code-Usage-Monitor adds live charts and "when will I hit my limit" predictions. Both /usage and /context shipped as built-in commands. Reddit documented multiple "agentic fork bomb" incidents - including a $3,800 overnight bill - that turned per-session spend caps into baseline governance.

    The bigger picture: Pawel Dolega's AI subscriptions are on borrowed time (Apr 26) makes the structural case that CPI thresholds will move. A $20 Pro plan currently burns $50-100 of compute; total enterprise LLM spend doubled in six months despite per-token prices falling (Jevons paradox); Anthropic pulled Claude Code from Pro and GitHub paused Copilot signups. Track CPI now, set per-session caps now, and assume re-pricing within 12 months.

  7. V1.0March 2026

    The metric arrived with the first edition, and it arrived with a number. Fifty cents per iteration was derived rather than picked: at an ITS of one to three it puts an agent-produced PR at $0.50 to $1.50 against a fully loaded engineer-hour of $50 to $100, and it shows precisely where the leverage disappears - a $3 to $5 iteration thrashing through eight attempts costs more than the human time it was meant to save.

Where does your team actually sit on this?

This guide describes one level of one area. Run the assessment to place your team across all 16 areas, see which gates you have passed, and get a report you can take to your stakeholders.

Start the assessment