v1.7October 1, 2026/covers September 2026

October 2026: System One

The cheap brain and the control plane. September split the agent's brain in two - fast, cheap decisions moved to decision models and routed models - while control moved into one layer: the gateway, managed settings and agent identity. 31 of 240 matrix items were rewritten; the structure is unchanged.

At a glance

Routing gets a third tier: a decision model decides, a cheap model executes, the frontier plans

Why: TypeSafe's Jev, a System One model that returns typed choices, scores and booleans instead of prose, reached Vercel, OpenRouter, Cloudflare and Databricks within ten days, with ~50 radar projects wiring it in for routing, compaction and safety gates. Its 444.6x cost claim is self-tested; the pattern is what the matrix adopts, not the multiplier.

Coding Agent Usage L4
Cheaper tokens made sessions longer, not bills smaller

Why: GPT-6 Luna set a new price floor and Opus 5.5 cut its price, yet Claude now works 3.3x longer per prompt. Ramp's data shows the bill bending only where firms set default policies that restrict frontier use, and Gemini 3.8 Flash cost 45% more per task at the same token price. Cost is measured per task, and the routing policy is a governance decision.

Metrics L3, AI Adoption Model L2 and L3
The AI gateway became the control plane, and an attack surface in the same month

Why: Kong AI Gateway 2.0 filters the tools each caller can see, and Claude Code now pins exact model versions with hard denies. On September 2 LiteLLM's MCP auth bypass reached the CISA KEV list: any bearer token yielded a fully authorised session. The layer that holds every key needs a patch SLA of its own.

Governance L3 and L4, MCP & Tool Integration L3 and L4
Every agent is a first-class identity with a named human owner

Why: Okta, CrowdStrike, GitHub and Google shipped agent identity, per-agent permissions and token kill switches, and the IETF adopted its first working-group baseline for agent authentication. The urgency is real: the first mass agent-run campaign breached 395 organisations and found Claude, Cursor and Gemini tokens on 5,871 machines.

Governance L2 to L5, Agent Runtime L2 to L5, Merge & Deploy L2 to L4
Governance on paper is not governance

Why: 98% of large US firms have formal AI policies and 47% skipped them for urgent deployments (EY). Only 4.4% of security rules in public CLAUDE.md files are backed by a technical control (Marmelab). Meta scrapped its token leaderboard. GitSpawn and Plugin4Shell ran code through git config and plugin marketplaces that no written rule would have stopped.

Context Engineering L3, Governance L5, MCP & Tool Integration L4

What changed in the model

Three shifts in what the model asks for. None changes the structure; each changes what counts as evidence.

Cost is measured per task, not per token

Token prices fell 20-50% at the top end in one month and the bills did not follow, because agents ran longer. Wherever the matrix asks for a cost figure, a price-per-token number no longer satisfies it; the unit is the cost of a finished task, and the routing policy behind it is owned centrally.

Agent identity runs through every perspective

Agents get their own principal, a named human owner, short-lived task-scoped tokens and a way to be revoked in seconds. These criteria now appear in Governance, Merge & Deploy, Agent Runtime, MCP and Observability at once, because an identity that only one area checks is not an identity.

A written rule counts only when something enforces it

A security rule in an instruction file, a policy with an urgent-deployment bypass or a governance document with no gateway behind it is scored as absent. The evidence the assessment looks for is the control: the gateway rule, the managed setting, the deny list, the revocation.

Every change, by area and level

2026-09 2026-10 - September taxonomy preserved unchanged.

Development

Coding Agent Usage

L2Autonomy is set increasingly by the organisation rather than the developer: GitHub's enterprise-managed Copilot agent permissions for shell commands, file access and network domains cannot be overridden by users
L3September model refresh: Claude Code on Opus 5.5 or Sonnet 5.5, Codex on GPT-6 Sol, with cheap and open-weight workers (GPT-6 Luna, DeepSeek V4.1 Flash, MiMo-V2.6, GLM-5.3, Kimi K3) for execution
L4Routing goes three-tier: a System One decision model decides (TypeSafe Jev - routing, compaction, safety gates in milliseconds), a cheap model executes, the frontier plans; re-costed monthly per task, not per token

Context Engineering

L3Context budgeting gets a hard cap (Uber: 400K tokens with auto-compaction); agent files are hand-written, because machine-generated context files did worse than none at 20%+ more cost
L3Any security rule written in CLAUDE.md is backed by a technical control - across public files, only 4.4% are

Delivery Management

CI/CD Pipeline

L5Verification moves before the PR opens and CI checks the evidence instead of re-running it: main-branch success fell to a five-year low of 70.8% across 28M workflows

Merge & Deploy

L2Agent commits come from a distinct bot identity, never a developer's account
L3Agents push with short-lived GitHub App or OIDC tokens and never hold deploy secrets
L4A green verdict still flows straight to production, but high-impact actions (token creation, production credentials, webhook edits) require a fresh human re-authentication (GitHub proof of presence)

Metrics

L2The token-leaderboard anti-pattern confirmed at scale: Meta scrapped its 85k-employee leaderboard and pulled AI usage from engineer performance reviews
L3Cost per iteration measured per task, not per token: Gemini 3.8 Flash kept its token price and still went from $0.40 to $0.58 per task, and cheaper tokens made Claude Code sessions 3.3x longer

Governance & Compliance

L2A deliberate default for new vendor features instead of letting them auto-enable (Copilot's "Default policy for new features")
L2Every agent has a named human owner, because 51% of organisations cannot say who owns their AI identities
L3The audit trail records the agent's own identity, registered as a distinct principal in the IdP (Entra Agent ID, Okta, Google Agent Identity), not a developer's token
L3All model traffic through an AI gateway: central keys with BYOK-only, per-team budgets, a model allowlist with exact-version pinning and hard denies; .git/config joins the mandatory-diff path
L4Provenance via gateway prompt IDs, prompt and response logs in the SIEM with retention beyond vendor defaults (Cursor keeps 30 days), agent access certified like human access
L4The AI gateway run as critical infrastructure with a patch SLA - LiteLLM's MCP auth bypass reached the CISA KEV list - plus DLP at the gateway and detection of unapproved agents (26% of large firms cannot)
L5Every agent action traces to one agent identity and one accountable human, and the urgent fast path runs through the same controls: 47% of firms with written policies skipped them for urgent deployments
L5RBAC per agent is task-scoped with no standing access, and new agents start on probation and earn scope from their track record

Organization

AI Adoption Model

L2Pilots track cost per task from day one - cheaper tokens lengthen sessions rather than shrink bills
L3A default-model policy set centrally: firms restricting frontier use cut spend per employee 9.7% while effective token prices fell 41% (Ramp AI Index)

Infrastructure

Agent Runtime & Sandboxing

L2Untrusted repositories opened with auto-run hooks disabled now includes git config: GitSpawn used core.fsmonitor to run attacker code on the agent's startup git status, outside the sandbox
L2Never a personal PAT or a shared long-lived key: the first mass agent-run campaign found Claude, Cursor and Gemini tokens on 5,871 machines and burned $600k of credits through one dashboard
L3Credentials injected at run time from a vault or broker, never in prompts, files or eval sandboxes - an injected evaluation sandbox handed over production keys for several providers
L4Assume escape, including over DNS: workload identity per agent (SPIFFE, WIF, Google Agent Identity) with task-scoped tokens that expire in minutes; Docker Cloud Sandboxes join the microVM options
L5Sender-constrained tokens (DPoP, mTLS) and every session of one agent revocable in seconds

MCP & Tool Integration

L3RBAC per MCP tool through a gateway that shows each caller only the tools it may use (Kong MCP bundling); clients registered via Client ID Metadata Documents, tokens bound to one server
L4Access granted centrally by the IdP (Okta Cross App Access / ID-JAG, MCP Enterprise-Managed Authorization) instead of per-user consent sprawl, with a kill switch that revokes an agent's live tokens at the gateway
L4Every server on a lifecycle from intake to deprecation with a pause switch; tool metadata watched at runtime (Deadbugz rewrote its tool descriptions after three calls); plugins pinned by full commit SHA, never a branch name (Plugin4Shell)

Build System

L2Packages an agent installs from docs or llms.txt verified against a registry allowlist first: 237 of 8,565 llms.txt files pointed at dead, typo'd or unregistered packages

Observability & Feedback Loop

L3Every tool call traced to the agent identity, the human it acts for and the token used
L4Anomalous agent behaviour (scope drift, credential reuse, unusual egress) revokes the agent's credentials automatically - a September training-sandbox escape got past partial monitoring because the automatic shutdown failed

The month in numbers

Context behind the edition. The figures that drove specific changes are cited with those changes above.

Updated Guides

31 of 240 matrix items were rewritten this edition, and no rows were added or removed. 38 guides were updated: the 31 whose items changed, plus 7 reworked around new September evidence.

Development
Agent in IDE (autonomy by written rule)

The rule is now set by the organisation: enterprise-managed agent permissions users cannot override

Development
CLI agents as primary

September model refresh: Opus 5.5, Sonnet 5.5, GPT-6 Sol, and a new price floor from GPT-6 Luna

Development
One-shot unattended agents

Three-tier routing: a System One model decides, a cheap model executes, the frontier plans

Development
Context budgeting

A hard context cap, hand-written agent files, and security rules backed by a technical control

Delivery
Production feedback into CI

Verification before the PR opens; CI checks evidence instead of re-running it

Delivery
CD pipeline with gates

Agent commits come from a distinct bot identity, never a developer's account

Delivery
Policy-based merge rules

Short-lived GitHub App or OIDC tokens; agents never hold deploy secrets

Delivery
Green verdict to production

High-impact actions require a fresh human re-authentication (proof of presence)

Delivery
Licenses vs usage rate

Meta scrapped its token leaderboard and pulled AI usage from performance reviews

Delivery
Cost per iteration

Per task, not per token - same token price, 45% more per task on Gemini 3.8 Flash

Delivery
Official AI tool policy

A deliberate default for new vendor features instead of letting them auto-enable

Delivery
Basic audit: who uses what

Every agent has a named human owner - 51% of organisations cannot say who owns theirs

Delivery
Minimum viable audit trail

Adds the agent's own identity as a distinct principal in the IdP

Delivery
Policy-as-code

All model traffic through an AI gateway: BYOK-only keys, budgets, exact-version pinning, hard denies

Delivery
Full provenance tracking

Gateway prompt IDs, SIEM retention beyond vendor defaults, agent access certification

Delivery
Automated compliance checks

The AI gateway as critical infrastructure with a patch SLA, after LiteLLM reached CISA KEV

Delivery
Self-documenting audit trail

One agent identity, one accountable human, and no fast path around the controls

Delivery
RBAC per agent

Task-scoped, no standing access; new agents start on probation

Infrastructure
Basic sandboxing

Git config is an execution surface too - GitSpawn ran through core.fsmonitor

Infrastructure
Scoped agent credentials

Never a personal PAT: agent tokens stolen from 5,871 machines in the first mass agent-run campaign

Infrastructure
Isolated agent environments

Credentials injected at run time, never in prompts, files or eval sandboxes

Infrastructure
MicroVM isolation by default

Assume escape, including over DNS; workload identity with task-scoped tokens

Infrastructure
Each agent on an isolated machine

Sender-constrained tokens; every session of one agent revocable in seconds

Infrastructure
RBAC per MCP tool

An MCP gateway shows each caller only the tools it may use

Infrastructure
One governed MCP gateway

Access granted centrally by the IdP, plus a kill switch for an agent's live tokens

Infrastructure
MCP governance

Runtime metadata watch (Deadbugz) and full-SHA pinning for plugins (Plugin4Shell)

Infrastructure
Basic build caching

Packages installed from docs or llms.txt checked against a registry allowlist

Infrastructure
Incident data for context

Every tool call traced to agent identity, human and token

Infrastructure
Self-healing known patterns

Anomalous agent behaviour revokes credentials automatically

Organization
Pilot metrics

Cost per task from day one, not cost per token

Organization
Standardized agent setup

A default-model policy set centrally - the lever that bent spend down at Ramp's top firms

Development
Agent instruction files in repo

llms.txt as a supply-chain path, and only 4.4% of CLAUDE.md security rules are enforced

Development
AI review agent as first pass

A second reviewer agent lowered success by 8% in Marmelab's harness study

Delivery
Shadow AI: private subscriptions

Personal agents are unmanaged credentials; EY and Deloitte shadow AI numbers

Delivery
EU AI Act awareness

No new EU guidance in September; the UK bill lets government order a stop on a specific model

Infrastructure
Production metrics to dashboards

The governed AI gateway as control plane, and the reason to patch it

Organization
Platform team owns AI tooling

The platform owns the AI gateway and the agent identity registry

Organization
AI-first development culture

Judge outcomes, not usage volume - Meta dropped AI usage from performance reviews

How the matrix is built

The shape

4 perspectives x 4 areas x 5 levels x 3 items = 240, and 240 is also the guide count. Every item owns exactly one guide and no guide is claimed twice. Levels run Assisted, Delegated, Systematic, Governed, Self-improving.

Two layers

The matrix is what we show: one sentence per item, in a practitioner's language. The maturity gates are what we assess: Must, Should, Prerequisites and Evidence per area and level, and they decide the workshop score. Both layers must say the same thing; a gap between them is a defect, not a nuance.

Rules for writing a criterion
  • Describe the capability. A vendor or a methodology belongs in parentheses as an example, never as the bar.
  • A threshold must not encode company size. A percentage of your own organisation scales; an absolute volume measures headcount.
  • L1 describes a floor a team actually stands on, never an absence. Deficit wording is blocked by a test.
  • Each capability has one owning area, so nobody scores twice for one practice.
How it is kept honest

Every edition is checked automatically against these rules before it ships: that the matrix and the assessment still say the same thing, that no capability is claimed by two areas, and that nothing is displayed which the assessment never actually scores. Previous editions are never edited, so a score from an earlier month stays comparable to what it meant at the time.

Known and deliberate

The level names promise a delegation and control arc, but CI/CD Pipeline and Build System still ladder on latency: their L4 criteria are wall-clock targets with no governance in them. Re-laddering those two areas would move existing workshop scores, so it is a decision about the model rather than a tidy-up.

What Didn't Change (and Why)

Matrix structure (5 levels, 4 perspectives, 16 areas) - Stable; 31 items were rewritten in place and no rows were added or removed.
Code Review, Testing Strategy, Knowledge Management, Team Structure, Tech Debt - No item changed; September's evidence landed on routing, gateways and identity rather than on review or testing practice.
MicroVM sandboxes as the runtime baseline - Still the baseline; identity and short-lived tokens were layered on top, and Docker Cloud Sandboxes joined the options.
Harness engineering as the named discipline - Held and reinforced: the same model scored 68-88% depending on which of eight harnesses ran it.
EU AI Act timeline - No new major guidance in September; Article 50 applies since Aug 2, high-risk duties stay deferred to Dec 2027 / Aug 2028.
Vendor multipliers as thresholds - Jev's speed and cost claims are self-tested and not independently reproduced; the matrix adopts the three-tier pattern, not the number.

Sources

TypeSafe AI: Introducing System One models

Jev: typed Choice / Score / Boolean decisions, $0.042 per 1M input tokens

Latent Space: Jev, System One models for Prod

The case for a decision tier below the frontier

TS2: the 445x claim is still self-tested

Why Jev's speed and cost multipliers are vendor-reported

Vercel: Jev on AI Gateway

A decision model reachable from a mainstream gateway within days

TechCrunch: GPT-6 Sol and Luna

Luna at $0.10/$0.50 sets a new API price floor; Sol becomes the Codex default

Anthropic: Opus 5.5 and longer coding sessions

Claude works 3.3x longer per prompt; context per request up 2.6x

Ramp AI Index, September 2026

Effective token price -41%; frontier share of tokens 53% to 45%

Uber: running a software factory efficiently

Agent requests 9.4x with spend flat since April - cheap subagents, 400K context cap

GitHub: Copilot Auto model selection tiers

Efficiency / Balance / Intelligence, routed per prompt

Kong AI Gateway 2.0 GA

MCP server bundling, identity principals, per-modality cost accounting

LiteLLM CVE-2026-59822 on CISA KEV

Any bearer token yields a fully authorised MCP session

Claude Code changelog

Exact-version model pinning, hard denies, gateway prompt-ID header

Okta: AI innovations at Oktane 2026

Agent-to-Agent Connections, agent access certifications, gateway kill switch

CrowdStrike: Agentic Identity Provider

Verifiable identity per agent, short-lived least-privilege credentials

GitHub: enterprise-managed agent permissions

Admins block, approve or allow shell, files and network; users cannot override

VentureBeat: AI agents breached 395 organisations

The first mass-exploitation campaign run by agents (GreyNoise)

Anthropic threat intelligence report, September

An injected eval sandbox handed over production API keys

The Hacker News: GitSpawn

Malicious git config runs code on the agent's startup git status

The Register: Plugin4Shell

Zero-click RCE through plugin marketplaces, bypassing SHA pinning

Pillar Security: Deadbugz

An MCP server that rewrites its tool metadata after three calls

EY: the AI governance gap

98% have policies, 47% skipped them for urgent deployments

Marmelab: State of AI Harness Engineering

Only 4.4% of CLAUDE.md security rules backed by a technical control

The Decoder: Meta drops AI usage from reviews

Token leaderboard for 85k employees scrapped after tokenmaxxing

Paul Stack: AI broke the assumptions behind CI

Main-branch success at a five-year low across 28M workflows

OSS AI policies study (arXiv)

281 policies: 83.3% permit AI, 48.8% require disclosure