October 2026: System One
The cheap brain and the control plane. September split the agent's brain in two - fast, cheap decisions moved to decision models and routed models - while control moved into one layer: the gateway, managed settings and agent identity. 31 of 240 matrix items were rewritten; the structure is unchanged.
At a glance
Why: TypeSafe's Jev, a System One model that returns typed choices, scores and booleans instead of prose, reached Vercel, OpenRouter, Cloudflare and Databricks within ten days, with ~50 radar projects wiring it in for routing, compaction and safety gates. Its 444.6x cost claim is self-tested; the pattern is what the matrix adopts, not the multiplier.
Why: GPT-6 Luna set a new price floor and Opus 5.5 cut its price, yet Claude now works 3.3x longer per prompt. Ramp's data shows the bill bending only where firms set default policies that restrict frontier use, and Gemini 3.8 Flash cost 45% more per task at the same token price. Cost is measured per task, and the routing policy is a governance decision.
Why: Kong AI Gateway 2.0 filters the tools each caller can see, and Claude Code now pins exact model versions with hard denies. On September 2 LiteLLM's MCP auth bypass reached the CISA KEV list: any bearer token yielded a fully authorised session. The layer that holds every key needs a patch SLA of its own.
Why: Okta, CrowdStrike, GitHub and Google shipped agent identity, per-agent permissions and token kill switches, and the IETF adopted its first working-group baseline for agent authentication. The urgency is real: the first mass agent-run campaign breached 395 organisations and found Claude, Cursor and Gemini tokens on 5,871 machines.
Why: 98% of large US firms have formal AI policies and 47% skipped them for urgent deployments (EY). Only 4.4% of security rules in public CLAUDE.md files are backed by a technical control (Marmelab). Meta scrapped its token leaderboard. GitSpawn and Plugin4Shell ran code through git config and plugin marketplaces that no written rule would have stopped.
What changed in the model
Three shifts in what the model asks for. None changes the structure; each changes what counts as evidence.
Token prices fell 20-50% at the top end in one month and the bills did not follow, because agents ran longer. Wherever the matrix asks for a cost figure, a price-per-token number no longer satisfies it; the unit is the cost of a finished task, and the routing policy behind it is owned centrally.
Agents get their own principal, a named human owner, short-lived task-scoped tokens and a way to be revoked in seconds. These criteria now appear in Governance, Merge & Deploy, Agent Runtime, MCP and Observability at once, because an identity that only one area checks is not an identity.
A security rule in an instruction file, a policy with an urgent-deployment bypass or a governance document with no gateway behind it is scored as absent. The evidence the assessment looks for is the control: the gateway rule, the managed setting, the deny list, the revocation.
Every change, by area and level
2026-09 2026-10 - September taxonomy preserved unchanged.
Development
Coding Agent Usage
Context Engineering
Delivery Management
CI/CD Pipeline
Merge & Deploy
Metrics
Governance & Compliance
Organization
AI Adoption Model
Infrastructure
Agent Runtime & Sandboxing
MCP & Tool Integration
Build System
Observability & Feedback Loop
The month in numbers
Context behind the edition. The figures that drove specific changes are cited with those changes above.
breached in the first mass-exploitation campaign run by AI agents, first compromise in 26 seconds
GreyNoise via VentureBeatof large US firms have formal AI governance policies / skipped them for urgent deployments
EYof security rules in public CLAUDE.md files are backed by a technical control
Marmelablonger Claude works per prompt, March to September, with context per request up 2.6x (vendor data)
Anthropicdrop in effective token price since the March peak, while the frontier share of tokens fell from 53% to 45%
Ramp AI Indexmain-branch success rate across 28M CI workflows, a five-year low, as throughput rose 59%
CircleCI via Paul StackUpdated Guides
31 of 240 matrix items were rewritten this edition, and no rows were added or removed. 38 guides were updated: the 31 whose items changed, plus 7 reworked around new September evidence.
How the matrix is built
4 perspectives x 4 areas x 5 levels x 3 items = 240, and 240 is also the guide count. Every item owns exactly one guide and no guide is claimed twice. Levels run Assisted, Delegated, Systematic, Governed, Self-improving.
The matrix is what we show: one sentence per item, in a practitioner's language. The maturity gates are what we assess: Must, Should, Prerequisites and Evidence per area and level, and they decide the workshop score. Both layers must say the same thing; a gap between them is a defect, not a nuance.
- Describe the capability. A vendor or a methodology belongs in parentheses as an example, never as the bar.
- A threshold must not encode company size. A percentage of your own organisation scales; an absolute volume measures headcount.
- L1 describes a floor a team actually stands on, never an absence. Deficit wording is blocked by a test.
- Each capability has one owning area, so nobody scores twice for one practice.
Every edition is checked automatically against these rules before it ships: that the matrix and the assessment still say the same thing, that no capability is claimed by two areas, and that nothing is displayed which the assessment never actually scores. Previous editions are never edited, so a score from an earlier month stays comparable to what it meant at the time.
The level names promise a delegation and control arc, but CI/CD Pipeline and Build System still ladder on latency: their L4 criteria are wall-clock targets with no governance in them. Re-laddering those two areas would move existing workshop scores, so it is a decision about the model rather than a tidy-up.
What Didn't Change (and Why)
Sources
TypeSafe AI: Introducing System One models
Jev: typed Choice / Score / Boolean decisions, $0.042 per 1M input tokens
Latent Space: Jev, System One models for Prod
The case for a decision tier below the frontier
TS2: the 445x claim is still self-tested
Why Jev's speed and cost multipliers are vendor-reported
Vercel: Jev on AI Gateway
A decision model reachable from a mainstream gateway within days
TechCrunch: GPT-6 Sol and Luna
Luna at $0.10/$0.50 sets a new API price floor; Sol becomes the Codex default
Anthropic: Opus 5.5 and longer coding sessions
Claude works 3.3x longer per prompt; context per request up 2.6x
Ramp AI Index, September 2026
Effective token price -41%; frontier share of tokens 53% to 45%
Uber: running a software factory efficiently
Agent requests 9.4x with spend flat since April - cheap subagents, 400K context cap
GitHub: Copilot Auto model selection tiers
Efficiency / Balance / Intelligence, routed per prompt
Kong AI Gateway 2.0 GA
MCP server bundling, identity principals, per-modality cost accounting
LiteLLM CVE-2026-59822 on CISA KEV
Any bearer token yields a fully authorised MCP session
Claude Code changelog
Exact-version model pinning, hard denies, gateway prompt-ID header
Okta: AI innovations at Oktane 2026
Agent-to-Agent Connections, agent access certifications, gateway kill switch
CrowdStrike: Agentic Identity Provider
Verifiable identity per agent, short-lived least-privilege credentials
GitHub: enterprise-managed agent permissions
Admins block, approve or allow shell, files and network; users cannot override
VentureBeat: AI agents breached 395 organisations
The first mass-exploitation campaign run by agents (GreyNoise)
Anthropic threat intelligence report, September
An injected eval sandbox handed over production API keys
The Hacker News: GitSpawn
Malicious git config runs code on the agent's startup git status
The Register: Plugin4Shell
Zero-click RCE through plugin marketplaces, bypassing SHA pinning
Pillar Security: Deadbugz
An MCP server that rewrites its tool metadata after three calls
EY: the AI governance gap
98% have policies, 47% skipped them for urgent deployments
Marmelab: State of AI Harness Engineering
Only 4.4% of CLAUDE.md security rules backed by a technical control
The Decoder: Meta drops AI usage from reviews
Token leaderboard for 85k employees scrapped after tokenmaxxing
Paul Stack: AI broke the assumptions behind CI
Main-branch success at a five-year low across 28M workflows
OSS AI policies study (arXiv)
281 policies: 83.3% permit AI, 48.8% require disclosure