Matrix/Development

Development

How developers work with AI day-to-day. From sidebar chat to fleet agents.

4capabilities20levels60practices60guides
The matrix · full map
Capability ↓
Maturity →
L1 · Stage 01
Assisted
L2 · Stage 02
Delegated
L3 · Stage 03
Systematic
L4 · Stage 04
Governed
Sweet spot
L5 · Stage 05
Self-improving
01·15 guides
Coding Agent Usage→
How your team uses AI coding assistants - from autocomplete to autonomous agent fleets
Autocomplete and a chat window
3 practices·3 guides→
An agent in the IDE, rules in the repo
3 practices·3 guides→
CLI agents become the primary interface
3 practices·3 guides→
Unattended agents, inside a written boundary
3 practices·3 guides→
A fleet ships faster than you can read
3 practices·3 guides→
02·15 guides
Context Engineering→
What information agents receive about your codebase, architecture, and conventions
The agent sees one open file
3 practices·3 guides→
CLAUDE.md tells it the basics
3 practices·3 guides→
Context is served, not scavenged
3 practices·3 guides→
The org pushes context to the agent
3 practices·3 guides→
Context maintains itself
3 practices·3 guides→
03·15 guides
Code Review & Quality→
How AI-generated code is reviewed, validated, and approved before merging
Humans review everything, slowly
3 practices·3 guides→
AI suggests, humans still decide
3 practices·3 guides→
Lint is architecture; AI takes first pass
3 practices·3 guides→
Green merges on policy, never without an owner
3 practices·3 guides→
Human eyes only on Red
3 practices·3 guides→
04·15 guides
Testing Strategy→
How tests are written, maintained, and validated in an AI-assisted workflow
Tests by hand, flakes by habit
3 practices·3 guides→
Agents write tests, humans own the oracle
3 practices·3 guides→
Requirements are the oracle, not the code
3 practices·3 guides→
A red test means a real defect
3 practices·3 guides→
The suite heals itself
3 practices·3 guides→
Climb the matrix

You don't have to figure this out alone.

Every level in this matrix has a path. Read the playbooks the teams that have climbed it wrote. Run the assessment with our consultants. Start where you are.

Live with Visdom

Book an AI Maturity Assessment session with your team.

We walk you through all four perspectives, score where you actually are, and leave you with a 90-day plan to climb in the dimensions that matter most.

Book an assessment →See what's included90-day plan - scored assessment - coaching
Author Commentary

The October 2026 zeitgeist is System One.

For two years the coding agent had one brain, and every decision went through it: which file to read, which tool to call, whether the context was full, whether the diff looked risky. In September that brain split. TypeSafe AI's Jev is a model that does not write prose at all - you hand it text and a schema, and it returns a typed choice, score or boolean with a calibrated probability, at $0.042 per million input tokens. Within ten days it sat behind four gateways, pydantic-ai had a model class for it, open-swe used it for automated model routing, and gpt-researcher swapped embeddings for it as a context selector. The shape that falls out is three-tier routing: a System One model decides, a cheap model executes, the frontier plans. GitHub made the same point from the product side with Copilot Auto's Efficiency / Balance / Intelligence tiers, routing per prompt. Treat the headline numbers with care, though: "193.6x faster, 444.6x cheaper" is TypeSafe's own benchmark, the company lists its own failure modes (counting, dates, irrelevant context), and there is no SLA. Put a decision model on your own eval set before you let it gate anything.

The second story is context, and it cuts against the instinct to write more of it. Uber's software factory post caps context at 400K tokens with auto-compaction and routes subagents to cheap models, which is how it held spend flat while agent requests grew 9.4x. Marmelab's State of AI Harness Engineering supplies the uncomfortable half: machine-generated context files did worse than having none, at 20%+ more cost, while human-written ones helped by about 4%. And only 4.4% of the security rules written in public CLAUDE.md files have a technical control behind them. A line in an instruction file that says "never touch production" is a request to the model. If it matters, it belongs in a deny rule, a sandbox or a credential the agent does not hold.

Which brings the autonomy question up a level. Last month the maturity signal was a written, version-controlled deny/ask ruleset in the repository. In September GitHub made enterprise-managed permissions for Copilot agent operations generally available: admins decide which shell commands, file operations and network domains an agent may use, and users cannot override them. Claude Code shipped exact-version model pinning in the same month. For an individual developer this feels like a loss of control. For the organisation it is the first time the rules are actually enforced, and that is the version of autonomy that survives an audit.

Other perspectives