Built by AI agents, checked against measurements.
This repo is written by AI coding agents working under an explicit contract kept in the repo itself. The human sets direction, approves risky changes and checks every claim against scripts/verify.py's output, never against an agent's own summary.
Two generations
Built on IBM Bob
RetiredIBM Bob authored the foundation and Phases 0–8: the engine, the demo-repo fixture, the mutation engine, compact MCP responses, and a .bob/ swarm of custom modes (Orchestrator, Test Writer, Critic, Gate, Publisher) with hooks that kept writes inside tests/.
12 Bob commits, ranked 1st by Bob commits in the hackathon.
Bob’s roles, as code
IBM Bob 2.0The writer and critic roles now run in watson_agent/ on watsonx.ai or Vertex AI, callable from the terminal or the Autofix button. The hooks became properties of the code: a write guard that rejects anything outside tests/, and a run report written after every fix. Claude Code authors the repo from Phase 13 on, under the same contract.
First live fix-loop success: 20.25% → 89.87% mutation score on demo-repo.
The session loop
Every task starts from a clean context and three short files. Re-reading them is cheaper than carrying a long session's dead ends forward.
- 1New task
- 2Fresh session
- 3Read CLAUDE.md · PENDING.md · LASTCONTEXT.md
- 4Branch
- 5Plan → code → tests
- 6scripts/verify.py
- 7PR with the verify output
- 8Human reviews and merges
- 9Update PENDING.md + LASTCONTEXT.md
- ↺ next task
The contract lives in the repo
| File | Role |
|---|---|
CLAUDE.md / AGENTS.md | Architecture rules, layer boundaries, the verification table and the actions that need human sign-off. Kept byte-identical so every agent reads the same rules. |
PENDING.md | The prioritized backlog, phase by phase. Items are checked off as they land, never deleted. |
LASTCONTEXT.md | Where things stand: what's done, what's mid-flight, decisions already made and gotchas still in effect. A new session reads this instead of re-deriving it from history. |
scripts/verify.py | The phase gates. A change isn't done without a PASS, pasted into the PR as the test plan. |
Verification over trust
Real results that looked fine and weren't, each found by checking a number against an independent one. Each now has a guard in the code.
Mutation score 100% (79/79), identical on two runs
Actuallypytest resolved through PATH to an environment without it, so every mutant run failed and counted as killed.
Guard nowSubprocesses use sys.executable; the check also asserts the known baseline, 16/79.
Mutation score 100% (79/79) after an AI fix run
ActuallyThe AI-written tests failed on the unmodified code, so every mutant inherited those failures.
Guard nowThe unmodified suite must pass before any mutant runs, or the engine refuses.
Accessibility: 0 violations
Actuallyaxe-playwright-python was never installed, so the check never ran.
Guard nowDeclared as a dependency; a check that can't run returns ok=false with the error.
Fix loop crashed mid-run
ActuallyThe model asked to read a file that doesn't exist, and the raw error escaped the loop.
Guard nowTool errors go back to the model as a normal result it can correct.
Who does what
The agent
- Plans the change from PENDING.md and the rules that apply
- Writes code and tests inside the layer boundaries
- Runs the matching verify.py check; no PASS, not done
- Reviews its own diff against the ask-first list
- Branch, Conventional Commit, PR with the verify output
The human
- Sets direction and priorities
- Approves anything on the ask-first list: demo-repo source, the write guard, mutation operators, published numbers
- Runs the steps that need real cloud access: IAM grants, secrets
- Reviews and merges every PR
Full write-up: AI_ASSISTED_DEVELOPMENT_FRAMEWORK.md · IBM_BOB_USAGE.md