AI-assisted development

Built by AI agents, checked against measurements.

This repo is written by AI coding agents working under an explicit contract kept in the repo itself. The human sets direction, approves risky changes and checks every claim against scripts/verify.py's output, never against an agent's own summary.

Two generations

1.0

Built on IBM Bob

Retired

IBM Bob authored the foundation and Phases 0–8: the engine, the demo-repo fixture, the mutation engine, compact MCP responses, and a .bob/ swarm of custom modes (Orchestrator, Test Writer, Critic, Gate, Publisher) with hooks that kept writes inside tests/.

12 Bob commits, ranked 1st by Bob commits in the hackathon.

2.0

Bob’s roles, as code

IBM Bob 2.0

The writer and critic roles now run in watson_agent/ on watsonx.ai or Vertex AI, callable from the terminal or the Autofix button. The hooks became properties of the code: a write guard that rejects anything outside tests/, and a run report written after every fix. Claude Code authors the repo from Phase 13 on, under the same contract.

First live fix-loop success: 20.25% → 89.87% mutation score on demo-repo.

The session loop

Every task starts from a clean context and three short files. Re-reading them is cheaper than carrying a long session's dead ends forward.

  1. 1New task
  2. 2Fresh session
  3. 3Read CLAUDE.md · PENDING.md · LASTCONTEXT.md
  4. 4Branch
  5. 5Plan → code → tests
  6. 6scripts/verify.py
  7. 7PR with the verify output
  8. 8Human reviews and merges
  9. 9Update PENDING.md + LASTCONTEXT.md
  10. ↺ next task

The contract lives in the repo

FileRole
CLAUDE.md / AGENTS.mdArchitecture rules, layer boundaries, the verification table and the actions that need human sign-off. Kept byte-identical so every agent reads the same rules.
PENDING.mdThe prioritized backlog, phase by phase. Items are checked off as they land, never deleted.
LASTCONTEXT.mdWhere things stand: what's done, what's mid-flight, decisions already made and gotchas still in effect. A new session reads this instead of re-deriving it from history.
scripts/verify.pyThe phase gates. A change isn't done without a PASS, pasted into the PR as the test plan.

Verification over trust

Real results that looked fine and weren't, each found by checking a number against an independent one. Each now has a guard in the code.

Looked like

Mutation score 100% (79/79), identical on two runs

Actually

pytest resolved through PATH to an environment without it, so every mutant run failed and counted as killed.

Guard now

Subprocesses use sys.executable; the check also asserts the known baseline, 16/79.

Looked like

Mutation score 100% (79/79) after an AI fix run

Actually

The AI-written tests failed on the unmodified code, so every mutant inherited those failures.

Guard now

The unmodified suite must pass before any mutant runs, or the engine refuses.

Looked like

Accessibility: 0 violations

Actually

axe-playwright-python was never installed, so the check never ran.

Guard now

Declared as a dependency; a check that can't run returns ok=false with the error.

Looked like

Fix loop crashed mid-run

Actually

The model asked to read a file that doesn't exist, and the raw error escaped the loop.

Guard now

Tool errors go back to the model as a normal result it can correct.

Who does what

The agent

  • Plans the change from PENDING.md and the rules that apply
  • Writes code and tests inside the layer boundaries
  • Runs the matching verify.py check; no PASS, not done
  • Reviews its own diff against the ask-first list
  • Branch, Conventional Commit, PR with the verify output

The human

  • Sets direction and priorities
  • Approves anything on the ask-first list: demo-repo source, the write guard, mutation operators, published numbers
  • Runs the steps that need real cloud access: IAM grants, secrets
  • Reviews and merges every PR

Full write-up: AI_ASSISTED_DEVELOPMENT_FRAMEWORK.md · IBM_BOB_USAGE.md