One deterministic engine, three thin adapters.
The engine measures and writes JSON; the CLI, the web API and the MCP server only call it. The AI layer sits beside the engine, never inside a measurement.
Architecture
Layering is a rule, not a preference: core.py has no web, MCP or CLI code, and every adapter goes through pipeline.py or core.py.
Terminalrepoguard analyze · fix · gateWebNext.js on Vercel → FastAPI on Cloud Run (HTTP, SSE, NDJSON)MCP clientrepoguard mcp · 9 tools over stdiopipeline.pyRuns the measurement steps in order, returns one resultwatson_agent/Fix loop: measure → write → critique → re-measurecore.pypytest, coverage, AST mutation, gaps, riskapi_check.pyFastAPI routes by AST, endpoint smoke testsvisual.pyScreenshots, pixel diff, console, accessibilityai_providers/watsonx.ai · Vertex AI, behind one ChatProviderrepoguard-out/*.jsonThe contract between agents: every step writes its result hereTarget repoRead for measurement; agents may write under tests/ onlyAI components. Everything else is deterministic.
Mutation engine
Written on the standard library's ast module (no mutmut or Stryker). It changes one node at a time, reruns the suite on a copy, and counts the mutants no test notices.
| Operator | Change | Catches missing tests for |
|---|---|---|
| Comparison | == ↔ !=, < ↔ >=, > ↔ <= | Off-by-one and boundary bugs |
| Arithmetic | + ↔ -, * ↔ / | Wrong formula, sign errors |
| Boolean | and ↔ or | Wrong condition logic |
| Constants | True ↔ False, n → n + 1 | Magic numbers, flags |
| Return | return x → return None | Results nobody asserts on |
| Raise | raise … removed | Error paths with no test |
Guardrails that keep the numbers honest
Baseline must pass first
Before mutating anything, the unmodified suite has to pass. Otherwise every mutant would count as killed and a broken suite would score 100%.
Same repo in, same numbers out
Each mutant runs on a temporary copy with PYTHONDONTWRITEBYTECODE=1, so stale bytecode can’t leak between runs. Two runs on demo-repo give 16/79 both times.
Pinned interpreter
Every pytest subprocess runs through sys.executable, never a bare pytest on PATH that could resolve to a different environment.
Write guard
The fix loop’s only write tool, write_test_file, rejects any path outside tests/. Autofix over HTTP also runs on a copy of the repo, so the original is never touched.
Fail loud, never fake
A check that couldn't run returns ok=false with the real error (missing credentials, missing dependency) instead of an empty result that looks like a pass.
AI fix loop
repoguard fix and the Autofix button run the same loop, in-process against the engine rather than through MCP.
- Measure. Full pipeline with mutation and endpoints.
- Prioritize. Up to 3 files, by the engine’s risk score.
- Write. The test-writer agent reads the source and writes a test through the guarded tool, and must run pytest before it finishes.
- Critique. A second agent with its own prompt reviews and corrects the test.
- Re-measure. The engine measures again; the report goes to
watson-evidence/.
| Provider | Select with | Status |
|---|---|---|
| IBM watsonx.ai | REPOGUARD_AI_PROVIDER=watsonx (default) | Live-verifiedThe default model calls tools unreliably |
| Google Vertex AI | REPOGUARD_AI_PROVIDER=vertex | Live-verifiedGemini 3: 20.25% → 89.87% on demo-repo |
HTTP API
What this dashboard calls. The backend is repoguard_engine/web/server.py.
| Method | Path | Does |
|---|---|---|
| GET | /api/analyze | Runs the pipeline: coverage, gaps, risk, optional mutation and gate |
| GET | /api/stream | Server-sent events with live progress for the same run |
| POST | /api/summary | Advisory AI prose over an already-measured result, plus the provider used |
| POST | /api/fix | Autofix: rate-limited (per-IP + global), one run at a time, NDJSON progress and engine before/after |
| GET | /api/projects/{slug}/… | Run history for the Results charts: trend, operators, risk, fix effect, survivors, flaky Phase 17 |
Deployment and CI
Frontend
web-next/ on Vercel; every merge to main redeploys. The backend URL is baked in at build time through NEXT_PUBLIC_REPOGUARD_API_BASE.
Backend
Docker image on Google Cloud Run, pushed by cd.yml through Workload Identity Federation (no stored keys). The identity is managed in Terraform.
Checks
ci.yml re-measures demo-repo and checks determinism; frontend-ci.yml lints and builds this app; infra-ci.yml validates Terraform.
Stack
Deeper reading: ARCHITECTURE.md · ARCHITECTURE-front.md · MULTICLOUD_AI.md