Technical

One deterministic engine, three thin adapters.

The engine measures and writes JSON; the CLI, the web API and the MCP server only call it. The AI layer sits beside the engine, never inside a measurement.

Architecture

Layering is a rule, not a preference: core.py has no web, MCP or CLI code, and every adapter goes through pipeline.py or core.py.

AI components. Everything else is deterministic.

Mutation engine

Written on the standard library's ast module (no mutmut or Stryker). It changes one node at a time, reruns the suite on a copy, and counts the mutants no test notices.

OperatorChangeCatches missing tests for
Comparison== ↔ !=, < ↔ >=, > ↔ <=Off-by-one and boundary bugs
Arithmetic+ ↔ -, * ↔ /Wrong formula, sign errors
Booleanand ↔ orWrong condition logic
ConstantsTrue ↔ False, n → n + 1Magic numbers, flags
Returnreturn x → return NoneResults nobody asserts on
Raiseraise … removedError paths with no test

Guardrails that keep the numbers honest

Baseline must pass first

Before mutating anything, the unmodified suite has to pass. Otherwise every mutant would count as killed and a broken suite would score 100%.

Same repo in, same numbers out

Each mutant runs on a temporary copy with PYTHONDONTWRITEBYTECODE=1, so stale bytecode can’t leak between runs. Two runs on demo-repo give 16/79 both times.

Pinned interpreter

Every pytest subprocess runs through sys.executable, never a bare pytest on PATH that could resolve to a different environment.

Write guard

The fix loop’s only write tool, write_test_file, rejects any path outside tests/. Autofix over HTTP also runs on a copy of the repo, so the original is never touched.

Fail loud, never fake

A check that couldn't run returns ok=false with the real error (missing credentials, missing dependency) instead of an empty result that looks like a pass.

AI fix loop

repoguard fix and the Autofix button run the same loop, in-process against the engine rather than through MCP.

  1. Measure. Full pipeline with mutation and endpoints.
  2. Prioritize. Up to 3 files, by the engine’s risk score.
  3. Write. The test-writer agent reads the source and writes a test through the guarded tool, and must run pytest before it finishes.
  4. Critique. A second agent with its own prompt reviews and corrects the test.
  5. Re-measure. The engine measures again; the report goes to watson-evidence/.
ProviderSelect withStatus
IBM watsonx.aiREPOGUARD_AI_PROVIDER=watsonx (default)Live-verifiedThe default model calls tools unreliably
Google Vertex AIREPOGUARD_AI_PROVIDER=vertexLive-verifiedGemini 3: 20.25% → 89.87% on demo-repo

HTTP API

What this dashboard calls. The backend is repoguard_engine/web/server.py.

MethodPathDoes
GET/api/analyzeRuns the pipeline: coverage, gaps, risk, optional mutation and gate
GET/api/streamServer-sent events with live progress for the same run
POST/api/summaryAdvisory AI prose over an already-measured result, plus the provider used
POST/api/fixAutofix: rate-limited (per-IP + global), one run at a time, NDJSON progress and engine before/after
GET/api/projects/{slug}/…Run history for the Results charts: trend, operators, risk, fix effect, survivors, flaky Phase 17

Deployment and CI

Frontend

web-next/ on Vercel; every merge to main redeploys. The backend URL is baked in at build time through NEXT_PUBLIC_REPOGUARD_API_BASE.

Backend

Docker image on Google Cloud Run, pushed by cd.yml through Workload Identity Federation (no stored keys). The identity is managed in Terraform.

Checks

ci.yml re-measures demo-repo and checks determinism; frontend-ci.yml lints and builds this app; infra-ci.yml validates Terraform.

Stack

Python ≥ 3.10pytest + coverage.pyast own mutation engineFastAPI + uvicorn, SSEMCP Python SDK, stdioPlaywright + Pillow, axewatsonx.ai default providerVertex AI GeminiNext.js 16 · React 19Docker Cloud RunGitHub Actions CI/CDTerraform WIF identity

Deeper reading: ARCHITECTURE.md · ARCHITECTURE-front.md · MULTICLOUD_AI.md