The project

Coverage tells you which lines ran. TestMind AI checks whether your tests would notice a bug.

It injects bugs into a Python repo on purpose, measures how many the test suite catches, and then has two AI agents write the tests that are missing. The AI decides what to test; the engine does all the measuring.

The gap coverage hides

The bundled demo-repo is a small shop app whose tests run most of the code but assert very little, on purpose.

Bugs caught (mutation score)20.25%→89.87%16 / 79 → 71 / 79 mutants killed
Line coverage60.3%→100%What most dashboards stop at
Tests5→71Passing, against unmodified source
API endpoints with a test1 of 7→7 of 7Found by parsing FastAPI routes

Measured by the engine on demo-repo. Before: its 5 deliberately weak tests, 60.3% coverage but only about 1 bug in 5 caught. After: the reference tests in docs/expected-after-tests/. A live run of the fix loop on Gemini 3 reached the same 89.87% (71/79). The 8 mutants still alive are equivalent: changes no test can observe.

What it does

Mutation testing

Injects one small bug at a time (flipped comparisons, swapped operators, changed constants, return None, removed raise) with its own AST engine and reruns the suite. A mutant that survives is a bug the tests would miss.

Writes the missing tests

A test-writer agent and a critic agent take the riskiest files, one at a time, and write tests through a tool that can only write under tests/. The engine then re-measures.

API checks

Finds a FastAPI app's routes by parsing its source, flags endpoints no test calls, and smoke-tests GET endpoints for 5xx errors.

Visual regression

Screenshots a running page with Playwright and diffs it pixel by pixel against a baseline, plus console errors and basic accessibility checks.

Risk ranking

Ranks files by the share of their lines no test runs, so the fix loop and the reader both start where the suite is weakest.

Quality gate

Pass/fail against a coverage threshold you choose, from the web UI, the CLI (repoguard gate) or CI.

How you use it

Four screens, in the order you'd use them. The first three work today on the public demo.

  1. Analyze

    Live

    Point it at a repo on the backend's machine. It runs the suite, coverage and, optionally, mutation testing.

    Open Analyze →
  2. Read the gaps

    Live

    Coverage next to mutation score, uncovered lines per file, the risk ranking and the gate result, on the same screen as the run.

  3. Autofix

    Token-gated

    From Analyze, with a token: the two agents write tests on a temporary copy of the repo, and you get the engine's before/after and the files they wrote.

  4. Track results

    Phase 17

    Stored runs per project: coverage vs. mutation over time, which bug kinds survive, and what each fix run changed.

    Open Results →

Three ways in

One engine behind all of them, so the numbers match wherever you read them.

Terminal

Measure, gate a CI job, or run the fix loop.

repoguard analyze ./demo-repo --mutation
repoguard gate ./demo-repo --threshold 80
repoguard fix ./demo-repo

Web

This dashboard on Vercel, calling the FastAPI backend on Cloud Run.

repoguard serve   # backend on :8000
npm run dev       # web-next on :3000

Any MCP client

Nine tools over stdio, compact responses by default, for Claude Code or any other agent.

repoguard mcp

What it will and won't do

The AI decides, the engine measures.No number on any screen comes from a model. Coverage, mutation score, risk and endpoint counts are engine output; AI text is labeled advisory.
Source code is the reference.The agents can only write under tests/. A new test that fails against the real code is treated as a wrong test, not a reason to change the code.