Brian Andrew Mulrooney

2026-09-21

Gap: Mutation tests ran locally for the wrong reasons

The project's full mutation-testing grid ran on the developer's laptop as part of every pre-merge check. It took over two hours, proved little that a one-second check did not already prove, and on the same evening a failing check on the hosted build service went unnoticed.

Gap: The full grid ran on the wrong machine at the wrong time

Mutation testing reintroduces each of the project's 69 historical bugs, one at a time, into a copy of the source and confirms that the test suite still catches it. The full grid is the strongest proof the suite has, and the slowest: 27 minutes on a rested laptop with 14 workers, and 2 hours 19 minutes when the laptop was already busy with other work.

A bar chart on a logarithmic time axis comparing three checks: the per-change drift check at about one second, the full grid at 27 minutes on a rested machine, and the full grid at 2 hours 19 minutes under a loaded session
The per-change check costs about a second; the full grid costs minutes to hours depending on how busy the machine is

The grid had been wired into the pre-merge checklist. Under a time-short merge it ran to completion anyway, and the wait was spent for nothing: the merge was already decided, and a failing result would have been a report to read later, not a reason to stop. The cheaper check that does belong before a merge is a drift check: it applies every mutation pattern to the current source and confirms each still matches a real line. That takes about a second and catches the case where a refactor silently retires a bug's guard.

The fix separates the two. The drift check stays in the per-change gate. The full grid becomes a job on the hosted build service that runs after every merge to the main line, on a self-hosted desktop the project already owns, with a four-hour ceiling and a one-job-at-a-time rule so a second merge queues behind the first instead of interrupting it. The hosted runners were priced first: their two-core machines would take hours per grid and their 14-gigabyte disks cannot hold the per-mutation copies of the source, at about 1.1 gigabytes each. The desktop runs the grid in one job at no cost.

A failing grid on that job is a report. The developer opens a fix when there is time; it never blocks a merge. The results ride as a downloadable run artifact rather than a commit, so the build service never writes to the repository.

Gap: A failing hosted check was missed at merge time

The same evening, the hosted build service reported a failure on the merge and nobody read it, because attention was on the local grid. The rule now is that the hosted check is the final read before a merge, and the local grid is run only when the developer asks for it.

At the time of writing, the self-hosted machine was registered and the job was in place, but the first full run after a merge had not yet been reviewed.

‹›