The first regression my pentest retest script caught was itself.

In April I swapped the zero-trust proxy in front of my personal finance app from one vendor to another. No application code changed. The retest did not care about application code. It decided whether a reflected CORS origin was accepted proxy behavior or a real application bug by looking for one vendor's response header, and that header stopped existing. The oracle was infrastructure-shaped, and the infrastructure moved out from under it.

That is the honest version of what a personal-scale security regression harness is. It is not a scanner. It is a small set of assertions about a system I already assessed once, and it decays at exactly the rate the system changes.

The application is Bevis: a private personal-finance and portfolio-intelligence app I built and maintain solo. The script is pentest-retest.sh. It replays the retest plan from a March 2026 assessment against the deployed app, covering eleven findings, each with a vulnerability class, a test ID, an exploit condition, and in most cases a CVSS v3.1 score. I run the same shape of assessment against my own code that I would run against any application I was asked to validate professionally. That is not extra work. It is the same craft pointed at a target with no client and no compliance driver.

The gap between professional security rigor and personal project shortcuts creates a specific blind spot. Practitioners who build enterprise security programs but skip the discipline on their own code miss the most direct feedback loop available. The one where you write the code, find the flaw, and fix the pattern. No client. No compliance driver. No one demanding the artifact. That absence of external accountability is the point, not the problem.

A fixed baseline replayed against a live deployment catches two things reliably. First, it confirms that fixed findings stay fixed: the script re-runs the original condition against the original surface and fails if the behavior returns. Second, it catches drift below the application. The platform and edge layers change on their own schedule, and a control verified once does not stay verified.[2] It also does a third thing I did not design for, which is fail loudly when the environment changes enough to invalidate the test itself. That turns out to be the most useful signal it produces.

The baseline misses more than it catches, and that matters. Business logic vulnerabilities have no recognizable HTTP-layer signature, so a harness that checks known conditions does not surface them unless someone writes an assertion for the specific behavior. Bevis handles portfolio state, per-record ownership across accounts, and precision-sensitive financial calculation. Testing any of those means writing assertions against intended behavior, such as one account never reading another's records, and I have not written them. Multi-step attack chains sit outside it for a related reason: individual requests appear benign, and the vulnerability lives in the sequence. This harness sends single requests and grades each response on its own; it never performs a sequence. Authentication-context flaws, such as whether a session that should have been invalidated remains usable, require session-state reasoning that this harness does not perform. Dependency-introduced vulnerabilities are also outside this script's reach: it verifies findings, it does not scan a dependency tree.

Cadence is the part I have not solved. Today the script runs when I run it: after a deployment group, after an infrastructure change, when I remember. Its header comment says it is designed for cron, and nothing schedules it. Weekly is the cadence I want, because weekly is close to the change rate of a solo project. Narrow enough that a failure is attributable to a specific week's work, wide enough that I do not start ignoring the output. Until it is wired to a scheduler, calling this a weekly control would describe intent rather than operation, and the distinction is the whole point of writing the results down.

The most recent committed report, from September 19, is a fair picture of what the harness produces, failure modes included. Twenty-two checks: fifteen passed, four skipped, and three failed. Three of the skips need access the script does not have from outside the identity layer; the fourth reviews a source file I had deleted with dead code in April. The three failures were the security-header checks, and every one of them was the oracle problem again. The script read response headers off whatever answered first, and what answers first for an unauthenticated request is the identity layer's redirect, not the application. The application sets all three headers; the harness was grading the wrong response. That was a real finding about the harness and a non-finding about the application, and the JSON report could not tell the difference on its own. Within a day I changed the harness instead of learning to read around it. The header checks now identify which layer answered before reading a single header, and skip with that layer named when it is not the application. A second change added an opt-in probe that authenticates as a narrowly scoped service identity with a short-lived token and reaches the deployed application behind the edge. The application refuses that caller by design, and its refusal carries its own headers. The verification run recorded on that change graded all three checks on the application's response: twenty-four checks, twenty passed, none failed, four skipped. It proves the deployed revision sets the headers; it does not prove the edge passes them through to a browser, and the evidence string names the response that was graded. A new check came with it: if that probe ever gets the application's page back instead of a refusal, the run fails. The harness found its own oracle problem, and the fix put the answer in the report instead of in my head.

Treat pentest-retest.sh as a guard rail, not assurance. Run periodic manual review of business logic paths, specifically state transitions and any operation involving financial calculation. Expand the baseline after each incident and each engagement cycle, adding new checks as the application evolves.