The first regression my pentest retest script caught was itself.
In April I swapped the zero-trust proxy in front of my personal finance app from one vendor to another. No application code changed. The retest did not care about application code. It decided whether a reflected CORS origin was accepted proxy behavior or a real application bug by looking for one vendor's response header, and that header stopped existing. The oracle was infrastructure-shaped, and the infrastructure moved out from under it.
That is the honest version of what a personal-scale security regression harness is. It is not a scanner. It is a small set of assertions about a system I already assessed once, and it decays at exactly the rate the system changes.
The application is Bevis: a private personal-finance and portfolio-intelligence app I built and maintain solo. The script is pentest-retest.sh. It replays the retest plan from a March 2026 assessment against the deployed app, covering eleven findings, each with a vulnerability class, a test ID, an exploit condition, and in most cases a CVSS v3.1 score. I run the same shape of assessment against my own code that I would run against any application I was asked to validate professionally. That is not extra work. It is the same craft pointed at a target with no client and no compliance driver.
The gap between professional security rigor and personal project shortcuts creates a specific blind spot. Practitioners who build enterprise security programs but skip the discipline on their own code miss the most direct feedback loop available. The one where you write the code, find the flaw, and fix the pattern. No client. No compliance driver. No one demanding the artifact. That absence of external accountability is the point, not the problem.
A fixed baseline replayed against a live deployment catches two things reliably. First, it confirms that fixed findings stay fixed: the script re-runs the original condition against the original surface and fails if the behavior returns. Second, it catches drift below the application. The platform and edge layers change on their own schedule, and a control verified once does not stay verified.[2] It also does a third thing I did not design for, which is fail loudly when the environment changes enough to invalidate the test itself. That turns out to be the most useful signal it produces.
The baseline misses more than it catches, and that matters. Business logic vulnerabilities have no recognizable HTTP-layer signature, so a harness that checks known conditions does not surface them unless someone writes an assertion for the specific behavior. Bevis handles portfolio state, per-record ownership across accounts, and precision-sensitive financial calculation. Testing any of those means writing assertions against intended behavior, such as one account never reading another's records, and I have not written them. Multi-step attack chains sit outside it for a related reason: individual requests appear benign, and the vulnerability lives in the sequence. This harness sends single requests and grades each response on its own; it never performs a sequence. Authentication-context flaws, such as whether a session that should have been invalidated remains usable, require session-state reasoning that this harness does not perform. Dependency-introduced vulnerabilities are also outside this script's reach: it verifies findings, it does not scan a dependency tree.
Cadence is the part I have not solved. Today the script runs when I run it: after a deployment group, after an infrastructure change, when I remember. Its header comment says it is designed for cron, and nothing schedules it. Weekly is the cadence I want, because weekly is close to the change rate of a solo project. Narrow enough that a failure is attributable to a specific week's work, wide enough that I do not start ignoring the output. Until it is wired to a scheduler, calling this a weekly control would describe intent rather than operation, and the distinction is the whole point of writing the results down.
The most recent committed report, from September 19, is a fair picture of what the harness produces, failure modes included. Twenty-two checks: fifteen passed, four skipped, and three failed. Three of the skips need access the script does not have from outside the identity layer; the fourth reviews a source file I had deleted with dead code in April. The three failures were the security-header checks, and every one of them was the oracle problem again. The script read response headers off whatever answered first, and what answers first for an unauthenticated request is the identity layer's redirect, not the application. The application sets all three headers; the harness was grading the wrong response. That was a real finding about the harness and a non-finding about the application, and the JSON report could not tell the difference on its own. Within a day I changed the harness instead of learning to read around it. The header checks now identify which layer answered before reading a single header, and skip with that layer named when it is not the application. A second change added an opt-in probe that authenticates as a narrowly scoped service identity with a short-lived token and reaches the deployed application behind the edge. The application refuses that caller by design, and its refusal carries its own headers. The verification run recorded on that change graded all three checks on the application's response: twenty-four checks, twenty passed, none failed, four skipped. It proves the deployed revision sets the headers; it does not prove the edge passes them through to a browser, and the evidence string names the response that was graded. A new check came with it: if that probe ever gets the application's page back instead of a refusal, the run fails. The harness found its own oracle problem, and the fix put the answer in the report instead of in my head.
Treat pentest-retest.sh as a guard rail, not assurance. Run periodic manual review of business logic paths, specifically state transitions and any operation involving financial calculation. Expand the baseline after each incident and each engagement cycle, adding new checks as the application evolves.
The Security Baseline as a Regression Instrument
A security baseline, in the context of regression testing, is a documented point-in-time snapshot of known vulnerability state. Without a baseline, each scan produces a list of findings with no reference point: you cannot determine whether a result represents a new vulnerability, a persisting old one, or a previously accepted risk.
The baseline in pentest-retest.sh is eleven findings from a single assessment, plus one monitoring probe added later. Each finding carries a vulnerability type, a test ID, the surface it appeared on, the exploit condition, and a CVSS v3.1 severity. The script does not scan Bevis generally. It probes the specific conditions that allowed each original finding to exist. This is the structural difference between continuous scanning and targeted regression: the former generates discovery, the latter generates verification.
The mechanics are deliberately unglamorous. The script is bash and curl. It sends requests with attacker-controlled origin headers and forged identity headers, inspects response headers and status codes, and classifies each result as pass, fail, or skip with an evidence string. Several checks are not network tests at all: they grep the working tree for patterns the assessment flagged, such as raw exception text being returned in an HTTP response rather than logged. One check opens a TLS connection to the origin and reads the certificate's expiry and subject alternative name. Every run writes a timestamped JSON report with per-test status and evidence, and exits non-zero if anything failed. That JSON is the artifact that makes two runs comparable.
Skips are a first-class result, not a failure to run. A large share of the baseline sits behind authentication, and the script tests from outside the zero-trust proxy. It cannot reach those endpoints as an authenticated user, so it records a skip with the reason rather than a false pass. Three of the twenty-two checks in the September 19 run landed there, and a fourth skipped because the code it reviews no longer exists. A harness that reported those as green would be worse than no harness.
Baseline selection follows professional practice in two ways. First, professional pentest retesting scopes the retest to previously reported findings rather than repeating the full engagement. Second, OWASP ASVS 5.0 Level 1 is the first-layer set of requirements every internet-exposed application should meet.[3] A baseline tracking authentication, session, authorization, error-handling, and configuration findings touches the Level 1 chapters; it does not cover them.
The categories most suited to automated baseline verification share a common property: deterministic, observable outputs at the HTTP response layer. Origin reflection, identity header handling, security header presence, information disclosure on unauthenticated surfaces, and input validation all fit this profile. Automated tools surface and retest these classes reliably because the signature is recognizable without application-context reasoning.
What the Automated Retest Catches
Automated retesting of a defined baseline reliably surfaces three problem categories.
Regression of fixed findings is the primary function. After remediating a finding, the retest replays the exact condition that triggered it and checks that the response no longer indicates the vulnerable behavior. The original finding provides the request, the response indicator, and the conditions that confirmed exploitability. Replaying that sequence against a fixed application produces a pass or fail result with high confidence.[4]
Configuration drift is a secondary benefit, and in practice it has been the primary one. NIST's configuration management guidance states the case plainly: planning and implementing secure configurations, then controlling change, is usually not sufficient to ensure that a system which was once secure remains secure, and monitoring is what identifies misconfigurations and unauthorized changes.[2] On a Cloud Run service the surfaces that drift are the ones nobody edits deliberately: the identity layer at the edge, the domain mapping and its certificate, response headers injected somewhere between the container and the browser. A check that probes for missing security headers or an unauthenticated surface returning more than it should catches these regressions regardless of whether the developer touched application logic.
Assumption failures in the harness itself are the category I did not plan for. When the proxy layer changed vendors, the checks that classified accepted-risk behavior by a vendor-specific response header stopped classifying it correctly. The fix was to detect the proxy layer generically, by any of the two vendors' marker headers or by an unauthenticated redirect status, to record which marker matched in the evidence string, and to accept that the vendor list will grow again. The lesson generalizes: any assertion that encodes a fact about the platform is a dependency on that platform, and it should fail visibly when the fact changes rather than quietly return the wrong verdict.
What the script does not catch is dependency-introduced vulnerability. It has no software composition analysis. Its only dependency-aware check confirms that a rate-limiting package is still declared, which verifies a remediation decision rather than a vulnerability. Dependency scanning is a separate control and should not be claimed by this one.
Beyond raw detection, consistent execution of pentest-retest.sh creates a structured feedback loop. A failure that narrows the search to one week's changes is a concrete, attributable signal. Early detection matters for cost: a vulnerability caught during development carries less remediation effort, and less exposure, than one discovered after deployment.[4]
What Automated Retest Cannot Reach
Automated regression against a fixed baseline has structural gaps that no tuning resolves.
Business logic vulnerabilities are the most significant gap. This harness, like a scanner, works from known patterns: it sends known-bad inputs and evaluates responses against known-bad indicators. Business logic flaws exploit the gap between what an application does and what it should do, not between what it does and a known attack pattern.[5] For Bevis, the business logic attack surface covers portfolio state transitions, precision handling in financial calculation, race conditions in concurrent writes, and ownership validation in a multi-account context. None of these has a signature this harness could match. Each one needs an application-specific assertion written against intended behavior, or a manual investigation that starts from the design.
Multi-step attack chains are out of reach for a related reason. This harness tests each endpoint in isolation and grades each response on its own. A chained attack, where two individually acceptable operations produce an exploitable outcome, shows up only in a test that performs the sequence and asserts the end state, and someone has to know the chain to write that test. In my workflow, finding the chains I have not written tests for is still manual investigation. Professional pentesting methodology calls out attack chain analysis as requiring manual engagement because no single request appears malicious; the vulnerability lives in the sequence.[6]
Authentication-context vulnerabilities require session-state reasoning. A script can confirm that an unauthenticated request to a protected surface is rejected. This harness does not reason about whether an authenticated session that should have been invalidated, after a state change or an account status change, remains valid due to improper session lifecycle management.
Coverage gaps from the vantage point are structural here rather than incidental. Testing from outside the identity layer means the script proves the perimeter holds and proves very little about what happens behind it. Input validation, authorization, and rate limiting are all better tested as an authenticated user, and none of them are, today. The honest reading of a green run is that the perimeter behaved and the code-review checks passed, not that the application is sound.
False negatives are a structural feature of this class of testing, not an edge case. Sources include insufficient coverage of the request surface, rate limiting that interrupts an active scan, input filtering that blocks obvious payloads while remaining bypassable by subtler ones, and context-dependent behavior that requires specific application state to manifest. A fixed baseline mitigates this by focusing on known paths, but it cannot self-expand to surface vulnerability classes outside its scope.
The right mental model for pentest-retest.sh is "regression guard," not "security assurance." It provides confidence that specific known findings behave as the remediation intended. It provides no assurance about business logic correctness, novel attack paths, or anything outside the baseline scope.
The Behavior Change Question
The most interesting question here is not technical. It is behavioral: does maintaining this loop actually change how I write code?
Two hypotheses apply. The accountability hypothesis holds that the script functions as an external constraint. A defined verification event, even a self-imposed one, means code written this week exists inside a window where it will be checked. The habit formation hypothesis holds that repeated exposure to security findings from personal code builds internalized security reasoning that applies before anything runs.
Research on the security mindset gives both a vocabulary, with nuance.[1] The study identifies three cognitive components: an unconscious monitoring habit, a conscious investigation process, and a contextual severity-evaluation capacity. The critical finding is that the monitoring component, the most valuable precursor to preventive coding behavior, is reinforced by the rewarding experience of finding and resolving actual security flaws in your own systems. A retest loop over your own code is my bet on that reinforcement cycle; the study did not test one.
The accountability mechanism works at the commit level. A failure from pentest-retest.sh narrows the search to whatever changed since the last passing run, which on a solo project is usually one deployment group. DevSecOps practitioners consistently emphasize visibility as a key driver of security behavior change: when developers see the direct relationship between their code and security outcomes, they adopt proactive practices more readily.[7]
The deeper claim, and it is still the bet rather than a finding, is that operating this loop over time internalizes security reasoning into the development process itself. The highest-severity finding in this baseline was a CORS configuration that reflected arbitrary origins while allowing credentials. Having watched an arbitrary attacker origin come back in a response header on my own application, with my own session semantics behind it, produced a constraint on every cross-origin decision I have made since. That is distinct from reading about CORS in a training module. The solo developer applying professional discipline to personal code is not responding to peer pressure, compliance mandates, or managerial accountability, the conventional levers of organizational security culture. The motivation is intrinsic, and my wager is that intrinsic motivation is the more durable kind because nothing external has to keep enforcing it. One developer's experience does not test that.
Even if the bet pays off, the benefit has a ceiling. This harness gives feedback only on the classes in its baseline, so whatever attention it builds will most likely cluster there. I have no evidence that watching a CORS check go red makes me more careful about ownership checks or calculation precision, and I do not assume it does. Business logic security gets its attention from deliberate design-time review and from the application-specific assertions I have not yet written. Outside the baseline, behavior depends on developer judgment.
False assurance risk compounds the ceiling effect. A clean run can create overconfidence, treating a passing retest as broader security validation than the scope warrants. This is a documented failure mode: teams mark issues as resolved by assumption rather than proof, and a clean result substitutes for threat modeling or design-level review. The scoring line in the script's own summary is deliberately worded as "all testable findings pass or are accepted risk," because "testable" and "accepted" are doing most of the work in that sentence.
Where the Script Sits in the Engagement
The baseline did not come from nowhere, and the validity of a retest depends entirely on how the findings it replays were originally discovered.
Bevis was assessed once, in March 2026, and the assessment was written up into the artifact set a structured engagement produces: scope, rules of engagement, recon notes, attack surface, threat model, a vulnerability register with CVSS scoring, a technical report, an executive report, a hardening plan, and a retest plan. pentest-retest.sh is the executable form of that last document. Every check in the script traces to a numbered test case in the retest plan, which traces to a numbered finding in the register.
The structure follows PTES at the bookends, from pre-engagement through post-exploitation and reporting,[8] and the findings are classified against the OWASP Web Security Testing Guide taxonomy. WSTG's stable release, v4.2, organizes tests into eleven categories: information gathering, configuration and deployment management, identity management, authentication, authorization, session management, input validation, error handling, weak cryptography, business logic, and client-side testing.[9] API testing is a twelfth category in the v5 draft, not in the stable release. Using a published taxonomy rather than ad hoc labels is what makes a personal engagement's findings comparable to anything else, including a later engagement against the same app.
I want to be precise about what is built and what is designed. The script is real and runs. The artifact set is real and versioned in the application repository. What does not exist yet is the thing that would make the front half of the engagement repeatable: a set of agent skills covering planning, recon, threat modeling, vulnerability analysis, exploitation, reporting, and post-engagement hardening, with human approval gates at scope, test plan, exploitation, and report. That is a design on paper and an open issue in my skills repository, not running software. The retest is the tail of a methodology that is currently executed by hand.
The approval gates are worth describing even as design, because they encode the part of the methodology that should not be automated. Scope, test plan, exploitation, and report are the four decision points where an automated system acting alone produces either a legal problem or a useless artifact. Even in a personal project where authorization is self-granted, the discipline of conscious scope definition matters. The retest script, operating without gates, is defensible precisely because it scopes itself to bounded verification of known findings against a target the existing baseline already authorizes.
OWASP ASVS 5.0 Level 1 is the benchmark I measure against, and the harness covers a fraction of it.[3] The 4.0 edition described Level 1 as verifiable by black-box testing; 5.0 dropped that framing and says meaningful verification needs access to the internals, which is what the skip column already shows. Level 2, which covers applications handling sensitive data such as this one, requires manual testing and code review on top. The standards framing gives the project a meaningful benchmark that is independent of any compliance requirement.
Comparing Approaches to Personal Application Security
Table 1 is a qualitative comparison of the retest approach against four alternatives for a solo-developed personal application with active development, on catch rate for known vulnerabilities, catch rate for novel vulnerabilities, operational cost, and behavior change potential.
| Approach | Known Vuln Catch Rate | Novel Vuln Catch Rate | Operational Cost | Behavior Change Potential |
|---|---|---|---|---|
On-demand pentest-retest.sh (current) |
Moderate - depends on remembering to run it | Low - scope-limited | Very low - one command | Low - no fixed verification event |
Scheduled pentest-retest.sh (intended) |
High - in-scope findings | Low - scope-limited | Low - automated, scheduled | Moderate - accountability plus habit loop |
| CI/CD security testing (per-commit) | High - pattern-matched issues | Low - same structural limits | Low to medium - pipeline cost | High - immediate feedback at code change |
| Periodic manual pentest | High - includes business logic | High - attacker reasoning applied | High - time-intensive | Low - too infrequent for habit formation |
| Hybrid: scheduled retest plus manual review | High - known vulns | High - novel vulns | Medium | Highest - both feedback loops active |
No single approach covers the full attack surface at low operational cost. The scheduled retest earns its place as the foundational continuous signal, and the gap between row one and row two in that table is a cron entry I have not written. Manual review, and eventually application-specific assertions, fill the business logic gap that this harness structurally cannot reach.
References
Schoenmakers, K., et al. "The security mindset: characteristics, development, and consequences." Journal of Cybersecurity, 2023. https://academic.oup.com/cybersecurity/article/9/1/tyad010/7147623
NIST. "SP 800-128: Guide for Security-Focused Configuration Management of Information Systems." 2011, updated 2019. https://csrc.nist.gov/pubs/sp/800/128/upd1/final
OWASP. "Application Security Verification Standard (ASVS) 5.0." 2025. https://owasp.org/www-project-application-security-verification-standard/
Mayhem Security. "3 Reasons Your Security Testing Tool Needs To Do Regression Testing." https://www.mayhem.security/blog/3-reasons-your-security-testing-tool-needs-to-do-regression-testing
Precursor Security. "Business logic vulnerabilities: what automated scanners miss in web applications." https://www.precursorsecurity.com/blog/business-logic-vulnerabilities-what-scanners-miss
CyCognito. "DAST vs Manual Pentesting vs Automated Pentesting: 5 Differences." https://www.cycognito.com/learn/application-security/dast-vs-manual-pentesting-vs-automated-pentesting/
Harness. "Continuous Security Monitoring DevSecOps." https://www.harness.io/harness-devops-academy/continuous-security-monitoring-devsecops
KirkpatrickPrice. "Stages of Penetration Testing According to PTES." https://kirkpatrickprice.com/blog/stages-of-penetration-testing-according-to-ptes/
OWASP Foundation. "OWASP Web Security Testing Guide v4.2." https://owasp.org/www-project-web-security-testing-guide/stable/