Learn · Pentesting concepts

Security Regression Testing: Catching Vulnerabilities That Come Back

Short answer: Security regression testing checks that vulnerabilities you already fixed stay fixed. The same pentest runs again on a schedule and each run's findings are compared with the last run's. A bug that a reverted commit, copied code or a new endpoint brought back is then labelled regressed, instead of getting lost among new findings.

What is a security regression?

A security regression is a vulnerability that was found, fixed, confirmed as fixed, and then found again later. The code moved on and the fix didn't come with it.

Security regression testing is repeated security testing of the same app that checks whether vulnerabilities fixed earlier have come back. Each run is compared with the one before, and a finding that was fixed and is then found again is marked as regressed.

It's the security version of a normal software regression, with one difference. A broken feature gets noticed by users within hours. A broken access check usually doesn't get noticed by anyone until someone goes looking, and that someone may not be on your side.

Why do fixed vulnerabilities come back?

Almost never because someone decided to undo the fix. They come back as a side effect of ordinary work.

CauseHow it happensTypical example
Reverted commitA release is rolled back, or a merge conflict is resolved in favour of the older branch, and the fix goes with itThe ownership check added to an order endpoint disappears when a hotfix branch is merged back
Copied codeSomeone copies an old handler or query as a starting point, from before the fixA new export feature is built from the pre-fix version of a search query, string concatenation included
New endpoint, old patternA new route is written the way the team used to write them/api/v2/orders/{id} gets the ownership check, and the new /api/v3/orders/{id} doesn't
Dependency downgradeA lockfile is regenerated, a version is pinned back to work around a bug, or a transitive dependency resolves to an older releaseA library version that fixed a template-escaping bug is replaced by the one before it
Config driftA new environment, a rebuilt gateway or a reset feature flag loses a setting that was part of the fixA rate limit set on the old load balancer isn't carried over to the new one

The last two are the hard ones, because the application code didn't change at all. A code review of the diff shows nothing. Only testing the running app does.

Why doesn't a one-off retest catch it?

A retest answers one question: is this fix in place today? That's worth knowing, and PCI DSS v4.0.1 asks for exactly that in Requirement 11.4.4, where exploitable findings are corrected and testing is repeated to verify the corrections.

But a regression happens after the retest. The retest passes in March, the hotfix merge in May takes the fix out, and nothing tests that endpoint again until next year's pentest. A retest is a point in time, and a regression needs a series.

Two more reasons a single retest misses regressions:

  • It only looks at the old finding. A retest checks the exact endpoint that was reported. The same bug on a new endpoint (the v3 route above) is a new finding to a retest, if it's looked at at all.
  • Nobody compares reports. Two PDFs from two different engagements, written by different testers with different titles for the same bug, don't line up by themselves. Someone has to do that by hand, and usually nobody does.

How does run-over-run comparison work?

The same pentest runs again, with the same scope, depth and test users, and every finding gets a stable identity so it can be matched between runs. Then each finding in the new run is compared with what the tool already knows.

LabelRuleWhat to do
NewNot seen in any earlier runTriage it like any fresh finding
Still openOpen in the last run and found againCheck that it's in someone's queue, and how long it's been there
FixedOpen in the last run, not found in this oneThe fix held for this run. Leave it in the history
RegressedFixed in an earlier run and found again nowTreat it as urgent. A fix came undone, so find out why

Some tools and reports call the third label "resolved". Barrion's dashboard says "fixed", so that's what the example below uses.

The matching is the part to ask vendors about. If a finding's identity depends on its title, a slightly different wording in the next run turns a regression into a "new" finding and the signal is gone. A stable key built from what the finding is (the endpoint, the parameter, the class of bug) holds up better.

Example: a run-over-run report

Here's what a weekly run looks like when it's compared with the week before. The app, dates and findings are made up, but the layout and the labels are the ones a scheduled Barrion run uses.

Run #14 vs run #13 · weekly scheduleExample
app.example.comStandard · run 2026-09-21
1 new1 still open1 fixed1 regressed
RegressedHigh
IDOR on GET /api/v2/orders/{id}
WSTG
WSTG-ATHZ-04
CWE
CWE-639
First seen
2026-08-10 (run #8)
Last seen
2026-09-21 (run #14)
Fixed in run #11, found again in this run. The test user with the Customer role read another customer's order with its own token. Request and response attached.
NewMedium
Reflected XSS in the q parameter on /search
WSTG
WSTG-INPV-01
CWE
CWE-79
First seen
2026-09-21 (run #14)
Last seen
2026-09-21 (run #14)
First seen in this run. The payload came back unencoded in the HTML response. Request and response attached.
Still openMedium
No rate limit on POST /login
WSTG
WSTG-ATHN-03
CWE
CWE-307
First seen
2026-08-17 (run #9)
Last seen
2026-09-21 (run #14)
Open since run #9. 50 failed logins in a row on the test account were neither slowed nor blocked.
FixedCritical
SQL injection in order lookup, GET /api/v2/orders?ref=
WSTG
WSTG-INPV-05
CWE
CWE-89
First seen
2026-08-10 (run #8)
Last seen
2026-09-14 (run #13)
Not found in this run. The same payloads now get a 400 response.
Example only. Made-up app, dates and findings, laid out the way a scheduled Barrion run compares its findings with the run before.

How to read it

  • Start with Regressed. The IDOR on /api/v2/orders/{id} was fixed in run #11 and is back in run #14, so something shipped between 14 and 21 September took the fix out. Check what merged in that week. That's usually faster than debugging the endpoint from scratch.
  • New is a normal finding. The reflected XSS on /search wasn't there before. Triage it like any other.
  • Still open shows age. "First seen" on the missing rate limit says it has been open since 17 August, six runs in a row. That's a prioritisation question for the team, and the date makes it visible.
  • Fixed means not found this time. The SQL injection didn't show up in run #14. "Last seen" records the last run that did find it, so if it comes back, you know exactly which window to look in.
  • WSTG and CWE say what kind of bug each one is. The WSTG ID points to the OWASP test case that covers it, and the CWE is the weakness class your tracker or a customer's questionnaire may ask for.

How to stop regressions as well as catch them

Testing finds a regression. It doesn't stop the next one. NIST's Secure Software Development Framework (SP 800-218, practice RV.3) asks teams to analyze the root causes of vulnerabilities so the same kind doesn't recur. In practice that means a few habits:

  1. Write a test for each fix. When you fix a finding, add an automated test to your own suite that fails if the bug returns, such as "user B gets 403 on user A's order". A reverted commit then breaks the build.
  2. Fix the pattern as well as the endpoint. If one handler lacked an ownership check, look for its siblings. Put the check in shared middleware so new routes get it by default.
  3. Review dependency changes. Treat a lockfile diff that downgrades a package like any other change and review it.
  4. Keep config in code. A rate limit that lives in a console setting is easy to lose in a rebuild. One in the repository comes along with everything else.
  5. Keep testing the running app. Tests in your suite cover the cases you thought of. Repeated pentests of the live app catch the ones you didn't, including the config and dependency changes a code review can't see.

How often should regression testing run?

As often as the app changes in ways that could undo a fix. For an app that deploys weekly, a weekly run is the natural rhythm. Teams that ship daily often pair a lighter run that goes only when the app has changed with a deeper monthly run that always goes. The trade-offs are in how often to pentest, and the wider model is covered in what continuous pentesting is.

What regression testing doesn't tell you

"Fixed" means a run didn't find the bug at its depth. It doesn't prove the bug can never be reached. And a regression can only be detected in something the test reaches. If a route is out of scope, or sits behind a role no test user has, a regression there goes unseen. Keep the scope and the test users up to date as the app grows.

How Barrion does it

Barrion's AI agents run full pentests of web apps and APIs across 8 testing areas, covering all 97 OWASP WSTG v4.2 test cases and the OWASP API Security Top 10. You can save a pentest and run it again on a schedule or on demand. You can choose whether a scheduled run goes every time, or only when your app has changed. You can also start a run from your CI/CD pipeline via the Barrion API, so a release gets tested soon after it ships.

  • Every run is compared with the previous one, and each finding is labelled new, still open, fixed or regressed. You can mark a finding as ignored and give a reason, so a colleague sees why.
  • New and regressed findings trigger an email.
  • Findings are checked against the live app before they're reported. Confirmed ones come with the request and response behind them. Anything we couldn't confirm is clearly marked and capped in severity.
  • Retests of found issues are free. From Standard up, a security engineer reviews the findings, and deeper tests come with a report signed off by that engineer.
  • Data is stored and hosted in Sweden, and AI processing runs in the EU.

See the schedule and the run comparison on continuous AI pentesting. Continuous programs are scoped to your apps, cadence and depth. Talk to us and we'll price it for your setup.

Sources

All sources checked 2026-09-26.

FAQ

Frequently asked questions

What is a regressed vulnerability?
A vulnerability that an earlier test found, a later test showed as fixed, and a newer test found again. It usually means a fix was reverted in a merge, code was copied from before the fix, a new endpoint reused an old pattern, or a dependency or config change undid it.
Is security regression testing the same as a retest?
No. A retest checks once that a specific fix is in place. Regression testing keeps checking on later runs, so a fix that comes undone after the retest is caught. A retest is a point in time. Regression testing needs a series of runs that are compared with each other.
How is a finding matched between two runs?
Each finding needs a stable identity that doesn't depend on how the report words it, for example the endpoint, the parameter and the class of bug. At Barrion each finding gets a stable key, and each run's findings are compared with the open and fixed keys from earlier runs.
What's the difference between fixed and resolved?
Nothing, in practice. Both mean a finding that was open in the last run wasn't found in this one. Barrion's dashboard uses fixed, and some of our pages and other tools say resolved.
Can unit tests replace security regression testing?
They're the first line. Add a test for every fix so a reverted commit breaks the build. But your tests only cover the cases you wrote down. They don't see a dependency downgrade, a lost gateway setting or the same bug on a new endpoint. Repeated pentests of the running app catch those.
Which Barrion plan includes regression tracking?
Scheduled pentests, and with them the run-over-run labels, are part of the Business plan. Continuous programs are scoped to your apps, cadence and depth, so talk to us at barrion.io/contact-sales and we'll price it for your setup.

Catch the fixes that come undone.

Put a pentest on a schedule and every run is compared with the last. Tell us about your apps and we'll scope a continuous program with you.