Always-on AI pentesting for your web apps and APIsAlways-on AI pentestingStart an AI pentest
Penetration testing

Pentesting in CI/CD: Run an AI Pentest on Every Deploy

Short answer: Put the pentest step after the deploy to staging, not in the build. The step starts a pentest and lets the pipeline carry on. A later job, or a person, gates the release on the findings. Barrion has REST API support for starting pentests, so any pipeline can do this, and a Light run is capped at 2 hours.

Why run a pentest on every deploy?

Because the time between a flaw going live and someone using it keeps shrinking. Google's Mandiant team measured the average time from a vulnerability's disclosure to its exploitation at 63 days in 2018 and 2019, 32 days in 2021 and 2022, and 5 days in 2023. Its M-Trends 2026 report puts the mean for 2025 at minus 7 days, which means that on average, exploitation started before a patch was out. VulnCheck counted 28.96% of the vulnerabilities first exploited in 2025 as used on or before the day their CVE was published, up from 23.6% in 2024.

Those figures are about CVEs in products like VPNs and firewalls. A bug in your own app never gets a CVE, so nobody publishes a clock for it. What carries over is the tooling, and it's getting faster at the kind of bug that lives in custom web apps.

Source and dateWhat it found
Google Threat Intelligence Group, AI Threat Tracker, May 2026The first case it's confident was an AI-developed zero-day planned for mass exploitation: a 2FA bypass in a web-based admin tool, caused by a hardcoded trust assumption. A logic flaw, not a pattern a scanner matches
Google Threat Intelligence Group, AI Threat Tracker, September 2026Attackers running multi-agent frameworks that "autonomously manage scanning pipelines", and one agent-enabled credential harvesting campaign planned, built and run in under six hours
Anthropic, AI-orchestrated espionage report, November 2025A state-sponsored group used AI for 80 to 90% of a campaign's work against about thirty targets, with a human stepping in at a handful of decision points. It succeeded in a small number of cases
Google Cloud (Mandiant), M-Trends 2026, March 2026Exploits were the most common way in for the sixth year running, at 32% of intrusions

So plan on an exposed flaw being found by automated tooling within days, not at next year's pentest. That's why we recommend testing the app surface daily and on every release, with the annual test as the compliance floor rather than the plan. If you ship once a month, a schedule is enough and you don't need a pipeline step. The cadence question on its own is covered in how often to pentest.

Where does the pentest step belong in a pipeline?

After the deploy to staging, running in parallel with everything else. A pentest needs a running app, and it takes hours, so it can't sit inside the build and it shouldn't hold one up.

Pipeline stageTypical timePentest step
Build and unit testsMinutesNo. There's no running app to test yet
Deploy to stagingMinutesStart the pentest here and move on
Integration and smoke testsMinutesRun them while the pentest works
Promote to productionSecondsGate here if you want a hard stop, or let it through and act on findings
After releaseOngoingDashboard schedules keep testing production

The pattern is asynchronous. Start the run and let the pipeline continue. Then decide what the result is allowed to block.

PatternWhat waitsWhen it fits
Start onlyNothing. Findings go to the dashboardYou deploy many times a day, or you're trying it out
Start, then gate the releaseA later job waits for the run and fails on critical or high findingsStaging runs ahead of a production promote, and a few hours' delay is fine
Start, flag, don't blockA later job posts the finding count to the job summaryYou want the signal in CI without holding releases

One limit shapes the setup. The target has to be verified in the Barrion dashboard first, the same check a dashboard run goes through. That means a stable environment such as your staging domain. A fresh preview URL per pull request won't pass it.

How do you start a pentest from your pipeline?

Barrion has REST API support for starting pentests, so any CI/CD pipeline can start one after a deploy. It works the same in GitHub Actions, GitLab CI, Jenkins, CircleCI or a plain shell script, because the step is your own. There's no Barrion GitHub Action or CI plugin, and you don't need one.

Once your staging domain is verified, set up API access from the dashboard, or contact us and we'll help you wire it into your pipeline.

How do you fail a build on new critical findings?

Each finding comes with a severity, so a later job can check for open critical or high findings and fail if there are any. Findings you've marked ignored in the dashboard don't count. So once someone has triaged an accepted risk, it stops failing builds.

What a pipeline check doesn't get in the API's first version is the run-over-run label. The new, still open, resolved and regressed labels come from dashboard schedules, which compare each run with the previous one and email you about new and regressed findings. So a pipeline gate answers "is anything critical or high open on staging right now". A schedule answers "what changed since last time".

In practice that means your first strict gate may fail on findings that were already there. Pick the policy to match where you are.

Gate policyThe job fails whenGood for
StrictAny critical or high finding is openApps with a triaged baseline, regulated products
Critical onlyAny critical finding is openThe first weeks, while older highs are worked down
Flag onlyNever. The count goes to the job summaryTeams deploying many times a day

How does it work alongside dashboard schedules?

They cover different things, and most teams want both. The pipeline step ties a test to a specific release on staging. Schedules keep testing production whether you deployed or not, and they're where regression tracking lives.

TriggerTargetLevelRuns
Every deploy, from the pipelineStagingLightAfter each deploy to staging
Daily scheduleProductionLightEvery day, or only when the app has changed
Monthly scheduleProductionDeepEvery time
Before a major releaseStagingStandard or DeepOn demand, from the dashboard or the pipeline

Each schedule has its own scope and depth. Schedules can run daily, weekly, monthly, quarterly, every six months, yearly, or on a custom rhythm, and each one runs every time or only when the app has changed. Change detection compares a snapshot of the start page's links and scripts, so a backend-only release can slip past it. A pipeline step closes that gap, because it's tied to the deploy itself and doesn't need to notice a change.

The background on scheduled testing is in what is continuous pentesting.

Is it safe to pentest from a pipeline?

Yes, with a few choices made up front. We recommend pointing per-deploy runs at staging, with test data and test accounts, and letting dashboard schedules cover production.

  • Agents test the way an attacker would, but they're non-destructive and rate-limited. They can still create test records, which is why staging is the safer target.
  • Scope is approved before any traffic, and only the verified target is tested.
  • Production is opt-in for pipeline runs. A start from the pipeline is refused unless you've allowed API runs for that target.
  • API access has its own credit budget. It applies on top of any spend cap on your own account, and the stricter one wins.
  • Runs in flight per account are capped. When you're at the cap, a new start is refused rather than queued, so decide whether your step should skip or try again later.
  • If a pipeline is cancelled, cancel the run too and you pay only for the work done.

What does a pentest from CI cost?

The same as one started from the dashboard, from the same credit balance. A run holds its level's credits, 400 for Light and 4,000 for Deep, and is charged for what it uses, with a minimum of 100 credits for a finished run. So a short run on a small staging app still costs 100. A failed run is free. Single-run prices are on the pricing page.

If you deploy twenty times a day, one pentest per deploy adds up fast and mostly retests the same code. Most teams run it on merges to the main branch, or on the last deploy of the day. Continuous programs across several apps are scoped to your apps, cadence and depth. Talk to us and we'll price it for your setup.

How Barrion does it

Barrion's AI agents run full pentests of web apps and APIs at five levels, from Light (400 credits, 3 agents) to Maximum. A coordinating agent splits the work across specialist agents that test in parallel (3 at Light, up to 100 at Maximum), covering 8 testing areas mapped to all 97 OWASP WSTG v4.2 test cases. Runs are capped by level: Light 2, Standard 4, Deep 8, Extended 12, Maximum 16 hours.

You can put a pentest on a schedule in the dashboard, start one from any CI/CD pipeline through the REST API, or start one on demand. Findings are checked against the live app before they're reported. One that's replayed and doesn't reproduce is dropped. One that can't be replayed is kept as a lower-confidence lead. From Standard level up, a security engineer reviews each report before release.

What we don't do: there's no Barrion GitHub Action or CI plugin, and the API has no webhooks yet, so your pipeline checks back on the run. It can't test ephemeral preview environments. And Barrion doesn't test internal networks, Active Directory, mobile apps, physical security or social engineering, and it isn't TLPT under DORA. Start a pentest to try the flow on your staging app first.

Sources

All sources checked 2026-09-26.

FAQ

Frequently asked questions

Is there a Barrion GitHub Action or CI plugin?
No. Barrion has REST API support for starting pentests, so your own pipeline step can start one after a deploy. That works the same in GitHub Actions, GitLab CI, Jenkins, CircleCI, Bitbucket Pipelines or a plain shell script. To set it up, use the dashboard or contact us.
How long does a pentest in a pipeline take?
Runs are capped by level: Light 2, Standard 4, Deep 8, Extended 12, Maximum 16 hours. That's why the pentest step should start the run and let the pipeline continue, with a later job or a person checking the result, rather than blocking the build.
Can I run it against preview environments?
Not ephemeral ones. The target has to be verified in the Barrion dashboard before any run, so the API tests stable environments such as a staging domain. A new preview URL for each pull request won't pass that check.
Can I pentest production from CI?
Only if you allow API runs for that target in the dashboard. It's off by default for every target. We recommend staging for per-deploy runs and dashboard schedules for production, where tests run at a rhythm you set.
Can the pipeline fail only on new or regressed findings?
Not in the API's first version. Findings come with their severity and ones you've marked ignored don't count, so you can fail on any open critical or high finding. The new, still open, resolved and regressed labels come from dashboard schedules, which email you about new and regressed findings.
Does the API send a webhook when a run finishes?
No. The first version has no webhooks, so your pipeline checks the run's status until it finishes. Once a minute is plenty.
What does a pentest from CI cost?
The same as a dashboard run, from the same credits. A run holds its level's credits and is charged for what it uses, with a minimum of 100 credits for a finished run, and a failed run is free. API access also has its own credit budget. Continuous programs are scoped to your apps, cadence and depth, so talk to us for a price.

Run your first pentest on staging.

Start a pentest yourself and check the flow on your staging app, then start the same run from your pipeline. For continuous testing across several apps, talk to us.