Short answer: Put the pentest step after the deploy to staging, not in the build. The step starts a pentest and lets the pipeline carry on. A later job, or a person, gates the release on the findings. Barrion has REST API support for starting pentests, so any pipeline can do this, and a Light run is capped at 2 hours.
Why run a pentest on every deploy?
Because the time between a flaw going live and someone using it keeps shrinking. Google's Mandiant team measured the average time from a vulnerability's disclosure to its exploitation at 63 days in 2018 and 2019, 32 days in 2021 and 2022, and 5 days in 2023. Its M-Trends 2026 report puts the mean for 2025 at minus 7 days, which means that on average, exploitation started before a patch was out. VulnCheck counted 28.96% of the vulnerabilities first exploited in 2025 as used on or before the day their CVE was published, up from 23.6% in 2024.
Those figures are about CVEs in products like VPNs and firewalls. A bug in your own app never gets a CVE, so nobody publishes a clock for it. What carries over is the tooling, and it's getting faster at the kind of bug that lives in custom web apps.
| Source and date | What it found |
|---|---|
| Google Threat Intelligence Group, AI Threat Tracker, May 2026 | The first case it's confident was an AI-developed zero-day planned for mass exploitation: a 2FA bypass in a web-based admin tool, caused by a hardcoded trust assumption. A logic flaw, not a pattern a scanner matches |
| Google Threat Intelligence Group, AI Threat Tracker, September 2026 | Attackers running multi-agent frameworks that "autonomously manage scanning pipelines", and one agent-enabled credential harvesting campaign planned, built and run in under six hours |
| Anthropic, AI-orchestrated espionage report, November 2025 | A state-sponsored group used AI for 80 to 90% of a campaign's work against about thirty targets, with a human stepping in at a handful of decision points. It succeeded in a small number of cases |
| Google Cloud (Mandiant), M-Trends 2026, March 2026 | Exploits were the most common way in for the sixth year running, at 32% of intrusions |
So plan on an exposed flaw being found by automated tooling within days, not at next year's pentest. That's why we recommend testing the app surface daily and on every release, with the annual test as the compliance floor rather than the plan. If you ship once a month, a schedule is enough and you don't need a pipeline step. The cadence question on its own is covered in how often to pentest.
Where does the pentest step belong in a pipeline?
After the deploy to staging, running in parallel with everything else. A pentest needs a running app, and it takes hours, so it can't sit inside the build and it shouldn't hold one up.
| Pipeline stage | Typical time | Pentest step |
|---|---|---|
| Build and unit tests | Minutes | No. There's no running app to test yet |
| Deploy to staging | Minutes | Start the pentest here and move on |
| Integration and smoke tests | Minutes | Run them while the pentest works |
| Promote to production | Seconds | Gate here if you want a hard stop, or let it through and act on findings |
| After release | Ongoing | Dashboard schedules keep testing production |
The pattern is asynchronous. Start the run and let the pipeline continue. Then decide what the result is allowed to block.
| Pattern | What waits | When it fits |
|---|---|---|
| Start only | Nothing. Findings go to the dashboard | You deploy many times a day, or you're trying it out |
| Start, then gate the release | A later job waits for the run and fails on critical or high findings | Staging runs ahead of a production promote, and a few hours' delay is fine |
| Start, flag, don't block | A later job posts the finding count to the job summary | You want the signal in CI without holding releases |
One limit shapes the setup. The target has to be verified in the Barrion dashboard first, the same check a dashboard run goes through. That means a stable environment such as your staging domain. A fresh preview URL per pull request won't pass it.
How do you start a pentest from your pipeline?
Barrion has REST API support for starting pentests, so any CI/CD pipeline can start one after a deploy. It works the same in GitHub Actions, GitLab CI, Jenkins, CircleCI or a plain shell script, because the step is your own. There's no Barrion GitHub Action or CI plugin, and you don't need one.
Once your staging domain is verified, set up API access from the dashboard, or contact us and we'll help you wire it into your pipeline.
How do you fail a build on new critical findings?
Each finding comes with a severity, so a later job can check for open critical or high findings and fail if there are any. Findings you've marked ignored in the dashboard don't count. So once someone has triaged an accepted risk, it stops failing builds.
What a pipeline check doesn't get in the API's first version is the run-over-run label. The new, still open, resolved and regressed labels come from dashboard schedules, which compare each run with the previous one and email you about new and regressed findings. So a pipeline gate answers "is anything critical or high open on staging right now". A schedule answers "what changed since last time".
In practice that means your first strict gate may fail on findings that were already there. Pick the policy to match where you are.
| Gate policy | The job fails when | Good for |
|---|---|---|
| Strict | Any critical or high finding is open | Apps with a triaged baseline, regulated products |
| Critical only | Any critical finding is open | The first weeks, while older highs are worked down |
| Flag only | Never. The count goes to the job summary | Teams deploying many times a day |
How does it work alongside dashboard schedules?
They cover different things, and most teams want both. The pipeline step ties a test to a specific release on staging. Schedules keep testing production whether you deployed or not, and they're where regression tracking lives.
| Trigger | Target | Level | Runs |
|---|---|---|---|
| Every deploy, from the pipeline | Staging | Light | After each deploy to staging |
| Daily schedule | Production | Light | Every day, or only when the app has changed |
| Monthly schedule | Production | Deep | Every time |
| Before a major release | Staging | Standard or Deep | On demand, from the dashboard or the pipeline |
Each schedule has its own scope and depth. Schedules can run daily, weekly, monthly, quarterly, every six months, yearly, or on a custom rhythm, and each one runs every time or only when the app has changed. Change detection compares a snapshot of the start page's links and scripts, so a backend-only release can slip past it. A pipeline step closes that gap, because it's tied to the deploy itself and doesn't need to notice a change.
The background on scheduled testing is in what is continuous pentesting.
Is it safe to pentest from a pipeline?
Yes, with a few choices made up front. We recommend pointing per-deploy runs at staging, with test data and test accounts, and letting dashboard schedules cover production.
- Agents test the way an attacker would, but they're non-destructive and rate-limited. They can still create test records, which is why staging is the safer target.
- Scope is approved before any traffic, and only the verified target is tested.
- Production is opt-in for pipeline runs. A start from the pipeline is refused unless you've allowed API runs for that target.
- API access has its own credit budget. It applies on top of any spend cap on your own account, and the stricter one wins.
- Runs in flight per account are capped. When you're at the cap, a new start is refused rather than queued, so decide whether your step should skip or try again later.
- If a pipeline is cancelled, cancel the run too and you pay only for the work done.
What does a pentest from CI cost?
The same as one started from the dashboard, from the same credit balance. A run holds its level's credits, 400 for Light and 4,000 for Deep, and is charged for what it uses, with a minimum of 100 credits for a finished run. So a short run on a small staging app still costs 100. A failed run is free. Single-run prices are on the pricing page.
If you deploy twenty times a day, one pentest per deploy adds up fast and mostly retests the same code. Most teams run it on merges to the main branch, or on the last deploy of the day. Continuous programs across several apps are scoped to your apps, cadence and depth. Talk to us and we'll price it for your setup.
How Barrion does it
Barrion's AI agents run full pentests of web apps and APIs at five levels, from Light (400 credits, 3 agents) to Maximum. A coordinating agent splits the work across specialist agents that test in parallel (3 at Light, up to 100 at Maximum), covering 8 testing areas mapped to all 97 OWASP WSTG v4.2 test cases. Runs are capped by level: Light 2, Standard 4, Deep 8, Extended 12, Maximum 16 hours.
You can put a pentest on a schedule in the dashboard, start one from any CI/CD pipeline through the REST API, or start one on demand. Findings are checked against the live app before they're reported. One that's replayed and doesn't reproduce is dropped. One that can't be replayed is kept as a lower-confidence lead. From Standard level up, a security engineer reviews each report before release.
What we don't do: there's no Barrion GitHub Action or CI plugin, and the API has no webhooks yet, so your pipeline checks back on the run. It can't test ephemeral preview environments. And Barrion doesn't test internal networks, Active Directory, mobile apps, physical security or social engineering, and it isn't TLPT under DORA. Start a pentest to try the flow on your staging app first.
Sources
All sources checked 2026-09-26.
How Low Can You Go? An Analysis of 2023 Time-to-Exploit Trends, Google Cloud (Mandiant), published 2024-10-15
M-Trends 2026, Google Cloud (Mandiant), published 2026-03-23
State of Exploitation 2026, VulnCheck, Patrick Garrity, published 2026-01-21
GTIG AI Threat Tracker: Adversaries Leverage AI for Vulnerability Exploitation, Augmented Operations, and Initial Access, Google Threat Intelligence Group, published 2026-05-11
GTIG AI Threat Tracker: From Prompting to Autonomy, Google Threat Intelligence Group, published 2026-09-08
Disrupting the first reported AI-orchestrated cyber espionage campaign, Anthropic, published 2025-11-13
Barrion product facts, facts page