Self-serve AI pentesting is now live
With Barrion's AI pentesting, AI agents test your web app or API the way a real attacker would. They map your attack surface and chain live requests to find what is genuinely exploitable, and findings are checked against your live app before they reach your report. Until now, kicking off a pentest meant booking a call with us to scope it first.
Now any Barrion account can set up and start an AI pentest directly, with no call required, and the same release deepened the engine underneath it. You define the scope, pick a level, pay for the run in credits, and follow it live. The full report comes with it.
What is new
This release changed two things.
The first is access. Setting up a pentest no longer needs a call. A wizard in your dashboard walks you through the target, the credentials, and anything you want left alone, and nothing is sent until you approve it.
The second is depth. This release extends the engine's multi-wave attack chaining and broadens its authenticated testing, so a run digs further into your app than it did before. Self-serve pentests are paid in credits before the run starts. The full report comes with the run. There is no second payment to read the findings. You can follow the run's progress in your dashboard, and the report has proof and remediation for every confirmed finding. That turns what used to be a booked engagement into something you start yourself.
How a run works
Each engagement moves through five phases inside an isolated Kali sandbox you can watch in real time:
- Scope and authorize: you set the target, its surfaces and credentials, and mark anything off-limits. The agent tests thoroughly inside the scope and stays quiet everywhere else.
- Recon and discovery: it maps hosts, endpoints, parameters, auth flows, and your stack before testing anything.
- Multi-wave testing: specialist agents work through injection, access control, authentication, API, and business-logic testing, carrying findings from one wave into the next.
- Reproduce before report: findings we can replay are sent again against your live app. If one doesn't reproduce, it's dropped. A finding we can't replay is labelled as an unconfirmed lead, never as a confirmed finding.
- Report and retest: you get a ranked report with the request and response behind each confirmed finding, remediation, and WSTG coverage as a downloadable bundle with a shareable PDF. Once you've shipped a fix, start a free retest from the dashboard. Each finding comes back as fixed, not fixed or inconclusive.
You can watch the run live in your dashboard. A run finishes within hours. From Standard level up, a security engineer reviews the report and it is released within one working day. A Light report is released as soon as the run completes.
Levels and what a run costs
You pick a level when you start a run. Every level tests every vulnerability class. A deeper level adds more agents working in parallel, more attack waves, more user roles and more expert review time.
| Level | Credits held | Agents | Expert review | Report released |
|---|---|---|---|---|
| Light | 400 | 3 | none | when the run completes |
| Standard | 1,000 | 5 | 1 hour | within one working day |
| Deep | 4,000 | 20 | 2 hours | within one working day |
| Extended | 10,000 | 50 | 4 hours | within one working day |
| Maximum | 20,000 | 100 | 8 hours | within one working day |
Credits cost €0.50 each as a top-up, so the most a run can cost ranges from €200 (Light) to €10,000 (Maximum). A run is charged for what it actually uses, with a minimum of 100 credits. A run that fails is free, and a cancelled run pays only for the work it got through.
Real pentesting, not a scan
A scanner matches patterns and gives you a list of things that might be wrong. This works differently. Each run spins up a per-engagement Kali box with the tools a manual tester reaches for, including sqlmap, nuclei, ZAP, katana, ffuf, dalfox, and jwt_tool, and the agent reasons about your specific app as it goes. It confirms a bug from the response itself, whether that is a blind SQL injection showing up as a timing difference or an endpoint returning 200 where it should return 403.
By the time a confirmed finding reaches your report it has already been reproduced, with the request, the response, the affected surface, and the matching OWASP WSTG and CWE references attached. There is far less to wade through, because findings that didn't reproduce were dropped, and the few we couldn't replay are clearly marked as leads.
What it looks at
Coverage maps to the OWASP Top 10, the OWASP API Security Top 10, and all 97 WSTG v4.2 test cases, each with a status in the report:
- Injection and web exploitation: SQL and NoSQL injection, XSS, command injection, SSRF, SSTI, XXE, CSRF, open redirects.
- Access control and authentication: broken access control, IDOR, privilege escalation, session and JWT flaws, multi-tenant isolation.
- API and business logic: BOLA and BFLA, mass assignment, excessive data exposure, rate-limiting gaps, business-logic abuse, file upload and path traversal, chained exploits.
The report ships with a full coverage matrix, so you can show an auditor or a customer exactly what ran and how each case came out.
Running it safely
Testing your own production app should not feel like a gamble, so every probe is rate-limited and non-destructive, and the agent only touches what you approved during setup. Testing can still create records, like test sign-ups or form submissions. If you would rather start on staging, point it there instead. It runs against any environment you control and authorize.
When to run it
Run it whenever you need a real pentest. That covers the yearly assessment for SOC 2, ISO 27001 or a customer security review, and the moments in between: when you ship a feature, before a launch, or any time a change makes you uneasy.
It works as a pentest on its own. If your security program also brings in manual testers, the report gives them a head start, since the OWASP Top 10, the API Security Top 10 and all 97 WSTG cases are already mapped, with proof on each confirmed finding. They can spend their hours on the business logic that needs human judgement.
Running it continuously
A single pentest tells you where the app stood on the day it ran. If you ship every week, that picture goes stale fast. So a pentest in Barrion doesn't have to be a one-off.
You can save a pentest and rerun it on a schedule: daily, weekly, monthly, quarterly, every six months, yearly, or a custom rhythm. One target can have more than one schedule. A common setup is a daily Light run to catch new issues quickly plus a monthly Deep run for depth.
Each schedule has its own scope and depth. Choose whether it runs every time, or only when your app has changed. For the second option, Barrion takes a snapshot of your live app's pages, links and scripts, and when it detects a change it runs the full pentest again at the depth you picked.
Every scheduled run is compared with the one before it. Each finding is labelled new, still open, resolved or regressed, so a vulnerability you fixed in March that comes back in June shows up as a regression rather than getting lost in a fresh list.
More on how this works: continuous pentesting and what agentic pentesting is.
Try it
Create a free account, open Pentesting in your dashboard, and set up your first run. Setup takes a few minutes, then the agent goes to work and you can follow the run in your dashboard until the report lands.
Running a larger estate, or need the scope and rules of engagement agreed up front? Continuous programs are scoped to your apps, cadence and depth. Talk to us and we'll price it for your setup. It's the same testing engine underneath either way.