Back to articles
Product
Updated Sep 26, 2026

Self-serve AI pentesting is now live

With Barrion's AI pentesting, AI agents test your web app or API the way a real attacker would. They map your attack surface and chain live requests to find what is genuinely exploitable, and findings are checked against your live app before they reach your report. Until now, kicking off a pentest meant booking a call with us to scope it first.

Now any Barrion account can set up and start an AI pentest directly, with no call required, and the same release deepened the engine underneath it. You define the scope, pick a level, pay for the run in credits, and follow it live. The full report comes with it.

What is new

This release changed two things.

The first is access. Setting up a pentest no longer needs a call. A wizard in your dashboard walks you through the target, the credentials, and anything you want left alone, and nothing is sent until you approve it.

The second is depth. This release extends the engine's multi-wave attack chaining and broadens its authenticated testing, so a run digs further into your app than it did before. Self-serve pentests are paid in credits before the run starts. The full report comes with the run. There is no second payment to read the findings. You can follow the run's progress in your dashboard, and the report has proof and remediation for every confirmed finding. That turns what used to be a booked engagement into something you start yourself.

How a run works

Each engagement moves through five phases inside an isolated Kali sandbox you can watch in real time:

  1. Scope and authorize: you set the target, its surfaces and credentials, and mark anything off-limits. The agent tests thoroughly inside the scope and stays quiet everywhere else.
  2. Recon and discovery: it maps hosts, endpoints, parameters, auth flows, and your stack before testing anything.
  3. Multi-wave testing: specialist agents work through injection, access control, authentication, API, and business-logic testing, carrying findings from one wave into the next.
  4. Reproduce before report: findings we can replay are sent again against your live app. If one doesn't reproduce, it's dropped. A finding we can't replay is labelled as an unconfirmed lead, never as a confirmed finding.
  5. Report and retest: you get a ranked report with the request and response behind each confirmed finding, remediation, and WSTG coverage as a downloadable bundle with a shareable PDF. Once you've shipped a fix, start a free retest from the dashboard. Each finding comes back as fixed, not fixed or inconclusive.

You can watch the run live in your dashboard. A run finishes within hours. From Standard level up, a security engineer reviews the report and it is released within one working day. A Light report is released as soon as the run completes.

Levels and what a run costs

You pick a level when you start a run. Every level tests every vulnerability class. A deeper level adds more agents working in parallel, more attack waves, more user roles and more expert review time.

LevelCredits heldAgentsExpert reviewReport released
Light4003nonewhen the run completes
Standard1,00051 hourwithin one working day
Deep4,000202 hourswithin one working day
Extended10,000504 hourswithin one working day
Maximum20,0001008 hourswithin one working day

Credits cost €0.50 each as a top-up, so the most a run can cost ranges from €200 (Light) to €10,000 (Maximum). A run is charged for what it actually uses, with a minimum of 100 credits. A run that fails is free, and a cancelled run pays only for the work it got through.

Real pentesting, not a scan

A scanner matches patterns and gives you a list of things that might be wrong. This works differently. Each run spins up a per-engagement Kali box with the tools a manual tester reaches for, including sqlmap, nuclei, ZAP, katana, ffuf, dalfox, and jwt_tool, and the agent reasons about your specific app as it goes. It confirms a bug from the response itself, whether that is a blind SQL injection showing up as a timing difference or an endpoint returning 200 where it should return 403.

By the time a confirmed finding reaches your report it has already been reproduced, with the request, the response, the affected surface, and the matching OWASP WSTG and CWE references attached. There is far less to wade through, because findings that didn't reproduce were dropped, and the few we couldn't replay are clearly marked as leads.

What it looks at

Coverage maps to the OWASP Top 10, the OWASP API Security Top 10, and all 97 WSTG v4.2 test cases, each with a status in the report:

  • Injection and web exploitation: SQL and NoSQL injection, XSS, command injection, SSRF, SSTI, XXE, CSRF, open redirects.
  • Access control and authentication: broken access control, IDOR, privilege escalation, session and JWT flaws, multi-tenant isolation.
  • API and business logic: BOLA and BFLA, mass assignment, excessive data exposure, rate-limiting gaps, business-logic abuse, file upload and path traversal, chained exploits.

The report ships with a full coverage matrix, so you can show an auditor or a customer exactly what ran and how each case came out.

Running it safely

Testing your own production app should not feel like a gamble, so every probe is rate-limited and non-destructive, and the agent only touches what you approved during setup. Testing can still create records, like test sign-ups or form submissions. If you would rather start on staging, point it there instead. It runs against any environment you control and authorize.

When to run it

Run it whenever you need a real pentest. That covers the yearly assessment for SOC 2, ISO 27001 or a customer security review, and the moments in between: when you ship a feature, before a launch, or any time a change makes you uneasy.

It works as a pentest on its own. If your security program also brings in manual testers, the report gives them a head start, since the OWASP Top 10, the API Security Top 10 and all 97 WSTG cases are already mapped, with proof on each confirmed finding. They can spend their hours on the business logic that needs human judgement.

Running it continuously

A single pentest tells you where the app stood on the day it ran. If you ship every week, that picture goes stale fast. So a pentest in Barrion doesn't have to be a one-off.

You can save a pentest and rerun it on a schedule: daily, weekly, monthly, quarterly, every six months, yearly, or a custom rhythm. One target can have more than one schedule. A common setup is a daily Light run to catch new issues quickly plus a monthly Deep run for depth.

Each schedule has its own scope and depth. Choose whether it runs every time, or only when your app has changed. For the second option, Barrion takes a snapshot of your live app's pages, links and scripts, and when it detects a change it runs the full pentest again at the depth you picked.

Every scheduled run is compared with the one before it. Each finding is labelled new, still open, resolved or regressed, so a vulnerability you fixed in March that comes back in June shows up as a regression rather than getting lost in a fresh list.

More on how this works: continuous pentesting and what agentic pentesting is.

Try it

Create a free account, open Pentesting in your dashboard, and set up your first run. Setup takes a few minutes, then the agent goes to work and you can follow the run in your dashboard until the report lands.

Running a larger estate, or need the scope and rules of engagement agreed up front? Continuous programs are scoped to your apps, cadence and depth. Talk to us and we'll price it for your setup. It's the same testing engine underneath either way.

Frequently asked questions

Q: What does self-serve AI pentesting mean?

A: You scope, authorize, and start a full AI penetration test yourself from the Barrion dashboard, with no sales call. This release also deepens the engine, so a run reaches further into your app, and you can go from setup to a running test in minutes.

Q: How much does it cost to start?

A: Credits. An Essential plan carries a set number every month, or you buy a top-up at €0.50 per credit. A run holds its level's credits (Light 400, Standard 1,000, Deep 4,000, Extended 10,000, Maximum 20,000) and is charged for what it uses, with a minimum of 100. The full report comes with the run and a retest after you patch holds no credits. If a report doesn't hold up, email contact@barrion.io and we return the credits. A failed run is free, and a cancelled run pays only for the work it got through. Larger estates can be scoped and quoted on a call.

Q: Will an AI pentest break my application or data?

A: It's built not to. Every probe is rate-limited and non-destructive, and you approve the scope and any off-limits paths before a single request is sent. Testing can still create records such as test sign-ups, so many teams run the first engagement against staging.

Q: How is this different from an automated scanner?

A: A scanner pattern-matches and hands you a list of maybes. Barrion actively tests the target the way a real attacker would, chains requests, and replays findings against the live app before reporting them. Findings that don't reproduce are dropped, and anything it can't replay is labelled as an unconfirmed lead, so you get proof-of-exploit rather than a pile of alerts.

Q: How long does a run take?

A: A run finishes within hours. From Standard level up, a security engineer reviews the report and it is released within one working day. A Light report is released as soon as the run completes. You can watch the run live in your dashboard while it works.

Q: Does this replace a human pentester?

A: No. It maps its results to the OWASP Top 10, the API Security Top 10 and all 97 WSTG cases, with proof on each confirmed finding, and from Standard level up a security engineer reviews the report. People are still better at deep, domain-specific business logic. Many teams run it as their ongoing pentest, continuously or on demand, and bring in a manual tester periodically for the paths that need judgement.

Q: Can I run it on a schedule?

A: Yes. You can save a pentest and rerun it daily, weekly, monthly, quarterly, every six months, yearly or on a custom rhythm, and choose whether each run happens every time or only when your app has changed. Each scheduled run labels findings new, still open, resolved or regressed compared with the previous run. Scheduled pentests are part of the Business plan.

A scan shows the surface. A pentest tests what gets in.

Passive scan

  • Reads what your app already exposes
  • Never logs in or submits a form
  • Cannot confirm what is exploitable

Active AI pentest

  • Tests your app the way an attacker would
  • Chains requests to confirm real exploits
  • Replays findings against your live app
  • Runs on a schedule or on demand, retests free

Probes for

  • SQL injection
  • Broken access control
  • IDOR
  • SSRF
  • Business-logic abuse

How AI pentesting works

Paid in credits. Free retests of found issues, and expert review from Standard up.

Secure your apps before
someone else finds the gaps.

Used by 5,000+ developers and engineering teams. Start with one pentest or put your apps on a schedule.