OWASP APTS: The Autonomous Penetration Testing Standard, Explained for Buyers
Short definition
The OWASP Autonomous Penetration Testing Standard (APTS) is a governance standard for platforms that run penetration tests autonomously: 173 tier-required requirements in eight domains that say how such a platform must stay in scope, stay stoppable and stay accountable.
Why this matters now
An autonomous pentesting platform takes offensive actions in production on its own, so the first procurement question is not what it finds but whether its AI can be kept in scope, stopped and audited. APTS gives buyers a shared way to ask that question, in the form of tiers, a vendor evaluation guide and a conformance claim template. It is also young (version 0.1.0), self-assessed and not a certification, so a buyer has to know what a claim of conformance does and does not prove.
Key points
- ▸OWASP Incubator project, version 0.1.0, published in April 2026 under CC BY-SA 4.0.
- ▸173 tier-required requirements (144 MUST, 29 SHOULD) in eight domains, plus 20 advisory practices outside the tiers.
- ▸Three tiers: Tier 1 has 72 requirements, Tier 2 has 157 cumulative, Tier 3 has 173 cumulative.
- ▸A tier claim needs every MUST met and every SHOULD met or justified in writing; there is no partial credit.
- ▸Tiers measure governance; autonomy levels L1 Assisted to L4 Autonomous measure independence. The standard says not to conflate them.
- ▸No certification body and no mandatory audit: claims are self-assessed or reviewed by a third party, and they are not EU AI Act compliance.
What APTS is, and what it is not
APTS describes itself as a governance standard for autonomous penetration testing platforms. It defines what these systems must do to operate safely, transparently and within defined boundaries, whether a vendor delivers them, a service provider operates them or an enterprise security team builds one to test its own organization.
It is not a testing methodology. It does not tell a tester how to test; it complements PTES, the OWASP Web Security Testing Guide and OSSTMM by addressing the problems that only exist when software, not a person, decides the next step: scope enforcement, safe autonomy, resistance to manipulation and accountability.
The basic facts a buyer should have on file:
- Status: OWASP Incubator project, which means an early-stage project rather than a mature flagship standard.
- Version: 0.1.0, with the standard's files first published in April 2026. Requirements can change between versions, which is why the standard asks contracts to cite a versioned identifier such as
APTS-v0.1.0-SE-001. - License: CC BY-SA 4.0, free to use and adapt in your own evaluation templates.
The eight domains
The 173 tier-required requirements are grouped in eight domains, each with its own prefix:
- Scope Enforcement (SE), 26 requirements: defining, validating and enforcing testing boundaries, including rules of engagement the platform can read.
- Safety Controls (SC), 20: impact classification, limits on blast radius, kill switches and rollback.
- Human Oversight (HO), 19: approval gates, dashboards, escalation and operator qualifications.
- Graduated Autonomy (AL), 28: the four autonomy levels and the obligations that grow with each one.
- Auditability (AR), 20: logging, decision trails and evidence integrity.
- Manipulation Resistance (MR), 23: resistance to prompt injection, adversarial input and attempts by a target to widen the scope.
- Supply Chain Trust (TP), 22: trust in AI providers, data handling, isolation between customers and disclosure of the foundation model.
- Reporting (RP), 15: finding validation, confidence and disclosure of what was not covered.
For a buyer the domains map neatly onto the questions a risk committee asks: will it stay in scope (SE), can we stop it (SC, HO), how much does it do on its own (AL), can we prove what it did (AR), can a target trick it (MR), who else sees our data (TP), and can we trust the findings (RP).
Tiers, MUST and SHOULD
APTS has three compliance tiers, each including the one below:
- Tier 1, Foundation (72 requirements): the platform will not test outside the agreed scope, can be stopped immediately, will not store or leak discovered credentials in plaintext, and keeps a basic audit trail. The standard positions it for supervised testing of non-critical systems.
- Tier 2, Verified (85 more, 157 cumulative): transparency about what the platform did and why, tamper-proof audit trails, formal incident response and findings that can be independently verified. Positioned for production environments and regulated industries.
- Tier 3, Comprehensive (16 more, 173 cumulative): the highest assurance bar, for critical infrastructure and fully autonomous (L4) operation.
Each requirement is marked MUST or SHOULD. Counted across the checklists, 144 are MUST and 29 are SHOULD. The claim rule is strict: a platform claims a tier only if it implements every MUST at that tier and all lower tiers with no deviation, and every SHOULD at those tiers is either implemented or covered by a documented justification in the conformance claim. An unimplemented MUST or an undocumented SHOULD deviation is a conformance gap. There is no partial credit, so "we meet 90% of Tier 2" is not a Tier 2 claim.
On top of the tiers sit 20 advisory practices, identified as APTS-<DOMAIN>-A0x. They are not counted toward any tier and do not affect conformance, but the standard recommends them for high-risk engagements, regulated industries and L4 operation, and suggests customers treat their adoption as a differentiator.
Tiers are not autonomy levels
APTS defines four autonomy levels in its Graduated Autonomy domain:
- L1 Assisted: the operator commands every action; the platform executes one technique per command.
- L2 Supervised: the platform chains techniques within one phase; the operator approves every phase boundary.
- L3 Semi-Autonomous: the platform runs complete chains inside pre-approved boundaries; the operator intervenes on exceptions.
- L4 Autonomous: the platform manages long multi-target campaigns; oversight moves to periodic review.
The standard is explicit that tiers and levels are distinct concepts and must not be conflated. A tier is a conformance posture, which requirements the platform satisfies. A level is an operating mode, how much the platform does on its own during an engagement. As a rule of thumb the standard gives, Tier 1 is generally suitable for L1 operation, Tier 2 covers what is needed for L2 and L3, and Tier 3 for L4.
For a buyer this means two separate questions, both in writing: which tier does the vendor claim, and at which autonomy levels will you actually run it. A Tier 1 claim from a platform you intend to run at L3 does not cover your use.
Who verifies a claim
APTS has no certification body, no mandatory third-party audit and no fee. Platforms are assessed against the requirements and the result is documented as a conformance claim. The standard does not prescribe who performs the assessment: internal self-assessment, independent internal review and external third-party assessment are all valid, and the choice is left to the reader.
The optional Conformance Claim Template shows what a serious claim contains: the APTS version, the claimed tier, the assessment method, the deployment model and autonomy levels covered (and any modules excluded), a disclosure of the foundation model behind the agents, and a table of every SHOULD the platform does not implement, with the justification and a review date. A claim that omits that table while deviating from a SHOULD is invalid under the template's own rule.
The standard then gives the customer three ways to check a claim, in increasing order of assurance: review the vendor's completed checklist and evidence; ask the vendor to demonstrate key safety controls such as the kill switch, scope enforcement and rate limiting; or run the optional Customer Acceptance Testing procedures, which cover 39 of the 173 requirements in a staging environment. Some requirements are behavioral and cannot be verified from documents alone, which is why the demonstration step matters.
Using APTS in an RFP
The Vendor Evaluation Guide is written for CISOs and procurement teams, and it translates into an RFP almost directly.
- Set the minimum tier before you talk to vendors. The guide says organizations in financial services, healthcare, critical infrastructure or any regulated industry should require Tier 2 as a minimum, and recommends Tier 3 for critical infrastructure and L4 operation.
- State the autonomy levels you will use and ask the vendor to confirm the claim covers them and your deployment model.
- Request the completed checklist and a conformance claim citing the version (for example
APTS-v0.1.0), the assessment method and the SHOULD deviations table. - Ask the seven screening questions from the guide: which tier do you claim; provide your completed assessment against the checklists; can you demonstrate your safety controls live; how does your kill switch work and can we test it; what happens to our data after the engagement; do you deploy agents or software on our infrastructure, and can they be removed without you; which AI models does the platform use and how do you track changes to them.
- Score the red flags the guide lists: no kill-switch demo; "we handle scope internally" instead of accepting your rules of engagement; no access to the audit trail; vague AI model governance; no data isolation between customers; no confirmed credential disposal; no incident notification timeline.
- Plan two to four weeks. Week 1: define the tier, send the checklists, review documentation. Week 2: presentation, evidence requests and live demonstrations. Weeks 3 and 4, optional: hands-on Customer Acceptance Testing in staging, then an accept, conditional or reject decision.
- Write it into the contract: the claimed tier and version, notice of material changes (a new AI provider, a new deployment model, a new class of testing), and re-evaluation after major platform changes, incidents or changes of autonomy level, as the guide recommends.
Record the outcome with the tier, the date and any conditions. That record is what an auditor or a regulator will ask for when they want to know how the tool was chosen.
What APTS does not do
- It is not EU AI Act compliance. The standard says so directly: APTS conformance does not constitute EU AI Act compliance, although many requirements, especially in Human Oversight, Auditability and Reporting, address concerns that overlap with the Act's obligations. Treat it as supporting evidence, not as a substitute.
- It is not a certification. A conformance claim is an operator-provided assessment unless you or a third party verify it.
- It does not replace regulated testing regimes. A platform's APTS tier says how the tool is governed. It does not say whether a given test satisfies DORA threat-led penetration testing, the CMMC Level 3 penetration testing requirement or any other rule; those have their own criteria. See automated penetration testing for how these tools relate to regulated pentests.
- It is not stable yet. Version 0.1.0 of an Incubator project will change. Pin the version in contracts and re-check a claim when a new version is published.
- It does not rank vendors on effectiveness. A Tier 3 platform can still find less than a Tier 1 one. APTS covers how safely and accountably a platform operates; whether it finds what matters in your environment is a separate evaluation, the one that Adversarial Exposure Validation programs are built around.
Where Zero Hunt stands
Zero Hunt has not yet published an APTS self-assessment and does not claim any APTS tier or conformance. A public self-assessment against the checklists is planned; until it exists, treat the points below as statements to verify in a demonstration, not as conformance.
The controls the site already describes bear on three APTS domains:
- Human oversight: every campaign runs at one of five autonomy levels that decide what the agents may do alone and what waits for a person. At the lowest level any active scan needs approval; only the highest, chosen explicitly, runs without approval gates inside its scope. Actions outside the chosen level wait in a review tab where execution needs a recorded consent and a signed record of who approved it. Operators can pause or stop any campaign, and the final verdict on every finding is theirs. See human in the loop.
- Scope: campaigns start from the scope you authorize, and an action approved in the review tab gets a fresh scope check before it runs.
- Auditability: every attack attempt is recorded in a SHA-256 hash chain per campaign with each entry signed with Ed25519, verifiable offline, and exported reports are ECDSA-signed.
Two points relate to the supply-chain questions APTS asks. The appliance runs on-premise with its own ZeroHunt Apex models and no external AI API, so no customer data leaves the perimeter. And Zero Hunt's five autonomy levels are its own scale: they have not been mapped to the APTS L1 to L4 levels, which the self-assessment will have to do. To see these controls working in your environment, request a demo.
Sources
- OWASP Autonomous Penetration Testing Standard, project page (OWASP)
- OWASP APTS repository: README, domains, tiers and license (GitHub)
- APTS Introduction: compliance tiers, verification model, EU AI Act note (GitHub)
- APTS Graduated Autonomy: autonomy levels and the tier and level mapping (GitHub)
- APTS Checklists: all 173 requirements by domain and tier (GitHub)
- APTS Vendor Evaluation Guide (GitHub)
- APTS Conformance Claim Template (GitHub)
- APTS Customer Acceptance Testing (GitHub)
Goes deeper
Want this against your environment?
Book a 30-minute scoping call — we will map this directly to your current compliance scope and threat profile.