← Learn
Playbook14 min read

Penetration Testing Requirements by Regulation: A Worldwide Guide (2026)

Short definition

A regulation-by-regulation map of who must run vulnerability assessments, penetration tests or threat-led red teaming, how often, who may test and what evidence to keep, with a detailed guide for each regime.

Why this matters now

Many regulated organizations answer to several of these regimes at the same time, and the texts differ on the point buyers most often get wrong. Only a few name the penetration test and fix its frequency; most require testing whose type and interval follow your own risk assessment, and some are guidance that supervisors compare you against. Knowing which is which decides both the testing budget and the evidence you must be able to show.

Key points

  • ▸Explicit rules name the test and a frequency: NYDFS §500.5 (annual, from inside and outside), PCI DSS 11.4 (annual and after significant changes), Saudi CSCC (every six months on critical systems) and CMMC Level 3 (annual).
  • ▸Risk-based rules make testing binding but let you set type and interval: NIS2 Art. 21, APRA CPS 234, Saudi ECC 2-11 and the HIPAA risk analysis; DORA Art. 24(6) adds a yearly floor for systems supporting critical or important functions.
  • ▸Supervisory guidance is not law but sets the benchmark: MAS TRM expects an annual pentest of internet-facing systems, and APRA CPG 234 and the NCSC CAF describe what good testing looks like.
  • ▸Threat-led red teaming is reserved for firms the regulators pick: DORA TLPT at least every three years and CBEST on request, both with strict rules on who may test.
  • ▸Few texts require an external tester for ordinary testing; most ask for qualified and independent testers, who can be internal. PCI ASV scans, CBEST providers, the TLPT threat intelligence provider and CAF third-party verification are the main exceptions.
  • ▸One program can serve several regimes if it keeps the same six records: scope, methodology, tester independence, findings, re-tests and sign-off.

Explicit obligation or risk-based expectation: the distinction that decides the budget

Regulations talk about penetration testing in three different ways, and the difference changes both what you must buy and what you must write down.

  • Explicit obligation: the text names the test and sets a frequency or a trigger. NYDFS §500.5(a)(1), PCI DSS 11.3 and 11.4, the Saudi CSCC for critical systems, CMMC Level 3 and, for the firms that are designated, DORA TLPT. If the interval is missed, no risk assessment makes up for it.
  • Risk-based obligation: testing or vulnerability management is binding, but the type and frequency follow your risk assessment. NIS2 Art. 21, the DORA testing programme (with a yearly floor on critical or important functions and weekly scans), APRA CPS 234, the HIPAA risk analysis and evaluation, Saudi ECC 2-11 (“periodically”) and, in Italy, the ACN measures for essential entities (“periodically and before go-live”).
  • Supervisory expectation: guidance that is not law but that supervisors compare you against, such as MAS TRM section 13, APRA CPG 234, the NCSC Cyber Assessment Framework, NIST SP 800-66r2 and HHS 405(d) HICP for healthcare.

The risk-based rules are not softer. When a regulation lets you choose the frequency, the examiner checks that the choice is written down, follows from your risk assessment and was actually applied. And one regime in this guide requires no testing at all: the SEC disclosure rules, which still expect the processes a company describes each year to exist.

Each item below links to a detailed guide that quotes the text clause by clause and cites its primary sources. Where the lists say “Who may test”, they describe what the text requires, not who is best placed to do it.

European Union: NIS2 and DORA

NIS2, with Implementing Regulation (EU) 2024/2690 and Italy's ACN measures. Detailed guide: NIS2 penetration testing and VA requirements.

  • Scope: essential and important entities under Directive (EU) 2022/2555. Regulation 2024/2690 binds a listed group of digital providers: DNS, TLD registries, cloud, data centers, CDNs, MSPs, MSSPs, online marketplaces, search engines, social networks and trust services.
  • What is required: vulnerability handling and disclosure, and policies to assess the effectiveness of measures (Art. 21(2)(e) and (f)); the Directive never names the penetration test. Regulation 2024/2690 requires a security testing policy and, where appropriate, vulnerability scans. In Italy, ACN measure ID.RA-01 requires essential entities to run vulnerability assessment and/or penetration testing on at least their relevant systems.
  • Frequency: risk-based; “at planned intervals” under 2024/2690; in Italy “periodically and in any case before go-live”, with the ACN measures due 18 months after the listing communication (October 2026 for entities listed in 2025, 31 July 2027 for those first listed in 2026).
  • Who may test: no text requires an external or certified tester. Regulation 2024/2690 asks for audit competence and independence for the independent review of the overall approach, not for each test.
  • Evidence: VA/PT reports describing each vulnerability and its impact (ID.RA-01), a board-approved vulnerability management plan (ID.RA-08), per-finding criticality and mitigation (2024/2690 point 6.5), remediation and re-test records.

DORA testing programme, Articles 24 and 25. Detailed guide: DORA penetration testing requirements beyond TLPT.

  • Scope: every financial entity under Regulation (EU) 2022/2554 except microenterprises, which test under a lighter risk-based approach (Art. 25(3)).
  • What is required: a risk-based testing programme; Art. 25(1) lists the tests it provides for, from vulnerability scans and source code reviews to penetration testing. The regulation does not require a penetration test of every system every year, but scans alone are hard to defend where exploitability is the question.
  • Frequency: at least yearly, appropriate tests on all ICT systems and applications supporting critical or important functions (Art. 24(6)); automated vulnerability scanning of those assets at least weekly (RTS 2024/1774, Art. 10(2)).
  • Who may test: independent parties, internal or external; internal testers need sufficient resources and no conflicts of interest (Art. 24(4)).
  • Evidence: the map from critical or important functions to systems, a test matrix, independence records, and a findings register with a validation step that confirms each fix (Art. 24(5)).

DORA threat-led penetration testing (TLPT), Articles 26 and 27. Detailed guides: TLPT under DORA and the DORA TLPT engagement playbook.

  • Scope: only the financial entities the authorities identify, using the criteria and default list of RTS 2025/1190, for example G-SII and O-SII credit institutions, central securities depositories, central counterparties and the largest payment and e-money institutions.
  • What is required: an intelligence-led red team test on live production systems supporting critical or important functions, with an active red team phase of at least 12 weeks, then a replay and a purple teaming exercise.
  • Frequency: at least every three years. From notification to attestation, the engagement realistically takes 12 to 18 months.
  • Who may test: the threat intelligence provider is always external; internal testers only with the authority's approval, and external testers every three tests; significant credit institutions under the SSM use external testers only. Providers must meet the experience, reference and insurance thresholds of RTS Art. 5(2).
  • Evidence: a scope specification approved by the management body, red team and blue team reports, a summary report, a remediation plan with root causes, and the authority's attestation (DORA Art. 26(7)).

United States, plus PCI DSS worldwide

NYDFS 23 NYCRR Part 500, §500.5. Detailed guide: NYDFS Part 500 penetration testing.

  • Scope: covered entities licensed or authorized under the New York Banking, Insurance or Financial Services Law. Entities with the §500.19(a) limited exemption are exempt from §500.5.
  • What is required: penetration testing from both inside and outside the information systems' boundaries; automated scans plus a manual review of systems the scans do not cover; monitoring for new vulnerabilities; timely remediation prioritized by risk.
  • Frequency: penetration test at least annually; scans at a frequency set by the risk assessment and promptly after any material system change.
  • Who may test: a qualified internal or external party. Part 500 does not define “qualified”, so document experience, independence and methodology.
  • Evidence: test reports showing both perspectives, scan records and remediation records behind the annual filing due 15 April, which the highest-ranking executive and the CISO sign; records are kept for five years.

HIPAA Security Rule. Detailed guide: HIPAA penetration testing requirements.

  • Scope: covered entities and business associates that handle electronic protected health information (ePHI).
  • What is required: today, an accurate and thorough risk analysis (§164.308(a)(1)(ii)(A)) and a periodic technical and nontechnical evaluation (§164.308(a)(8)); the rule never names penetration testing or scanning. The January 2025 HHS proposal would add both.
  • Frequency: none fixed in the current rule. The proposal would require scans at least every six months and a penetration test at least every 12 months; as of 30 September 2026 it is not final, and the latest regulatory agenda lists final action for July 2027.
  • Who may test: no rule requires an external tester. The proposal asks for a qualified person, and HHS 405(d) HICP allows qualified internal staff or external partners.
  • Evidence: inventory, risk analysis, scan and test results, remediation and re-tests, kept for six years (§164.316(b)). HHS must consider recognized security practices in place for the previous 12 months when setting penalties (Public Law 116-321).

PCI DSS v4.0.1, Requirements 11.3 and 11.4 (card payments, worldwide). Detailed guide: PCI DSS 11.4 penetration testing requirements.

  • Scope: entities that store, process or transmit payment card account data. Whether and how they validate compliance is set by the payment brands and acquirers, not by PCI SSC.
  • What is required: internal and external vulnerability scans (11.3); internal and external penetration tests under a documented methodology (11.4.1 to 11.4.3); repeated testing to verify corrections (11.4.4); segmentation tests (11.4.5).
  • Frequency: scans at least every three months and after significant changes; penetration tests at least every 12 months and after significant changes; segmentation tests every 12 months, or every six months for service providers (11.4.6).
  • Who may test: a qualified internal resource or external third party with organizational independence, not necessarily a QSA or ASV. The quarterly external scans must come from a PCI SSC Approved Scanning Vendor.
  • Evidence: the methodology, 12 months of test reports with tester qualifications, four quarters of passing ASV scans and internal scans, remediation and re-test records, and change records; results kept for at least 12 months.

CMMC, 32 CFR Part 170. Detailed guide: CMMC Phase 2 assessment evidence.

  • Scope: defense contractors whose information systems process, store or transmit FCI or CUI. From 10 November 2026 (Phase 2), DoD intends to require Level 2 (C3PAO) status as a condition of award in applicable solicitations.
  • What is required: at Level 2 (NIST SP 800-171 Rev. 2), vulnerability scanning (3.11.2), remediation according to risk (3.11.3) and periodic control assessments (3.12.1); Level 3 adds penetration testing (3.12.1e).
  • Frequency: “periodically” means an interval you define of no more than one year; Level 3 penetration testing at least annually or when significant security changes are made.
  • Who may test: the contractor runs the testing; a C3PAO assesses Level 2 and DCMA DIBCAC assesses Level 3. For 3.12.1e, NIST expects a team with the right skills that is objective in its assessment.
  • Evidence: final-form artifacts, hashed and kept for six years (§170.17(c)(4)), including the defined frequencies, scan results, plans of action and test reports.

SEC cybersecurity disclosure rules. Detailed guide: SEC Form 8-K Item 1.05 playbook.

  • Scope: companies that report under the Securities Exchange Act of 1934.
  • What is required: no testing. A Form 8-K Item 1.05 filing within four business days of determining that an incident is material, and an annual description of cybersecurity risk management under Regulation S-K Item 106.
  • Frequency: none for testing; the Item 106 description appears in every annual report.
  • Who may test: not addressed. Item 106 asks whether the company engages assessors, consultants, auditors or other third parties.
  • Evidence: if the annual report says the company runs regular penetration tests, keep the dated reports, remediation records and board reporting that show it.

United Kingdom: CBEST and the NCSC CAF

CBEST, the Bank of England, PRA and FCA intelligence-led assessment. Detailed guide: UK CBEST and CAF penetration testing.

  • Scope: firms and financial market infrastructures that the regulators select in the supervisory cycle, that request a CBEST with the regulator's agreement, or that are asked for one after an incident.
  • What is required: an intelligence-led penetration test of the live systems behind important business services, followed by a remediation plan that the regulator tracks. CBEST is not pass/fail.
  • Frequency: when requested, with no fixed cycle. A CBEST averages 9 to 12 months, including about 14 weeks of penetration testing.
  • Who may test: CBEST-accredited threat intelligence and penetration testing providers that are CREST members, led by certified individuals.
  • Evidence: scope specification, provider accreditation and certification records, test and risk management plans, the final report, the remediation plan and the updates sent to the regulator.

NCSC Cyber Assessment Framework (CAF) v4.0. Same detailed guide.

  • Scope: organizations whose management of cyber risk to essential functions is assessed, by themselves or by independent assessors such as regulators. The NCSC has no regulatory powers, so ask your regulator whether and how the CAF applies to you.
  • What is required: outcome B4.d (vulnerability management) expects regular testing to fully understand vulnerabilities and prompt mitigation; A2.c (assurance) expects methods chosen with their limits understood, including the risks of penetration testing in operational environments.
  • Frequency: “regularly” and “recently”, with no fixed interval.
  • Who may test: your own team for regular testing; “achieved” on B4.d also requires verification by third-party testing. The NCSC recommends CHECK testers for UK government (HMG) organisations.
  • Evidence: your record of exposure to known vulnerabilities, tracking and mitigation dates, the results of your own testing, third-party test reports, and the reasoning behind each assurance method.

Asia-Pacific: Singapore and Australia

Singapore: MAS Technology Risk Management Guidelines, section 13. Detailed guide: MAS TRM penetration testing requirements.

  • Scope: financial institutions supervised by MAS. The Guidelines are not law, but MAS considers how far an institution observes their spirit; for banks, the binding Notice FSM-N06 requires patch timeframes that follow the risk of each vulnerability.
  • What is required: regular vulnerability assessment with a minimum scope; penetration testing on production with proper safeguards; black-box and grey-box testing for online financial services; adversarial attack simulation exercises; a remediation process with timeframes by severity.
  • Frequency: systems directly accessible from the Internet at least once a year and whenever they undergo major changes (13.2.4); other testing at a frequency set by criticality and exposure.
  • Who may test: no particular kind of tester is specified; the test has to deliver the in-depth evaluation that 13.2.1 asks for.
  • Evidence: VA records, PT reports with test type and production safeguards, the inventory of internet-facing systems with test dates, exercise records and the remediation register.

Australia: APRA Prudential Standard CPS 234. Detailed guide: APRA CPS 234 security testing requirements.

  • Scope: all APRA-regulated entities, including ADIs, general insurers, life companies, private health insurers and RSE licensees, with their non-operating holding companies.
  • What is required: a systematic testing program for information security controls, binding since 1 July 2019. CPS 234 never names the penetration test; APRA's guidance CPG 234 lists penetration tests and red team tests among the techniques.
  • Frequency: commensurate with threats, criticality, consequences, untrusted exposure and change (para 27), with the program reviewed at least annually (para 31). CPG 234 expects a sufficient set of controls tested at least annually and controls exposed to untrusted environments tested throughout the year.
  • Who may test: appropriately skilled and functionally independent specialists (para 30). No external tester is required; CPG 234 reads independence as testers without operational responsibility for the controls they test.
  • Evidence: the program and its reasoning, test reports with tester roles, independence records, findings and escalations. A material control weakness that cannot be fixed in a timely manner must be notified to APRA within 10 business days (para 36).

Middle East: Saudi Arabia

Saudi Arabia: NCA Essential Cybersecurity Controls (ECC-2:2024) and Critical Systems Cybersecurity Controls (CSCC). Detailed guide: Saudi NCA ECC penetration testing requirements.

  • Scope: government agencies and their affiliated companies and entities, inside and outside the Kingdom, and private entities that own, operate or host Critical National Infrastructure. The CSCC adds controls for systems the organization deems critical.
  • What is required: periodic vulnerability assessment with severity-based remediation (ECC 2-10); penetration tests covering at least all externally provided services and their technical components (ECC 2-11). For critical systems, the CSCC extends the scope to all internal and external services.
  • Frequency: “periodically” under the ECC. Under the CSCC, vulnerability assessment at least monthly and penetration tests at least every six months on critical systems.
  • Who may test: the ECC is silent and the CSCC asks for a qualified team. ECC 1-8 requires an independent review and audit of the controls by parties other than the cybersecurity department.
  • Evidence: approved requirements documents, the list of critical systems, VA records that show the monthly cadence, penetration test reports with scope, date and team, and the independent audit results.

The UAE and other Gulf states are not covered in this guide yet.

How to build one testing program that satisfies several regimes

Most of the regimes above ask for the same six things in different words. Build the program around them once, then layer each regime's frequency and tester rules on top.

  1. Scope from one inventory. Every regime starts from a list: relevant systems (ACN), systems supporting critical or important functions (DORA), the cardholder data environment and critical systems (PCI DSS), systems that touch ePHI (HIPAA), systems directly accessible from the Internet (MAS TRM), critical systems (CSCC). Keep one asset inventory, tag each asset with the regimes that apply, and write down every exclusion and the reason for it.
  2. A written methodology. PCI DSS 11.4.1 requires one by name, Regulation 2024/2690 asks for tests under a documented methodology, and TLPT follows the RTS. Cover testing from inside and outside (NYDFS, PCI DSS), black-box and grey-box testing (MAS TRM), rules of engagement and the safeguards for production systems.
  3. Frequency set by the strictest regime that applies. Where an asset falls under several regimes, the shortest interval wins: for example weekly scans on DORA critical-function assets, quarterly scans under PCI DSS, monthly assessment and six-monthly tests on Saudi critical systems, and at least an annual penetration test under NYDFS and PCI DSS. Record trigger events too: significant or material changes call for new scans or tests under PCI DSS, NYDFS and MAS TRM.
  4. Independence on record. For every test, who ran it, their qualifications, and why they are independent of the systems under test. Where a regime names an accredited or external party (ASV scans, CBEST providers, the TLPT threat intelligence provider, CAF third-party verification), that party is not optional.
  5. One findings register. Exploitability, risk-based priority, owner, target date, and any accepted risk with its approval. DORA Art. 24(5), NYDFS §500.5(c), MAS TRM 13.6 and CMMC 3.12.2 all ask for this record.
  6. Re-test, then sign-off. A finding closes when the re-test passes (PCI DSS 11.4.4, DORA Art. 24(5)). Sign-off looks different in each regime: a board-approved vulnerability management plan (ACN), a TLPT scope approved by the management body (DORA), the annual filing signed by the highest-ranking executive and the CISO (NYDFS), escalation to the Board or senior management (APRA), the annual affirmation (CMMC). Keep the approvals with the evidence.

Retention follows the same rule: the longest period wins. Six years for HIPAA documentation and CMMC artifacts, five years for NYDFS records, at least 12 months for PCI DSS test results.

Where continuous autonomous pentesting on an on-premise appliance fits

Zero Hunt is an autonomous AI red team for networks and infrastructure that runs on an on-premise appliance, on private AI, with a human in the loop. Across all of these regimes it does the same job: it produces testing evidence in the months between the tests that regulators and assessors require from named parties. For the category itself, see automated penetration testing.

  • Continuous black-box and gray-box testing. Campaigns from outside and with standard-user access run once, daily, weekly, monthly or on a custom schedule, and after changes, so trigger-based tests do not wait for an engagement slot.
  • Findings proven exploitable, then re-tested. Each finding records whether it was exploited in your environment, which is the input risk-based remediation needs, and each fix can be re-tested with the same proof.
  • Evidence mapped to 34 frameworks. They include NIS2, DORA and TIBER-EU, NYDFS Part 500, HIPAA, PCI DSS 4.0.1, CMMC, the SEC rules, MAS TRM, APRA CPS 234, CBEST, UK CAF and Saudi NCA, with exportable compliance reports. The engine does not map to Italian ACN measure codes, and some of its control references do not follow the regulators' numbering, so tie reports to paragraphs in your own documentation.
  • Human approval of risky steps. Five autonomy levels define what waits for an operator, and exploit proofs of concept can be held for review. See human in the loop.
  • Records that hold up. Every attack attempt is an Ed25519-signed entry in a SHA-256 hash chain per campaign, verifiable offline, and exported reports are ECDSA-signed.
  • No customer data leaves the appliance. The models run on it with no external AI service, so testing does not add a third party that you would have to assess; in air-gapped mode it also stops downloads from public sources. See on-premise AI red team.

What it does not replace: the PCI ASV, a QSA or a C3PAO, CBEST-accredited providers, TLPT testers and the external threat intelligence provider, the third-party verification that CAF B4.d asks for, the independent reviews required by Regulation 2024/2690 and Saudi ECC 1-8, or the qualified, independent tester that several regimes name. It does not make your team independent and does not decide whether you comply. To check which of these regimes apply to you and where your evidence has gaps, book a 30-minute readiness call.

Sources

Goes deeper

Want this against your environment?

Book a 30-minute scoping call — we will map this directly to your current compliance scope and threat profile.