← Learn
Playbook11 min read

UK penetration testing requirements — CBEST for financial firms and the NCSC Cyber Assessment Framework

Short definition

How the Bank of England's CBEST intelligence-led testing and the NCSC Cyber Assessment Framework v4.0 treat penetration testing and vulnerability management: who they apply to, who may test, and what evidence counts.

Why this matters now

The NCSC released CAF v4.0 on 4 August 2025, and its vulnerability management outcome separates organisations that regularly test from those that also verify their understanding with third-party testing. In financial services, CBEST remains the regulators' instrument for intelligence-led testing of live systems at the firms they select, run only by accredited providers. Neither framework sets an annual pentest rule, so the reasoning behind your own testing cadence is part of what gets assessed.

Key points

  • ▸CBEST is requested by the regulators as part of the supervisory cycle (the list is agreed by the PRA and FCA), agreed at a firm's own request, or requested after an incident.
  • ▸CBEST tests live systems behind important business services; the threat intelligence and penetration testing providers must be CBEST-accredited and CREST members, with certified individuals leading.
  • ▸A CBEST takes around 9 to 12 months, including about 14 weeks of penetration testing, and ends with a remediation plan that the regulator tracks.
  • ▸CAF v4.0 B4.d: “achieved” means regular testing to fully understand vulnerabilities, verified with third-party testing, and prompt mitigation of announced vulnerabilities.
  • ▸CAF v4.0 A2.c: assurance methods chosen with an understanding of their limits, including the risks of penetration testing in operational environments.
  • ▸The NCSC has no regulatory responsibilities; organisations should ask their regulator whether and how the CAF applies to them.

Two instruments with different jobs

The UK has no single penetration testing rule. Two frameworks matter to most regulated buyers, and they answer different questions.

  • CBEST is a supervisory assessment for the financial sector. Since 2014 it has been part of the collective toolkit of the Bank of England, the Prudential Regulation Authority (PRA) and the Financial Conduct Authority (FCA) to assess the cyber resilience of firms and financial market infrastructures (FMIs). It is an intelligence-led penetration test of the systems behind important business services, done at the regulators' request, by accredited providers.
  • The Cyber Assessment Framework (CAF) is the NCSC's method for assessing how well an organisation manages cyber risk to its essential functions. It sets outcomes, not tests, and it is used by organisations themselves or by independent assessors, including regulators.

A bank may face both: CBEST from its supervisors, and CAF-style outcomes wherever its essential functions are assessed against them. Neither replaces the ordinary vulnerability management and testing that both assume you already run.

CBEST: who, when and who may test

The PRA's CBEST Implementation Guide, 2024 edition, sets out the process.

When CBEST happens. A firm or FMI carries out a CBEST when (1) the regulator requests it as part of the supervisory cycle, from a list the PRA and FCA agree regularly in line with thematic focus and supervisory strategy; (2) the firm asks to run one as part of its own cyber resilience programme and the regulator agrees; or (3) an incident leads the regulator to request one to support remediation and validation. The Bank of England describes it as improving the resilience of systemically important firms and, through them, of the wider financial system.

Who tests. Four parties are involved: the regulator, the firm's Control Group, a threat intelligence service provider (TISP) and a penetration testing service provider (PTSP). Both providers must be CBEST-accredited, through an accreditation process undertaken by the Bank of England, and must be members of CREST, which acts as the CBEST accreditation and certification body. Accredited providers must employ certified individuals: a CREST Certified Threat Intelligence Manager for the TISP, and a CREST Certified Red Team Manager or Cyber Scheme Red Team Manager with CREST Certified Red Team Specialists for the PTSP.

What is tested. An intelligence-led penetration test, using manual and automated techniques, against the systems that underpin each in-scope important business service, covering the end-to-end processes and systems unless otherwise agreed. Testing is on live systems, so the PTSP prepares a risk management plan, and the Control Group can order a temporary halt at any time. Secrecy is kept inside the Control Group; informing system owners or the SOC is listed as manipulation of the process.

How long. The regulators find that a CBEST averages 9 to 12 months: Initiation around 6 weeks, Threat Intelligence around 10 weeks, Penetration Testing around 14 weeks, and Closure around 4 weeks.

What follows. CBEST is not pass/fail. The firm drafts a remediation plan, agrees it with the regulator, and the regulator tracks its implementation, typically over six to 12 months or longer. The regulators also publish anonymised thematic findings; the 2024 guide notes the value of simulating highly privileged internal attackers, such as malicious insiders and supply chain attacks, and the importance of a strong foundation of cyber hygiene.

The CAF v4.0: the outcomes that involve testing

CAF v4.0 was released on 4 August 2025. It has four objectives, A to D, and 14 principles, each broken into contributing outcomes with indicators of good practice marked not achieved, partially achieved or achieved. Assessments can be done by the organisation itself or by an independent external entity, such as a regulator or an NCSC-assured provider acting for one. Two outcomes are about testing.

B4.d Vulnerability management (principle B4, System security). To be achieved, all of these must be true:

  • You maintain a current understanding of the exposure of your essential functions to publicly known vulnerabilities.
  • Announced vulnerabilities for all software packages, network and information systems supporting essential functions are tracked, prioritised and mitigated (for example by patching) promptly.
  • You regularly test to fully understand the vulnerabilities of the systems that support your essential functions and verify this understanding with third-party testing.
  • You actively maximise the use of supported software, firmware and hardware.

The partially achieved level already requires regular testing, but not the third-party verification, and allows temporary mitigations for vulnerabilities that are not externally exposed. Not achieved includes not having recently tested to verify your understanding of vulnerabilities, and not mitigating externally exposed vulnerabilities promptly. The B4 guidance lists regular vulnerability and security assessments, for example penetration tests and vulnerability scans, among the effective methods, and asks operators to consider their approach to testing live operational technology carefully; assurance can also come from non-operational environments or laboratory testing of components.

A2.c Assurance (principle A2, Risk management). To be achieved, you validate that security measures are effective and remain effective, choose appropriate assurance methods knowing their strengths and limitations, can justify your confidence to a third party who can verify it, and remedy deficiencies found by assurance activities in a timely and effective way. Among the not-achieved indicators: applying assurance methods without appreciating their limits, such as the risks of penetration testing in operational environments, and assuming assurance because there have been no known problems.

The NCSC is explicit about its role: it has no regulatory responsibilities, and organisations subject to cyber regulation should consult their regulators to learn whether they should use the CAF to meet regulatory requirements.

What the NCSC says about penetration testing

The NCSC's penetration testing guidance explains how the CAF expects testing to be used, and it is blunt:

  • Penetration testing should be viewed as a method for gaining assurance in your vulnerability assessment and management processes, not as a primary method for identifying vulnerabilities. The NCSC compares it to a financial audit of processes your own team runs every day.
  • Ideally you know what the testers will find before they find it; their report should improve your internal processes.
  • A test only validates that systems are not vulnerable to known issues on the day of the test, and it is not uncommon for a year or more to pass between tests.
  • Third-party tests should be performed by qualified and experienced staff only, and the NCSC recommends that UK government (HMG) organisations use testers and companies in the CHECK scheme.
  • Vulnerability risk assessment and mitigation is a business process and should not be wholly outsourced to the test team.

Read together with B4.d, the model is clear: continuous internal vulnerability management does the finding, and periodic third-party testing verifies that it works.

What is explicit and what is left to you

Explicit:

  • CBEST, when requested: accredited TISP and PTSP, certified leads, live testing of systems behind important business services, a remediation plan tracked by the regulator.
  • CAF B4.d at achieved: regular testing to understand vulnerabilities, verified by third-party testing, and prompt mitigation of announced vulnerabilities.
  • CAF A2.c: assurance methods chosen with an understanding of their limits, and deficiencies remedied in a timely way.
  • For HMG organisations, the NCSC's recommendation to use CHECK testers.

Not written in these texts:

  • A fixed frequency. CBEST happens when the regulator requests it; the CAF says “regularly” and “recently”.
  • An accreditation requirement for ordinary testing outside CBEST. The CAF asks for third-party verification at the achieved level of B4.d; it does not name a scheme.
  • A rule that internal or automated testing is enough on its own for B4.d at achieved. The third-party verification is a separate indicator.

If you are subject to a sector regulator, what it asks for, and which CAF profile it expects, is the thing to confirm first.

The evidence to keep

  • For CBEST: scope specification, provider accreditation and certification records, test and risk management plans, the final penetration test report, the remediation plan and the updates you send the regulator until each action is closed.
  • For B4.d: your record of exposure to publicly known vulnerabilities, tracking and prioritisation of announced vulnerabilities with mitigation dates, the schedule and results of your own regular testing, and the third-party test reports that verify it.
  • For A2.c: the assurance methods you chose for each essential function and why, including how you handled the risks of testing operational environments, and the remediation of deficiencies they found.
  • Unsupported technology: the list, the temporary mitigations and the migration plan, since B4.d looks at both.

How continuous autonomous pentesting on an on-premise appliance fits

An autonomous AI red team for networks and infrastructure is not a CBEST provider, is not accredited, and does not supply threat intelligence. Tests you run yourself are also not the third-party verification that B4.d asks for at the achieved level. What it supports is the part the NCSC says should do the finding: your own regular testing. For the category, see automated penetration testing.

  • Regular testing between third-party tests. Black-box campaigns from outside and gray-box campaigns with standard-user access, on a schedule and after changes, so you know what a third-party tester or a CBEST team is likely to find before they arrive.
  • The inside view. Gray-box campaigns from a standard account show how far lateral movement goes today, the question CBEST's thematic findings point to with insider and supply chain scenarios.
  • Care with operational environments. Scope validation, an emergency stop and five autonomy levels define what the agents may do alone and what waits for an operator, and exploit proofs of concept can be held for review, which is the kind of control A2.c expects when testing live systems. See human in the loop.
  • Framework mapping and records. CBEST and UK CAF are among the 34 frameworks the engine maps findings to, with exportable compliance reports; tying them to B4.d and A2.c stays in your documentation. Every attack attempt is an Ed25519-signed entry in a SHA-256 hash chain per campaign, verifiable offline.
  • Findings stay on site. The models run on the appliance with no external AI service, and no customer data leaves it. See on-premise AI red team and the on-premise pentest platform buyer's guide.

Sources

Goes deeper

Want this against your environment?

Book a 30-minute scoping call — we will map this directly to your current compliance scope and threat profile.