AI Strategy, Cybersecurity, Compliance Automation & Microsoft 365 Managed IT for Security-First Financial Institutions | ABT Blog

Microsoft 365 Penetration Testing for Financial Institutions

Written by Justin Kirsch | Thu, Sep 17, 2026

Every year your institution buys a penetration test, and every year the report comes back and goes in the examination file. The scope statement usually reads something like this: external network perimeter, internal network, the public website, the online banking portal, and a phishing simulation against staff. The tester finds a few things. You fix them. The finding gets closed.

Now look at where an attacker actually ends up when a bank, a credit union, or a mortgage company loses money. Not the perimeter. Somebody's mailbox. A session token lifted through a fake sign-in page. A forwarding rule nobody saw. An application consent granted in about four seconds by an employee who thought they were opening a shared document. All of that lives inside the Microsoft 365 tenant, and the tenant is frequently the one environment the penetration test never touched.

That gap is not a failure of diligence by the people who bought the test. It comes from a real structural problem: Microsoft 365 is a shared platform that Microsoft operates, so the usual penetration testing playbook does not apply to it cleanly, and the rules about what you may do to it are set by Microsoft rather than by your tester. The result is that the most valuable part of your environment gets assessed with the wrong instrument, or with none at all.

$3.05B
Business email compromise losses reported to the FBI's Internet Crime Complaint Center in 2025, out of $20.877 billion in total reported losses across 1,008,597 complaints. Reported to IC3 across all victim types, not financial institutions specifically.

What a penetration test is, in your regulator's words

It helps to start with the definition your examiner is working from, because it is narrower and more honest than the way the word gets used in sales conversations.

A penetration test subjects a system to real-world attacks selected and conducted by the testers. A penetration test targets systems and users to identify weaknesses in business processes and technical controls. The test mimics a threat source's search for and exploitation of vulnerabilities to demonstrate a potential for loss.

Section IV.A.2(b), Penetration Tests

Two things follow from that definition, and both of them get lost in practice.

The first is that a penetration test is a demonstration, not an inventory. It proves that a particular path from A to B is walkable. It does not produce a complete list of everything wrong with your environment, and it was never designed to. A tester who finds one way in usually stops looking for the other six, because the finding has been demonstrated and the clock is running.

The second is that the FFIEC says the quiet part out loud in the very next breath. The booklet notes that some tests "focus on only a subset of the institution's systems and may not accurately simulate a determined threat actor." That is the regulator, in its own guidance, telling you that a narrow scope produces a narrow assurance. Whether your scope was narrow is a question about your scope statement, not about the quality of the tester.

Notice also what the guidance does not say. It sets no universal annual cadence. It says the frequency and scope "should be a function of the level of assurance needed by the institution and determined by the risk assessment process." Your testing calendar is supposed to be an output of your risk assessment. If your risk assessment says the Microsoft 365 tenant is where your member and borrower data lives, and your testing calendar never looks at it, those two documents are contradicting each other, and an examiner reading both will notice.

What your regulator actually requires

This is where precision matters, because the requirement genuinely differs depending on who supervises you, and a lot of published advice blurs the two.

Two different rulebooks, depending on who examines you

Banks and federally insured credit unions. Your obligation runs through your prudential regulator and the interagency guidance in the FFIEC IT Examination Handbook. The expectation is a risk-based testing and assurance program combining self-assessments, penetration tests, vulnerability assessments, and audits, with coverage, depth, and independence appropriate to the risk. There is no fixed federal "annual penetration test" mandate in that guidance. There is an expectation that you can explain why your testing program is proportionate to your risk.

Mortgage companies, mortgage brokers, and non-federally-insured credit unions. You are a financial institution under the FTC's jurisdiction, and the Safeguards Rule is prescriptive where the FFIEC guidance is principles-based. Under 16 CFR 314.4(d) you either implement continuous monitoring, or you conduct annual penetration testing plus vulnerability assessments, "including system-wide scans every six months designed to test for publicly-known security vulnerabilities." Testing is also required whenever there are material changes to your operations or business arrangements.

Worth being careful with the second half of that. The Safeguards Rule applies to financial institutions under FTC jurisdiction that are not subject to another regulator's enforcement authority under section 505 of the Gramm-Leach-Bliley Act. Banks and federally insured credit unions have prudential regulators, so they sit outside it. That does not mean they carry a lighter obligation. It means their obligation is written differently, and it is judged by an examiner applying professional judgment rather than by a clock.

For an institution that owns a mortgage subsidiary, both rulebooks can apply inside the same organization, to different entities, on different schedules. That is worth mapping before your next exam rather than during it. If you have never worked through whether the FTC's definition reaches a particular entity, our walkthrough of who counts as a financial institution under the Safeguards Rule covers the scope question on its own.

What Microsoft permits you to test

Here is the part that surprises people who have only ever scoped tests against infrastructure they own. You do not get to point an arbitrary testing methodology at Microsoft 365, because Microsoft 365 is not yours. Microsoft owns and operates the platform. You own a tenant on it.

Microsoft publishes rules of engagement governing security testing of its cloud services, and they are permissive in one direction and firm in the other. Testing within your own tenant is explicitly encouraged. Microsoft's guidance states that all encouraged testing activities "must be performed within your own tenant or assets for which you have explicit authorization." The old regime where you filed a notification form and waited for approval is gone. You can test what is yours.

What Microsoft prohibits includes, and this list is not exhaustive, so read the current rules before you scope anything:

  • Denial-of-service testing, in any form, against the platform.
  • Accessing customer or Microsoft data, or testing customer systems, without explicit permission. Your tenant only. Not a neighbor's, not Microsoft's.
  • Using, accessing, or retrieving credentials or other secrets that are not your own.
  • Network-intensive fuzzing or automated testing that generates excessive traffic.
  • Phishing or social engineering attacks targeting Microsoft employees. Phishing your own staff, with your own authorization, is a different matter entirely and is a normal part of a security program.
  • Post-compromise or post-exploit actions such as enumerating internal networks and files, or dumping secrets.

Read that list next to a standard penetration testing methodology and the friction becomes obvious. A large share of what a tester normally does after gaining a foothold, the enumeration, the lateral movement, the credential harvesting, is off the table on a platform Microsoft runs for thousands of other tenants at the same time. A tester who respects these rules produces a thinner report on your tenant than on your network. A tester who ignores them creates a contractual and legal problem for you, not for themselves.

So the honest conclusion is not that Microsoft 365 cannot be tested. It is that the penetration test is the wrong primary instrument for it, and reaching for a bigger penetration test is reaching for the wrong tool harder.

A traditional penetration test scope and the Microsoft 365 tenant configuration layer overlap far less than most scope statements assume.

The layer the test usually skips

The attacks that empty accounts at financial institutions are not, for the most part, exploitation of unpatched network services. They are identity attacks, and identity attacks land in a layer that sits above the network and below the application: your tenant's configuration.

Consider the adversary-in-the-middle phishing pattern. An employee clicks a link, lands on a proxy that looks exactly like the Microsoft sign-in page, and enters their credentials. The proxy relays them to Microsoft in real time. Microsoft prompts for multifactor authentication. The employee approves it, because from where they are standing everything looks correct. The proxy captures the resulting session token, and the attacker replays that token. Your multifactor authentication control worked perfectly and the attacker is inside anyway, because they did not defeat the control, they waited for the employee to satisfy it and then stole the receipt.

Tier-1 Cloud Solution Provider (CSP) ABT Partner Insight

Microsoft counted 10.7 million business email compromise attacks in the first quarter of 2026. Tycoon2FA, one of the adversary-in-the-middle phishing kits sold as a service, surged 44 percent in February and then fell 15 percent in March, after Microsoft's Digital Crimes Unit disrupted its infrastructure with Europol and industry partners. It did not stop. Its domains moved registration instead.

ABT manages Microsoft 365 tenants for more than 750 credit unions, banks, and mortgage companies as a Tier 1 Microsoft Cloud Solution Provider, which means the controls that decide whether one of these attacks succeeds are the same controls we configure and monitor on a daily basis. Every item in the tenant configuration layer described below is part of that baseline.

A network penetration test does not see any of this, and it is not supposed to. The controls that actually decide whether that attack succeeds are all tenant settings: whether Conditional Access requires a compliant or hybrid-joined device, whether phishing-resistant authentication methods are enforced for the accounts that matter, whether token protection is configured, whether sign-in frequency is bounded for high-risk sessions, whether legacy authentication protocols are still reachable, and whether risk-based policies can actually fire.

That last one deserves a note, because it is the failure that is hardest to see from the portal. Risk-based Conditional Access policies depend on Microsoft Entra ID P2, which is not included in Microsoft 365 Business Premium. An institution can write a perfectly good risky-sign-in policy, see it listed as enabled in the portal, and have it enforce for nobody, because the licensing that powers the signal was never purchased. The policy exists. The control does not. If you are on Business Premium, our breakdown of what E3, E5, and Business Premium actually include covers where that line falls. On Business Premium the path to P2 is the Microsoft Defender Suite for Microsoft 365 Business Premium add-on, which is what we recommend to every Business Premium institution we manage, because risk-based policies and privileged identity management both depend on it.

What a network penetration test reaches

  • Perimeter firewall and VPN exposure
  • Unpatched services on internal hosts
  • Web application flaws on sites you host
  • Network segmentation between branches and the core
  • Wireless and physical access paths
  • Whether staff click a simulated phishing link

What it typically never touches

  • Conditional Access policy logic, exclusions, and enforcement
  • Application consent grants and OAuth permissions
  • Mailbox forwarding and transport rules
  • Administrative role assignments and standing privilege
  • Guest and external sharing configuration
  • Audit logging coverage and retention
  • Whether the licensing behind a policy makes it enforceable

None of the items in the right column are vulnerabilities in the usual sense. There is no CVE for an exclusion somebody added during a 2023 outage and never removed. There is nothing to patch. These are configuration facts, and configuration facts are found by reading configuration, not by attacking it. We have written separately about how Conditional Access exclusions accumulate and why nobody reviews them.

Configuration assessment is a different instrument

If the penetration test proves a path is walkable, a configuration assessment answers a different question: does the environment match a defined, defensible baseline, and where does it diverge?

There is a credible free starting point for this, and more institutions should know it exists. The Cybersecurity and Infrastructure Security Agency runs the Secure Cloud Business Applications project, which publishes secure configuration baselines for Microsoft 365. Alongside the baselines, CISA publishes ScubaGear, described in its own words as "a no-cost assessment tool that verifies M365 tenant configuration alignment to the policies described in SCuBA's secure configuration baselines." It is not restricted to federal agencies. CISA states it has made the tool and baselines "available to all agencies and private sector organizations seeking security improvements."

Run it and you will get a report of where your tenant diverges from a published federal baseline. That is a real artifact, produced by a named authority, that you can hand an examiner.

Be equally clear about what it is not. ScubaGear reads configuration and reports on alignment. It does not remediate anything. It does not monitor continuously, so its output is true for the moment it ran and begins aging immediately. It does not know which divergences are deliberate and documented at your institution and which are accidents. And it does not translate a finding into the language your examiner uses, which is the translation work that usually consumes the most time.

M365 Guardian is the operating model ABT uses to close that gap: a defined configuration baseline for the Microsoft 365 tenant, continuous monitoring for drift away from it, and an evidence trail written in the terms an examiner asks about. Where ScubaGear gives you a point-in-time divergence list against a federal baseline, Guardian keeps the baseline enforced between assessments and tells you when something moves. Guardian MxDR sits on top of that for the detection and response layer, because a configuration that was correct on Tuesday is not a control that noticed Thursday.

Penetration test

  • Question: can an attacker get from here to there?
  • Output: demonstrated paths, with evidence
  • Cadence: periodic, risk-based or annual
  • Blind spot: everything it did not have time or scope to try

Configuration assessment

  • Question: does the tenant match a defensible baseline?
  • Output: divergence list against a named standard
  • Cadence: point in time, ideally repeated
  • Blind spot: whether a compliant configuration is actually being attacked

Managed detection and response

  • Question: is something happening right now?
  • Output: detections, investigations, containment
  • Cadence: continuous
  • Blind spot: whether the configuration would hold against an attack nobody has launched yet

Those three answer different questions, and an assurance program needs all three answers. The common failure is buying one and filing it as though it settled the other two.

Independence, stated plainly

The FFIEC is specific about what makes a test independent, and the sentence is worth reading closely because it cuts in a direction most vendors would rather not discuss.

To be considered independent, testing personnel should not be responsible for the design, installation, maintenance, and operation of the tested system, or the policies and procedures that guide its operation.

Section IV.A.3, Independence of Tests and Audits

Read that against how financial institutions actually buy technology. The provider who knows your tenant best is almost always the provider who built it, and that is precisely the provider the guidance says should not be the one testing it. The conflict is structural, it is not an accusation about anyone's integrity, and pretending it is not there is how it ends up in a finding.

This applies to us, and to any provider who both builds and reviews your tenant

ABT manages Microsoft 365 tenants for financial institutions. Where we designed, configured, and operate your tenant, our people are not independent of it, and a report we write about our own configuration is not the independent test your assurance program needs. Use our assessment for what it is good for, which is finding and fixing what is actually wrong, and keep an independent party in the program to test the result.

Two practical notes follow. First, independence is a property of personnel and responsibility, not of a logo on a report cover. A large firm that both runs your environment and tests it with the same team has an independence problem; a firm with genuine separation may not. Ask who specifically performed the test and what else they are responsible for in your environment. Second, the level of independence required is a judgment your management makes and documents, because the guidance says so explicitly: management should determine the level of independence required of the test. An examiner is more interested in whether you reasoned about it than in which answer you reached.

The three instruments answer different questions, and the regulatory driver for each one differs too. Document which instrument answers which question, and the independence of each.

A testing program that holds up at exam time

Practical steps, in the order we would take them with an institution that has a clean penetration test report and an unexamined tenant.

Read last year's scope statement, not last year's findings

Find the sentence that defines what was in scope. If Microsoft 365, Microsoft Entra ID, or "cloud services" appears nowhere in it, you have located the gap in one minute.

Establish which rulebook governs each legal entity

The bank or credit union follows the FFIEC risk-based expectation. A mortgage subsidiary follows 16 CFR 314.4(d), which means continuous monitoring, or annual penetration testing plus six-month scans. Write down which path each entity is on.

Run a configuration assessment against a named baseline

CISA's ScubaGear and the SCuBA baselines cost nothing and produce a citable artifact. Do this before commissioning more testing, because it tells you what you are dealing with.

Verify that your policies can actually enforce

For each Conditional Access policy, confirm the licensing behind it, the exclusions on it, and that its grant controls are not empty. A policy that grants nothing blocks nothing.

Confirm your audit logging would survive an investigation

Check what is captured and how long it is kept against how long it would take you to discover a quiet mailbox compromise. Retention shorter than discovery time means the evidence is gone before you look, which is the trap we covered in our guide to Microsoft 365 audit log retention.

Add the tenant to next year's independent test scope, explicitly

Name Microsoft Entra ID, Exchange Online, SharePoint Online, and Teams in the scope statement, and require the tester to work inside Microsoft's rules of engagement.

Record the independence reasoning in the file

Who tested, what else they are responsible for, and why management considered that level of independence appropriate. The FFIEC asks management to make this determination, so make it on paper.

If your institution is working through this and wants the tenant side covered properly, our note on continuous monitoring covers what changes when point-in-time assessment becomes ongoing, and our Microsoft 365 incident response guide covers what you do on the day the tenant tells you something is wrong.

Key Takeaway

A clean penetration test report tells you that the paths someone tested were not walkable on the day they tested them. It tells you nothing about the tenant configuration layer where identity attacks actually land, and on Microsoft 365 the rules that govern testing make that layer hard to reach with a penetration test at all. Assess the configuration against a named baseline, test independently, monitor continuously, and document why each instrument was chosen.

See what your last penetration test left out

ABT manages Microsoft 365 for more than 750 credit unions, banks, and mortgage companies as a Tier 1 Microsoft Cloud Solution Provider. An M365 Guardian assessment reads your tenant configuration against a defensible baseline and reports what an examiner would ask about, in the language they use. It is the find-and-fix instrument, and it is built to sit alongside your independent test rather than in place of it, so your file shows both.

Frequently Asked Questions

Yes, within limits Microsoft sets. Microsoft encourages security testing performed inside your own tenant or against assets you have explicit authorization to test, and pre-approval is no longer required. Microsoft's rules of engagement prohibit a number of activities, including denial-of-service testing, accessing data or systems you do not own, retrieving credentials that are not yours, generating excessive automated traffic, phishing Microsoft employees, and post-exploitation actions such as enumerating internal networks or dumping secrets. That list is not exhaustive, so review the current rules before scoping any engagement.

No. The FFIEC Information Security booklet treats testing frequency as risk-based rather than fixed, stating that the frequency and scope of a penetration test should be a function of the level of assurance needed by the institution and determined by the risk assessment process. Examiners expect a testing and assurance program combining self-assessments, penetration tests, vulnerability assessments, and audits with appropriate coverage, depth, and independence, and they expect management to be able to explain why the chosen cadence is proportionate to the institution's risk.

Generally no for banks and federally insured credit unions, because the Rule applies to financial institutions under FTC jurisdiction that are not subject to another regulator's enforcement authority under section 505 of the Gramm-Leach-Bliley Act, and those institutions have prudential regulators. It does apply to mortgage companies, mortgage brokers, and non-federally-insured credit unions. Institutions outside the Rule are not free of testing obligations; their requirements come from FFIEC interagency guidance and their own regulator instead, judged on risk rather than on a fixed schedule.

Under 16 CFR 314.4(d), a covered financial institution either implements continuous monitoring of its information systems, or conducts annual penetration testing together with vulnerability assessments that include system-wide scans every six months designed to test for publicly-known security vulnerabilities. Testing is also required whenever there are material changes to operations or business arrangements. Continuous monitoring and the annual-test path are alternatives, so an institution that genuinely monitors continuously is not separately required to run the annual test.

It depends on personnel separation rather than on the company name. The FFIEC states that to be considered independent, testing personnel should not be responsible for the design, installation, maintenance, and operation of the tested system, or the policies and procedures that guide its operation. A provider that both operates your environment and tests it with the same people has an independence problem. A provider with genuine separation between the operations team and the testing team may not. The guidance also places the decision with management, which should determine and document the level of independence required for each test.

Start with CISA's Secure Cloud Business Applications project, which publishes secure configuration baselines for Microsoft 365 along with ScubaGear, a no-cost assessment tool that verifies tenant configuration alignment to those baselines. CISA has made the tool and baselines available to private sector organizations, not only federal agencies. Treat the output as a starting point: it reports divergence from a baseline at a moment in time, it does not remediate anything, it does not monitor continuously, and it does not distinguish a deliberate documented exception from an accident.

Justin Kirsch

Co-Founder & CEO, Access Business Technologies

Justin Kirsch has been building and defending Microsoft cloud environments for financial institutions since 1999, and has sat on both sides of the examination table as institutions worked through FFIEC, NCUA, and FTC Safeguards testing expectations. As Co-Founder and CEO of Access Business Technologies, the largest Tier-1 Microsoft Cloud Solution Provider primarily dedicated to financial services, he helps more than 750 credit unions, banks, and mortgage companies assess what their security testing actually covered and close what it missed.