SOC 2 penetration test: scope, timing and the report

Updated 9 min read soc2pentest

A penetration test bought for a SOC 2 program has two readers who never meet: a service auditor looking for evidence, and a customer's security reviewer looking for reassurance. Scope, dates and structure are what decide whether either of them can use it.

Scope comes from the system description, not from a price list

A SOC 2 report describes a system: what the service does, where its boundaries are, which subservice organizations sit inside it, and what commitments you have made to customers. Everything the examination touches is defined by that description, and your test scope should be traceable to the same document.

The practical test is simple. If a reviewer holds the system description in one hand and the test scope statement in the other, can they see which parts of the described system were examined and which were not? If not, the report is interesting but it is not evidence, because nobody can tell what it covers.

Write the scope statement before asking for quotes. A tester quoting against "our SaaS platform" will price a different engagement from one quoting against "the multi-tenant web application, its public and partner APIs, the customer-facing admin console and the AWS account hosting them, excluding the marketing site and the corporate network".

What belongs in scope, and what usually does not

Typical scope decisions for a SOC 2 program
ComponentUsually in scopeWhy
The multi-tenant applicationYesTenant isolation is the question your customers are actually asking, and nothing else in the evidence pack answers it.
Public and partner APIsYesOften broader than the interface, and frequently missing the authorization checks the interface enforces.
Administrative consolesYesPrivilege escalation into an admin role is the highest-impact finding available in most products.
Cloud account configurationUsuallyIdentity policy, storage exposure and network segmentation carry more real risk than most application findings.
CI/CD and source controlConsiderA pipeline that can deploy to production is production. Include it if it is in the system description.
Corporate IT and workstationsRarelyRelevant to your risk assessment, but usually outside the boundary of the described service.
Marketing websiteNoOutside the system, and it consumes budget that belongs on the product.
Subservice organizationsNoYou cannot test your cloud provider. Their controls are covered as complementary subservice organization controls and by their own assurance report.
Scope exclusions are as much a part of the evidence as inclusions. State them in the report rather than letting a reader infer them.

Production or a copy?

Test the environment your customers use, or an environment you can demonstrate is identical to it. A staging environment with a different identity provider, relaxed network rules and synthetic data does not answer questions about production, and an experienced reviewer will spot the substitution in the scope statement.

Where production testing genuinely carries unacceptable risk – a payment flow, a medical workflow, a system with irreversible side effects – the honest approach is a documented hybrid: destructive classes of testing on a mirrored environment, everything else against production, with the split explained in the report. A documented compromise is defensible; an undocumented one is a finding waiting to be raised.

Authenticated testing is the point

An unauthenticated external test tells you what an anonymous internet user can reach. For a SaaS product, that is the least interesting question available, because the interesting attacker is a paying customer.

Provide accounts. At minimum two separate tenants and, within one of them, one account per privilege level. The findings that decide whether your report is useful – reading another tenant's data, escalating from a read-only role, reaching an administrative function through an API that the interface hides – can only be produced with credentials in hand.

Where the test lands in the calendar

The two report types are defined by what they opine on. In the international standard for service organization assurance the distinction is set out in the report titles themselves: a type 1 is a "report on the description and design of controls at a service organization", and a type 2 is a "report on the description, design and operating effectiveness of controls at a service organization". The AICPA suite follows the same split.

That single difference drives the schedule. A type 1 is a point in time: the test needs to exist and be reasonably current when the report is dated. A type 2 covers a period, and evidence has to belong to that period. A test performed before the observation window opened describes a system that the examination is not looking at.

  • Run readiness and remediation first. Testing an environment you already know is unfinished converts your own backlog into report findings, each of which then has to be tracked and retested at report prices.
  • Schedule the test so that testing, remediation and retest all fall inside the observation window for a type 2.
  • Leave four to six weeks between the test and the end of the window. Remediation and retest take longer than anyone plans, and a critical finding fixed after the window closes is a critical finding that was open throughout it.
  • For a first type 2, three months is the shortest window most auditors will accept. Twelve is what a mature enterprise buyer expects to see on the second cycle.

What the report has to contain

Most penetration test reports are written for engineers. A report bought for a SOC 2 program has to survive two other readers as well, and the structure below is the minimum that lets it.

Report sections, and who each one is for
SectionWhat it must containWho reads it
Scope statementNamed systems, environments, account types and tenants; the exclusions; the traceability back to the system descriptionThe auditor, first. The customer's reviewer, first.
DatesWhen testing ran, not only when the report was issued, and the date of any retestEveryone. This is the field that decides whether the report is current.
MethodologyThe approach, the tooling class, and what was manual rather than automatedThe reviewer deciding whether this was a scan with a cover page.
Severity scaleThe scale, its definitions, and how it relates to the risk scale used elsewhere in your evidenceThe auditor, who has to reconcile "high" here with "high" in your risk register.
FindingsReproduction detail sufficient to verify the fix, and evidence of exploitation rather than an assertionYour engineers, and any reviewer who doubts the finding.
Remediation statusWhat was fixed, what was accepted, and by whom the risk was acceptedThe auditor, who will look for an accepted risk that contradicts your own policy.
Retest evidenceIndependent confirmation that the fix works, with its own dateThe customer's reviewer. This is the section they check first.
Executive summaryA plain account a procurement reader can act on without over- or under-reactingThe person who decides whether the deal moves.
The retest section is the one most reports omit and the one that most often decides whether a report closes a review.

Retest: the difference between a claim and evidence

A finding marked "remediated" on the strength of a client email is a claim. A finding retested by the same firm, with a date and a short note on what was verified, is evidence. The cost difference is a few hours; the credibility difference is the whole point of buying an independent test.

It also matters to the examination. The service auditor is looking at whether your vulnerability handling actually closes things, over the period. A test report with an open critical finding and no follow-up is, read carefully, evidence against the control it was bought to support.

What you may hand to a customer

Three documents, in decreasing order of detail.

  • The full report, under a non-disclosure agreement, to a customer whose review genuinely needs it. Most enterprise reviewers will accept this and many require it.
  • A redacted report, where reproduction detail for unfixed findings is removed but the scope, dates, methodology, severity distribution and retest evidence remain. Redaction that removes the scope statement or the dates is not redaction, it is a leaflet.
  • An attestation letter: a short signed statement from the testing firm naming the scope, the dates, the methodology and the fact that findings were remediated and retested, with no finding detail at all. This is the document to attach to a questionnaire before an NDA is in place.

The same logic exists on the SOC side. A SOC 2 report is restricted to specified parties who understand the system, which is why it travels under agreement. A SOC 3 addresses the same subject matter in far less detail and is a general use report that can be freely distributed, so if a customer simply wants something to file, ask your CPA firm to issue one from the same examination.

The questionnaire questions one test closes

A properly structured report is the cheapest way to answer a whole block of a security questionnaire. Typically it settles: whether independent testing is performed and how often; who performs it; whether testing covers the production environment; whether authenticated and multi-tenant testing is included; how findings are risk-rated; the remediation timescales by severity; whether fixes are verified; and when the most recent test was performed.

That is eight to ten questions, answered by attaching one document. It is worth structuring the report so that a reviewer can find each answer without reading the finding detail, because the reviewer will not read the finding detail.

Testing again after a change

The annual cadence is a convention, not a rule, and it is the wrong trigger on its own. The right trigger is significant change to the described system: a new tenancy model, a new identity provider, a migration between cloud accounts, a new public API surface.

The EU's own technical rules for NIS2 entities describe testing in exactly these terms rather than as a calendar item. Point 6.5 of the Annex to Commission Implementing Regulation (EU) 2024/2690 requires entities to establish "the need, scope, frequency and type of security tests" from a risk assessment, to test "according to a documented test methodology", to "document the type, scope, time and results of the tests, including assessment of criticality and mitigating actions for each finding", and to "apply mitigating actions in case of critical findings". A risk-driven retest after a major change is precisely what that describes, and it reads well in an evidence pack.

For the underlying question of why a test is expected at all, see does SOC 2 require a penetration test?. For who signs which document at the end of it, see who can issue a SOC 2 report in the EU.

Sources

  1. ISAE 3000 (Revised), Assurance Engagements Other than Audits or Reviews of Historical Financial Information IAASB · 2013 The conforming amendments to ISAE 3402 define the type 1 and type 2 report titles used here: description and design, versus description, design and operating effectiveness.
  2. Promises of 'fast and easy' threaten SOC credibility Journal of Accountancy · 2026 SOC 2 reports are restricted to specified parties; template reports and the consequences of a report a business partner rejects.
  3. SOC 3 – audit and assurance topic page AICPA & CIMA SOC 3 addresses the same subject matter as SOC 2 in less detail and is a general use report that can be freely distributed.
  4. SOC engagements: Ethics risks with tool providers Journal of Accountancy · 2026 The service auditor must obtain sufficient, appropriate evidence in all circumstances.
  5. Commission Implementing Regulation (EU) 2024/2690 EUR-Lex · 2024 Annex point 6.5 on risk-based security testing, documented methodology, recorded results and mitigation of critical findings.
  6. 2017 Trust Services Criteria (with revised points of focus, 2022) AICPA & CIMA · 2023 The criteria describe control outcomes and are published for use in attestation or consulting engagements.

Questions

Related questions

How long does a SOC 2 penetration test take?

For a single multi-tenant application with its APIs, plan one to two weeks of testing, a few days of reporting, then remediation on your side and a short retest. The elapsed time is dominated by remediation, not by testing, which is why scheduling it late in an observation window is a mistake.

Does the whole cloud environment have to be in scope?

Only the part inside your system description. Cloud configuration review is usually worth including because identity policy and storage exposure carry real risk, but your provider's own infrastructure is a subservice organization and is covered by their assurance report, not by your test.

Can we share the penetration test report with customers?

Yes, and it is often the fastest way to close a vendor review. Share the full report under an NDA, a redacted version where reproduction detail is removed but scope, dates and retest evidence remain, or an attestation letter where no finding detail is appropriate.

What if the test finds something critical just before the window closes?

Fix it, retest it and document the timeline honestly. An auditor can work with a critical finding that was identified, escalated and closed inside the period. What is difficult to work with is one that was identified and left open, or one that appears only in an email chain.