
Yes, SOC 2 applies to privately hosted and air-gapped AI systems, and auditors already know how to test them. There’s no AI-specific Trust Services Criteria yet, but examiners map the same five criteria onto your model lifecycle and expect physical, cryptographic proof: attributable inference logs, an immutable model registry, RBAC/ABAC access records, encryption and key-rotation evidence, physical access logs, and a tested disaster recovery plan. Everything below breaks that down into a working checklist, lifecycle controls, and a realistic audit timeline.
TL;DR:
- Air-gapped AI systems must produce cryptographic and physical evidence, such as signed updates and access logs, aligned with the five Trust Services Criteria.
- Evidence collection for on-prem or air-gapped environments often requires 8 to 12 weeks due to the need to generate internal infrastructure logs and test DR plans.
- Model versioning and deployment change management require immutable registry entries, sign-off tickets, and detailed rollback procedures to satisfy auditors.
- Physical and transfer controls, including data diodes, signed updates, and controlled manual transfers, are essential for proving environment isolation.
- Effective SOC 2 readiness involves designing architecture with built-in traceability, tamper-evident logs, and clear scope boundaries from the outset, minimizing retrofitting efforts.
Table of Contents
- What SOC 2 for AI Actually Requires From Air-Gapped Systems
- A Readiness Checklist Mapped to Trust Services Criteria
- Controls Mapped to Every Stage of the AI Model Lifecycle
- Air-Gapped Architecture: What Auditors Want to See
- Timeline: Assembling Your Evidence Package
- Who Owns What: Governance Auditors Expect to See
- Vendor and Third-Party Risk for AI Components
- Data Privacy in Training and Inference Pipelines
- Testing for Fairness, Bias, and Explainability
- Real Examples: Model Lifecycle Controls in Practice
- Setting Scope Boundaries for Air-Gapped Environments
- Traceability Is the Real Audit Currency
- Get Your Air-Gapped AI Deployment Audit-Ready
- Sources
- FAQ
What SOC 2 for AI Actually Requires From Air-Gapped Systems
SOC 2 doesn’t care whether your inference server sits in AWS or in a locked server room with no internet connection. Auditors evaluate AI deployments against the same five Trust Services Criteria they’ve always used: security, availability, processing integrity, confidentiality, and privacy. What changes for air-gapped and on-prem systems is who produces the evidence. A cloud vendor hands you a shared responsibility matrix and a pre-built compliance dashboard. Your own infrastructure hands you nothing until you build the collection pipeline yourself.
That’s the trap most enterprise teams fall into. They assume “we control everything, so we’re automatically fine,” then discover mid-audit that nobody exported six months of firewall logs or documented who approved the last three model deployments. SOC 2 criteria are implementation-agnostic: bare-metal and colocation environments are fully compatible with a clean report, but only if you design the evidence architecture on purpose instead of hoping a cloud-style audit trail appears by accident.
One misconception worth killing early: labeling an AI system “out of scope” because it doesn’t touch payment data is a common and losing move. Auditors consider a model in scope the moment it processes customer information or influences a business decision, which covers nearly every enterprise deployment worth having a SOC 2 conversation about in the first place.

A Readiness Checklist Mapped to Trust Services Criteria
Work through this by control family, not by task list. Each line below pairs the criterion with the specific artifact an auditor will ask to see.
- Security (CC6.1, CC6.6): Export RBAC/ABAC role assignments and MFA enforcement logs for every account with model or infrastructure access.
- Change management (CC8.1): Pull the last six months of deployment tickets showing approval, rollback plan, and who signed off on each model push.
- Monitoring (CC7.2): Confirm inference logs stream to a centralized SIEM with tamper-evident storage, not just local disk.
- Availability: Schedule and document a disaster recovery test with backup restoration timestamps.
- Physical security: Gather badge logs, CCTV retention records, and visitor sign-in sheets for the facility housing your hardware.
- Privacy/confidentiality: Document data retention and deletion policies for both training data and inference logs.
Some of this is a fast win. Turning on SIEM export and enforcing MFA can happen this week. Other items, like a documented DR test with restoration evidence, need four to eight weeks of lead time because you have to schedule the test, run it, and let the results settle into a report an auditor can review.
Statistic Callout: On-premise SOC 2 preparation routinely runs longer than cloud-based audits because teams have to generate infrastructure evidence, like DR test results and environmental control logs, from scratch rather than pulling a report from a cloud vendor’s trust portal.
Controls Mapped to Every Stage of the AI Model Lifecycle
Auditors don’t evaluate “the AI system” as one blob. They walk the lifecycle stage by stage and ask for evidence at each handoff.
- Training and data provenance. Keep a dataset manifest listing source, collection date, and access permissions for every training set. Restrict who can pull raw training data, and log every access event.
- Model registration. Every model version needs an immutable registry entry: SHA hash, timestamp, and the name of whoever approved it for deployment. If your registry allows silent overwrites, that’s a finding waiting to happen.
- Pre-deployment validation. Retain the test results, benchmark scores, and sign-off ticket proving someone checked the model before it went live, not after.
- Deployment and change control. Each push needs a ticket documenting the rollback plan and the approver’s name, satisfying the same change-management criteria auditors apply to any production system.
- Inference logging. Every request needs a record tying together user identity, input hash, model version, and retrieval context, which is the minimum decision lineage auditors expect for processing integrity.
- Retirement and sanitization. When a model gets pulled, document the sanitization method and confirm associated keys were rotated or destroyed.
Pro Tip: Build one evidence pipeline that tags each artifact for multiple frameworks at once. A single immutable prompt log, properly instrumented, can satisfy SOC 2 monitoring criteria and double as HIPAA or CMMC accountability evidence without duplicating the collection work.
Air-Gapped Architecture: What Auditors Want to See
An air-gapped system has to prove a negative: that nothing left the environment. That’s harder to demonstrate than “here’s our firewall config,” so auditors lean on a specific set of artifacts.
- Data diode documentation. If you use a one-way transfer device to move data or model updates into the environment, keep the hardware spec sheet and configuration proof showing it physically cannot pass traffic back out.
- Signed update verification. Every model or software update entering the enclave should carry a cryptographic signature, checked and logged before installation.
- Controlled sneakernet records. For manual transfers, log the transfer ticket, the checksum verification, and the name of the approving authority for each physical handoff.
- Centralized SIEM off the inference host. Moving logs off the model server and into write-once storage is how you prove logs weren’t altered after the fact, which matters enormously once an auditor starts asking pointed questions.
Pro Tip: Keep an append-only “audit provenance ledger” that records every model import and update event with verified hashes. It’s a small addition to your pipeline that turns a vague “we trust our process” claim into something an auditor can independently check.
Timeline: Assembling Your Evidence Package
Budget 8 to 12 weeks for a first SOC 2 Type 1 engagement on an on-prem or air-gapped stack, longer than most cloud-native teams need because you’re generating infrastructure evidence internally rather than exporting a vendor’s existing report.
- Weeks 1 to 2: Finalize system scope and boundary documentation; identify which models, data stores, and network segments fall inside the audit.
- Weeks 3 to 5: Stand up the model registry and confirm SIEM ingestion is capturing inference logs correctly.
- Weeks 6 to 8: Run quarterly access reviews and collect sign-off tickets for every privileged account.
- Weeks 9 to 10: Execute and document a disaster recovery test with backup restoration proof.
- Weeks 11 to 12: Package evidence pairs and walk through sample selection with your auditor before fieldwork starts.
Auditors typically sample five to ten user accounts per control area rather than reviewing every account, so prioritize completeness over volume. Hand over evidence in pairs: a config export alongside the signed review ticket that shows someone actually looked at it.
| Common pitfall | Why it fails the audit |
|---|---|
| Assuming a “phantom” cloud-vendor policy covers your on-prem stack | Auditors want evidence from your own environment, not a policy document referencing infrastructure you don’t use |
| Incomplete visitor logs at the data center | Physical access evidence gaps are one of the most frequent on-prem findings |
| Logs stored only on the inference host | Local-only logs are trivially alterable and fail tamper-evidence testing |
Who Owns What: Governance Auditors Expect to See
Assign a named owner for models, training data, and infrastructure, then document who signs each change-management ticket. Separate the people who write deployment code from the people who approve it going live, and log every break-glass access event with a reason code. Set a review cadence and stick to it: quarterly access reviews for all privileged accounts, monthly vulnerability triage for anything touching Tier 1 model infrastructure. Auditors read cadence gaps as governance gaps, even when the underlying control is technically sound.
Vendor and Third-Party Risk for AI Components
Most air-gapped AI stacks still depend on outside components, base models, fine-tuning frameworks, hardware vendors, even the consultants who set up your enclave. Each one needs a documented risk assessment before it enters your environment, not after.
Ask for the vendor’s own SOC 2 report or equivalent attestation where one exists, and document what you found even if the answer is “no report available, compensating controls applied instead.” For open-source base models pulled from public repositories, log the source, checksum, and version at the time of ingestion. That record becomes part of your model registry entry and closes a gap auditors specifically look for: an untraceable model origin.
Vendor contracts should spell out data handling terms explicitly, especially for any component that touches training data before it enters your air-gapped boundary. If a third-party fine-tuning service ever had access to your data, that access needs an entry in your provenance ledger and a termination or access-revocation record once the engagement ends. Treat every external dependency, hardware, software, or human, as a line item in your third-party risk register, reviewed on the same cadence as your internal access reviews.
Data Privacy in Training and Inference Pipelines
Training data and inference data create different privacy obligations, and auditors expect you to treat them separately. Training data usually sits in bulk for extended periods, so retention policy and access restriction matter most: who can query the raw dataset, how long it’s kept, and whether personally identifiable fields were minimized or masked before training began.
Inference data is different. Every prompt and completion your system generates is itself a record that needs a retention and deletion policy, not just a log line that accumulates forever. Decide upfront how long inference logs live in your SIEM before archival or deletion, and document that decision as a formal policy rather than an informal default. If your AI system processes regulated data categories, health records, financial details, government identifiers, that decision needs its own sign-off separate from your general logging policy.
Document data flow boundaries clearly: where training data enters, where it’s transformed, and where inference requests and responses are stored. For air-gapped systems, this documentation does double duty. It supports SOC 2 confidentiality criteria and gives you the evidence trail to prove data never crossed the air gap in either direction.

Testing for Fairness, Bias, and Explainability
SOC 2 doesn’t have a fairness criterion by name, but processing integrity examiners increasingly ask how you validate that a model behaves consistently and predictably. Keep bias testing reports and drift detection dashboards as standing evidence, refreshed on a set schedule rather than produced once at launch and forgotten.
Document your testing methodology: what benchmark datasets you use, what disparity thresholds trigger a review, and who signs off when a model passes or fails that review. If your model informs a decision with material consequences, credit approvals, hiring screens, security flagging, retain enough explainability documentation to answer “why did the model produce this output” for at least a sample of past decisions.
Drift monitoring belongs in the same evidence package. A model that performed well at deployment can degrade as real-world inputs shift, and an auditor evaluating processing integrity wants to see that you’re watching for that degradation, not just trusting the initial validation results indefinitely.
Real Examples: Model Lifecycle Controls in Practice
A model registry entry that satisfies an auditor looks like this: a SHA-256 hash of the model weights, a timestamp, the name of the engineer who trained it, and a separate approval field signed by someone other than the trainer. That separation of duties is what turns a spreadsheet into audit evidence.
For change management, a deployment ticket should show the model version being replaced, the rollback procedure if the new version underperforms, and a link back to the pre-deployment validation report. Pairing those three documents together is what auditors mean when they talk about evidence pairing: a config export alone proves nothing without the ticket showing someone reviewed it.
For retirement, sanitization evidence should include the deletion method used on any retained copies and confirmation that associated API keys or access tokens were rotated. A model that’s “retired” but still has a live key floating in a config file is a finding, not a formality.
Setting Scope Boundaries for Air-Gapped Environments
Scope definition is where most air-gapped SOC 2 engagements go sideways before fieldwork even starts. Draw the boundary around every system that processes, stores, or transmits data relevant to the trust criteria you’re pursuing, including the data diode or transfer mechanism itself, since that’s the control proving isolation.
Exclude systems only when you can document why they’re genuinely outside the data flow, not because excluding them makes the audit smaller. A staging environment that occasionally receives production data copies belongs in scope even if it’s labeled “test.” Document your boundary decisions in a scope memo reviewed with your auditor before evidence collection starts, so nobody discovers a scope disagreement halfway through fieldwork.
Include your update mechanism, whatever moves new models or patches into the air-gapped environment, as an explicit part of the boundary. That’s usually the single component auditors scrutinize hardest, since it’s the one place where the “air gap” claim gets tested directly.
Traceability Is the Real Audit Currency
Every enterprise deployment we’ve studied that struggled through a SOC 2 audit had the same root problem: they built the AI system first and thought about evidence collection later. Retrofitting immutable logging onto a production model registry is possible, but it’s slower and messier than designing the registry with SHA hashes and approval fields from day one.
The friction point that comes up again and again in enterprise audits is inference logging that exists but isn’t attributable, logs that show something happened without showing who caused it or why. That gap is exactly what air-gapped and enclave architecture needs to close, since there’s no cloud vendor logging layer to fall back on.
— John Ezzell, Founder
Get Your Air-Gapped AI Deployment Audit-Ready
Certain providers build sovereign AI deployments where audit-readiness is designed in, not bolted on afterward. That means immutable model registries, tamper-evident inference logging routed to your SIEM, and disaster recovery testing built into the deployment plan from the start, so your compliance team isn’t reverse-engineering evidence collection six months after go-live.

Every deployment integrates into your existing infrastructure with model governance and change-management controls already mapped to the artifacts auditors ask for. If you’re procuring a privately hosted or air-gapped AI system and need SOC 2-aligned controls from day one rather than retrofitted later, Forge’s deployment and integration solutions are built specifically for that requirement. Request a readiness assessment to see where your current architecture stands against the evidence checklist above.
Sources
- SOC 2 for AI companies — what auditors expect
- AI agents and SOC 2 readiness (Teleport blog)
- SOC 2 readiness for bare metal SaaS — practitioner guidance
FAQ
Does SOC 2 Apply to Air-Gapped AI Systems?
Yes. SOC 2’s five Trust Services Criteria are implementation-agnostic, so air-gapped and on-prem systems qualify as long as you build the evidence collection yourself rather than relying on a cloud vendor’s report.
How Long Does SOC 2 Prep Take for On-Prem AI?
Plan for roughly 8 to 12 weeks for a first Type 1 engagement, longer than typical cloud timelines because teams have to generate infrastructure evidence internally, including DR tests and physical access records.
What Evidence Do Auditors Need for AI Model Changes?
Auditors want an immutable model registry entry with a SHA hash, timestamp, and approver name for every version, paired with a deployment ticket showing rollback plans and sign-off, satisfying standard change-management criteria.
Can Forge Help With SOC 2 Readiness for AI?
Forge designs air-gapped and on-prem AI deployments with model governance, tamper-evident logging, and DR testing built into the architecture, giving compliance teams the evidence trail auditors expect from the start. Details on services are available on the Forge solutions page.
What Counts as Scope for an Air-Gapped SOC 2 Audit?
Scope includes every system that processes, stores, or transmits relevant data, including the update or transfer mechanism used to move models into the environment. Staging systems that occasionally handle production data belong in scope too, regardless of their label.