
A practical framework identifying and prioritizing the top security risks in AI datacenter infrastructure, covering hardware, networking, management planes, and supply chain, with mitigation strategies for providers and customers.

FORGE - Harden the metal beneath the model.
🌐 Website: https://forge-framework.io
AI datacenters are being built faster than they are being secured.
The rapid expansion of datacenters, GPU clouds, and specialized compute environments has introduced security risks across hardware, networking, storage, orchestration, identity, management planes, and physical operations. Many of these risks resemble traditional datacenter or cloud security issues, but modern data centers and AI infrastructure changes their severity: systems originally designed for trusted operators are now supporting high-value, multi-tenant workloads from unrelated customers.
The Top 10 Data Centers & AI Infrastructure Security Risks provides a practical framework for identifying, prioritizing, and reducing the most important security risks in the infrastructure layer that powers AI. The framework defines the most critical failure modes in AI infrastructure and helps translate them into concrete security requirements.
This framework focuses on the security of AI infrastructure and the datacenters that house it: the physical hardware, networking fabrics, management planes, orchestration systems, storage systems, and operational environments on which AI workloads run.
It does not focus on AI models themselves or application-layer risks such as prompt injection, insecure agent behavior, model abuse, or model-level evaluation. Those risks are addressed by frameworks focused on other parts of the AI stack, including the OWASP Top 10 for LLM Applications, MITRE ATLAS, the NIST AI Risk Management Framework, and ISO/IEC 42001. Together, these resources give practitioners a more complete picture of AI security, from governance and application-layer risks down to the underlying compute infrastructure.
Some risks in this document also exist in traditional datacenter and cloud environments. They are included here because AI infrastructure makes them materially more severe: shared high-value compute, complex accelerator clusters, dense management layers, and multi-tenant operations can turn ordinary infrastructure weaknesses into more severe security risks because the likelihood and impact of any incident are much higher than in traditional enterprise-level software and service deployments.
Neo-cloud providers can benefit from the framework as a practical guide to the AI infrastructure threat landscape, attack surface review, environment hardening, security prioritization, and maturity demonstration.
AI infrastructure customers can benefit from the framework as a practical guide for procurement, security reviews, contractual requirements, provider comparison, and assessing resilience against realistic tenant-to-infrastructure compromise.
Hybrid Cloud security teams can benefit from this framework as a practical guide for configuring, maintaining, and securing their on-premise footprint against modern attacks, which can quickly pivot between environments.
Built to evolve with the field, and kept open so the whole AI and security community can use it, challenge it, and improve it.
We would like to thank and acknowledge all experts which took part in reviewing and validating this document.
Have feedback or want to contribute? [email protected]
FORGE domains define the evaluation lens: the infrastructure areas where AI security risk lives. The Risk Matrix below maps individual risks into these domains.
| Domain | Name | Description |
|---|---|---|
| F | Fleet integrity | Trust in the hardware, firmware, software artifacts, images, dependencies, and supply paths that make up the AI infrastructure fleet. |
| O | Operations & management planes | Privileged systems used to control, automate, and administer AI infrastructure, including BMCs, schedulers, orchestration, automation, and admin tooling. |
| R | Resource isolation | The boundaries that separate tenants, workloads, execution environments, and reused infrastructure across shared AI systems. |
| G | Grid | The networking fabrics and facility systems that connect AI clusters and keep them powered, cooled, and operational. |
| E | Evidence & exposure management | The evidence customers need to understand provider security maturity, architectural scope, exposed services, and patch velocity. |
FORGE IDs are ordered by severity, from highest to lowest. The matrix groups each risk by domain and shows its likelihood, impact, and detection difficulty.
| ID | Domain | Risk | Risk Level | Likelihood | Impact | Detection Difficulty |
|---|---|---|---|---|---|---|
| FORGE-01 | F | Hardware & Firmware Integrity Compromise | Critical | Medium | Severe | High |
| FORGE-02 | G | Network & Interconnect Vulnerabilities | Critical | Medium | Severe | High |
| FORGE-03 | R | Unsafe Multi-Tenant Isolation and Resource Reuse | Critical | Low | Severe | Very High |
| FORGE-04 | O | Insecure Out-of-Band Management Plane | Critical | Medium | High | Very High |
| FORGE-05 | F | AI Infrastructure Supply Chain Compromise | Critical | High | High | High |
| FORGE-06 | G | Insecure Facility & Datacenter Management Systems | High | Low | High | Very High |
| FORGE-07 | R | Insecure Data and Artifact Handling | High | High | High | Medium |
| FORGE-08 |
Each risk follows a consistent structure: Definition, Description, Impact and Failure Modes, Prevention and Mitigation Strategies (for providers and for customers), Attack Scenarios, and References.
The rate at which the modern datacenter and AI infrastructure is being built is heavily outpacing the ability to secure it. When this infrastructure, GPU clusters, training pipelines, high-performance networking, and inference endpoints, is compromised, the blast radius is unlike conventional compute. Attackers gain access to proprietary models representing hundreds of millions of dollars, the power to poison foundational training data, and persistent footholds in the most highly privileged environments available.
This dynamic is reinforced by a structural market imbalance: GPU scarcity gives providers outsized leverage. When demand for accelerated compute far outstrips supply, customers often cannot choose their provider based on security posture – they take what is available. Providers face little market pressure to invest in security maturity, and customers accept risk they would not tolerate in conventional cloud. This imbalance is the backdrop against which every risk in this document should be read. This market pressure is becoming more dangerous as AI-enabled security research and exploitation capabilities compress the timeline for defenders. Weaknesses that might once have remained obscure for years may now be discovered, chained, and operationalized much faster.
At the same time, AI-enabled security research and exploitation capabilities are compressing the timeline for defenders: weaknesses that might once have remained obscure for years may now be discovered, chained, and operationalized much faster.
AI infrastructure sits at the intersection of cloud computing, high-performance computing, and physical datacenter operations. It uses familiar components such as servers, storage, networks, schedulers, management planes, and identity systems, but combines them in ways that change the security model.
Traditional HPC environments were often designed for trusted users, research communities, or internal operators. Modern AI infrastructure increasingly supports commercial, high-value, multi-tenant workloads from unrelated customers. As a result, assumptions that were acceptable in trusted or single-organization environments can become serious security risks when applied to shared AI infrastructure.
Securing AI infrastructure and data centers will require collaboration across cloud providers, hardware and networking vendors, security companies, system integrators, and customers. This is especially important in areas such as high-performance networking, east-west visibility, segmentation, management-plane protection, and infrastructure security controls.
Attackers have already begun targeting the infrastructure, supply chains, and datacenter ecosystems that support advanced AI and HPC workloads. In some cases, the objective is direct theft of sensitive research, models, or engineering data, in others, it is espionage, prepositioning, or the ability to disrupt strategically important compute environments. Recent reporting illustrates both patterns:
These are early examples. The attack surface is likely to expand as the stack matures, and this document will evolve alongside it.
The risk is best understood by looking at potential attackers. The threat model spans far more than the classic "outside attacker," and the controls in this document are calibrated against the full set:
The risks in this framework focus on AI infrastructure-specific or AI infrastructure-amplified failure modes. They do not replace the need for strong enterprise security controls. In practice, many attacks against AI infrastructure may begin through familiar enterprise compromise paths, including weak or missing MFA, phishing, credential theft, SaaS compromise, and weak identity controls.
These paths should be treated as cross-cutting initial access vectors. A compromised identity, endpoint, SaaS account, CI/CD system, or administrative credential can provide the foothold needed to reach several of the risks described in this framework. These sections therefore focus on what can happen once an attacker reaches, or can influence, the AI infrastructure environment itself.
Candidate risks were drawn from:
Each candidate was evaluated against four dimensions, Likelihood, Impact, Exploitability, and Detection Difficulty, with the highest-scoring risks selected for inclusion.
This is a living document. AI infrastructure is evolving rapidly, and so are the attacks against it. Future revisions will add new risks, refine existing ones, and incorporate lessons learned from the field. Feedback and proposed additions from practitioners are welcome and will be credited upon publication of the new version.
Content in this repository is licensed under CC BY-NC-SA 4.0.
| E |
| Certification Gaps & Provider Transparency Failures |
| High |
| Medium |
| High |
| Medium |
| FORGE-09 | O | Insecure Operational Infrastructure Services | High | High | Medium | Medium |
| FORGE-10 | E | Vendor Embargo Gaps & Patch Velocity Failures | Medium | Medium | Medium | Low |