© Eclypsium, Inc.
The Growing Challenge
The global AI arms race is accelerating, and AI infrastructure is now critical infrastructure, vital to national security and defense. AI infrastructure is uniquely challenging to secure. It processes high volumes of sensitive data on complex, nested racks of GPUs, CPUs, DPUs, BMCs, networking components, and more. A compromise in any component can lead to loss of data, leaking of intellectual property, and poisoning of model weights, causing significant harm to any nation’s battle for AI supremacy.
The urgency of proactive cybersecurity for AI infrastructure could not be higher.
The Infrastructure Gap in AI Cyber Risk Guidance
The discussions of cyber risk in AI have typically focused on the model and application layer, with little to no emphasis on hardware and infrastructure. This gap is present across the industry. The current pace of AI infrastructure development has pushed security far down the priority stack.
On July 2, 2026, U.S. President Donald Trump signed Executive Order 14409, Promoting Advanced Artificial Intelligence Innovation and Security ordering federal agencies to prioritize cybersecurity across defense and civilian agencies. While the EO is a move in the right direction for AI security, it does not emphasize AI infrastructure and hardware level security. But as new vulnerabilities and exploits against AI hardware are discovered, neoclouds and enterprises operating AI infrastructure will need to pursue greater security at the hardware level.
Why the Hardware Layer Is Exposed in AI Data Centers
The economics and nested complexity of AI infrastructure create risk conditions that do not exist in a normal enterprise fleet.
Computers Within Computers Expand The Attack Surface
AI hardware contains layers of interconnected devices with their own management interfaces, privileges, access management, and connectivity to GPUs, storage, and memory inside AI infrastructure. Most device inventories and access management systems do not account for the deeply embedded devices, yet vulnerable components like BMCs and DPUs can grant an adversary total control over compromised AI infrastructure.
Hardware That Changes Hands
In a neocloud or shared research environment, the same physical GPU infrastructure serves one customer after another. Bare metal assets customized by one tenant, or altered by an attacker who rented time on that hardware, persists into the next tenancy unless someone verifies it. Manual verification takes days of otherwise billable rack time.
Blind Trust in the Supply Chain
AI servers are assembled from components sourced several layers down the supply chain, under intense demand pressure. Provenance is hard to confirm, gray-market inventory is common, and counterfeit or substituted parts are difficult to catch by inspection or asset tag.
Monitoring Gaps Adversaries The Advantage
EDR operates at the OS and application layer and does not cover networking devices at all. Vulnerability scanners produce a flood of findings dependent on version strings and known CVEs, with minimal visibility into components and firmware. Neither reads GPU firmware, verifies BMC integrity, or checks whether a UEFI image matches a known-good hash. An attacker at that layer can persist through a rebuild, disable security controls, manipulate device behavior, or brick expensive hardware outright.
Two existing NIST documents offer a starting point for thinking about AI infrastructure security.
- The NIST AI Risk Management Framework (AI RMF) helps organizations manage risk across the AI lifecycle through four functions: Govern, Map, Measure, and Manage. Most of the framework addresses the AI model itself, its data, its behavior, and its impact on people. But every AI system runs on physical infrastructure, and that infrastructure carries its own risk. GPUs, AI accelerators, servers, and the network devices that connect them arrive with vulnerabilities, misconfigurations, and, in some cases, tampered or counterfeit components introduced anywhere in the supply chain.
- NIST SP 800-223 highlights several risks facing all High Performance Compute (HPC) environments that are especially urgent and under-addressed in the AI infrastructure context. These are:
- Attacks on critical hardware components and software manipulation to gain unauthorized access.
- Rapidly changing infrastructure for HPC firmware and hardware components, increasing supply chain risk
- Compute Node Sanitization challenges, such as validating firmware between task runs on shared compute infrastructure.
Eclypsium Delivers Proactive Security for AI Infrastructure
Eclypsium verifies and hardens the components other security controls can’t, and detects hidden threats deep inside AI infrastructure. To learn more about how the Eclypsium Infrastructure Assurance platform works, visit our AI Infrastructure Platform page.






