Trending
ASC cable break causes connectivity disruption between Perth and Singapore Damac and Vodafone Türkiye triple Izmir data center project budget DCD News: August review Google’s European energy and infrastructure principle joins Mistral Compute NEC signs MoU with UNDP to support conservation, climate action through use of AI O2 Business launches satellite D2D service in UK Sponsored: From alarm overload to action: How AI is improving service performance today What does sustained AI growth mean for data center fiber infrastructure? Qatar’s Meeza signs 8MW lease with global hyperscaler Stulz launches new air cooling systems for Edge data centers Bitcoin hosting firm Bitari files for $30m Nasdaq IPO Palantir selects Nebius as sovereign AI infrastructure provider Fujitsu develops diamond-spin quantum computer that incorporates tin-vacancy centers Corvex targets 8MW of AI cloud capacity by end of 2026 Sponsored: From training to inference: The AI economic turning point

The Data Center’s Hidden Attack Surface: Why OT Security Can’t Wait

Why OT Security Is Critical to Prevent Data Center Outages. 6 Min Read. Operational technology in data centers is a growing attack surface.

Securing OT systems is vital to prevent outages and ensure resilience.Getty Images. Data center resilience is moving into the spotlight, and security is a core part of that story. They are increasingly being treated as critical infrastructure because a dominant share of economic growth and security operations now depend on them..

But too often, the coverage stops at the perimeter. Inside the fence line, a separate and underexamined risk is compounding: the operational technology (OT) and cyber-physical systems (CPS) that keep these facilities running. Uninterrupted power supplies (UPS), power distribution units (PDUs), cooling systems, environmental sensors, and building management platforms are now routinely connected, often through remote access tools and legacy protocols that were never designed for today’s threat landscape..

These are not IT systems, and most organizations do not govern them as such. When this layer is compromised, it often does not look like a breach. It looks like an outage.

As AI-driven demand accelerates reliance on connected operational infrastructure, the scale of what is riding on these systems has changed, too. The facilities that train and run AI models are now strategic assets, and the cost of an outage scales with everything that now depends on them..

Related:Build an Autobahn for Zero-Day Defense with Strategic Gateways. Your Facility Is a Cyber-Physical Ecosystem. Most operators intuitively understand the physical layout of a data center: white space where the IT load lives, and gray space, where facility power and HVAC live.

The challenge is that organizational structures have not evolved as fast as technology. Facilities teams speak in voltage and cooling. Security teams speak in packets and vulnerabilities..

Traditional security frameworks were not built for this environment, and neither side has the tools to translate the other’s risks into operational action. The result is a gap in ownership, visibility, and prioritization. That gap is where outages are born..

It also changes what an incident looks like. Corporate IT security is largely built around confidentiality. In a facility environment, the primary consequence is often disruption.

It can mean unexpected shutdown behavior, false readings and delayed alarms, or cooling degradation that triggers thermal throttling. In the worst case, a localized failure, such as a rack-level power event, can cascade into upstream facility systems.. This distinction should change how leaders measure security outcomes.

The question is not solely “are we preventing intrusions?” but “can we sustain safe operations during one?” That is the core of resilience.. Related:From Data Centers to Models: White House Targets AI Risks at the Source. The Stakes Are Higher Than Most People Realize.

One reason this conversation is lagging is that we have not fully internalized the knock-on effects of consequential failure.. When one critical sector goes down, others degrade quickly. A hospital, for example, depends on more than clinicians and buildings.

It needs medical records, transportation, cold-chain storage for medicines, and stable power for environmental conditions. When those enabling systems fail, care often continues but becomes manual, slower, and riskier.. Now apply that lens to data centers.

As AI becomes a dependency for how businesses operate, how supply chains run, and how national security systems analyze and act, data center uptime becomes an economic security and public safety issue, not just a service availability issue.. The challenge is not simply that there are more data centers.

It’s that they’re being designed, built, and commissioned faster than ever before, creating conditions where governance, security validation, and operational maturity struggle to keep pace.. Hypergrowth Is Creating New Risk. Hypergrowth means governance is struggling to keep pace with newly acquired sites, rapidly deployed devices, and compressed deployment timelines.

When organizations break ground on a new data center every three months, speed introduces challenges across the supply chain and operational quality.. Related:Anthropic’s Project Glasswing Tackles AI Security Challenges in Data Centers. Supply Chain: The technologies powering and cooling modern data centers are evolving rapidly.

Many new vendors, including startups and suppliers from semi-trusted geographies, may lack mature secure software development life cycles (SDLC) and manufacturing processes. As these technologies become embedded in critical infrastructure, resilient CPS programs are essential for managing third party risk..

Quality: The market rewards speed, and billions of dollars can hinge on meeting aggressive deployment timelines. As contractors, third parties, and internal teams work faster than ever, speed-driven mistakes become more common. Real-world examples include misconfigured network segmentation (VLANs), exposed ports, and default credentials left unchanged..

Speed has become a competitive advantage, but it has also become a security liability. Without governance keeping pace, organizations risk building tomorrow’s critical infrastructure on today’s operational blind spots.. What The Research Is Telling Us.

Recent research shows how directly these vulnerabilities can map to operational disruption. A recent report found near-maximum-severity flaws in the network interfaces used to manage UPS systems. An authenticator bypass combined with a remote code execution condition could, in the worst case, enable an attacker to bypass login controls and issue commands that interrupt power to protected loads.

In a data center, that is not an inconvenience. It is a potential facility-scale outage event.. Other findings revealed a chain of vulnerabilities in a widely deployed HVAC controller, including a path to root-level remote access without authentication and exposure of sensitive facility data.

In combination, this type of chain can give an attacker influence over cooling systems that directly determine whether high-density workloads remain stable. Cooling is not a background system. It is an uptime dependency, and as rack densities rise, thermal margin shrinks..

The pattern matters more than the specific risks. In both cases, the equipment that most IT security teams rarely think about turned out to be exactly the equipment that determines whether a facility stays online.. Closing the Gap.

Many cyber-physical devices were designed for uptime and maintainability in an era of limited connectivity and a very different threat model. Today, connectivity is a business requirement; remote maintenance is commonplace, and these systems often sit entirely outside traditional IT vulnerability management.

Facilities teams own the vendor relationship but lack a workflow to track firmware risk. Security teams know how to operationalize vulnerability management, but do not have these assets in inventory.. Start with asset visibility across both gray space and white space, so you know what exists, where it sits, what it talks to, and which protocols and firmware versions it runs.

Map those assets to operational purpose so remediation and compensating controls prioritize what is most critical to uptime. From there, make segmentation a first-class uptime control by restricting communications to what is required, limiting management traffic to authorized paths, and separating control networks from business networks so a compromise cannot become a facility-wide event.

Apply the same discipline to remote access by enforcing strong authentication, least privilege, and auditable sessions for vendors and contractors.. Finally, treat detection and response as cyber-physical disciplines by correlating security events with operational anomalies and ensuring that playbooks anticipate the manipulation of monitoring and control systems, with the goal of preventing incidents from becoming outages..

Why This Is an Executive Issue. Underwriters increasingly expect auditable proof of cyber-physical controls that preserve operational resilience. That means visibility into CPS assets, segmentation that limits blast radius, governed remote access, and reporting that shows progress over time.

These expectations increasingly mirror the rigor applied to other critical infrastructure sectors.. That shifts OT security from a technical debate into a business one. If uptime is the product, cyber-physical resilience is part of product quality..

The industry is right to focus on physical threats and perimeter security. But the systems keeping the lights on, the servers cool, and the power flowing deserve the same scrutiny. Because when this layer fails, the headlines rarely say “breach.” They say “outage.”.

About the Author

 

Join the conversation

Your email address will not be published. Required fields are marked *