AI-driven demand continues to fuel the need for more reliable and resilient data centers. As data centers are now as mission-critical as aircraft, it is time to rethink the traditional redundancy architectures—and this requires moving beyond the static, over-provisioned approach.

Recent studies demonstrate new redundancy strategies to address these challenges. These include leveraging redundant data center systems to generate revenue by supporting the grid. Data centers are prosumers; they both consume and produce power. Similarly, shifting the cooling redundancy model from idle standby to active operation and optimising server workloads to improve efficiency have been demonstrated. This article reviews these recent advancements across three key data center pillars: cooling, power, and server-level redundancy, and argues that smarter implementation is more impactful than more hardware.

Cooling-level redundancy

The end goal of data center cooling is to keep the IT equipment operational without interruption. In a study cited by Cho et al. (2024), cooling system failures account for 51 percent of downtime, making it the most common cause of interruption in data centers. Hence, it is imperative to strike a balance between maintaining the reliability of data center cooling systems and reducing energy consumption.