Presentation
Distributed Modular Digital Twin Network for High-Performance and Reliable Data Centers
DescriptionHigh performance computing (HPC) workloads are driving rack power densities beyond 100 kW, creating unprecedented stress on data center cooling and power systems. Conventional CFD-based digital twins provide high-fidelity design optimization but are too computationally intensive and rigid for operational use. We present the first physics-constrained Distributed Modular Digital Twin Network (DMDTN), designed for real-time performance evaluation, load prediction, and fault detection. Each subsystem (e.g., cooling, power, IT load) is represented by an AI-driven surrogate model, interconnected through conservation laws and coordinated via a distributed message bus. This modular design preserves physical consistency while enabling scalability and rapid adaptability. Using synthetic datasets, DMDTN achieved ~60% lower prediction error (RMSE 172 vs. 450) and more than 2× faster training (201 vs. 442 seconds) than a monolithic model, while maintaining robustness under stress. DMDTN complements CFD by enabling accurate, real-time operational management of HPC data centers.

Event Type
Research and ACM SRC Posters
TimeThursday, 20 November 20258:00am - 5:00pm CST
LocationSecond Floor Atrium
Archive
view
