data center career en · 2026 · Technical Research Note
00. From Disaster Recovery to Data Center Infrastructure
Keeping eight years of disaster recovery experience while learning the physical systems that sustain continuous service
A structured learning path from disaster recovery into power, cooling, facility monitoring, commissioning, and data center project delivery.
Research edition · 8 chapters · 6 minute read · Updated 2026-08-25
I graduated in Computer Science and Technology in 2018 and have worked in disaster recovery since then. Most projects discuss recovery point objectives (RPOs), which limit acceptable data loss, and recovery time objectives (RTOs), which limit how long a service may remain unavailable. My work has concentrated on servers, storage, networks, replication, and failover. Other engineers handled the power and cooling systems beneath them.
That division is less distinct than it appears. A second copy of the data cannot keep a service available if the uninterruptible power supply (UPS) exhausts its batteries before the generators start. A redundant application cluster still depends on cooling capacity after a chiller fails. A recovery plan also depends on monitoring that identifies a facility fault before it affects both the primary system and its recovery path.
I started filling this gap on August 25, 2026. The plan runs through the end of the year at two to three hours per day. Four months will not make me an electrical or mechanical designer. The target is narrower: read a medium-scale data center design, understand the dependencies among construction, commissioning, operations, and disaster recovery, and defend a three-data-center design under technical review.
What carries over and what is missing
The previous eight years still count. Active-active services, remote replication, RPOs, RTOs, failover exercises, and incident procedures already belong to data center business continuity. The missing side is the physical infrastructure that keeps the computing equipment running: electrical distribution, cooling, fire protection, facility monitoring, and the delivery process that turns drawings into an operational site.
This series follows those dependencies. It begins with the complete facility, then studies power and cooling. Once that foundation is clear, it moves to data center infrastructure management (DCIM), building management systems (BMS), acceptance testing, and integrated commissioning. The final articles place every system back into a two-city, three-data-center recovery design.
Part I: the complete facility
Article 01 follows a project through planning, design, procurement, construction, commissioning, handover, and operations. It identifies the responsibilities of the owner, design institute, general contractor, vendors, supervision team, and commissioning authority.
Article 02 starts at a server rack and traces its dependencies across compute, storage, network, electrical power, cooling, fire protection, physical security, and monitoring.
Article 03 explains N, N+1, 2N, and Tier I through Tier IV. The useful question is not the label. It is whether planned maintenance or one equipment failure interrupts the supported service.
Part II: electrical power
Article 04 follows electricity from the utility supply through switchgear, transformers, low-voltage distribution, UPS equipment, row distribution, and rack power distribution units. A UPS uses stored energy to sustain the load between a utility failure and the arrival of an alternate source.
Article 05 examines rectifiers, inverters, static bypasses, maintenance bypasses, batteries, and the operational window created by battery runtime.
Article 06 covers generators, automatic transfer switches (ATS), and static transfer switches (STS). An ATS transfers between power sources; an STS provides faster transfer for supported electrical paths.
Article 07 compares N+1 and 2N power paths by injecting utility, UPS, and generator failures and identifying remaining single points of failure.
Part III: cooling and efficiency
Article 08 traces server heat through air, water, computer-room cooling equipment, chillers, pumps, and cooling towers until it reaches the outdoor environment.
Article 09 studies hot aisles, cold aisles, containment, and airflow. The same nominal cooling capacity can produce different server inlet temperatures when airflow is poorly controlled.
Article 10 examines high-density GPU racks and liquid cooling. Liquid cooling moves heat with a fluid near the rack or processor and changes power density, piping, monitoring, and maintenance requirements.
Article 11 uses power usage effectiveness (PUE), the ratio of total facility energy to IT equipment energy. PUE describes facility-wide efficiency but does not by itself establish the efficiency of an application or rack.
Part IV: monitoring and commissioning
Article 12 separates DCIM, BMS, and power and environmental monitoring. DCIM manages data center capacity, assets, and infrastructure operations. BMS supervises building systems. Power and environmental monitoring collects equipment and room conditions and raises alarms.
Article 13 explains how SNMP, Modbus, and BACnet connect servers, meters, UPS equipment, cooling systems, and building controllers. It then studies alarm severity, aggregation, and automated responses.
Article 14 distinguishes factory acceptance testing (FAT), site acceptance testing (SAT), equipment startup, system testing, and integrated systems testing.
Article 15 designs a utility-failure test with prerequisites, risk controls, observation points, and acceptance criteria for the UPS, generators, switching, cooling, monitoring, and IT load.
Article 16 builds a failure library covering power, cooling, networking, storage replication, monitoring loss, fire alarms, and human error. Every scenario records detection, service impact, response, recovery, and acceptance criteria.
Part V: returning to disaster recovery
Article 17 revisits RPO, RTO, mean time to repair (MTTR), and mean time between failures (MTBF). These measures describe different subjects and should not be collapsed into one availability number.
Article 18 designs a financial-services deployment across a primary site, a metropolitan recovery site, and a remote recovery site. Business tiers, replication, networks, power, cooling, and monitoring appear in one dependency model.
Article 19 creates failover and failback exercises with separate acceptance criteria for technical transition, business availability, and data consistency.
Article 20 builds a small DCIM demonstration with synthetic temperature, humidity, UPS load, rack power, PUE, alarm, replication, RPO, and RTO data.
Part VI: converting the work into career evidence
Article 21 rewrites eight years of disaster recovery work in terms understood by data center employers while preserving the actual scope of each project.
Article 22 prepares for solution architect, delivery manager, commissioning manager, and technical project manager interviews. Answers must be supported by the complete project rather than memorized definitions.
Article 23 reviews the first applications and interviews. Job descriptions and unanswered questions become inputs to the learning plan instead of waiting until every subject feels complete.
What should exist by the end of the year
Time spent is not the acceptance criterion. By December 31, 2026, this work should have produced:
- A complete data center systems map.
- Basic electrical and cooling diagrams.
- A library of at least twenty failure scenarios.
- An integrated utility-failure test plan.
- A three-data-center infrastructure and disaster recovery design.
- A DCIM demonstration driven by synthetic data.
- A role-specific resume and project brief.
- A recorded thirty-minute design presentation.
The artifacts must support one another. Redundancy in the design should appear in the tests. Risks found in testing should enter the failure library. DCIM alarms should correspond to equipment states. Resume claims should be visible in the design and demonstration. The result will be a reviewable engineering record rather than a list of completed courses.
Statement: If no specific statement in the content, the copyright belongs to sshipanoo . Reprint please indicate the link of this article.
(The content is authorized with CC BY-NC-SA 4.0 protocol)
Title:00. From Disaster Recovery to Data Center Infrastructure
