- Why You Need a Data Center Cooling Emergency Response Plan
- Step 1: Confirm the Failure and Assess Risk
- Step 2: Activate Temporary Airflow Measures
- Step 3: Bring in Portable Cooling or Spare Units
Why You Need a Data Center Cooling Emergency Response Plan
A CRAC unit can stop without warning. A tripped breaker, a worn belt, or a failed compressor—and the room starts heating up. In a typical server room with 10 racks drawing 5 kW each, that's 50 kW of heat with nowhere to go. Without cooling, temperatures can rise several degrees per minute, depending on room volume and IT load. That's why a data center cooling emergency response plan isn't a luxury. It's a necessity.
The plan should be simple enough to execute under stress. You don't want people fumbling through manuals while alarms are blaring. You want clear roles, clear actions, and clear thresholds. Let's walk through what a practical response looks like.

Step 1: Confirm the Failure and Assess Risk
First, verify the alarm. Sometimes a sensor glitch or a brief power dip triggers a false positive. Check the CRAC unit's display or controller. Is it actually off? Is the fan spinning? Is the compressor running? If the unit is down, note the time—this becomes your baseline for rate of temperature rise.
Next, assess the risk. How much IT load is in the room? What's the current inlet temperature? If you have redundant cooling, can the remaining units handle the load? For example, if you have two 30 kW CRAC units and the IT load is 40 kW, one unit can't cool the room alone. You're already in trouble. If the load is 20 kW, you have margin—but not for long.
- Check the CRAC controller for error codes and log the time of failure.
- Measure inlet temperatures at the top of racks—hot spots appear first.
- Review the cooling redundancy: is there an N+1 unit that can start automatically?
- Note the room's rate of temperature rise—this tells you how much time you have.
Step 2: Activate Temporary Airflow Measures
Time is short. If the room is heating up faster than 1°C per minute, you need to buy time. Open the floor tiles if you have raised-floor cooling. That lets cold air escape closer to the racks. Position portable fans to push air across hot spots. But be careful—fans can recirculate hot air if placed wrong. Aim them to pull cool air from the floor or from the nearest operating CRAC.
In many server rooms, the issue is not total heat but uneven distribution. A single rack with high-density equipment might be cooking while the rest of the room is fine. Check the aisle containment—are there gaps in the cold aisle? Seal them temporarily with plastic sheeting or even cardboard. It's not pretty, but it works.

Step 3: Bring in Portable Cooling or Spare Units
If the room continues to heat, you need more cooling capacity. Portable air conditioners or spot coolers can help, but they have limits. A typical portable unit might provide 5–10 kW of cooling, which is often not enough for a whole room. Use them to target hot spots—place them near the racks with the highest inlet temperatures.
If you have a spare CRAC unit on site, now's the time to roll it in. But swapping a unit takes hours, not minutes. You need to disconnect the failed unit, move the spare into place, connect power and coolant lines, and test it. That's a job for trained technicians. In the meantime, your temporary measures are keeping the room alive.
Some facilities have pre-installed connections for portable cooling units. If you don't, consider adding them during your next upgrade. A simple quick-connect fitting on the chilled water loop can save hours during an emergency.
Step 4: Reduce IT Load
Sometimes the fastest way to reduce heat is to shut down non-critical workloads. If you have a virtualized environment, you can migrate VMs to other locations or simply power off development servers. In a colocation facility, you might need to contact customers and ask them to reduce load. That's not fun, but it's better than a full outage.
Know your load priorities in advance. Which applications are truly business-critical? Which can be down for an hour? Document this in your data center cooling emergency response plan. When the alarm goes off, you don't want to be making those decisions on the fly.
- Identify non-critical workloads and shut them down first.
- Migrate virtual machines to other hosts if possible.
- If you have multiple server rooms, move high-density workloads to a room with spare cooling.
- Communicate with IT stakeholders before shutting down anything—they may have dependencies you don't know about.
Step 5: Know Your Safe Shutdown Thresholds
There's a point where you must stop the IT equipment, even if it means downtime. Most server manufacturers specify a maximum inlet temperature—often 35°C to 40°C (95°F to 104°F). But running at those limits for long is risky. Components can degrade, and fans spin at full speed, drawing more power.
Set your shutdown threshold lower than the hardware limit. For example, if the server spec says 40°C, start shutting down at 35°C. That gives you a buffer. And don't forget humidity—if the room gets too dry, static discharge can damage equipment. Most CRAC units control humidity, so when they fail, you might see humidity drop. If it falls below 20% RH, that's another reason to shut down.
| Parameter | Typical Threshold | Action |
|---|---|---|
| Server inlet temperature | 35°C (95°F) | Begin load shedding |
| Server inlet temperature | 40°C (104°F) | Shut down non-critical IT |
| Room humidity | Below 20% RH | Shut down if sustained |
| Rate of temperature rise | >2°C per minute | Immediate shutdown of all IT |
These numbers are examples—check your equipment specs and set your own thresholds. The key is to decide in advance, so you don't hesitate when the time comes.

Step 6: Communicate Clearly and Document Everything
During an incident, communication is as important as technical action. Who needs to know? The facilities manager, the IT operations team, and possibly senior management. If you're in a colocation facility, the provider's NOC (Network Operations Center) should be on your call list. And if you have customers, they need to know if their services are at risk.
Designate a single incident commander. That person coordinates all actions and communications. Everyone else reports to them. This avoids chaos and conflicting instructions. Use a conference bridge or a group chat to keep everyone informed. And log everything—who did what, when, and what the temperatures were. That log will be invaluable for the post-incident review.
Step 7: Diagnose the Root Cause and Restore Cooling
Once the immediate threat is under control, you need to find out why the CRAC unit failed. Check the error codes, inspect the belts, look for refrigerant leaks, and test the electrical connections. Was it a simple power trip? A worn belt that snapped? A failed compressor motor? The diagnosis will determine how long the repair takes.
If you have a maintenance contract with a provider like VERHI, call them right away. They can dispatch a technician with the right parts. If you're handling it in-house, make sure you have spare parts on hand—belts, filters, contactors, and maybe a spare fan motor. Waiting for parts is the worst part of any repair.
Before restarting the CRAC unit, verify that the power supply is stable. If the failure was caused by a power surge, you might need to check the UPS and the electrical panel. Also, inspect the condenser coils—if they're clogged with dirt, the unit will fail again soon after restart.
Step 8: After the Event—Review and Improve
Once cooling is restored and temperatures are back to normal, the work isn't over. Hold a post-incident review within a few days. What went well? What didn't? Were the thresholds correct? Did the communication plan work? Did anyone hesitate because they weren't sure what to do?
Update your data center cooling emergency response plan based on what you learned. Maybe you need better monitoring, or more portable fans, or a different shutdown order. Perhaps you should add a redundant CRAC unit, or improve your preventive maintenance schedule. The goal is to make the next failure less disruptive.
Also, check your PUE (Power Usage Effectiveness) after the incident. If you had to run fans at full speed or use portable cooling, your PUE likely spiked. That's expected, but it's worth documenting to show the financial impact of downtime.
Frequently Asked Questions
What is the first thing to do when a CRAC unit fails?
Confirm the failure, note the time, and check the rate of temperature rise. If the room is heating up fast, activate your emergency plan immediately—starting with temporary airflow and load reduction.
How hot can a server room get before equipment shuts down?
Most servers can handle inlet temperatures up to 35-40°C (95-104°F), but running at those limits for long is risky. Set your shutdown threshold lower, like 35°C, to give yourself a buffer.
Can portable air conditioners save a server room?
They can help with hot spots, but a typical portable unit only provides 5-10 kW of cooling—often not enough for a whole room. Use them to target the hottest racks while you work on a more permanent fix.
How often should I test my emergency cooling plan?
At least twice a year. Run a simulated failure drill to see how your team responds. You'll likely find gaps in your plan that you can fix before a real emergency.
What maintenance prevents CRAC failures?
Regular checks of belts, filters, refrigerant levels, and electrical connections. Also, keep condenser coils clean and replace worn parts before they fail. A good preventive maintenance schedule reduces the chance of sudden breakdowns.
Need help building or testing your cooling emergency plan? VERHI's engineers can review your current setup and recommend practical steps to improve resilience. Talk to us today.
Based on VERHI's field experience with precision air conditioning systems and data center cooling failures.
