Data Center Cooling Emergency Response Plan: What to Do When a CRAC Unit Fails

Article snapshotSmart cooling
2 min readEstimated time
10 sectionsWhat’s inside
  • Why You Need a Data Center Cooling Emergency Response Plan
  • Step 1: Confirm the Failure and Assess Risk
  • Step 2: Activate Temporary Airflow Measures
  • Step 3: Bring in Portable Cooling or Spare Units

Why You Need a Data Center Cooling Emergency Response Plan

A CRAC unit can stop without warning. A tripped breaker, a worn belt, or a failed compressor—and the room starts heating up. In a typical server room with 10 racks drawing 5 kW each, that's 50 kW of heat with nowhere to go. Without cooling, temperatures can rise several degrees per minute, depending on room volume and IT load. That's why a data center cooling emergency response plan isn't a luxury. It's a necessity.

The plan should be simple enough to execute under stress. You don't want people fumbling through manuals while alarms are blaring. You want clear roles, clear actions, and clear thresholds. Let's walk through what a practical response looks like.

Server room temperature rising after CRAC failure

Step 1: Confirm the Failure and Assess Risk

First, verify the alarm. Sometimes a sensor glitch or a brief power dip triggers a false positive. Check the CRAC unit's display or controller. Is it actually off? Is the fan spinning? Is the compressor running? If the unit is down, note the time—this becomes your baseline for rate of temperature rise.

Next, assess the risk. How much IT load is in the room? What's the current inlet temperature? If you have redundant cooling, can the remaining units handle the load? For example, if you have two 30 kW CRAC units and the IT load is 40 kW, one unit can't cool the room alone. You're already in trouble. If the load is 20 kW, you have margin—but not for long.

  • Check the CRAC controller for error codes and log the time of failure.
  • Measure inlet temperatures at the top of racks—hot spots appear first.
  • Review the cooling redundancy: is there an N+1 unit that can start automatically?
  • Note the room's rate of temperature rise—this tells you how much time you have.

Step 2: Activate Temporary Airflow Measures

Time is short. If the room is heating up faster than 1°C per minute, you need to buy time. Open the floor tiles if you have raised-floor cooling. That lets cold air escape closer to the racks. Position portable fans to push air across hot spots. But be careful—fans can recirculate hot air if placed wrong. Aim them to pull cool air from the floor or from the nearest operating CRAC.

In many server rooms, the issue is not total heat but uneven distribution. A single rack with high-density equipment might be cooking while the rest of the room is fine. Check the aisle containment—are there gaps in the cold aisle? Seal them temporarily with plastic sheeting or even cardboard. It's not pretty, but it works.

Portable fan directed at server rack during cooling failure

Step 3: Bring in Portable Cooling or Spare Units

If the room continues to heat, you need more cooling capacity. Portable air conditioners or spot coolers can help, but they have limits. A typical portable unit might provide 5–10 kW of cooling, which is often not enough for a whole room. Use them to target hot spots—place them near the racks with the highest inlet temperatures.

If you have a spare CRAC unit on site, now's the time to roll it in. But swapping a unit takes hours, not minutes. You need to disconnect the failed unit, move the spare into place, connect power and coolant lines, and test it. That's a job for trained technicians. In the meantime, your temporary measures are keeping the room alive.

Some facilities have pre-installed connections for portable cooling units. If you don't, consider adding them during your next upgrade. A simple quick-connect fitting on the chilled water loop can save hours during an emergency.

Step 4: Reduce IT Load

Sometimes the fastest way to reduce heat is to shut down non-critical workloads. If you have a virtualized environment, you can migrate VMs to other locations or simply power off development servers. In a colocation facility, you might need to contact customers and ask them to reduce load. That's not fun, but it's better than a full outage.

Know your load priorities in advance. Which applications are truly business-critical? Which can be down for an hour? Document this in your data center cooling emergency response plan. When the alarm goes off, you don't want to be making those decisions on the fly.

  1. Identify non-critical workloads and shut them down first.
  2. Migrate virtual machines to other hosts if possible.
  3. If you have multiple server rooms, move high-density workloads to a room with spare cooling.
  4. Communicate with IT stakeholders before shutting down anything—they may have dependencies you don't know about.

Step 5: Know Your Safe Shutdown Thresholds

There's a point where you must stop the IT equipment, even if it means downtime. Most server manufacturers specify a maximum inlet temperature—often 35°C to 40°C (95°F to 104°F). But running at those limits for long is risky. Components can degrade, and fans spin at full speed, drawing more power.

Set your shutdown threshold lower than the hardware limit. For example, if the server spec says 40°C, start shutting down at 35°C. That gives you a buffer. And don't forget humidity—if the room gets too dry, static discharge can damage equipment. Most CRAC units control humidity, so when they fail, you might see humidity drop. If it falls below 20% RH, that's another reason to shut down.

ParameterTypical ThresholdAction
Server inlet temperature35°C (95°F)Begin load shedding
Server inlet temperature40°C (104°F)Shut down non-critical IT
Room humidityBelow 20% RHShut down if sustained
Rate of temperature rise>2°C per minuteImmediate shutdown of all IT

These numbers are examples—check your equipment specs and set your own thresholds. The key is to decide in advance, so you don't hesitate when the time comes.

Technician monitoring temperature and humidity sensors

Step 6: Communicate Clearly and Document Everything

During an incident, communication is as important as technical action. Who needs to know? The facilities manager, the IT operations team, and possibly senior management. If you're in a colocation facility, the provider's NOC (Network Operations Center) should be on your call list. And if you have customers, they need to know if their services are at risk.

Designate a single incident commander. That person coordinates all actions and communications. Everyone else reports to them. This avoids chaos and conflicting instructions. Use a conference bridge or a group chat to keep everyone informed. And log everything—who did what, when, and what the temperatures were. That log will be invaluable for the post-incident review.

Step 7: Diagnose the Root Cause and Restore Cooling

Once the immediate threat is under control, you need to find out why the CRAC unit failed. Check the error codes, inspect the belts, look for refrigerant leaks, and test the electrical connections. Was it a simple power trip? A worn belt that snapped? A failed compressor motor? The diagnosis will determine how long the repair takes.

If you have a maintenance contract with a provider like VERHI, call them right away. They can dispatch a technician with the right parts. If you're handling it in-house, make sure you have spare parts on hand—belts, filters, contactors, and maybe a spare fan motor. Waiting for parts is the worst part of any repair.

Before restarting the CRAC unit, verify that the power supply is stable. If the failure was caused by a power surge, you might need to check the UPS and the electrical panel. Also, inspect the condenser coils—if they're clogged with dirt, the unit will fail again soon after restart.

Step 8: After the Event—Review and Improve

Once cooling is restored and temperatures are back to normal, the work isn't over. Hold a post-incident review within a few days. What went well? What didn't? Were the thresholds correct? Did the communication plan work? Did anyone hesitate because they weren't sure what to do?

Update your data center cooling emergency response plan based on what you learned. Maybe you need better monitoring, or more portable fans, or a different shutdown order. Perhaps you should add a redundant CRAC unit, or improve your preventive maintenance schedule. The goal is to make the next failure less disruptive.

Also, check your PUE (Power Usage Effectiveness) after the incident. If you had to run fans at full speed or use portable cooling, your PUE likely spiked. That's expected, but it's worth documenting to show the financial impact of downtime.

Frequently Asked Questions

What is the first thing to do when a CRAC unit fails?

Confirm the failure, note the time, and check the rate of temperature rise. If the room is heating up fast, activate your emergency plan immediately—starting with temporary airflow and load reduction.

How hot can a server room get before equipment shuts down?

Most servers can handle inlet temperatures up to 35-40°C (95-104°F), but running at those limits for long is risky. Set your shutdown threshold lower, like 35°C, to give yourself a buffer.

Can portable air conditioners save a server room?

They can help with hot spots, but a typical portable unit only provides 5-10 kW of cooling—often not enough for a whole room. Use them to target the hottest racks while you work on a more permanent fix.

How often should I test my emergency cooling plan?

At least twice a year. Run a simulated failure drill to see how your team responds. You'll likely find gaps in your plan that you can fix before a real emergency.

What maintenance prevents CRAC failures?

Regular checks of belts, filters, refrigerant levels, and electrical connections. Also, keep condenser coils clean and replace worn parts before they fail. A good preventive maintenance schedule reduces the chance of sudden breakdowns.

Need help building or testing your cooling emergency plan? VERHI's engineers can review your current setup and recommend practical steps to improve resilience. Talk to us today.

V
VERHI Editorial Team
Precision cooling, UPS and data center infrastructure content team
Reviewed by VERHI Technical Editorial Review on 2026-09-02

Based on VERHI's field experience with precision air conditioning systems and data center cooling failures.

Information can change. Confirm time-sensitive details with the official provider or your trip/technical advisor before making plans.

Similar Posts

Leave a Reply

Your email address will not be published. Required fields are marked *