Skip to main content

    How to Survive (and Prevent) an IT Emergency

    By Joseph HolkoMay 26, 2026Updated August 11, 2026Business Continuity7 min read

    The worst part of a serious outage usually isn't the failure itself. It's the first hour afterward, when everybody who wants to help starts helping. Someone reboots the server to see if that clears it. Someone else finds a copy of a folder on a shared drive and restores it over the top of live data. A manager tells the office it'll be back shortly, because that's what he hopes. None of it is malice. But by the time the person who can actually fix it has the full picture, there are two problems instead of one, and nobody can say for certain what the first one was. Most of how that hour goes was decided months earlier, on ordinary days when nothing was broken.

    So the first move in a real outage is to stop long enough to name one person who makes the calls, and to get everyone else to agree they'll check with that person before touching anything. Who is that at your company at 6:40 on a Tuesday morning, when the office manager can't get into email and your IT contact hasn't picked up yet?

    Then map it before you explain it. What is actually down, and for whom? One application or all of them, one site or the whole company? I once watched two hours disappear into a server that was fine, because nobody had checked whether the internet circuit was up. Scope is cheap to establish. Getting it early also tells you within minutes rather than hours whether this is a failure or an intrusion. Those get handled differently. If there's any chance it's the second, isolate the machine and don't wipe it, because what you erase at 7am is what your insurer and your attorney ask for at 4pm.

    Somebody should be writing this down as it happens, with times. Nothing formal, just a running note of what was observed, what was tried, and what changed after. It costs one person's attention.

    You can hear a good provider in the shape of the communication. Somebody takes ownership by name. They tell you plainly what they know and what they don't, and commit to a next update at a stated time, even when that update is nothing more than still working on it. The bad version sounds more confident. Fast certainty about a cause in the first ten minutes is usually somebody guessing out loud to make you feel better, and it ages badly.

    Everything I've described so far is execution, and execution runs on what you already prepared. A named decision-maker, a plan short enough to read, an honest list of what the business actually needs, backups that restore, and systems somebody has been maintaining. None of it can be assembled while the phones are ringing.

    You can watch that play out in the incident numbers. Verizon's 2025 breach report splits small and mid-size businesses from large organizations across just over three thousand incidents, and ransomware shows up in 88% of the SMB breaches against 39% at the large ones. Their explanation for the gap runs to a single sentence: "It is simply a bonus for the attacker that SMBs are less likely to have up-to-date and readily available backups than a large organization."1 Nothing in that sentence is about the first hour. The same attackers hit both groups.

    Start with the plan, because it's the cheapest thing to fix. Most recovery plans we're asked to review share one flaw. They live on the file server. If the outage is the file server, the plan is inside the outage. Ours run about two pages and exist on paper and on somebody's phone. Two pages is the point, because nobody reads a forty-page binder at 6:40 in the morning. Who to call at the provider, at the carrier, at the insurer, at the bank, in what order, and who tells the staff. If your file server went down this afternoon, where is your copy?

    The plan is short because the list underneath it is short. Every business has a handful of systems it genuinely can't run without, and the handful is smaller than people expect. Internet and email, plus the one or two line-of-business applications where the work actually happens. Everything else can wait a day, and knowing which is which decides what gets restored first when you can't restore everything at once. It also tells you how much downtime you can absorb before it costs real money. How long can your team work on paper before you lose a customer or miss payroll? That's a business decision for somebody who signs checks.

    Which brings us to backups, where the distance between what people believe and what's true is widest. A backup job that completes proves only that data was written somewhere. It doesn't prove that the data reads back, that you're keeping it long enough, or that a full restore finishes inside the window your business can survive. NIST's control catalog treats testing restorability as a control of its own, CP-9(1), separate from the control that requires performing backups at all.2 If a completed job demonstrated restorability, nobody would have needed to write it.

    Ask a room of owners how fast they could recover and you'll get confident answers. Kaseya surveyed just over three thousand IT professionals for its 2025 continuity report, and more than 60% of their organizations believed they could be back inside a day. Only 35% could.3 Nearly all of them had backups running. Most of that 25-point difference got discovered on the worst day of somebody's year.

    Part of the reason is that backups stopped being a bystander in these events. Now attackers go after them first. In a study of nearly three thousand organizations that had been hit, 94% of victims said the attackers tried to compromise their backups, and 57% of those attempts worked. The ones whose backups were destroyed paid roughly eight times as much to recover, and 26% were fully back inside a week, against 46% of the ones whose backups survived.4 So there are two things worth putting to your provider. Whether at least one copy sits somewhere your own admin credentials can't reach it, and when somebody last restored a real system end to end and timed it. If nobody can give you a date, you don't have a tested backup.

    The last piece is maintenance, and it predicts more than any of the others. Most of the emergencies we get called into start with maintenance somebody deferred. A firewall running firmware from three years ago. A VPN appliance nobody has owned since the technician who set it up left.

    The good data on this comes from data-center operators, who are far larger than anyone reading this, so take it as direction rather than as your own numbers. Across the operations Uptime Institute surveyed for 2026, configuration and change management failures caused 41% of major network outages and hardware failure another 34%, while malicious attack accounted for 9%. For system and software outages it was software faults at 53% and configuration at 51%, against 15% for attack. And 92% said human error contributed to their most recent significant outage, usually staff not following procedures that already existed.5 Those are organizations with dedicated teams and written runbooks. If drift and skipped procedure still take them down at that rate, a company where one person handles all of it between three other jobs isn't doing better.

    So who patches your firewall, and how would you find out if they stopped? If you're not sure how to ask that without it turning into an argument, our MSP Performance Scorecard is one way to structure the conversation. The answer that comes out of it is often to fix the process with the provider you already have.

    Want to know where you actually stand?

    A Technology Confidence Assessment is an independent review of your backup posture, recovery planning, and infrastructure resilience. You get a written read on where the real exposure is, and what to address first.

    Book your assessment

    None of this makes you immune. Things break, and somebody still clicks the thing. What changes is the size of the event. The businesses that come through a serious outage in half a day and the ones that lose a week are rarely separated by what they did in the first hour. The hour everyone remembers was decided in the months nobody remembers.

    Get Professional Guidance

    Schedule a free Technology Confidence Assessment to get personalized recommendations for your business.

    Book your assessment