Pages

▼

Backup, Recovery & Data Protection

🧑🏻‍🎓 AL Academy Masterclass

Backup, Recovery & Data Protection

Everyone has backups. Almost nobody has restores. Here is the difference, and why it decides whether a bad day becomes a catastrophe.


The checkmark that lies

Walk into most organizations and ask whether they back up their data, and the answer is a confident yes. Ask when they last restored from one of those backups, and the room goes quiet. That gap - between a job that reports success and data you can actually get back - is where businesses die.

A backup job's green checkmark proves exactly one thing: that a write happened. It says nothing about whether the data is complete, uncorrupted, decryptable, or recoverable inside the time you have. Backups fail silently in a dozen mundane ways - an excluded folder nobody noticed, a database file copied while it was open and now useless, a chain of incremental backups with one broken link, an encryption passphrase that walked out the door with a departed admin. None of these trip an alarm. All of them are discovered at the worst possible moment, with the original data already gone.

So the first principle of data protection is uncomfortable and absolute: if you have never restored it, you do not have a backup. You have a job that appears to succeed.

Two numbers before any tool

The instinct, when someone says "we need backups," is to go shopping for a product. Resist it. Backup design starts with two numbers, agreed with the business and written down for each system.

The first is the Recovery Point Objective - how much data, measured in time, you can afford to lose. If losing a day's work is survivable, a nightly backup is fine. If losing fifteen minutes is not, you need frequent snapshots or continuous replication. RPO tells you how often to back up.

The second is the Recovery Time Objective - how long you can be down before service is restored. RTO tells you how fast and by what method you must recover. A four-hour RTO means detection, decision, restore, and validation all have to fit inside four hours. If your restore takes nine, you do not have a four-hour RTO no matter what the policy document claims.

These two numbers justify every later decision: the schedule, the tooling, the storage tiers, the budget. A backup plan without an RPO and RTO is just hope on a timer.

The rule, hardened for the ransomware era

The structural backbone of any serious plan is the 3-2-1 rule: three copies of your data, on two different media types, with one copy off-site. Three copies mean no single failure leaves you at zero. Two media types avoid a shared failure mode. The off-site copy survives the fire, the flood, the theft, and the ransomware blast that reaches everything on the local network.

The modern world adds two more digits: 3-2-1-1-0. The extra one is a copy that is offline or immutable - air-gapped, or written to storage that cannot be altered or deleted within a retention window, even by an administrator. The zero means zero errors, proven by testing restores. That immutable copy is the single most important upgrade of the past decade, because attackers now hunt and encrypt reachable backups first. A victim with clean backups does not pay the ransom, so the backups become the target.

This is also why two comforting things are not backups. RAID is not a backup - it protects against a disk dying, not against deletion, corruption, ransomware, or a controller that scrambles the whole array at once. And file-sync is not a backup - OneDrive, Dropbox, and Google Drive cheerfully replicate your deletions and your ransomware encryption to every copy, fast.

The blind spots

Even teams with a solid on-premises backup routine tend to share the same gaps.

The biggest is SaaS. Microsoft 365 and Google Workspace run a shared-responsibility model: the provider keeps the service available, but your data is your responsibility. Their recycle bins and retention windows are short and are not designed as a backup. A deleted mailbox or a ransomware-synced document library can be gone permanently. If your business runs on cloud collaboration tools, you need to back that data up into storage you control.

The second is bare-metal recovery. Teams faithfully back up the data and never test rebuilding the whole machine - then discover mid-disaster that they never captured the boot configuration, the drivers, or the exact partition layout. Data without the means to run it is not a recovered service.

The third is simply scope drift: a new server, a new share, a new SaaS app comes online and never gets onboarded into the backup system. Every environment change needs a scope review.

Make recovery boring

The fix for all of this is unglamorous and it works: rehearse the restore until it is routine. Schedule restore tests - monthly for critical systems, quarterly for the rest - with a named owner. Restore to an isolated location, never over production. Verify integrity, not just that files exist: open them, run a database consistency check, compare checksums and record counts. Time the restore against the RTO. Log the result, and treat a failure as a real incident.

Then write the part that recovers the business, not just the data: a disaster-recovery runbook that someone who is not you, under pressure, at three in the morning, can follow. Trigger and roles. Priority order driven by RTO. Exact commands, hostnames, where the credentials and the encryption keys live, and the dependency order - restore the database before the application server. How to validate each service. Who to tell, and how, while the network may be down.

Do all of this and "we have backups" stops being a hope you repeat at the audit and becomes a fact you have proven with your own hands. The checkmark will still be green. The difference is that now it tells the truth.

This article accompanies the free Backup, Recovery & Data Protection masterclass at AL Academy. Workshop, PDF handbook and curated resources: alouatiq.com/academy.
backupdisaster-recoverydata-protectionransomwarerestic

No comments:

Post a Comment