Systems & databases · 22 August 2026 · 7 min read

Backups that actually restore: a disaster-recovery playbook for healthcare databases

A backup you have never restored is a hope, not a plan. For healthcare data — where both downtime and data loss carry real consequences — disaster recovery has to be something you have rehearsed, not something you assume works. Here is how to build database DR you can actually trust.

Start with RPO and RTO

Two numbers drive every other decision. RPO (Recovery Point Objective) is how much data you can afford to lose, measured in time — five minutes, an hour, a day. RTO (Recovery Time Objective) is how long you can afford to be down. A patient-facing clinical database might demand an RPO of minutes and an RTO of under an hour; an internal reporting database can tolerate far more. Agree these with the business first, because they decide how much you need to spend and which techniques apply.

Backups: full, incremental, and point-in-time

Scheduled full and incremental backups are the floor. What gets you a tight RPO is point-in-time recovery (PITR) — continuously archiving the write-ahead log (PostgreSQL WAL, MySQL binlog, or SQL Server transaction log) so you can restore to any moment, including the second before a bad change. Backups should be automated, encrypted, and stored off-site or cross-region, never only on the same host or in the same availability zone as the database.

Replication is not a backup

This is the most expensive misunderstanding in the field. A read replica or a hot standby protects you against a hardware or availability-zone failure — it does not protect you against a bad DELETE, a buggy migration, or ransomware, because those changes replicate faithfully to the standby in seconds. You need both: replication for availability, and independent, immutable backups for recovery. Treat them as separate controls solving separate problems.

Test the restore — on a schedule

The only backup that counts is one you have restored. Run game-day restores on a schedule: spin up a fresh instance from backup, restore to a target point in time, and verify the data is complete and consistent. Measure the actual RTO you achieve — teams are routinely surprised that a restore they assumed took twenty minutes takes four hours. Finding that out during a drill is cheap; finding it out during an outage is not.

Encryption, access, and retention

Backups are copies of your most sensitive data, so they deserve the same protection as production: encryption at rest, least-privilege access to the backup store, and separate credentials so a compromised app account can't also delete your backups. Set retention to match your policy and any HIPAA or contractual requirements, and consider immutable/object-lock storage so backups can't be tampered with or deleted during an incident.

Failover and runbooks

Technology aside, DR is a procedure. Write the runbook: who declares a disaster, who promotes the standby or triggers the restore, how the application is repointed, and how you verify service is healthy again. Rehearse it so the steps are muscle memory at 2 a.m., not a document nobody has opened. A tested runbook turns a crisis into a checklist.


Not sure your backups would restore? We design and test database backup and DR for SQL Server, PostgreSQL, and MySQL across cloud and on-prem. See our database work or talk to an engineer →