I told you so.

Restore drills

Backing up two servers so they can be rebuilt, and what rebuilding them found.

Most of what I run lives on two servers. One carries the web sites and most of the services. The other is a dedicated box in Germany that holds DNS and a few databases. If either one died tonight, I wanted to know how long it would take to get it back, and I wanted to have actually done it once.

So I built a tool for it, called hostrepo, and then spent two days rebuilding both servers from their backups on throwaway machines.

🔺 What gets saved

A server is more than its files. Every night, each server gathers everything that makes it itself into one directory:

  • configuration, web roots, home directories, and anything on disk that no installed package accounts for
  • a dump of every Postgres cluster, a snapshot of each Redis that keeps data, and a consistent copy of every SQLite database
  • the full list of installed packages and versions, the user accounts, the services that were running, the ports that were open, and the firewall

restic encrypts that directory and uploads it. restic deduplicates, so a night where little changed costs very little.

🔺 Where it goes

The backup lands in a US bucket. That bucket cannot be locked, because restic has to delete old data when it prunes. So the provider replicates it to a second bucket in the EU, which is locked for 30 days in a mode nobody can override, including me. Deletions do not replicate, so the EU copy only ever grows.

Each server also pulls the other one’s backup every morning. It holds that copy without the password, so to the server keeping it, the copy is noise. If one server and the US bucket were both gone, the other server still has a complete, encrypted copy.

🔺 4-2-3-1-0

The old rule for backups is 3-2-1: three copies of the data, on two kinds of media, one of them offsite. It was written for a failed disk and a building fire. The newer 3-2-1-1-0 adds two numbers: one copy that cannot be changed or deleted, and zero errors when you test a restore.

This setup meets that and goes a little further. Written the same way, it is 4-2-3-1-0:

  • 4 copies. The live server, the copy on the other server, the US bucket, and the EU bucket.
  • 2 kinds of media. Server disks and object storage.
  • 3 offsite. Both buckets, and the peer copy, which sits on another continent with another provider.
  • 1 immutable. The EU bucket, under a lock that nobody can shorten.
  • 0 errors. The restores are drilled until they pass, and checked again each quarter.

How each number is kept:

  • The nightly backup uploads to the US bucket, encrypted and deduplicated, and prunes old snapshots there.
  • The storage provider replicates every new file to the EU bucket, where it is locked for 30 days in compliance mode. Deletions are not replicated, so pruning in the US never reaches the EU copy. A weekly check confirms every US file is present in the EU with the same checksum, and that each server’s newest snapshot is locked.
  • Each morning, each server pulls the other’s backup from the US bucket. The pull only ever adds files, is checked against the bucket’s checksums without downloading anything twice, and cannot be read by the server holding it.
  • The drills cover the zero. A monthly check reads back a different twelfth of the stored data, so every byte is read once a year.

The extra copy and the drills are the parts that matter most now. A locked copy is what survives ransomware, which deletes backups first. A tested restore is what catches silent problems, like a database dump that loads nothing.

The count leaves one question open: whether a single stolen key can reach more than one copy. If one key can delete three of the four, the real count is closer to two. Copy math looks the same either way. That is decided by how narrowly each key is scoped.

🔺 Bit rot

Copies can rot. A disk returns a flipped bit, a file is cut short, and nobody notices until the day it is needed. restic helps here more than it seems to: every file in a restic repository, apart from one small config file, is named after the SHA-256 of its own contents. Hashing a file and comparing the result with its name is a complete check, with no other record to consult.

On top of that, every file gets a par2 parity set at 10 percent. The sets for the bucket are kept in the bucket, and each peer copy keeps its own on its own disk. A damaged file is repaired from its parity, or fetched again from another copy, and a repair only counts when the file hashes to its name again. That means a repaired file is always the exact original. The damaged copy is kept aside as evidence.

The checks run on their own schedule:

  • Each night the backup writes parity for the day’s new files, reading each one back from the bucket first. That also catches damage on the day it happens.
  • If the repository check fails, the backup repairs the files it names and checks again.
  • Each peer copy hashes every file every night, and checks a seventh of its parity sets, so every set is looked at once a week. Parity can rot too, and par2 on its own does not notice when it has.
  • A restore repairs anything damaged before it reads it.

A full drill downloads everything, so it runs quarterly. Every night there is also a lite drill that downloads nothing. The storage provider keeps a SHA-1 for every file and returns it in a directory listing, and I record each file’s SHA-1 once its bytes have been proven good. The lite drill compares the two for every file in the bucket, and hashes every file of the copy on the other server. Each server’s backups are drilled this way every night by the other server. Anything this tooling uploads carries both hashes: the SHA-1, which the provider checks against the bytes it receives, and the SHA-256, stored with the file.

I tested it by overwriting a few bytes in the middle of one file in a live peer copy. The next check found it, rebuilt it from parity, and set the damaged version aside.

Building this found another problem. One server’s nightly copy of the other’s backup had failed that morning. An earlier fix had copied some files in the bucket onto themselves, which changed their timestamps but not a byte, and the copy tool refused to continue because the dates no longer matched. It now compares content instead of dates.

🔺 What a drill is

A drill is one command. It rents a fresh virtual machine, restores the newest backup onto it, and compares the result with the live server: which services run, how many rows every database table holds, and what every hosted site returns. Then it deletes the machine.

The copy is isolated. It must not join the private network, send mail, post to the fediverse, or run its own backups. A second copy of a server doing those things would fight the real one. The drill masks those services before anything starts.

The restore itself goes in a fixed order. User accounts come first, with their original ids, then every package, then the files, then ownership, then the databases, then the services.

🔺 What the drills found

restic check said the backups were fine every time. The first restores were not fine.

Installing packages created their system accounts with new ids. The restored files still carried the old numeric owners, so services could not read their own data. Accounts are now created before any package installs.

The databases restored nothing and reported no errors. The dumps were readable only by root, so Postgres could not load them. The check counted errors, found none, and passed. It now counts databases.

Services came back on the wrong settings. Installing a package starts its service with the package defaults, and the restore only started services that were stopped, so those kept the defaults. It now restarts everything that was running on the old server.

One drill hung for three and a half hours, waiting for Redis to stop. Ubuntu’s Redis service has no stop timeout. The drill was meant to take twenty minutes. Every service action in the restore now gives up after two minutes, and the restore as a whole has a limit too.

An account id collided. On the German server, Redis runs as user 100. On a fresh Ubuntu install, 100 is the system logger. The restore handled that for files, but the log directories came back owned by the logger, and Redis could not write its log. The next drill proved the fix.

DNS would not start on a copy. The German server is the hidden primary for my zones: two public nameservers answer the world and pull their data from it. Its DNS server was told to listen on the server’s public address, and it refuses to start if any listen address is missing, which it always is on a copy. One setting tells it to skip missing addresses. A restored copy now comes up with every zone loaded, ready to take over as primary.

A drill also found a bug on the live server. A duplicate entry in the web server’s config meant its next restart would fail and take down every site on it. Nothing had restarted it yet. It is fixed.

🔺 Where it stands

The main server was drilled again today, under the stricter rule. The first run failed, for two reasons. fail2ban would not start on the copy: the restore brought back the log directories but not the log file one of its jails watches, and fail2ban refuses to run without it. On a real replacement that would have left the server without its brute-force protection. The restore now recreates every log file the old server had, empty, with its owner and mode. The drill also fetched the live sites through public DNS, so a name behind a CDN, or one served by the other server, was compared with something else entirely. Both sides are now fetched directly from their own server. The rerun passed: 66 of 66 services, 15 of 15 databases identical, and 99 of 100 sites identical, the one difference expected and explained.

The German server took three drills today. The first failed: Redis and DNS did not start, and three sites did not match. The second was cut short: the drill script was edited while the drill was running it, and bash reads a script from disk as it goes, so the run fell over halfway. Drills now run from a frozen copy. The third passed. Every database matched, and every service and site either matched or had a written reason not to.

Those reasons are the other change. The drill used to tolerate a couple of differing sites and simply list services that were not running. Now each server’s configuration names every difference a drill should expect, with the reason next to it: a RAID monitor has no RAID to watch on a virtual machine, and a site that proxies over the private network cannot answer while the drill keeps that network down. Anything without a reason fails the drill.

Each drill downloads the whole backup, a few gigabytes, so I run them quarterly and after big changes. Between drills, a monthly check reads a different twelfth of the stored data.

Every one of those restore bugs passed the backup check.

🔺 Live status

The results are public at https://backups.itys.net. Every backup, copy, lite drill, full drill and replica check leaves a record, and the page is rebuilt from them every hour. Each backup is a dot: green for completed with no further action, blue for completed with a repair, yellow for succeeded with warnings, red for failed. The servers appear as Server A and Server B.

🔺 The code

hostrepo is published under the BSD 2-Clause license. It is bash, restic, rsync, rclone and par2, written for Ubuntu 26.04 hosts, with B2 for storage and Linode for drills. docs/OPERATING.md covers setup, daily use, restoring by hand, a real replacement, drills and the pass rule, bit rot detection and repair, the lite check, and the status page. hosts/example.conf documents every per-host setting.