Backups and standby site

Standby site: three schemes and what to check every quarter

Contents

For managers and administrators whose main servers are in Ukraine or in the office, or who rely on a single server in Poland. You will choose one of three standby site schemes and get a quarterly checklist.

Why a second site#

A backup answers the question “will the data survive”. A standby site answers “where and when will we be working again”. RAID, redundant power and the data centre’s network links protect against a failed disk or line, but not against losing the whole site. We know this first-hand: the United DC data centre in Ukraine was destroyed by a strike on 23 September 2026 (our story). That is why copies and the standby server belong in another country.

No scheme gives absolute guarantees. It reduces data loss and downtime to a level you have chosen and tested in practice.

Three schemes#

First the manager answers two questions: how much recent data the company is prepared to re-enter, and how long it can go without its accounting system.

SchemeWhat it needsHow much data you may loseHow long recovery takes
“Cold”: only copies in another countryEncrypted scheduled copies, a tested restore, a written planEverything changed since the last copyThe longest: get a server, install the system, restore the data from copies
“Warm”: a standby server with regular restores or replicationA second server in another country that receives restored copies on a schedule or asynchronous replication; manual switchoverEverything since the last restore; with replication, usually only the latest transactionsNoticeably shorter: the server is already running; you decide, switch DNS and check the result
“Hot”: continuous replication and fast switchoverThe same, plus continuous replication of the database and files, monitoring, a rehearsed switchover; for automatic switchover, a third witness nodeThe least: in synchronous mode confirmed transactions are already on both serversThe shortest, but the hardest to set up and maintain

Actual figures depend on your scheme: the backup schedule, database size, the link and how well the switchover has been rehearsed. The only way to learn them is a timed switchover drill. For the “cold” scheme add server delivery time: within 12 or within 72 hours of payment for servers in stock — the catalogue shows which for each configuration.

The “warm” scheme with a manual switchover usually suits a small company best: a person decides, not automation that a brief loss of connectivity can mislead.

What you will need#

  • A second server in another country. If the main servers are in Ukraine or in the office, take a server in Poland: the response time from Kyiv is about 15 ms. If the primary is already in Poland, take the standby in Germany or France: a separate data centre with separate power and network links; from Kyiv about 35 and 40 ms respectively. The figures are our measurements from Ukrainian providers’ networks, October 2026; they depend on the provider and route and are higher on mobile internet.
  • A private network between countries on the main lines: replication runs over it, so the database port need not be opened to the internet. Support connects it on request; the setup is in the article on the private network and VLAN.
  • A VPN between the office and both servers, set up beforehand: WireGuard between the office and the server.

Software licences for the standby server are not covered here.

Steps#

1. Names and DNS with a low TTL#

Users should connect by name, for example app.example.com, not by IP address: then a switchover is a change to one A record. Set its TTL to 300 seconds beforehand. Lowering the TTL during an outage is too late: the old value is already cached. DNS hosting must not depend on either server. Check the record:

dig +noall +answer app.example.com

The second field is the remaining TTL in seconds; it must not exceed the value you set. On Windows Resolve-DnsName app.example.com shows the same.

2. Database replication#

The example is PostgreSQL 18 streaming replication on Debian or Ubuntu (PGDG packages, cluster 18/main); on other systems paths and service names differ. The addresses 10.10.0.1 (primary) and 10.10.0.2 (standby) are example private-network addresses. Install the same major version on the standby beforehand (the postgresql-18 package). On the primary, create a role for replication:

sudo -u postgres createuser --replication --pwprompt replicator

In /etc/postgresql/18/main/postgresql.conf set the private address and cap the WAL a replication slot may retain: without a cap the primary’s disk fills up if the standby is unreachable for long. 20 GB is an example; past the limit, the replica has to be rebuilt.

listen_addresses = 'localhost,10.10.0.1'
max_slot_wal_keep_size = 20GB

Add this line to /etc/postgresql/18/main/pg_hba.conf:

host  replication  replicator  10.10.0.2/32  scram-sha-256

Warning. Restarting PostgreSQL drops every user connection — do it outside working hours. Do not open port 5432 to the internet: replication should run over the private network or a VPN, and the firewall should allow this port only from the standby’s address.

sudo systemctl restart postgresql@18-main

pg_basebackup copies only the data directory; the standby keeps its own files in /etc/postgresql/18/main. Before starting the replica, carry the primary’s settings over to them: max_connections must not be lower than on the primary, or the replica will not work.

Warning. Run the commands below only on the standby server: they move its current data directory aside (renamed, not deleted). To roll back, stop the service, remove the new main directory, move main.old back and start the service.

sudo systemctl stop postgresql@18-main
sudo mv /var/lib/postgresql/18/main /var/lib/postgresql/18/main.old
sudo -u postgres pg_basebackup -h 10.10.0.1 -U replicator -D /var/lib/postgresql/18/main -R -X stream -C -S standby1 -P
sudo systemctl start postgresql@18-main

-R creates the standby.signal file and writes the connection settings, password included, to postgresql.auto.conf; -C -S standby1 create a replication slot. Replication is asynchronous: the standby lags slightly, so the latest transactions can be lost. Synchronous mode (synchronous_standby_names) removes that loss, but every write waits for confirmation from the other country, and without the standby, writes on the primary hang. See the PostgreSQL documentation.

MS SQL Server has two built-in mechanisms. Log shipping: SQL Server Agent jobs back up the transaction log, copy the backup to the standby server and restore it there; there is no automatic switchover, so this is a “warm” scheme. Always On availability groups: the log is sent continuously, synchronously or asynchronously; on Windows they need a failover cluster (WSFC), and Standard edition offers only basic groups — one database and two replicas. A file-based BAS database cannot be synchronised reliably while users are working in it: move it with scheduled copies, as in the article on database backups.

3. File synchronisation#

Warning. The --delete option removes from the standby server everything that is not on the primary. Check the paths and trailing slashes, and run the command with --dry-run first: together with -v it only shows the changes.

rsync -av --delete --dry-run /srv/files/ admin@10.10.0.2:/srv/files/

If the list of changes is right, remove --dry-run and run the command on a schedule. On Windows robocopy with the /MIR switch does the same — with the same caveat.

4. Action plan and access list#

  • Who decides to switch over, and who stands in for that person.
  • The steps in order: isolate the old primary, bring the standby database into service, change DNS, check sign-in, notify users.
  • The access list: servers, DNS panel and domain registrar, VPN, databases, backup storage and encryption keys, support contacts.
  • The plan is kept in a password manager and on paper, not only on a server that may be gone.

5. A switchover drill#

Warning. After pg_promote() the standby becomes an independent server and cannot be turned back into a replica with a single command. From then on the old primary must not accept writes, or the two databases will diverge. To go back, make the old primary a replica again (pg_basebackup or pg_rewind) and switch over once more.

  1. Pick a time outside working hours, warn the users, take a fresh backup.
  2. Stop the applications on the primary, make sure the replica is not lagging, and stop PostgreSQL on it: sudo systemctl stop postgresql@18-main.
  3. On the standby, bring the database into service with the command under this list.
  4. Change the A record to the standby’s address and wait for the TTL to pass.
  5. Ask users to sign in and do their usual work. Note the time from the start to “we are working”.
  6. Return to the primary and note what went wrong.
sudo -u postgres psql -c "SELECT pg_promote();"

How to check the result#

On the primary, check replication and the slot:

sudo -u postgres psql -x -c "SELECT client_addr, state, sent_lsn, replay_lsn, replay_lag FROM pg_stat_replication;"
sudo -u postgres psql -c "SELECT slot_name, active, wal_status FROM pg_replication_slots;"

Expected: state is streaming, active is t, wal_status is reserved or extended. On the standby:

sudo -u postgres psql -c "SELECT pg_is_in_recovery(), now() - pg_last_xact_replay_timestamp() AS delay;"

The first field is t, the second a small value. If nobody is working on the primary, delay grows — that is normal.

What to check once a quarter#

  • Restore from a copy: restore the latest copy on the standby or into a test database and ask the accountant to find the most recent documents.
  • A switchover drill by the plan, with the actual time written down; compare it with the previous quarter.
  • Replication: state, lag, an active slot, free disk space on both servers.
  • DNS: the TTL has not changed, at least two people have access to the DNS panel and the registrar, the domain is paid for.
  • Network: the VPN from the office and from home to the standby works, firewall rules match on both servers.
  • Versions: the OS, the database server and the accounting platform are updated to the same versions on both servers; TLS certificates have not expired.
  • Access and the plan: the keys and passwords on the list work; access for people who have left is closed; phone numbers, responsible people and steps in the plan are up to date; invoices for both servers are paid.

Common mistakes#

  • The standby server is in the same data centre, and the copies are on the same server.
  • Replication or file synchronisation is treated as a backup. An accidental deletion, or encryption of data by a virus, reaches the standby within seconds too, so every scheme needs copies with history: the 3-2-1 rule.
  • The standby server is less protected than the primary: RDP or the database port is open to the whole internet. See secure RDP and basic Linux server hardening.

What next#