← Back to Writing

High Availability Is Not Recovery

When I started PgSentry, I had a compressed idea of database reliability: run PostgreSQL replicas, add automatic failover, and the database is safe.

That idea was not completely wrong. It was hiding several different problems behind one word.

A replica can take over when the primary machine disappears. It cannot tell whether a valid SQL statement was a human mistake. It cannot recover an earlier version of a row after the same deletion reaches every node. It cannot make an ambiguous client response unambiguous. Even stronger commit durability has an availability cost that appears exactly when the cluster is unhealthy.

PgSentry became my attempt to separate those problems and make each one fail in front of me.

The first milestone looked reassuring

The lab runs real PostgreSQL processes across a reproducible VM topology. Patroni controls PostgreSQL leadership, an etcd quorum stores coordination state, and HAProxy gives clients one stable write endpoint. Under normal operation there is one writable primary and two read-only streaming replicas.

That architecture initially felt like the finished idea. If the primary stopped, Patroni could promote an eligible replica and HAProxy could route new connections to it. I could stop thinking about a particular database hostname and instead connect through one address.

Then I began injecting failures.

The workload always went through HAProxy. It did not cheat by discovering the new leader and connecting directly. Each attempted write received a run ID, a sequence number, timing information, and an outcome classification. After each fault, the harness restored the stopped service, VM, or firewall rule and required the entire cluster to become healthy before the next experiment.

This made the difference between an architecture diagram and an operational claim obvious. “Patroni supports failover” was no longer enough. I wanted to know what my client observed while failover happened.

A failed connection did not tell me what happened to the transaction

The most useful classification in the failure harness was not success or failure. It was:

ambiguous outcome

A connection can disappear after PostgreSQL commits but before the client receives confirmation. From the client's perspective, the request failed to return successfully. From the database's perspective, the row may already exist.

I had treated connection errors as if they were transaction results. Testing primary loss forced me to stop doing that.

The application now has to decide how to retry. If it blindly repeats a non-idempotent operation, it may apply the business action twice. If it assumes every lost response means the transaction committed, it may invent state that does not exist.

High availability moved the write endpoint to another node. It did not recover the missing knowledge at the client boundary.

The partition tests made ownership more important than reachability

Stopping a database process is the clean failure. A network partition is harder because different parts of the system continue running with incomplete views.

PgSentry isolates database nodes from etcd, partitions replication traffic, removes DCS quorum, and tests etcd peer loss. The invariant I cared about was not that every node remained reachable. It was that the accepted trials never observed more than one writable PostgreSQL primary.

That changed how I viewed etcd. It was not an extra service placed beside PostgreSQL for the sake of an HA stack. Its quorum determines whether the coordination layer can safely change leadership. A minority that cannot commit is inconvenient, but letting disconnected sides independently claim ownership would be worse.

The same experiments exposed another boundary. HAProxy gives clients stable routing, but HAProxy itself is a failure domain in this lab. PostgreSQL can be healthy and correctly coordinated while the client endpoint is unavailable. Database health, leader ownership, and client reachability are three separate signals.

Synchronous replication did not give me a free safety switch

I then compared three explicit durability policies: asynchronous, synchronous, and synchronous-strict operation.

In asynchronous mode, the primary can acknowledge after its own durable work without waiting for another PostgreSQL node to flush the transaction. That favors write availability, but it leaves a window in which acknowledged WAL has not reached a standby.

In synchronous mode, Patroni selects a synchronous standby and an acknowledged commit waits for the required WAL flush. This narrows the durability exposure, but the behavior still depends on which standby is eligible and which policy is active at commit time.

Synchronous-strict mode made the trade-off impossible to ignore. When no eligible synchronous standby remained, bounded writes stayed blocked until a standby returned. The system refused to silently fall back to ordinary asynchronous acknowledgements.

higher write availability
        ↕
stronger acknowledgement durability

Before running that experiment, “turn on synchronous replication” sounded like a straightforward improvement. Afterward, I understood it as a decision about which failure behavior is acceptable. Waiting is safer for one kind of loss and worse for availability. There is no configuration that removes the trade-off.

The accepted finite trials did not lose an acknowledged row, but that observation is not a universal RPO-zero guarantee. Simultaneous storage loss, arbitrary partitions, operator mistakes, and failures outside the lab remain outside the claim.

The DELETE that every replica handled correctly

The experiment that changed my thinking most was not a node failure. It was a successful SQL statement.

I created a marker row, established a named recovery point, wrote another marker after that point, and then deleted the protected row through the normal HAProxy endpoint. Streaming replication propagated the deletion to both replicas.

primary:   row deleted
replica 1: row deleted
replica 2: row deleted

Nothing in the HA system malfunctioned. PostgreSQL generated WAL for the deletion. The replicas replayed it. Patroni still saw a healthy cluster. HAProxy still had a valid primary.

The latest state was precisely the state I did not want.

This is where my original idea of replicas as “backups that are ready to take over” finally broke. A replica protects a current database lineage against certain machine failures. It is designed to reproduce the primary's accepted history, including destructive history.

Recovery needed a different source of truth

PgSentry uses pgBackRest for a physical base backup and continuous WAL archiving. The base backup provides a consistent starting point. Archived WAL extends the recoverable history beyond the instant that backup was taken.

For point-in-time recovery, I restored into an isolated PostgreSQL instance on a separate port. It was deliberately not a Patroni member and not an HAProxy backend. PostgreSQL replayed archived WAL only through the named recovery target.

The verification checked both sides of the boundary:

  • The row deleted from the live primary and both replicas existed again.
  • The marker written after the recovery point did not exist.

That second check mattered. Finding the old row would not prove that replay stopped at the intended history. Recovery is not merely starting PostgreSQL from copied files; it is selecting and verifying a specific database state.

I also tested latest-state recovery. It reconstructed data written after the full backup, which proved that archived WAL—not only the base backup—was necessary. A backup inventory that looks healthy is insufficient if required WAL segments never reached the archive.

Why I restored away from the cluster

Restoring over the only live cluster would have mixed investigation with destruction. If I selected the wrong target or discovered a missing archive segment halfway through, I would have damaged the environment I was trying to recover.

The isolated restore let me inspect the recovered state before deciding what should happen next. It also forced a useful identity distinction: the restored instance shared the PostgreSQL system lineage, but it was not automatically a Patroni member and did not inherit the live routing role.

choose target
restore base backup
replay required WAL
verify expected rows
verify excluded rows
only then plan reintegration

Monitoring had to describe the failure class

By this point, a single “database down” alert no longer seemed useful enough. A cluster can have no primary, degraded replication, lost etcd quorum, unavailable HAProxy routing, a failed WAL archive command, or a stale backup. Those conditions require different responses.

PgSentry therefore checks infrastructure metrics and semantic database state, with runbooks attached to distinct alerts. I exercised alerts through real faults and verified that they resolved after recovery. The goal was not a dashboard full of green panels. It was an alert that narrowed the operator's next question.

What I would not claim from the lab

PgSentry is a controlled, single-host laboratory. Its measurements are not a production SLA, and the accepted experiments do not establish guaranteed RTO or universal RPO. The backup repository is not geographically separate, immutable, or ransomware-resistant. HAProxy and the monitoring stack are not highly available.

Those limitations do not weaken the lesson. They define where the evidence stops.

What actually changed for me

I no longer use “database reliability” as if it describes one mechanism.

replication keeps another current copy
consensus protects leadership decisions
failover selects a writable survivor
routing gives clients a stable endpoint
durability policy defines acknowledgement behavior
backup and WAL preserve recoverable history
PITR selects an earlier state
monitoring identifies the current failure

Each mechanism owns a different failure class. Sometimes they even pull in opposite directions: stricter acknowledgement durability can reduce write availability, and perfect replication can spread a destructive mistake faster.

The assumption I started with was that enough replicas made the database safe. The assumption I have now is narrower and more operational: before calling a database reliable, I need to name the failure, identify which mechanism addresses it, run the failure, and verify the recovery state.

High availability matters. It just is not recovery.

Related project: PgSentry case study →