← All projects

Database Reliability · Completed lab

PgSentry

A PostgreSQL reliability lab that separates high availability, acknowledgement durability, routing, and historical recovery.

Why I built it

I started PgSentry with a compressed model: replicas plus automatic failover meant the database was safe. Fault injection showed that primary ownership, client reachability, acknowledged durability, and recoverable history are different properties.

Architecture

Patroni controls PostgreSQL leadership, an mTLS etcd quorum stores coordination state, and HAProxy provides one client endpoint. Terraform and libvirt create the VM topology. pgBackRest stores physical backups and archived WAL; Prometheus, Alertmanager, and Grafana observe system and semantic database state.

PostgreSQL · Patroni · etcd · HAProxy · pgBackRest · Terraform

The failures I wanted to reproduce

Network and DCS partitions test the invariant that no accepted trial observes multiple writable primaries. Async, sync, and sync-strict experiments expose the availability cost of stronger acknowledgement rules. A valid DELETE was allowed to replicate to every standby; an isolated point-in-time restore then replayed archived WAL only to a named recovery target and recovered the deleted row.

What I actually tested

  • primary and Patroni process loss through the normal HAProxy endpoint
  • DCS quorum changes and targeted network partitions
  • async, synchronous, and synchronous-strict write behavior
  • HAProxy loss and client-visible ambiguous outcomes
  • latest-state restore and point-in-time recovery after destructive data loss

Boundary and trade-off

The measurements belong to a controlled single-host lab, not an SLA or universal RPO/RTO claim. The backup repository is not geographically separate or immutable, and HAProxy and monitoring are not themselves highly available.

Related writing