Platform Reliability · Completed lab
Aegis
A GitOps platform organized around proving controls under failure: build, secure, observe, break, recover, and record evidence.
Why I built it
Aegis stopped being a Kubernetes configuration exercise when trusted components disagreed with reality. The useful work became the evidence: which control blocked a fault, which one missed it, and how the architecture changed afterward.
Architecture
Flux owns desired state across disposable kind and persistent single-node K3s environments. The delivery path scans images, attaches an SBOM, signs with keyless Cosign, and verifies identity through Kyverno. Cilium policies, SOPS-encrypted secrets, Prometheus SLOs, Authentik identity, and PostgreSQL recovery cover separate boundaries.
Kubernetes · Flux · Kyverno · Cilium · Cosign · Prometheus
Controls I tested by breaking things
Unsigned and wrong-signer images produced distinct admission failures. A signed and admitted release contained a real algorithmic latency regression; supply-chain controls passed, while an SLO detected behavior and recovery proceeded through Git. Destructive PostgreSQL and Authentik tests proved that Flux rebuilds objects but backups restore rows. A Cilium Gateway datapath failure survived configuration checks, so live packet evidence led the home environment to use a small Flux-owned nginx proxy instead.
What I actually tested
- unsigned and wrong-identity image admission
- network-policy lateral-movement attempts
- SLO detection and Git rollback of a signed bad release
- destructive PostgreSQL and Authentik recovery
- host reboot, replacement-host reconstruction, and datapath control experiments
Boundary and trade-off
The evidence is bounded to development environments. Home K3s is single-node; monitoring is not HA; backups lack PITR; rollback is deliberate rather than automatic; and recovery keys remain concentrated on one workstation.