That was not enough. The question I wanted the project to answer was narrower: when I introduce the failure a control is supposed to stop, does the outcome actually change? That turned the work into a repeating loop: build, secure, observe, break, recover, and keep the command output that proves what happened.
Admission produced two different failures
The release path builds an owned Go service, scans the image, attaches an SBOM, signs through GitHub Actions OIDC, and records an immutable digest. Kyverno verifies the expected signing identity before admission.
I tested an actually unsigned image rather than reading the policy and declaring success. Admission rejected it because no signature existed. Then I signed an image from an unrelated workflow. That image was not unsigned; it had the wrong identity. Kyverno rejected it with a subject mismatch. The distinction matters because presence of a signature and trust in its signer are separate properties.
This also separated responsibilities in the pipeline. Flux image policy selects what Git records as desired. Kyverno independently decides whether Kubernetes may admit that digest. Neither one establishes that the application behaves correctly after it starts.
The trusted release that was still bad
The clearest failure was a release containing an intentionally inefficient Fibonacci implementation. The image came from trusted CI, passed scanning, carried its SBOM and expected signature, and was admitted. Kubernetes considered its pods healthy. The application returned successful responses.
It was still a bad release because latency crossed the service objective under the experiment load. Prometheus caught behavior that provenance and availability checks were not designed to judge. Recovery went through Git: stop image automation, restore the known-good digest, let Flux reconcile, fix the implementation, and add a regression test before resuming promotion.
My earlier mental model treated supply-chain gates as a long funnel ending in a safe workload. The experiment replaced that with independent questions:
Was this artifact built by the identity I trust?
May this artifact enter the cluster?
Is the admitted application behaving acceptably?
Can I recover if the answer changes?
A passing answer at one boundary does not imply a passing answer at the next.
Network policy needed traffic evidence
For lateral movement, I ran probes from an unrelated namespace toward Authentik's PostgreSQL and application ports. I captured the allowed flows before enforcing the policy, then repeated the probes and observed denied traffic while legitimate peer relationships continued to work.
That before-and-after pattern prevented a weak test. A timeout alone is ambiguous: the service might already be broken, DNS could be wrong, or the probe might target the wrong port. Establishing that the same path worked before enforcement made the later denial useful evidence.
Git rebuilt the database perfectly—and the row was gone
GitOps is good at reconstructing declared objects. It can restore a StatefulSet, Service, PVC declaration, policy, and encrypted configuration. It cannot restore mutable rows that were never committed to Git.
I deleted the storage behind a PostgreSQL workload and let Flux do exactly what it promises. It reconstructed a working but empty database. Only an encrypted off-host database backup recovered the original rows. I repeated the distinction with Authentik: after destroying its PostgreSQL volume, the platform and schema returned, but an API-created test identity was absent. Restoring the database brought back that same identity and its properties.
The useful result was not “backup script succeeded.” The restore happened after deliberate destruction, and I checked for absence before restoring. Without that negative check, I could accidentally prove that the data never disappeared.
The architecture had to diverge
I initially wanted the disposable kind cluster and persistent home K3s host to use the same Cilium Gateway ingress path. On the home host, every request through that path returned an error even though the configuration looked healthy.
Packet traces showed the backend responding to connections tagged with Cilium's ingress identity, but the handshake did not complete. I ran an ordinary nginx proxy against the same application paths as a control. That path completed the requests, which narrowed the failure to the Gateway datapath rather than the application or general host networking.
The home environment now uses a small Flux-owned nginx proxy while the disposable environment retains Cilium Gateway. I prefer one architecture when the operational assumptions match. Here, forcing symmetry would have preserved a cleaner diagram and a broken data path.
What I actually tested
Aegis records live rejection of unsigned and wrong-signer images, lateral-movement probes before and after policy, detection and Git rollback of a signed bad release, destructive PostgreSQL and Authentik recovery, host restart and replacement-host reconstruction, and ingress control experiments. The repository keeps the runbooks, commands, and captured results alongside the configuration.
The tests are bounded to two development environments: a disposable local cluster and a single-node home server. The monitoring stack is not highly available. Backups do not provide point-in-time recovery. Rollback remains a deliberate operator action, and important recovery keys are concentrated on one workstation. Those are constraints, not footnotes.
What changed for me
I now read an infrastructure claim as a proposed experiment. “Images are verified” suggests an unsigned image and a wrong signer. “Network isolation works” suggests a known-good connection followed by the same probe under policy. “GitOps recovers the platform” suggests deleting both declared objects and undeclared data, then observing what each recovery source can actually reconstruct.
The tools still matter, but their names are not the result. The result is a changed failure outcome and enough evidence to explain why it changed.
Related project: Aegis case study →