← All projects

Distributed Systems · Completed study

ForgeCI

A distributed CI control plane built in Go to study durable coordination across runners, servers, GitHub, and content-addressed storage.

Why I built it

I built ForgeCI after realizing that executing a shell command is the least surprising part of CI. The system has to preserve job ownership and exact source identity while workers disappear, servers restart, and external API responses are lost.

Architecture

A YAML pipeline becomes a validated DAG persisted in PostgreSQL. Ready jobs are claimed by registered remote runners. Each run refers to an immutable source snapshot in a dedicated content-addressed store; artifacts and cache use separate stores, while logs are persisted in ordered chunks.

Go · PostgreSQL · Docker · GitHub Apps · content-addressed storage

The ownership problem

Every remote operation is authorized by the exact run, job, runner, lease ID, generation, and expiration. That fence applies to source downloads, logs, artifacts, cache access, heartbeats, and completion. A stale runner may keep computing, but it cannot mutate current state. Native GitHub delivery is persisted before processing; Check creation is reconciled when an API response is lost, and newer pull-request revisions supersede older work.

What I actually tested

  • multiple PostgreSQL workers competing for ready jobs
  • stale and expired workers attempting completion and data transfer
  • exact event revision execution from an immutable snapshot
  • runner loss and conservative server-restart recovery
  • duplicate webhook delivery and GitHub Check reconciliation

Boundary and trade-off

ForgeCI marks uncertain remote work ABORTED after runner loss or server recovery. It does not automatically retry a generic shell job because the job may already have produced an external side effect. Retry policy remains deliberately out of scope.

Related writing