← Back to Writing

A CI System Is a Lease Protocol with Shell Commands Attached

Executing a shell command was never the difficult part of ForgeCI. The difficult part was deciding which machine still had the right to say what happened.

A remote runner can claim a job, lose contact with the server, and keep executing. The server may conclude that the runner is gone. If the old runner later returns, accepting its result can overwrite a newer decision or attach output to work it no longer owns.

A timeout does not revoke computation

Expiring a row in PostgreSQL changes control-plane state. It cannot stop instructions already running elsewhere. That gap is why a lease needs more than an expiration timestamp.

Authentication is not ownership

ForgeCI identifies ownership with the run, job, runner, lease ID, generation, and expiration. Source download, artifact transfer, log append, heartbeat, and completion all require the exact current tuple. A bearer token authenticates a runner, but does not prove that runner owns this job now.

authenticated runner
+ current lease identity
+ unexpired generation
= permission to mutate job state

Fencing belongs on every side effect

Checking ownership only at completion leaves holes. A stale worker could still append misleading logs, upload an artifact later consumed by another job, or fetch source after cancellation. ForgeCI applies the same ownership test across every runner route.

Each lease gets an isolated workspace keyed by run, job, and lease. Before execution, the runner streams the immutable source archive to a temporary file, verifies the blob, extracts it under bounded rules, recomputes the logical manifest, and compares that digest with the leased source identity. Editing the original checkout after submission cannot change the run.

The boundary is PostgreSQL

The scheduler does not rely on two workers politely avoiding one another. Claiming work, validating the lease, and changing terminal state meet at database transactions with conditional updates. That made the database more than storage: it is the serialization point the HTTP handlers and runners must share. If a condition is missing from one mutation, the protocol is incomplete even when the happy path passes.

The test I care about

I wanted to know whether the ownership tuple was doing real work, so I leased a job and tried to complete it three ways. First I kept the runner and generation but substituted a different lease ID. PostgreSQL rejected the completion. Then I used the real lease ID with the wrong generation. That was rejected too. Only the exact current runner, lease ID, and generation could move the job to PASSED.

current lease: runner A · lease L · generation 1

runner A · wrong lease · generation 1  → rejected
runner A · lease L    · generation 999 → rejected
runner A · lease L    · generation 1   → accepted

The runner-protocol tests repeat that boundary for heartbeat, source, logs, artifacts, cache, and completion. The expired-job test is deliberately more conservative than automatic reassignment: it moves uncertain work to ABORTED. That is the behavior the current repository proves.

Immutable source is part of execution correctness

A valid owner can still execute the wrong tree if submission points at a mutable checkout. ForgeCI records Git revision identity separately from its canonical source digest, publishes a deterministic snapshot into source CAS, and makes the runner verify both the compressed blob and extracted manifest.

Ownership answers who may mutate the result. The snapshot answers what source that owner is allowed to execute. The completion is credible only when both identities still match.

Recovery refuses to guess

There is an uncomfortable limitation in the current design: if the server restarts with a remotely running job, ForgeCI marks that uncertain work ABORTED. It does not automatically reassign it.

That is conservative because a generic shell step may have performed an external side effect. Re-executing it could be worse than stopping. Retry semantics need an explicit policy; they cannot be inferred from a missing heartbeat.

SCM delivery work is different. Its database lease can expire and be reclaimed because processing is designed around durable request identity and idempotent boundaries. GitHub Check delivery is also reconciled later if an API response is lost.

What I deliberately do not retry

ForgeCI reclaims SCM delivery work because that path has durable identity and bounded retry semantics. It does not automatically reassign an uncertain shell job after runner loss. The old attempt may have changed something outside ForgeCI; executing it again would be a guess.

What changed

I started with a picture of CI as parsing YAML and running commands in dependency order. Remote runners changed the center of the design. The real system is an ownership protocol wrapped around execution: immutable inputs define what should run, fenced leases define who may report it, and durable reconciliation repairs external status.

The shell commands are still there. They are simply the least surprising part.

Related project: ForgeCI case study →