// failure lab / operating evidence

Reliability is learned
after the diagram.

Incidents and near misses from systems I operate, reconstructed without heroics: what users saw, which layer actually failed, what made recovery safe, and which control changed afterward.

METHOD

Evidence before action

FOCUS

Failure propagation

RECOVERY

Safety before speed

OUTPUT

A stronger platform

Failure Lab 002Incident

The proxy was healthy. The service had disappeared.

A Kubernetes node loss stranded Gitea’s RWO volume, blocked its replacement pod, emptied its Service endpoints, and surfaced externally as a Caddy 503.

HTTP 5030 ready endpoints14h stuck terminationRecovered to HTTP 200
Replay the failure →
Failure Lab 001Near miss

The change worked. The process was still risky.

A remote Authentik change over Tailscale exposed the difference between reaching a control plane and having an independent path to recover it.

Identity control planeRemote accessUnproven rollbackOperating gate added
Replay the failure →