go-docker-orphan-reconciliation
Root cause: sandbox containers are long-lived "sleep-infinity" docker containers. When the server dies, nothing removes them, and since there was no startup reconciliation, the next boot leaks them forever. There was also no way to even enumerate them (no label-filtered list on DockerClient).
The fix has three parts, in a 3-package module (uhlp.dev/sandbox, Go 1.25, docker SDK v28.5.2):
dock — ListContainers(label-filter) on the DockerClient interfaceThe interface is extended with a label-filtered list, backed by the docker SDK's ContainerList + filters.Arg label key, and a decoupled ContainerSummary{ID,Labels} type so the manager never touches SDK types:
// dock/dock.go
type ContainerSummary struct {
ID string
Labels map[string]string
}
type DockerClient interface {
// ListContainers returns running containers that carry the given label
// key (presence filter). Exited containers are excluded.
ListContainers(ctx context.Context, labelKey string) ([]ContainerSummary, error)
RemoveContainer(ctx context.Context, id string) error
}
func (c *Client) ListContainers(ctx context.Context, labelKey string) ([]ContainerSummary, error) {
if c == nil || c.api == nil {
return nil, ErrNoDocker
}
list, err := c.api.ContainerList(ctx, container.ListOptions{
All: false, // only running containers
Filters: filters.NewArgs(filters.Arg("label", labelKey)),
})
if err != nil {
return nil, fmt.Errorf("dock: list containers with label %q: %w", labelKey, err)
}
summaries := make([]ContainerSummary, 0, len(list))
for _, ctr := range list {
summaries = append(summaries, ContainerSummary{ID: ctr.ID, Labels: ctr.Labels})
}
return summaries, nil
}
sandbox — Manager.ReconcileOrphansThe core of the fix (sandbox/manager.go). Every container the manager creates carries uhlp.sandbox.pool=<poolID>; orphans are containers carrying that label whose pool is no longer live:
const PoolLabel = "uhlp.sandbox.pool"
func (m *Manager) ReconcileOrphans(ctx context.Context) error {
if m.docker == nil {
return nil // nil-docker no-op: docker disabled, nothing to reconcile
}
// RLock for the ENTIRE pass: while held, GetOrCreatePool (write lock)
// blocks, so a pool cannot be created between the orphan list and the
// removals — a container can never be deleted from a pool that comes
// to life mid-reconciliation.
m.mu.RLock()
defer m.mu.RUnlock()
containers, err := m.docker.ListContainers(ctx, m.labelKey)
if err != nil {
return fmt.Errorf("sandbox: list orphan candidates: %w", err)
}
var errs []error
for _, ctr := range containers {
if ctxErr := ctx.Err(); ctxErr != nil {
errs = append(errs, ctxErr)
break
}
poolID := ctr.Labels[m.labelKey]
if poolID == "" {
continue // not ours
}
if _, live := m.pools[poolID]; live {
continue // skip containers owned by live pools
}
if err := m.docker.RemoveContainer(ctx, ctr.ID); err != nil {
errs = append(errs, fmt.Errorf("sandbox: remove orphan %s: %w", ctr.ID, err)) // best-effort
continue
}
slog.Info("reconcile: removed orphaned sandbox container", "id", ctr.ID, "pool", poolID)
}
return errors.Join(errs...) // failures joined, never aborts the rest
}
GetOrCreatePool takes the write lock (unchanged semantics, same mutex), which is what makes the RLock-during-removal race-proof:
func (m *Manager) GetOrCreatePool(ctx context.Context, id string) (*Pool, error) {
m.mu.Lock()
defer m.mu.Unlock()
if p, ok := m.pools[id]; ok { return p, nil }
p := &Pool{ID: id}
m.pools[id] = p
return p, nil
}
app — composition root, WARN-continueWired into framework startup so reconciliation runs before serving but never blocks or fails boot:
func Compose(ctx context.Context, cfg Config) (*sandbox.Manager, error) {
var dk dock.DockerClient
if cfg.DockerEnabled {
c, err := dock.NewClientFromEnv()
if err != nil { return nil, err }
dk = c
}
mgr := sandbox.NewManager(dk)
if err := mgr.ReconcileOrphans(ctx); err != nil {
slog.Warn("startup orphan reconciliation failed; continuing", "error", err) // WARN-continue
}
return mgr, nil
}
sandbox/manager_test.go)The mock records every RemoveContainer call, supports per-container error scripts, and exposes a blockable removal hook for the lock-ordering test:
| Test | Verifies |
|---|---|
TestReconcileOrphans_NilDockerNoOp |
nil docker ⇒ no-op, nil error, no panic; pool bookkeeping still works |
TestReconcileOrphans_RemovesOrphansSkipsLivePools |
exact removal set: live-pool + unlabeled containers skipped, exactly [c-orphan-1, c-orphan-2] removed |
TestReconcileOrphans_BestEffortErrorsJoin |
two removals fail ⇒ joined error contains both failures, yet c-ok is still removed |
TestReconcileOrphans_RLockBlocksGetOrCreatePool |
while a removal is in flight, GetOrCreatePool blocks (100 ms probe) and completes only after reconciliation releases the RLock |
All checks run in `~/go-docker-orphan-reconciliation`: - **`go build ./...`** — clean (docker SDK v28.5.2 fetched from proxy, compiles). - **`go vet ./...`** — clean. - **`gofmt -l .`** — clean (no unformatted files). - **`go test -race -count=1 -v ./...`** — 4/4 PASS: ``` === RUN TestReconcileOrphans_NilDockerNoOp --- PASS (0.00s) === RUN TestReconcileOrphans_RemovesOrphansSkipsLivePools --- PASS (0.00s) === RUN TestReconcileOrphans_BestEffortErrorsJoin --- PASS (0.00s) === RUN TestReconcileOrphans_RLockBlocksGetOrCreatePool --- PASS (0.10s) ok uhlp.dev/sandbox/sandbox 1.114s ``` - **Race detector repeated 10× on the concurrency test** (`-race -count=10 -run RLock`) — stable, no flakes, no data races. **Edge cases covered:** docker-disabled (nil client) startup; orphan vs live-pool ownership discrimination via the label value; containers with the marker label but no pool value; partial removal failures (best-effort, `errors.Join`, continued reconciliation); context cancellation mid-pass; and the GetOrCreatePool/ReconcileOrphans lock race that motivated the RLock design.
{"model": "deepseek-v4-flash", "problem_class": "go-docker-orphan-reconciliation", "result": "passed", "tests": 4}