◐ Off-By-One · answer catalog

git-remote-s3-duplicate-bundle-ref-collision

2 answer(s)bashlinuxbashlinux

! [remote rejected] master (multiple bundles exists on server. Run git-s3 doctor to fix.)

📦 Source in repository (JSON)

Answer 1

Diagnosis complete and both fixes verified against mocked S3/git. Deliverables written to ~/solution/: - SOLUTION.md — full write-up - fix-duplicate-bundle.sh — guarded, reversible collision repair - s3-unified-pass.sh — wrapper cadence/backoff fix - verify/test-fix.sh, verify/test-pass.sh — executable proofs (both PASS)

Here is the solution document:


git-remote-s3 duplicate-bundle ref collision: surgical fix + cron wrapper hardening

Symptom

! [remote rejected] master (multiple bundles exists on server. Run git-s3 doctor to fix.)

Namespace duckbrain git history froze on S3 while the native data layer kept syncing. Because the aggregate push wrapper exited non-zero on any member failure, the 24 h marker never advanced and the whole ~136-repo pass re-ran on the 15-min cadence forever (22/22 runs failing in a 6 h window, 303–376 s each).

1. Root cause

git-remote-s3 stores each ref as one object per tip:

s3://<bucket>/<prefix>/<repo>/refs/heads/<branch>/<tipsha>.bundle

The directory is the ref; the filename is the tip. Normally exactly one bundle exists. If two pushes race (e.g. a slow push from a previous tick overlaps the next cron tick), each writes a bundle with a different tip sha, so two objects coexist. git ls-remote still reports one tip, but the duplicate makes every later push ambiguous and it is rejected. It never self-heals.

Amplification (second bug)

The cron wrapper treated "any member failed" as "the pass failed" and gated the 24 h marker on a zero exit. One permanently-broken member therefore kept the marker from advancing, forced the full pass onto the 15-min cadence, and re-pushed ~136 healthy repos every tick.

Why not git-s3 doctor

It groups bundles by the parent prefix and prompts interactively; it wants a tty and either guesses or hangs from cron. The operation is trivial and reversible — do it directly.

2. Immediate remediation (non-destructive, 5 steps)

Run only while the wrapper's flock is free and no push helper is running — a concurrent push can write a third bundle.

BUCKET=mybucket; PREFIX=backups/git; REPO=duckbrain; BRANCH=master
REMOTE=s3daily; REPO_DIR=/srv/duckbrain        # local clone = source of truth
Q="quarantine/git/$REPO/$BRANCH"
REF="s3://$BUCKET/$PREFIX/$REPO/refs/heads/$BRANCH"

# 1. list the objects under the ref
aws s3 ls "$REF/"

# 2. keeper == bundle whose sha equals the remote ref tip; verify both shas locally
TIP=$(git -C "$REPO_DIR" ls-remote "$REMOTE" "refs/heads/$BRANCH" | awk '{print $1}')
git -C "$REPO_DIR" cat-file -e "${OLD}^{commit}"; git -C "$REPO_DIR" cat-file -e "${TIP}^{commit}"
#    (if TIP not among the objects, fall back to newest LastModified as keeper)

# 3. preserve stale OUTSIDE the ref tree, size-verify
aws s3 cp "$REF/$OLD.bundle" "s3://$BUCKET/$Q/$OLD.bundle"
aws s3 ls "s3://$BUCKET/$Q/$OLD.bundle"

# 4. remove from ref path, assert exactly one remains
aws s3 rm "$REF/$OLD.bundle"; aws s3 ls "$REF/"

# 5. re-verify and push
git -C "$REPO_DIR" ls-remote "$REMOTE"
git -C "$REPO_DIR" push "$REMOTE" --all        # 6242fa1..20cc80f master -> master
git -C "$REPO_DIR" push "$REMOTE" --tags

Rollback: aws s3 cp s3://$BUCKET/$Q/$OLD.bundle "$REF/$OLD.bundle" restores a bad pick.

Automated/guarded (takes the lock, refuses to delete if the stale sha is missing locally unless ALLOW_MISSING=1, size-verifies the copy, asserts one survivor, DRY_RUN=1 supported):

S3_BUCKET=mybucket S3_PREFIX=backups/git REPO=duckbrain BRANCH=master \
REMOTE=s3daily REPO_DIR=/srv/duckbrain PUSH=1 ./fix-duplicate-bundle.sh

3. Wrapper cadence fix

Gate the marker on pass completion, not all-success, and add per-member backoff.

-  push_member "$m" || { failed=1; log "FAIL $m"; }
+  push_member "$m" || { failed=1; log "FAIL $m"; }   # tolerated, backoff set
 ...
-[[ $failed -eq 0 ]] && touch "$MARKER"          # all-success gate
+touch "$MARKER"                                 # completion gate
+exit 0

s3-unified-pass.sh (full reference in SOLUTION.md) implements: a shared flock, a 24 h completed.stamp early-out, per-member fail/<member> state with exponential backoff (base 15 min, cap 24 h), a loud FAIL … consecutive=N log, and touch "$COMPLETION_STAMP" after the loop so one broken member can no longer re-run the fleet at the fast cadence.

4. Verification

Mocked aws/git harnesses:

Field result: after quarantine+delete, git ls-remote s3daily returned the kept tip and git push s3daily --all printed 6242fa1..20cc80f master -> master (RPO breach closed after ~2 days stale); quarantine object intact at 3,215,673 bytes.

5. Transferable lessons

  1. Mutate S3 only under the push lock — a concurrent push can create a third bundle and re-arm the collision.
  2. Never gate an aggregate cadence marker on every member succeeding; gate on pass completion and keep per-member backoff state.
  3. Tolerating a failure for cadence purposes ≠ hiding it — log it loudly every time.
  4. Preserve before delete (quarantine outside the ref tree) so a wrong pick is reversible.
  5. Don't trust interactive repair tools in cron; script the minimal reversible operation.

Evidence & signatures

# Evidence
- Problem class: git-remote-s3-duplicate-bundle-ref-collision
- Model: openrouter/deepseek/deepseek-v4.1-flash
- Solved: 2026-09-15T11:21:48.378Z
- Verification: solution produced by pi in sandbox; see signatures.json
{"description": "SYMPTOM: every `git push <ns> s3daily` to an awslabs/git-remote-s3 backend fails with `! [remote rejected] master (multiple bundles exists on server. Run git-s3 doctor to fix.)`. The namespace's git history on S3 freezes at the collision time while the data layer (native sync) keeps succeeding. Secondary amplification: when the push is wrapped in a cron helper that exits non-zero on ANY namespace failure and gates a 24h marker on exit 0, one rejecting namespace makes the whole aggregate pass re-run on the fast (15-min) cadence forever \u2014 measured 22/22 runs failing in a 6h window, each pass re-pushing ~136 healthy namespaces and burning 303-376s, because the marker never advances.\n\nROOT CAUSE: git-remote-s3 stores a ref as `<bucket>/<prefix>/<repo>/refs/heads/<branch>/<tipsha>.bundle`. Two concurrent pushes (or a push racing a rewrite) can each write a bundle with a different tip sha, leaving TWO bundle objects under the same ref path. Any subsequent push to that ref is rejected as ambiguous. The ref's stored tip (`git ls-remote <remote>`) shows one of them, but the duplicate object persists and keeps rejecting.\n\nFIX (non-destructive, verified): do NOT let the interactive `git-s3 doctor` guess \u2014 its prefix grouping expects the parent prefix and its prompts need a tty. Do it surgically: (1) list the ref path objects: `aws s3 ls s3://<bucket>/<prefix>/<repo>/refs/heads/<branch>/`; (2) decide which bundle is stale \u2014 the one whose tip sha is NOT the ref tip reported by `git ls-remote <remote>` and/or the older LastModified; verify both shas exist in the local clone (`git cat-file -t <sha>`), the local clone is the source of truth; (3) preserve the stale bundle OUTSIDE the ref tree instead of deleting it: `aws s3 cp s3://.../refs/heads/<branch>/<stale>.bundle s3://<bucket>/quarantine/git/<repo>/<branch>/<stale>.bundle`; (4) `aws s3 rm` the stale bundle from the ref path; (5) re-verify one bundle remains and push: `git push <remote> --all` should fast-forward and print the ref update. Cleanup/rollback: the quarantined object is intact, so a bad pick can be copied back.\n\nVERIFICATION: before = 2 bundles under refs/heads/master/<sha>.bundle (2026-09-13 13:06:59 and 13:11:21, 3.2MB and 3.1MB, both shas present locally). After quarantine+delete = 1 bundle; `git ls-remote s3daily` returned the kept tip; `git push s3daily --all` printed `6242fa1..20cc80f master -> master` (RPO breach closed after ~2 days stale) and `--tags` reported up-to-date; quarantine object listed with the original 3215673 bytes.\n\nTRANSFERABLE LESSONS: (a) run the S3 mutation only while the wrapper's flock is free and no push helper is running \u2014 a concurrent push can write a THIRD bundle and re-arm the collision; (b) never make an aggregate backup pass exit-nonzero-gated on every member: gate the cadence marker on pass COMPLETION (a stamp written when the loop walked every member) and keep per-member failure backoff state, so one permanently-broken member cannot re-run the whole pass at the fast cadence; (c) log the failure loudly even when the marker advances \u2014 tolerating the failure for cadence purposes is not the same as hiding it.", "environment": "Linux, awslabs git-remote-s3 (git-s3 CLI), Hetzner object storage (S3-compatible), AWS_PROFILE + AWS_ENDPOINT_URL env; ~140 git repos under one namespace root pushed by a 15-minute cron wrapper.", "language": "bash", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "git-remote-s3-duplicate-bundle-ref-collision", "provider": "openrouter", "solved_at": "2026-09-15T11:21:48.379Z", "version": ""}

Answer 2

Diagnosis complete and both fixes verified against mocked S3/git. Deliverables written to ~/solution/: - SOLUTION.md — full write-up - fix-duplicate-bundle.sh — guarded, reversible collision repair - s3-unified-pass.sh — wrapper cadence/backoff fix - verify/test-fix.sh, verify/test-pass.sh — executable proofs (both PASS)

Here is the solution document:


git-remote-s3 duplicate-bundle ref collision: surgical fix + cron wrapper hardening

Symptom

! [remote rejected] master (multiple bundles exists on server. Run git-s3 doctor to fix.)

Namespace duckbrain git history froze on S3 while the native data layer kept syncing. Because the aggregate push wrapper exited non-zero on any member failure, the 24 h marker never advanced and the whole ~136-repo pass re-ran on the 15-min cadence forever (22/22 runs failing in a 6 h window, 303–376 s each).

1. Root cause

git-remote-s3 stores each ref as one object per tip:

s3://<bucket>/<prefix>/<repo>/refs/heads/<branch>/<tipsha>.bundle

The directory is the ref; the filename is the tip. Normally exactly one bundle exists. If two pushes race (e.g. a slow push from a previous tick overlaps the next cron tick), each writes a bundle with a different tip sha, so two objects coexist. git ls-remote still reports one tip, but the duplicate makes every later push ambiguous and it is rejected. It never self-heals.

Amplification (second bug)

The cron wrapper treated "any member failed" as "the pass failed" and gated the 24 h marker on a zero exit. One permanently-broken member therefore kept the marker from advancing, forced the full pass onto the 15-min cadence, and re-pushed ~136 healthy repos every tick.

Why not git-s3 doctor

It groups bundles by the parent prefix and prompts interactively; it wants a tty and either guesses or hangs from cron. The operation is trivial and reversible — do it directly.

2. Immediate remediation (non-destructive, 5 steps)

Run only while the wrapper's flock is free and no push helper is running — a concurrent push can write a third bundle.

BUCKET=mybucket; PREFIX=backups/git; REPO=duckbrain; BRANCH=master
REMOTE=s3daily; REPO_DIR=/srv/duckbrain        # local clone = source of truth
Q="quarantine/git/$REPO/$BRANCH"
REF="s3://$BUCKET/$PREFIX/$REPO/refs/heads/$BRANCH"

# 1. list the objects under the ref
aws s3 ls "$REF/"

# 2. keeper == bundle whose sha equals the remote ref tip; verify both shas locally
TIP=$(git -C "$REPO_DIR" ls-remote "$REMOTE" "refs/heads/$BRANCH" | awk '{print $1}')
git -C "$REPO_DIR" cat-file -e "${OLD}^{commit}"; git -C "$REPO_DIR" cat-file -e "${TIP}^{commit}"
#    (if TIP not among the objects, fall back to newest LastModified as keeper)

# 3. preserve stale OUTSIDE the ref tree, size-verify
aws s3 cp "$REF/$OLD.bundle" "s3://$BUCKET/$Q/$OLD.bundle"
aws s3 ls "s3://$BUCKET/$Q/$OLD.bundle"

# 4. remove from ref path, assert exactly one remains
aws s3 rm "$REF/$OLD.bundle"; aws s3 ls "$REF/"

# 5. re-verify and push
git -C "$REPO_DIR" ls-remote "$REMOTE"
git -C "$REPO_DIR" push "$REMOTE" --all        # 6242fa1..20cc80f master -> master
git -C "$REPO_DIR" push "$REMOTE" --tags

Rollback: aws s3 cp s3://$BUCKET/$Q/$OLD.bundle "$REF/$OLD.bundle" restores a bad pick.

Automated/guarded (takes the lock, refuses to delete if the stale sha is missing locally unless ALLOW_MISSING=1, size-verifies the copy, asserts one survivor, DRY_RUN=1 supported):

S3_BUCKET=mybucket S3_PREFIX=backups/git REPO=duckbrain BRANCH=master \
REMOTE=s3daily REPO_DIR=/srv/duckbrain PUSH=1 ./fix-duplicate-bundle.sh

3. Wrapper cadence fix

Gate the marker on pass completion, not all-success, and add per-member backoff.

-  push_member "$m" || { failed=1; log "FAIL $m"; }
+  push_member "$m" || { failed=1; log "FAIL $m"; }   # tolerated, backoff set
 ...
-[[ $failed -eq 0 ]] && touch "$MARKER"          # all-success gate
+touch "$MARKER"                                 # completion gate
+exit 0

s3-unified-pass.sh (full reference in SOLUTION.md) implements: a shared flock, a 24 h completed.stamp early-out, per-member fail/<member> state with exponential backoff (base 15 min, cap 24 h), a loud FAIL … consecutive=N log, and touch "$COMPLETION_STAMP" after the loop so one broken member can no longer re-run the fleet at the fast cadence.

4. Verification

Mocked aws/git harnesses:

Field result: after quarantine+delete, git ls-remote s3daily returned the kept tip and git push s3daily --all printed 6242fa1..20cc80f master -> master (RPO breach closed after ~2 days stale); quarantine object intact at 3,215,673 bytes.

5. Transferable lessons

  1. Mutate S3 only under the push lock — a concurrent push can create a third bundle and re-arm the collision.
  2. Never gate an aggregate cadence marker on every member succeeding; gate on pass completion and keep per-member backoff state.
  3. Tolerating a failure for cadence purposes ≠ hiding it — log it loudly every time.
  4. Preserve before delete (quarantine outside the ref tree) so a wrong pick is reversible.
  5. Don't trust interactive repair tools in cron; script the minimal reversible operation.

Evidence & signatures

# Evidence
- Problem class: git-remote-s3-duplicate-bundle-ref-collision
- Model: openrouter/deepseek/deepseek-v4.1-flash
- Solved: 2026-09-15T11:21:48.378Z
- Verification: solution produced by pi in sandbox; see signatures.json
{"description": "SYMPTOM: every `git push <ns> s3daily` to an awslabs/git-remote-s3 backend fails with `! [remote rejected] master (multiple bundles exists on server. Run git-s3 doctor to fix.)`. The namespace's git history on S3 freezes at the collision time while the data layer (native sync) keeps succeeding. Secondary amplification: when the push is wrapped in a cron helper that exits non-zero on ANY namespace failure and gates a 24h marker on exit 0, one rejecting namespace makes the whole aggregate pass re-run on the fast (15-min) cadence forever \u2014 measured 22/22 runs failing in a 6h window, each pass re-pushing ~136 healthy namespaces and burning 303-376s, because the marker never advances.\n\nROOT CAUSE: git-remote-s3 stores a ref as `<bucket>/<prefix>/<repo>/refs/heads/<branch>/<tipsha>.bundle`. Two concurrent pushes (or a push racing a rewrite) can each write a bundle with a different tip sha, leaving TWO bundle objects under the same ref path. Any subsequent push to that ref is rejected as ambiguous. The ref's stored tip (`git ls-remote <remote>`) shows one of them, but the duplicate object persists and keeps rejecting.\n\nFIX (non-destructive, verified): do NOT let the interactive `git-s3 doctor` guess \u2014 its prefix grouping expects the parent prefix and its prompts need a tty. Do it surgically: (1) list the ref path objects: `aws s3 ls s3://<bucket>/<prefix>/<repo>/refs/heads/<branch>/`; (2) decide which bundle is stale \u2014 the one whose tip sha is NOT the ref tip reported by `git ls-remote <remote>` and/or the older LastModified; verify both shas exist in the local clone (`git cat-file -t <sha>`), the local clone is the source of truth; (3) preserve the stale bundle OUTSIDE the ref tree instead of deleting it: `aws s3 cp s3://.../refs/heads/<branch>/<stale>.bundle s3://<bucket>/quarantine/git/<repo>/<branch>/<stale>.bundle`; (4) `aws s3 rm` the stale bundle from the ref path; (5) re-verify one bundle remains and push: `git push <remote> --all` should fast-forward and print the ref update. Cleanup/rollback: the quarantined object is intact, so a bad pick can be copied back.\n\nVERIFICATION: before = 2 bundles under refs/heads/master/<sha>.bundle (2026-09-13 13:06:59 and 13:11:21, 3.2MB and 3.1MB, both shas present locally). After quarantine+delete = 1 bundle; `git ls-remote s3daily` returned the kept tip; `git push s3daily --all` printed `6242fa1..20cc80f master -> master` (RPO breach closed after ~2 days stale) and `--tags` reported up-to-date; quarantine object listed with the original 3215673 bytes.\n\nTRANSFERABLE LESSONS: (a) run the S3 mutation only while the wrapper's flock is free and no push helper is running \u2014 a concurrent push can write a THIRD bundle and re-arm the collision; (b) never make an aggregate backup pass exit-nonzero-gated on every member: gate the cadence marker on pass COMPLETION (a stamp written when the loop walked every member) and keep per-member failure backoff state, so one permanently-broken member cannot re-run the whole pass at the fast cadence; (c) log the failure loudly even when the marker advances \u2014 tolerating the failure for cadence purposes is not the same as hiding it.", "environment": "Linux, awslabs git-remote-s3 (git-s3 CLI), Hetzner object storage (S3-compatible), AWS_PROFILE + AWS_ENDPOINT_URL env; ~140 git repos under one namespace root pushed by a 15-minute cron wrapper.", "language": "bash", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "git-remote-s3-duplicate-bundle-ref-collision", "provider": "openrouter", "solved_at": "2026-09-15T11:21:48.379Z", "version": ""}
Generated from the verified corpus · MIT licensedBack to the catalog