Skip to content

feat(coordinator): reap orphaned base snapshots (image-refresh storage leak) - #439

Merged
nikhilunni merged 1 commit into
mainfrom
base-snapshot-reaper
Jun 25, 2026
Merged

nikhilunni merged 1 commit into
mainfrom
base-snapshot-reaper

Conversation

@nikhilunni

Copy link
Copy Markdown
Contributor

What

Adds a coordinator sweeper that reaps orphaned per-image base snapshots — the storage leak behind the Jun-23/24 jump in GCS live-object bytes.

Why

Every enabled-image re-bake/RefreshImage captures a fresh base snapshot and swaps enabled_images.base_snapshot_id to it (upsert_enabled_image's ON CONFLICT). The prior base row is left dangling: session_id IS NULL, pointed to by no enabled_images row. Nothing deleted it:

So each refresh permanently leaked one base snapshot's chunks — 20–32 GB for the heavy dogfood images (dev-brain, dev-engrams). Prod currently holds 23 orphaned bases (~150 GB logical) dating back 9 days, none in any GC queue.

How

prune_orphan_base_snapshots — the session_id IS NULL mirror of prune_session_snapshots:

DELETE FROM snapshots s
WHERE s.session_id IS NULL
  AND s.created_at < NOW() - $grace
  AND s.id NOT IN (SELECT base_snapshot_id FROM enabled_images WHERE base_snapshot_id IS NOT NULL)
  • Skips bases referenced by any enabled_images row, including soft-deleted (their chunk lineage is intentionally still pinned, ADR 0021 P1.8).
  • Bumps chunk_generation in the same TX (GC-barrier symmetry). The existing chunk-GC (ADR 0016 Phase C) and snapshot-blob-GC (ADR 0028 addendum) sweeps then reclaim the now-unpinned chunks and portable snapshots/<id>/ blobs — this reaper deletes nothing in BlobStorage directly.
  • The base_snapshot_id FK (REFERENCES snapshots(id), no ON DELETE) is a hard backstop against deleting an in-use base even if the predicate regressed.
  • Spawned beside checkpoint_retention; grace via ENGRAM_BASE_SNAPSHOT_RETENTION_HOURS (default 24).

A sweeper (not an inline delete in RefreshImage) so it both drains the existing 23-orphan backlog and prevents future leaks, per the "scanner drives transitions" convention.

Test

Live-PG test wired into the existing Postgres-gated CI lane (checkpoint_reconcile_live_pg): a superseded orphan past grace is reaped; the current enabled-image base and a fresh orphan within grace are kept; session snapshots are untouched. Validated against a real Postgres locally (just check green: fmt + clippy -D warnings + 1270 workspace tests).

Also fixes the now-stale snapshot_blob_pin_set doc comment that asserted base rows are never deleted.

Scope notes

  • No ADR — this is an incremental extension of the existing ADR-0016/0028 GC machinery, not a new decision. The broader idle-session GC the orphan investigation surfaced is a separate follow-up (different row population: session_id IS NOT NULL).
  • No admin trigger — matches checkpoint_retention (its closest sibling); the 24h sweeper drains the backlog automatically. For immediate prod reclaim we can run the equivalent one-shot SQL out of band.

Roll

Coord-only; auto-rolls on merge. The 23 existing orphans age out within one grace window (24h) after deploy.

🤖 Generated with Claude Code

…e leak)

Each enabled-image re-bake/refresh captures a fresh per-image base
snapshot and swaps `enabled_images.base_snapshot_id` to it
(`upsert_enabled_image`'s ON CONFLICT), leaving the PRIOR base row
dangling: `session_id IS NULL`, referenced by no `enabled_images` row.
Nothing deleted it — `prune_session_snapshots` (checkpoint retention) is
`session_id IS NOT NULL` only — and an orphan base keeps pinning its own
disk+memory chunks via pin-set sources #3/#4
(`list_recoverable_snapshot_{disk,memory}_manifests`, which filter on
`recoverable = TRUE` with no session predicate). So every image refresh
permanently leaked one base snapshot's chunks (20-32 GB for the heavy
dogfood images). Prod had 23 such orphans (~150 GB logical) going back 9
days, none in any GC queue.

Add `prune_orphan_base_snapshots`, the `session_id IS NULL` mirror of
`prune_session_snapshots`: delete base rows older than
`ENGRAM_BASE_SNAPSHOT_RETENTION_HOURS` (default 24) that are referenced by
no `enabled_images.base_snapshot_id` (live OR soft-deleted — soft-deleted
lineage is intentionally still chunk-pinned, ADR 0021 P1.8), bumping
`chunk_generation` in the same TX (GC-barrier symmetry). The existing
chunk-GC (ADR 0016 Phase C) and snapshot-blob-GC (ADR 0028 addendum)
sweeps then reclaim the now-unpinned chunks and portable `snapshots/<id>/`
blobs — the reaper deletes nothing in BlobStorage directly. The
`base_snapshot_id` FK (REFERENCES snapshots(id), no ON DELETE) is a hard
backstop against ever deleting an in-use base.

Spawned beside `checkpoint_retention`. A sweeper (not an inline
delete-old-base in RefreshImage) so it both drains the existing backlog
and survives restarts, per the "scanner drives transitions" convention.

Also fixes the now-stale `snapshot_blob_pin_set` doc that claimed base
rows are never deleted.

Test (live-PG, wired into the existing Postgres-gated CI lane): a
superseded orphan past grace is reaped; the current (enabled-image) base
and a fresh orphan within grace are kept; session snapshots are
untouched. The sibling checkpoint-retention test's template is pinned
recent so the new global reaper can't collect it under local parallel
runs (CI serializes the lane).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@nikhilunni
nikhilunni merged commit 3479ee6 into main Jun 25, 2026
17 checks passed
@nikhilunni
nikhilunni deleted the base-snapshot-reaper branch June 25, 2026 15:33
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant