Skip to content

chore: stop an interrupted repack from slowing every write - #1439

Merged
MicBun merged 1 commit into
mainfrom
clear-interrupted-repack
Sep 30, 2026
Merged

MicBun merged 1 commit into
mainfrom
clear-interrupted-repack

Conversation

@MicBun

@MicBun MicBun commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

tn_vacuum runs pg_repack --all every 50,000 blocks, under exec.CommandContext. When the node shuts down during a run, the context kills pg_repack with SIGKILL, which skips the cleanup pg_repack runs when interrupted. Its repack_trigger stays on the table it was repacking and copies every later write into a repack.log_<oid> table, and nothing removes it. The next --all run then repacks those log tables as well, so the triggers chain and one write becomes many.

One mainnet sentry had collected 53 leftover triggers and 96 log tables holding 226 GB, in a 255 GB database. A write to primitive_event_type there cascaded into 19 log tables, and a write to primitive_events into 17. The replication publication covers every table, so the node decoded all of those rows for every block. On 26-Sep one auto_digest block pushed it past the 30 s the node then allowed for a block's commit ID, and it failed that block on every retry for three days. A second sentry had 37 triggers and 48 GB of log tables. The leader, and a node run by another operator, had none.

What changed

Leftovers are cleared before each run. Before starting pg_repack, the mechanism counts repack_triggers and the tables in the repack schema. A fresh extension owns no tables there, so anything it finds was left by an interrupted run. It then drops the extension with CASCADE and creates it again in one transaction, which removes the triggers, the log tables and their types. lock_timeout is 5 s, because block execution queues behind that lock. If the lock is not free, the run is reported as failed and pg_repack does not start on top of the leftovers. The next scheduled run tries again.

A cancelled pg_repack gets SIGINT. cmd.Cancel sends SIGINT, and WaitDelay kills the process only after 30 s. pg_repack 1.5.3 handles SIGINT by cancelling its query and calling repack.repack_drop from its exit handler (on_interrupt in pgut.c, repack_cleanup_callback in pg_repack.c). It installs no handler for SIGTERM. The extension's Close does not wait for the run, so on a node shutdown this is best effort. The check before each run is what makes sure nothing stays behind.

Nothing in the repack schema is consensus state. The leader has none of it, and the sentry cleared by hand has applied more than 14,000 blocks since then with no app hash mismatch.

Tests

  • TestClearInterruptedRepack (kwiltest, on the kwil-postgres image nodes run). It leaves an interrupted run on a table the way pg_repack does, by running the create_pktype, create_log and create_trigger statements from pg_repack's own repack.tables view. Then it does the same on that run's log table, so the triggers chain, and checks that one insert lands in both log tables. After clearing, no repack_trigger and no table in repack remain, the extension exists again, and the table keeps its rows and takes writes. A log table left without its trigger is cleared too. With nothing left behind, the extension is not dropped.
  • TestRunInterruptsPgRepackOnCancel: a stand-in pg_repack records SIGINT when its run's context is cancelled.
  • TestRunSkipsPgRepackWhenLeftoversCannotBeCleared: when clearing fails, the run fails and pg_repack does not start.

Seven mutations, seven caught: SIGKILL instead of SIGINT, the clearing error ignored, clearing never called, leftovers counted but not dropped, the extension not created again, the extension dropped when nothing was left, and tables in repack not counted.

go test ./extensions/tn_vacuum/... and go test -tags kwiltest ./extensions/tn_vacuum/... pass. go vet (both tags) and golangci-lint with the repo's config are clean.

Not covered

A node that already carries leftovers is cleared at its next scheduled run after it upgrades, not at startup. --all is unchanged; once this lands, the repack schema holds no tables when a run starts, so pg_repack no longer repacks its own log tables.

There is no Problem issue for this in node, so the PR carries no closing keyword.

Summary by CodeRabbit

  • Bug Fixes
    • Interrupted database repacking is now cleaned up before a new repack begins, helping prevent leftover database objects from blocking maintenance.
    • If cleanup fails, repacking does not start and the operation reports a failure.
    • Cancelling a repack now interrupts the running process and allows time for cleanup before forceful termination.

@MicBun MicBun self-assigned this Sep 29, 2026
@holdex

holdex Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

Time Submission Status

Member # Time Running Total Status Last Update
MicBun 4h ✅ Submitted Sep 29, 2026, 11:40 PM

Submit or update total time with:

@holdex pr submit-time 2h

Add time on top of previous submission with:

@holdex pr add-time 1h30m

See available commands to help comply with our Guidelines.

@coderabbitai

coderabbitai Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Advanced

Run ID: e7d408d0-7eb5-4c3b-81e0-1c1c640b5691

📥 Commits

Reviewing files that changed from the base of the PR and between 9e7ef82 and e8f8c41.

📒 Files selected for processing (3)
  • extensions/tn_vacuum/mechanism_repack.go
  • extensions/tn_vacuum/mechanism_repack_pg_test.go
  • extensions/tn_vacuum/mechanism_repack_test.go

Included review availability: This review used your included allowance. Your plan provides up to 1 included review per hour; 0 remain after this review.


📝 Walkthrough

Walkthrough

The pg_repack mechanism now clears interrupted-run leftovers before starting the command. It returns a failed report if cleanup fails. Context cancellation sends SIGINT to the command and allows a 30-second wait period.

Changes

pg_repack lifecycle

Layer / File(s) Summary
Leftover cleanup and preflight
extensions/tn_vacuum/mechanism_repack.go, extensions/tn_vacuum/mechanism_repack_pg_test.go, extensions/tn_vacuum/mechanism_repack_test.go
The mechanism counts non-internal repack_trigger triggers and ordinary tables in the repack schema. If either count is nonzero, it drops and recreates the extension in a transaction with a 5-second lock timeout. Cleanup errors return a failed report before command startup. PostgreSQL integration tests cover untouched extensions, chained log tables, and a remaining log table without a trigger.
Command cancellation
extensions/tn_vacuum/mechanism_repack.go, extensions/tn_vacuum/mechanism_repack_test.go
The command sends SIGINT when its context is canceled and uses a 30-second WaitDelay. Tests verify that the fake executable receives SIGINT and exits.

Priority: ➖ Normal

Estimated code review effort: 3 (Moderate) | ~20 minutes

Change: Bug fix

Sequence Diagram(s)

sequenceDiagram
  participant Caller
  participant Run
  participant Cleanup as clearInterruptedRepack
  participant DB as PostgreSQL
  participant Repack as pg_repack
  Run->>Cleanup: Check and clear leftover repack objects
  Cleanup->>DB: Count triggers and repack tables
  DB-->>Cleanup: Return counts
  Cleanup->>DB: Recreate extension if leftovers exist
  Cleanup-->>Run: Return cleanup result
  Run->>Repack: Start command if cleanup succeeds
  Caller->>Run: Cancel context
  Run->>Repack: Send SIGINT
Loading

Merge Risk: ⚪ Minimal · up to e8f8c

The change clears interrupted-repack leftovers before starting new work and allows graceful cancellation. No actionable merge-blocking issue remains; merge after normal checks.

Security Architecture Review

Security architecture risk: 🟡 Moderate · up to e8f8c

The recovery design addresses interrupted maintenance and prevents new work after cleanup fails. Its database-wide cleanup nevertheless assumes exclusive ownership that is enforced only within one worker. Concurrent maintenance and database permissions remain unresolved; no attacker-reachable exploit was established.

Retained concerns

  • Medium · reliability · inferred: The new recovery predicate treats matching database artifacts as interrupted without establishing database-wide exclusive ownership. Another active pg_repack run can produce the same artifacts, but the scheduler serializes only one Extension instance. If independent maintenance overlaps and the required locks become available, CASCADE cleanup could remove its artifacts or disrupt its execution. The five-second lock timeout limits waiting but does not establish lifecycle ownership. Deployment overlap and active-run locking remain unverified.
Security review details

Security Blast Radius

  • inferred — The destructive path targets the selected PostgreSQL database and pg_repack's dependency closure. Intended artifacts include triggers on source tables and tables in schema repack. No cross-database or cross-node cascade is shown; the complete production dependency closure remains unknown.

Trust Boundaries and Controls

  • observed — The examined production caller derives database identity and credentials from local service configuration. Cleanup and command execution use the same selected configuration. Documentation already describes elevated database privileges, but it does not establish deployed grants or whether untrusted roles can create objects matching the cleanup predicate.

Resilience and Maintainability Implications

  • observed — Cleanup errors stop new maintenance work, lock waits are bounded, and scheduling advances after attempted execution rather than retrying on every block. These controls contain repeated failures within one worker, but do not coordinate independent processes sharing a database.

Hardening Proposals

  • proposed — Make exclusive database ownership an explicit recovery precondition. Where independent maintenance is supported, use coordination honored by every maintenance initiator, verify artifact provenance, and validate cleanup against an active repack and restricted production-role permissions.
🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 28.57% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 14 functions across 3 files. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly summarizes the main change: preventing interrupted pg_repack runs from leaving leftovers that slow later writes.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
  • Fix all pre-merge checks with AI
✨ Finishing Touches 💡 1
📝 Generate docstrings 💡
  • Commit to this branch
  • Create a new PR
🧪 Generate unit tests (beta)
  • Commit to this branch
  • Create a new PR

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@MicBun

MicBun commented Sep 29, 2026

Copy link
Copy Markdown
Contributor Author

@holdex pr submit-time 4h

@MicBun
MicBun merged commit aadacd8 into main Sep 30, 2026
8 checks passed
@MicBun
MicBun deleted the clear-interrupted-repack branch September 30, 2026 13:01
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant