Serialize database destruction with concurrent opens - #787
Conversation
There was a problem hiding this comment.
Code Review
This pull request introduces robust database lifecycle management for RocksDB JS bindings. It implements a timed-wait mechanism (lifecycleWaitSeconds) for open, destroy, and shutdown operations to prevent concurrent lifecycle conflicts. It also introduces a "quarantine" state for database paths when a native close, flush, compaction, or physical directory cleanup fails, preventing subsequent opens until the cleanup is retried via destroy() or shutdown(). Additionally, it ensures that in-flight operations (like backups and checkpoints) are safely awaited before destruction, and that thread-affine N-API references are cleaned up safely. There are no review comments, so I have no feedback to provide.
📊 Benchmark Resultsget-sync.bench.tsgetSync() > random keys - small key size (100 records)
getSync() > sequential keys - small key size (100 records)
ranges.bench.tsgetRange() > small range (100 records, 50 range)
realistic-load.bench.tsRealistic write load with workers > write variable records with transaction log
transaction-log.bench.tsTransaction log > read 100 iterators while write log with 100 byte records
Transaction log > read one entry from random position from log with 1000 100 byte records
worker-put-sync.bench.tsputSync() > random keys - small key size (100 records, 10 workers)
worker-transaction-log.bench.tsTransaction log with workers > write log with 100 byte records
Results from commit 3842f2f |
|
Reviewed Re-review of the one new commit since
Also checked and cleared: the dropped Verification: full suite 55 files, 773 passed / 1 skipped / 0 failed; targeted destroy + ranges 64/64 including the new One merge-ordering note, not a defect in this PR: — |
…GetCount against concurrent close finishClose() cleared compactCancelRequested right after the operationsInFlight drain, but an async compact() releases its OperationGuard at setup handoff and is not awaited until the closables sweep — so it can still be running after the drain returns, and clearing the token there left it able to stall teardown (and every concurrent open on the path) indefinitely. Keep the token armed for finishClose()'s whole duration instead, and have the close-time compact-on-close pass opt out via a new compactRange() `cancellable` param rather than relying on the shared flag being cleared. Transaction::GetCount now takes an OperationGuard and checks isClosing() before scanning: without it, finishClose()'s drain can return immediately and the closables sweep can roll back the transaction while the count scan is parked between rows, reading freed memory. Carries the in-progress PR #787 lifecycle repair plan describing the fuller atomic-admission fix these two changes are a first slice of. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013CvCKcGKG7vyhRts1Mv4Wh
f24ef7a to
5b459e4
Compare
…GetCount against concurrent close finishClose() cleared compactCancelRequested right after the operationsInFlight drain, but an async compact() releases its OperationGuard at setup handoff and is not awaited until the closables sweep — so it can still be running after the drain returns, and clearing the token there left it able to stall teardown (and every concurrent open on the path) indefinitely. Keep the token armed for finishClose()'s whole duration instead, and have the close-time compact-on-close pass opt out via a new compactRange() `cancellable` param rather than relying on the shared flag being cleared. Transaction::GetCount now takes an OperationGuard and checks isClosing() before scanning: without it, finishClose()'s drain can return immediately and the closables sweep can roll back the transaction while the count scan is parked between rows, reading freed memory. Carries the in-progress PR #787 lifecycle repair plan describing the fuller atomic-admission fix these two changes are a first slice of. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013CvCKcGKG7vyhRts1Mv4Wh
5b459e4 to
542b058
Compare
|
Reviewed Re-review of the one new commit since
Also re-verified at object-code level:
— |
…GetCount against concurrent close finishClose() cleared compactCancelRequested right after the operationsInFlight drain, but an async compact() releases its OperationGuard at setup handoff and is not awaited until the closables sweep — so it can still be running after the drain returns, and clearing the token there left it able to stall teardown (and every concurrent open on the path) indefinitely. Keep the token armed for finishClose()'s whole duration instead, and have the close-time compact-on-close pass opt out via a new compactRange() `cancellable` param rather than relying on the shared flag being cleared. Transaction::GetCount now takes an OperationGuard and checks isClosing() before scanning: without it, finishClose()'s drain can return immediately and the closables sweep can roll back the transaction while the count scan is parked between rows, reading freed memory. Carries the in-progress PR #787 lifecycle repair plan describing the fuller atomic-admission fix these two changes are a first slice of. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013CvCKcGKG7vyhRts1Mv4Wh
542b058 to
3cdd9f9
Compare
…GetCount against concurrent close finishClose() cleared compactCancelRequested right after the operationsInFlight drain, but an async compact() releases its OperationGuard at setup handoff and is not awaited until the closables sweep — so it can still be running after the drain returns, and clearing the token there left it able to stall teardown (and every concurrent open on the path) indefinitely. Keep the token armed for finishClose()'s whole duration instead, and have the close-time compact-on-close pass opt out via a new compactRange() `cancellable` param rather than relying on the shared flag being cleared. Transaction::GetCount now takes an OperationGuard and checks isClosing() before scanning: without it, finishClose()'s drain can return immediately and the closables sweep can roll back the transaction while the count scan is parked between rows, reading freed memory. Carries the in-progress PR #787 lifecycle repair plan describing the fuller atomic-admission fix these two changes are a first slice of. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013CvCKcGKG7vyhRts1Mv4Wh
107b216 to
ab5cd96
Compare
cb1kenobi
left a comment
There was a problem hiding this comment.
Please rebase with main and resolve the merge conflicts.
…GetCount against concurrent close finishClose() cleared compactCancelRequested right after the operationsInFlight drain, but an async compact() releases its OperationGuard at setup handoff and is not awaited until the closables sweep — so it can still be running after the drain returns, and clearing the token there left it able to stall teardown (and every concurrent open on the path) indefinitely. Keep the token armed for finishClose()'s whole duration instead, and have the close-time compact-on-close pass opt out via a new compactRange() `cancellable` param rather than relying on the shared flag being cleared. Transaction::GetCount now takes an OperationGuard and checks isClosing() before scanning: without it, finishClose()'s drain can return immediately and the closables sweep can roll back the transaction while the count scan is parked between rows, reading freed memory. Carries the in-progress PR #787 lifecycle repair plan describing the fuller atomic-admission fix these two changes are a first slice of. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013CvCKcGKG7vyhRts1Mv4Wh
…ng, destroy flush waste - CloseDB and the iterator finalizer detached from `closables` before close()/Reset(), leaving a concurrent destroy()/shutdown() sweep unable to see (and wait for) a handle/iterator still draining async work or mid-Next() -- a real use-after-free window on the shared rocksdb::DB. Detach after close() returns instead; closeMutex/iteratorMutex already serialize a foreign close arriving in that window. - DBHandle::close() released `logRefs` napi_refs gated on a std::thread::id equality check, the same recycled-pthread-id hazard invariant 18 already fixed for TransactionHandle. Add DBRegistry::ReleaseLogRefsByEnv, wired into the env cleanup hook like CloseTransactionsByEnv, so a dying env's logRefs are emptied while it is still alive -- before its thread id could ever be reused against the stale guard. - OpenDB's "still closing"/"retry in progress" waits parked on one matching descriptor's condition variable but re-scanned the whole path in their predicate, so two descriptors closing on one path (e.g. a writable and a secondary) could leave an opener asleep for the full lifecycleWaitSeconds even after the path was free. Track the specific selected entry instead. - finishClose(destroying=true) still ran a full flush + close-time compaction before the files are unlinked -- wasted I/O, and with allow_write_stall's default a destroy() of a stalled database could hang indefinitely while holding the destroyingPaths gate. Skip both when destroying. - Add a fixture covering a lifecycleWaitSeconds timeout actually firing and the path recovering afterward (previously only config validation was tested). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01122SfiNfCtiLgWvQTD6wZ7
…eal two-descriptor regression test - fork-lifecycle-timeout.mts raced the shutdown retry claim against the main thread's open() with no barrier, so under load the open could win and throw the wrong error instead of timing out -- a flake, not a proof. Expose `closeRetrying` on registryStatus() and poll for it before racing the open. - fork-compact-cancel-destroy.mts lost its discriminating power once destroy() skips compactOnClose (previous commit): with nothing left to block on compactMutex, an early vs. late cancellation arm became timing-indistinguishable. Drive it through shutdown() instead, which still runs compactOnClose. Verified: reverting the early arm now makes the fixture fail again (7.7s vs the ~500ms bound), confirming this restores the fixture's purpose. - Added a genuine two-descriptor regression test for the OpenDB condition-variable/predicate fix (db_registry.cpp:603/:655): a writable and a read-only descriptor quarantine, then retry concurrently under one shutdown() call, with an opener racing in. Verified against the pre-fix predicate: 4/4 runs stalled to the full 8s deadline instead of the expected ~3s, confirming this catches the regression the prior test explicitly could not. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01122SfiNfCtiLgWvQTD6wZ7
shutdown() is process-wide, not scoped to one path, and its worker call could still be re-scanning for anything left to close after the racing open() call returned but before it posted shutdownResult -- a handle opened in that window is not safe to read from or hold onto (shutdown() could sweep and force-close it too). Keep the timing-only open (the actual regression proof) but defer the data-preservation check to a fresh open, strictly after the worker's shutdown() call and the worker itself are both confirmed done. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01122SfiNfCtiLgWvQTD6wZ7
Review follow-up on the `shutdown()`-throws thread: the behavior is kept
(db.close() throws the identical quarantine error, and three of the four
throw sites are lifecycle timeouts that `database:closeFailed` cannot
carry), but the only place it was written down was the README. The
exported symbol now carries the same contract, including why a
`process.on('exit')` listener has to wrap the call: a throw from an exit
listener skips every listener registered after it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…eams
Independent pre-push review round 1 on the rebased head (codex + gemini +
harper-domain) found five real items:
- `DBHandle::open()` now drops the previous lifecycle's `logRefs`. The
owner-thread guard that makes a foreign close napi-safe (AGENTS.md
invariant 18) also means a cross-env `destroy()`/`shutdown()` leaves the
cache populated, so a reopened handle handed `useLog()` back a
`TransactionLog` whose store `weak_ptr` pointed at the unregistered store
of the closed lifecycle. Only `addEntry` re-resolves, so every read
accessor reported an empty log — `getLogFileSize()` returned 0 instead of
31 in the new fixture, which fails 1/1 without the fix.
- `registryStatus()`'s destroy-cleanup tombstone branch omitted
`transactionDetails`, which `RegistryStatusDB` declares non-optional; a
monitor reading `entry.transactionDetails.length` threw on exactly the
entry shape that only appears when a physical destroy failed.
- The three test-delay seams this PR added (`ROCKSDB_JS_BACKUP_DELAY_MS`,
`ROCKSDB_JS_DESTROY_DELAY_MS`, `ROCKSDB_JS_CLOSE_RETRY_DELAY_MS`) still
called `::getenv` from a libuv worker / arbitrary teardown thread. They
are snapshotted in `initializeTestSeams()` now, like the fault flags
beside them.
- `fork-shutdown-retry.mts` raced the worker's retry claim: the worker posts
before calling `shutdown()`, so the parent's open could hit the still
quarantined entry. It now polls `closeRetrying` first and re-reads data
from a fresh handle after the worker is done — the two barriers the
two-descriptor fixture already had.
- `fork-compact-cancel-async.mts` had no `'message'` listener attached
across `await outcome`, so a `{destroyed:true}` posted in that window was
dropped and the fixture would hang to its timeout.
Declined, with reasons in the PR body: making `isClosing()` a relaxed load
(the reviewer's own estimate is that it is dwarfed by `iterator->Next()`,
and the flag is read under mutexes elsewhere), and the aggregate
comment-narration nit.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
CI turned red on macOS (Bun + Deno) with four quarantine tests failing on
`/private/var/...` vs `/var/...`: a tombstoned `registryStatus()` entry and
every `database:closeFailed` event reported the registry key's resolved
identity, so a caller matching either against the path it opened did not
recognize it. That is exactly what AGENTS.md invariant 19 forbids
("returning only the resolved identity breaks callers that match paths
against the spelling they supplied") and what `registryStatus()` already
does correctly for a live descriptor.
- `emitCloseFailures()` reports `descriptor->path`.
- `DBRegistryEntry::reportedPath` remembers the opening caller's spelling so
a destroy-cleanup tombstone — which has no descriptor left to ask — can
still report it, in `registryStatus()` and in its own emit.
Not caused by the rebase (the pre-rebase head carried identical code and was
green on the older macos-26-arm64 runner image); the new image's TMPDIR made
the latent mismatch reachable. `test/destroy.test.ts` now covers it on every
platform with an explicit symlink instead of depending on macOS's `/var`
link — verified to fail with `expected undefined to be true`, the same
assertion macOS CI produced, when either report site is reverted.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…tombstone Round 3 of the independent pre-push review (codex) found the one gap in the previous commit: a second `destroy()` after a failed physical cleanup starts with `reportedPath = identityPath` and its scan skipped the existing tombstone, because that entry has no descriptor. So `registryStatus().path` kept reporting the opened spelling while the retry's `database:closeFailed` reverted to the resolved identity — the two disagreeing is the same defect one step later. The scan now falls back to a matching entry's remembered spelling, still preferring a live descriptor's. The symlink test asserts the event path on the retry as well, and fails with `expected [ …(2) ] to match object [ …(2) ]` without the fallback. Also dropped from that round, verified rather than assumed: Gemini's blocker on `checkpoint.cpp` leaking `operationsInFlight` when admission or queueing fails — `CheckpointInFlightClaim` is the RAII equivalent of `BackupInFlightClaim` and `handedOff` is only set after both succeed (`src/binding/database/checkpoint.cpp:69-76,114-115,206`). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
CI's remaining SIGSEGV (`fork-destroy-open.mts`, one Deno/Bun job per run since the rebase) is a real native use-after-free on the JS thread, not a flake. Reproduced locally at 2/60 by running that fixture 6-way concurrently, and gdb put it exactly here: __strlen_avx2 v8::String::NewFromUtf8 napi_set_named_property rocksdb_js::DBRegistry::RegistryStatus `registryStatus()` holds `databasesMutex`, which covers the registry map but not a descriptor's own maps. It walked `descriptor->columns` — guarded by `columnsMutex` — while a cross-env `destroy()`'s `finishClose()` cleared that map from the worker thread, so `name.c_str()` pointed into a freed map node and `napi_set_named_property()` `strlen()`ed it. `locks.size()` was read unguarded the same way (a count, so a torn read rather than a fault). The column summary is now snapshotted under `columnsMutex` (plus the per-CF `userSharedBuffersMutex` for its buffer count) and the JS values are built after releasing it — holding it across the N-API calls would risk a finalizer re-entering the same non-recursive mutex on this thread, which is why `transactions` already had this shape under `txnsMutex`. `locks.size()` is read under `locksMutex`. Neither adds a lock-order inversion: `OpenDB` and `CollectWriteBufferManagerInventory` already establish `databasesMutex → columnsMutex`, `getUserSharedBuffer` is a leaf, and no `locksMutex` region reaches the registry. Pre-existing on `main` (its `registryStatus()` walks `columns` unguarded too); this PR is what makes it reachable, because `destroy()` now force-closes every descriptor and its own fixtures poll `registryStatus()` across that window. `test/fixtures/fork-registry-status-column-race.mts` makes it deterministic rather than leaving it to CI luck: a worker churns `dropSync()` against the shared descriptor while the main thread polls `registryStatus()`, with the new `ROCKSDB_JS_REGISTRY_STATUS_COLUMNS_DELAY_MS` seam parking the walk per column family so an erase lands inside it. Column names run past libstdc++'s 15-char small-string buffer so the erase frees a separate heap allocation (an SSO name usually survives the free intact and hides the bug). 5/5 abort without the fix, 3/3 clean with it; the concurrent `fork-destroy-open.mts` stress went 2/60 → 0/132. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TMokk8DsJyGwpHz4Lmts85
…iming Round 5 of the independent pre-push review (codex, seconded by the domain adjudicator) caught `test/fixtures/fork-registry-status-destroy-race.mts` shipping unreferenced: it was the first attempt at the `registryStatus()` column-walk regression test and does not reproduce (0/7 against the unfixed walk), because a single close-time `columns.clear()` rarely lands inside a walk and a small-string column name usually survives the free intact. The drop-churn fixture that replaced it does reproduce deterministically (5/5), and it covers the same thing — the fix is the snapshot in the walk, not anything per-mutator, so reverting that snapshot fails the churn fixture too. Also reworded the tombstone branch's comment, which claimed it fills "every non-optional field of RegistryStatusDB": `userSharedBuffers` is declared non-optional and set by no branch, and `columnFamilies` is typed `string[]` against an object. Both predate this PR and are noted rather than fixed here. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TMokk8DsJyGwpHz4Lmts85
Co-Authored-By: GPT-5 Codex <noreply@openai.com>
Co-Authored-By: GPT-5 Codex <noreply@openai.com>
Co-Authored-By: GPT-5 Codex <noreply@openai.com>
5e8803d to
8784294
Compare
Co-Authored-By: GPT-5 Codex <noreply@openai.com>
Co-Authored-By: GPT-5 Codex <noreply@openai.com>
cb1kenobi
left a comment
There was a problem hiding this comment.
This was a big one, but it looks great! I love the OperationGuard!
Brings the branch up to `main` (through #866) and reconciles it with #850, which landed the deferred column-family reclamation work in the same lifecycle files. Nine files conflicted; the resolutions keep both sides' intent and, where the two branches had grown the same mechanism twice, keep one: - `DBRegistry::OpenDB`: #850's outer reclaim-retry loop now wraps this branch's destroy/quarantine/close-retry wait loop, and the retiring- generation reclaim runs inside the column block. The handle is still adopted, cancellation-cleared and attached under `databasesMutex` (invariant 28), so `DBHandleParams` stays unused and is removed. - `Transaction::Commit`: main's `PendingTransactionCommitState` guard subsumes this branch's hand-rolled unwind, so the async-work refusal and queue-failure paths reject and let the guard release the claim, delete the async work and restore the handle to `Pending` — instead of `admitAsyncWorkOrReject`/`queueAsyncWorkOrReject`, whose `delete state` cannot coexist with the guard's ownership. `finishCommitCompletion()` became `completion->finish()` in b7a23da. - `DBDescriptor::finishClose`: keeps this branch's `cancelBlockingWork()` sweep first and its `closeWorkersStopped` retry gate, with main's newer per-env `CommitCompletion::release()` pass inside it. - `registryStatus()`: keeps this branch's value-only snapshot (which already fixes the `columns`/`transactions` races #850 addressed inline) and widens `TxnSummary::id` to the 64-bit transaction id from #853. - `ROCKSDB_JS_COMMIT_EXECUTE_DELAY_MS`: one seam, at main's call site (which also sets the `…DelayActive` flag), read from this branch's `initializeTestSeams()` snapshot rather than `::getenv`. - `drop.test.ts`: drops this branch's poisoned-environment assertions — #850 refuses the commit at admission, so #726 no longer poisons the database and main's "still writable" expectations are the correct ones. - `db_handle.cpp`: `commitCompletion.reset()` moves into `open()` beside `releaseLogRefs()`, both owner-thread-only pre-adoption resets. - AGENTS.md: main's three new invariants become 23/24/25 and this branch's three become 26/27/28; stale cross-references on both sides (including `unregisterColumnFamily` → `retireColumnFamily`) updated. Verification: `pnpm check` clean, `pnpm test:native` 234/234, `pnpm test` 1045 passed / 10 skipped. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Pre-push review (Claude advisory leg) flagged that `DBHandle::open()` read `this->descriptor->identityPath` after `DBRegistry::OpenDB()` returned, which contradicts invariant 28's "no handle shared-pointer field may be read after the registry lock drops". OpenDB already publishes every other descriptor-backed field under `databasesMutex`; `identityPath` now joins them. Not a live defect — a foreign close resets `descriptor` only on the owning thread, which is the thread inside `open()` — but the code and the invariant it cites have to agree, and this is the cheaper side to change. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Pre-push review (harper-domain, kept from cursor-grok's pass): `Transaction::GetCount` bound `txnDbHandle` as a reference to `TransactionHandle::dbHandle`, null-checked it, and then dereferenced `->descriptor` to build its OperationGuard. `TransactionHandle::close()` resets that member from whichever thread drives a forced teardown, with no owner-thread check, so a foreign destroy()/shutdown() landing between the two is a null dereference. A local shared_ptr copy closes the window and pins the handle for the guard's lifetime. Also drops a comment in `Database::CatchUpWithPrimary` that still described `DBHandle::close()`'s async-work drain as bounded with its failure ignored; this branch made that drain untimed (invariant 17). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…ebase main's own #787 landed invariants 26-28 in the same slot this branch's purge invariant occupied (25), so the rebase conflict resolution moved purge to 29 and its sibling databaseFlushed invariant (which collided with main's new 26) to 30. Fixed the three stale in-code "invariant 25" cross-references this displaced (AGENTS.md, transaction-log-reader.ts, transaction_log_store.cpp), following the whole-tree re-grep lesson already recorded in this PR's Findings. Dispatch-Task: pr-maint-68bd037170be49b2a01e194607ecae40 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
…g to read it (#820) * Release a purged segment's mapping instead of pinning it A purge unlinks the segment, which removes one link to an inode whose bytes a retired segment never changes again; a reader's MemoryMap is the other link, so the entries it mapped are still exactly the committed history and stay readable. The bug in HarperFast/harper#2337 was never that those bytes were served — it was that the mapping was never released, so the purge reclaimed no space: 16 MiB of a deleted .txnlog resident until restart. TransactionLog._currentLogBuffer, the fast path over the already-weak _logBuffers cache, held a strong reference and is only refreshed by query(), so a reader that calls query() once and next() forever (harper's audit subscription) froze it on whatever segment was current then. It is now a WeakRef: the mapping goes at the next GC once the iterator holding it moves on, with no purge-time invalidation and no cross-handle signalling. Also here, because they are the same reclaim path: - nextReadableLogBuffer() skips a run retention deleted when an iterator advances. Stopping at the hole stopped the iterator permanently, since every later poll stopped in the same place. _findPosition(0) names the oldest survivor, so a purged prefix costs one native call rather than a probe per segment, and only a run the store no longer has is skipped: a segment it still knows is merely unmappable for now, so iteration stops and retries. - readableExtent() bounds a read of a purged segment by its mapping, since the store reports no size for a segment it has forgotten. The 0 it reports dropped every entry the reader had not reached yet — including entries appended after it last polled, which the writer's overlay extension made visible in that same mapping. - removeFile() uses the non-throwing std::filesystem::remove overloads on both platforms; a Windows sharing violation used to unwind a C++ exception through the N-API purge boundary. - A segment that vanished between the purge's scan and its unlink is forgotten from sequenceFiles the way the scan forgets an already-missing one, and a segment that could not be deleted for a real reason is reported once per purge run via log.warn instead of silently stalling retention. Refs HarperFast/harper#2337 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Gate the txn.state rewrite on the file's existence, not the stream's is_open() databaseFlushed() keeps the flushed-state stream open across flushes, and a stream describes a descriptor, not a pathname: once txn.state (or the whole store directory) is unlinked, every write lands in the orphaned inode while getLastFlushedPosition(), which reads by path, returns the {0,0} sentinel and retention never advances. The pathname is now checked before the unchanged-position shortcut, since a flush resolving to the already-recorded position must still restore a missing file. The directory is recreated the way getLogFile() does, isClosing is re-checked under flushedStateMutex so a concurrent destroy cannot be resurrected, and the reopen is in-place (in | out) rather than truncating: the 8-byte record is overwritten whole, and a truncating reopen after a failed write would erase the last durable position before a retry that can fail again. The creating open is taken only after the file is verified absent. This runs on RocksDB's flush thread, where an escaping exception ends the process, so the whole rewrite sits behind a catch-all and every failure is reported once via log.warn, leaves lastWrittenFlushedPosition untouched, and is retried on the next flush. Hardening rather than a live bug: only purgeLogs({ destroy: true }) removes the directory today, and Harper does not call it in production. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Fix AGENTS formatting Co-Authored-By: GPT-5 Codex <noreply@openai.com> * Correct the oxfmt scope claim that misdiagnosed this branch's CI failure AGENTS.md said oxfmt formats "TS/JS/JSON only" and does "not touch C++ or Markdown". It does format Markdown, including AGENTS.md — that file is what the failing `Check` job named, and renumbering its ordered invariant list is what oxfmt objected to. Believing the doc is why the failure was first read as a hand-fixable numbering slip. Also record the trap the doc hid: a `pull_request` build formats the merge commit, so a branch whose own `fmt:check` is green fails CI whenever it and main have each appended an invariant and the numbers collide. Verified: oxfmt scans .md and skips .cpp entirely, renumbers `19, 19, 20` to `19, 20, 21`, and leaves lazy `1.` numbering alone. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GePs22dNhggr8XKDc1DThG * Resolve the next readable segment by successor, and stop a read recreating one Two independent defects in the purge-advance path, both found by the pre-push review of this branch. nextReadableLogBuffer() asked _findPosition(0) for "what comes after N". That walks backward from the current sequence and stops at the first gap, so it names the bottom of the contiguous run ending at the current segment. That is the oldest survivor only when the deletions form a single prefix. With a survivor between two holes - a purge({all}) that continued past a segment it could not unlink, or segments deleted out of band and registered that way at load - it lands past the survivor, and that segment's committed entries are never yielded to the reader. Silently: the skip has no signal. _nextLogId() (TransactionLogStore::nextSequenceAfter, sequenceFiles.upper_bound) answers the question actually being asked, in one O(log n) lookup. The probe is a bounded loop rather than one hop, because a registered successor can be absent too - unlinked out of band, or by another process's retention, before this process's purge run forgets it - and stopping at the first one that will not map is the same permanent wedge. It terminates because _nextLogId() strictly increases and is capped at the latest sequence. Separately, resolving a registered-but-closed segment opened it, and TransactionLogFile::open() creates (O_RDWR | O_CREAT, and OPEN_ALWAYS on Windows). Probing a segment that discovery registered but that is no longer on disk therefore recreated it as a header-only ghost that the next startup registers again. openIfPresent() skips a definite absence, using the same reasoning ensureExtent() already documents - only a definite absence skips, since a stat that errors leaves the extent unresolved - and is meaningful because dataSetsMutex is held across the check and the open. All three paths that resolve a segment go through it: getLogFileSize(), getMemoryMap(), and findPositionByTimestamp()'s backward walk, which skips to the previous sequence instead of opening. The walk's outcome is unchanged, since an absent file yielded position 0 and continued anyway; only the ghost goes away. An already-open segment is unaffected by any of this: its handle still describes the file, and on POSIX its unlinked inode is still exactly the committed history this branch exists to keep serving. The regression covers the successor lookup against a two-hole registry, a registered successor that is itself absent, and the absence of resurrection across all three resolve paths. It deliberately does not drive an iterator end to end: findPositionByTimestamp() keeps the backward-walk shape, so after a restart with holes no entry point positions a reader below one - measured _findPosition(0) = 5 and startFromLastFlushed = 5 on a 1/3/5 layout - and that shape is only reachable for an iterator already live when the holes appear. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GePs22dNhggr8XKDc1DThG * Bound the resync scan by the written extent, and descend holes when positioning Two follow-ups approved after review, both regressions this branch introduced or left behind. The resync bound at src/transaction-log-reader.ts:133 fell back to the raw mapping length when the store reports no extent for a purged segment. Invariant 11 forbids exactly that, and prior behaviour was dataEnd = 0 (no scan), so this branch introduced the failure rather than inheriting it: findResyncPosition tries every start offset, so a mapped-capacity bound byte-scans the whole pre-extended map on the JS thread, finds nothing (zeros never satisfy frameFits), and returns undefined - reporting a recoverable mid-log break as a torn tail, the harper#2016 amputation the invariant exists to prevent. It now uses the cached extent or readableExtent(), which walks the frames to the end-of-entries marker when the store has forgotten the segment. That does not cover every case, and the gap is worth naming: endOfEntries() deliberately returns the whole mapping when framing is broken, so that it never becomes the thing that decides a corrupt frame ends the log. For a purged segment whose framing is broken - the case that reaches corruptFrame in the first place - the bound is therefore still the mapped capacity. Every other purged-segment case is fixed and none is made worse, but choosing a bound when no authoritative extent exists (the store has forgotten the segment and the frame walk is defeated by the break) is a design question left open. findPositionByTimestamp() carried the same defect just fixed in nextReadableLogBuffer(), and a worse consequence. It stepped with sequenceFiles.find(--sequenceNumber), so the walk ended at the first missing sequence: after out-of-band deletion a reader asking for timestamp 0 silently received only the newest contiguous run while every older survivor sat registered, on disk, with a valid extent. It now descends by map order, and tracks the next registered sequence above the entry being examined so the two "the timestamp belongs further up" exits name a segment that exists rather than sequenceNumber + 1, which a hole may have removed. Only a segment that actually opened becomes that tracked sequence: both exits hand it back as a position to read from, so a registered-but-absent one would send the reader to a file that is not on disk. Fixing positioning is what makes the end-to-end case testable at all: before it, no entry point could put a reader below a hole, which is why the earlier regression asserted the successor primitive instead. The test now covers a three-wide gap - 2 and 4 never registered, 3 registered but absent - asserting query({start: 0}) yields [1, 5] where the old lookup started at 5 and yielded [5], plus a timestamp past every segment so the other exit is exercised rather than only the position-zero path. Its read order is load-bearing and says so: reading a segment's extent opens it, and an open handle keeps reporting the real size after an unlink (invariant 20), so nothing may touch a segment before it is meant to be gone. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GePs22dNhggr8XKDc1DThG * fix: preserve txnlog purge read boundaries Carry the append-owned readable extent with retained memory maps and invalidate stale flushed-state correlations when destructive purge empties a live store. Co-Authored-By: GPT-5 Codex <noreply@openai.com> * fix: harden txnlog purge safeguards Co-Authored-By: GPT-5 Codex <noreply@openai.com> * test: close txnlog review gaps Co-Authored-By: GPT-5 Codex <noreply@openai.com> * fix: close txnlog pathname races Use atomic no-create opens for read probes and verify flushed-state writes through their pathname so purge cannot leave ghost segments or stale retention state. Co-Authored-By: GPT-5 Codex <noreply@openai.com> * test: assert only the portable purge-vs-mapping contract The Windows-only "converge on a purge refused by a live mapping" test asserted that a live reader mapping refuses the unlink there. Windows CI showed both outcomes for that test across this branch's heads (refused at 9a8606a, removed at c8e790a) with no change to the mapping's lifetime in between, so the outcome of any single purge run is not something to assert on that platform. Replace it with a cross-platform test of the contract the code actually implements: the reader keeps every entry it mapped, and retention reclaims the segment once nothing maps it — immediately where the unlink lands, on the next run where it did not. It reads the mapping reference after the purge so V8 cannot collect it first, which is what made the old test's premise unverifiable. The three POSIX-only tests that assert a first-run deletion now say why they are skipped on Windows. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CVAGTbmYzkKxvuH4HQjXKd * test: force the refused-unlink path, and lock the error it reports The purge's refused-unlink branch (segment stays registered, one warn line per run, next run reclaims it) had no test that reached it: POSIX always unlinks and the Windows outcome is not predictable. An unwritable store directory fails the unlink with EACCES, which is the same branch a sharing violation takes, so the contract — including end-to-end delivery of the `log.warn` line — is now covered deterministically where permissions apply. `lastRemoveError` is a plain std::error_code written under fileMutex; the purge read it unlocked, which can tear its value/category pair against a concurrent retirement of the same file. Read it through a locked accessor. Also trims the `nextReadableLogBuffer` header to the two constraints the code cannot state itself; the rest restated invariant 22 verbatim. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CVAGTbmYzkKxvuH4HQjXKd * docs: renumber duplicate AGENTS.md invariant 23 to 24 The rebase onto main moved this branch's purge invariant from 22 to 23 (main appended its own 22, "queued unlock callback"), which collided with this branch's existing invariant 23 ("databaseFlushed persists and verifies txn.state by pathname"). Bump the latter to 24. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs: point purge-invariant cross-references at 23, not 22 Three comments still cited "invariant 22" for the purge/read-coherence rule after the rebase renumbered it to 23 (main's own new invariant 22, queued unlock callback, now occupies that number). Caught by the independent pre-push review on the rebased head. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix: preserve committed txnlog bounds across rotation Co-Authored-By: GPT-5 Codex <noreply@openai.com> * docs: renumber AGENTS.md invariants 23-24 to 24-25 after rebase main's own rebase-picked invariant 23 (deferred column-family drops, #850) now occupies the number this branch's purge invariant held. Move purge to 24 and databaseFlushed to 25, and fix the two source comments and the one AGENTS.md self-reference that cited the old number 23 for the purge invariant. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Dispatch-Task: pr-maint-68bd037170be49b2a01e194607ecae40 * fix: close the last two creating-open ghost-recreate paths, and drop a hot-path refcount Pre-push review (round 30, codex+cursor-grok+gemini+harper-domain) found two more places this PR's own non-creating-open fix missed: - TransactionLogStore::load() opened both the surviving current segment and each older segment scanned for a recovery boundary with the creating open(). A file deleted out of band between the directory scan and that open (this same load()) would silently come back as a header-only ghost, exactly the resurrection this PR closed on the read path (openIfPresent/openExisting). Both call sites now use openExisting() and treat a genuine absence like any other non-fatal open failure already handled there. - TransactionLogFile::publishReadableExtentLocked() runs on every append (the commit hot path) and copied a shared_ptr<MemoryMap> to read one field, paying two atomic refcount ops for no ownership need. Now takes memoryMap.get() and only falls back to frozenMapCache.lock() (which has no raw-pointer equivalent) when there is no live memoryMap. The refused-unlink-drops-current-segment-mapping finding from the same round is real (three independent legs converged on it) but reachable only through purgeLogs({ destroy: true }), which Harper does not run in production, and touches the destroy path this PR has deliberately left alone since the September 2 scope correction — recorded in the PR body for the human reviewer instead of fixed here. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Dispatch-Task: pr-maint-68bd037170be49b2a01e194607ecae40 * docs: correct frozen-map ownership in README, trim review-narration comments README claimed native code holds every file's map until purge/close; frozen (rotated) files are only weakly cached (frozenMapCache) and survive solely through JS Buffer references, independent of purge or close. Also corrects stats.memory.activeMaps: it counts only the current file's strongly-held map, not a frozen one kept alive by JS. Trims three comments added by the last commit that narrated this PR's own history ("this PR closed on the read path") instead of stating the invariant itself. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Dispatch-Task: pr-maint-68bd037170be49b2a01e194607ecae40 * docs: renumber duplicate AGENTS.md invariant 25, fix stale invariant-24 refs Rebasing onto main (which landed its own new invariant 24, the two process-wide clocks note) collided with this branch's purge invariant, also numbered 24. Resolved by keeping main's 24 and bumping the purge invariant to 25 — but that collided with this branch's own databaseFlushed invariant, already numbered 25 from an earlier rebase's renumbering. Bumped databaseFlushed to 26 and fixed the three "invariant 24" cross-references (AGENTS.md, transaction-log-reader.ts, transaction_log_store.cpp) that meant the purge invariant. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Dispatch-Task: pr-maint-68bd037170be49b2a01e194607ecae40 * docs: qualify frozen-map weak-ownership claim as POSIX-only The README's memory-map section said a frozen (rotated/purged) file's map is weakly held and excluded from stats.memory.activeMaps unconditionally. That's only true on POSIX. On Windows, TransactionLogFile::getMemoryMapLocked() never applies the weak-for- frozen optimization -- every frozen read re-creates the mapping and re-pins it strongly for the file's life (by design, per the comment above that function), and getStats() counts memoryMap into activeMaps regardless of platform. Surfaced by this rebase's pre-push review. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Dispatch-Task: pr-maint-68bd037170be49b2a01e194607ecae40 * docs: renumber purge invariant 25->29, databaseFlushed 26->30 after rebase main's own #787 landed invariants 26-28 in the same slot this branch's purge invariant occupied (25), so the rebase conflict resolution moved purge to 29 and its sibling databaseFlushed invariant (which collided with main's new 26) to 30. Fixed the three stale in-code "invariant 25" cross-references this displaced (AGENTS.md, transaction-log-reader.ts, transaction_log_store.cpp), following the whole-tree re-grep lesson already recorded in this PR's Findings. Dispatch-Task: pr-maint-68bd037170be49b2a01e194607ecae40 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs: fix a stale invariant number and a wrong function name in comments Pre-push review round 1 (codex) caught two factually wrong references that predate this rebase and were never corrected by earlier renumbering passes: - test/transaction-log.test.ts:3112 cited "invariant 20" (now the transaction-timestamp invariant) for a property that invariant 29 (the purge/mapping invariant) actually documents. - src/transaction-log-reader.ts's corruptFrame() docstring named `getLogFileSize` as the store-mutex-taking call being avoided, but the caller resolves the extent via `readableExtent()` (a lock-free atomic accessor), not `getLogFileSize`. Dispatch-Task: pr-maint-68bd037170be49b2a01e194607ecae40 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs: correct attribution in corruptFrame docstring Pre-push delta review caught that my prior fix (404b19ef) misattributed the readableExtent() resolution to "the caller" -- corruptFrame() itself calls it (conditionally, on a logBuffer.size cache miss) at line 134 for the readUncommitted branch. Describe the actual cache-then-native-fallback mechanism instead. Dispatch-Task: pr-maint-68bd037170be49b2a01e194607ecae40 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs: renumber purge invariant 30->30, databaseFlushed 30->31 after second rebase main advanced again mid-review (PR #868, "Keep user shared buffers for the life of the column family") and landed its own new invariant 29, colliding with this branch's purge invariant. Kept both (main's 29 stays, purge moves to 30, databaseFlushed to 31) and re-grepped the whole tree for stale "invariant 29" cross-references left behind by the collision, per the lesson already recorded in this PR's Findings. Dispatch-Task: pr-maint-68bd037170be49b2a01e194607ecae40 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs: trim redundant corruptFrame docstring Round 3 (domain) flagged the docstring as restating the cache-then- native-fallback pattern already visible in the code two lines below (logBuffer.size ?? readableExtent(logBuffer)) plus the existing inline comment above it. Cut it down to the one-line purpose statement. Dispatch-Task: pr-maint-68bd037170be49b2a01e194607ecae40 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> Co-authored-by: GPT-5 Codex <noreply@openai.com> Co-authored-by: Chris Barber <chris@harperdb.io>
Summary
Rebases this PR onto latest
main. While this was in flight,mainindependently landed an equivalent-but-different fix for the same destroy-vs-open race (2c2066aa,84157bf9): a resolved-identity, per-entry gate (DBKey{path, readOnly, secondaryPath}) instead of this PR's original path-onlydestroyingPathsset. Per Kris's direction, this rebase adopts main's gate rather than porting the PR's original design: the first commit skips the PR's now-duplicate commit and adapts every dependent commit onto main's per-entry/cross-key conditions, and the second carries forward the PR's quarantine/retry/compaction-cancellation work and its test fixtures on top of that.destroyingPathsitself is re-derived (not ported) as a separate, lock-free-compatible gate for the window between "registry entries erased" and "physical files deleted" — physical deletion runs withoutdatabasesMutexheld (it's I/O-bound), so main's design alone left that window ungated; this was found by test failures against the carried-forward fixtures, not assumed from the PR.This is the rocksdb-js root-cause fix for Harper PR #2169: it moves a lifecycle invariant Harper worked around in JavaScript (locking root opens) into the native registry, makes teardown failures observable/recoverable via quarantine +
database:closeFailed, and adds cancellation tokens so a manual compaction cannot block teardown indefinitely.Merged onto
mainthrough #866, consolidated with #850 — 2026-09-21This note supersedes the rebase and verification statements below.
mainmoved fromfb4ba092to
dcca4ed3; the load-bearing change in that range is#850, Defer physical column-family drops behind admitted commits,
which rewrote the same lifecycle files this PR does. The branch is brought up to date with a
merge, not a rebase: its history already contains one merge from
main,mainrestructureddb_registry.cppby 430 lines, and replaying 16 commits over that would have meant resolving thesame conflict sixteen times in the most delicate file here.
Nine files conflicted (20 hunks). Every resolution kept both sides' intent, and where the two
branches had grown the same mechanism twice, one of them is now gone.
Kept one mechanism instead of two
main'sPendingTransactionCommitStateguard already does everything this PR's admission/queue-failure paths did by hand — mark the
state completed, delete the async work, drop the refs, and (which the hand-rolled version did
not do) put the handle back to
Pending. Both failure paths inTransaction::Commitnowreject and return,
letting the guard unwind, rather than calling
admitAsyncWorkOrReject/queueAsyncWorkOrReject,whose
delete statecannot coexist with the guard's ownership. Those two helpers remain for theeight
database.cpp/backup.cpp/checkpoint.cppcall sites that have no such guard.finishCommitCompletion()becamecompletion->finish()inb7a23da2and is spelled that way now.registryStatus()locking. Defer physical column-family drops behind admitted commits #850 added per-map locking inline around the N-API construction;this PR's value-only snapshot, taken under each owning mutex and built into JS values after
every unlock, already covers the same race and additionally avoids the descriptor pin that can
strand a last-handle purge. The snapshot wins;
TxnSummary::idis widened to the 64-bittransaction id from fix(transaction): widen transaction ids to 64 bits so log writes survive past 2^31 #853.
ROCKSDB_JS_COMMIT_EXECUTE_DELAY_MS. Two copies existed —main's at the top of the executecallback (which also sets the
…DelayActiveflag its fixtures poll) and this PR's snapshot-readcopy inside the
elsearm. One remains, atmain's call site, reading this PR'sinitializeTestSeams()snapshot instead of::getenvon a worker thread.DBHandleParamsis deleted. Invariant 28 exists because returning it and attaching inDatabase::Open()left a teardown gap; the struct had been left behind unused.Reconciled, both sides kept
DBRegistry::OpenDB—Defer physical column-family drops behind admitted commits #850's outer reclaim-retry
for (;;)now wraps this PR's destroy-gate / quarantine /close-retry wait loop, and the retiring-generation reclaim sits inside the column block where
Defer physical column-family drops behind admitted commits #850 put it. Handle adoption, cancellation reset and attachment still happen under
databasesMutexbefore the function returns (invariant 28).DBDescriptor::finishClose— this PR'scancelBlockingWork()sweep stays first and itscloseWorkersStoppedretry gate stays, withmain's newer per-envCommitCompletion::release()pass inside it. Defer physical column-family drops behind admitted commits #850'sretryPendingReclaims(true)and theretiring-generation handle destruction run where Defer physical column-family drops behind admitted commits #850 placed them, ahead of
db.reset().db_handle.cpp—commitCompletion.reset()joinsreleaseLogRefs()inopen(); both areowner-thread-only resets of the outgoing descriptor's cached state, done before adoption.
Dropped, because #850 fixed the underlying bug
drop.test.ts's pessimistic-cross-family case no longer asserts a poisoned environment, a failingclose(), and adestroy()recovery. #850 refuses the commit at admission withERR_COLUMN_FAMILY_DROPPED, so #726 nolonger latches a background error and
main's "still writable afterwards" expectations are thecorrect ones.
Invariant renumbering
maingrew three invariants (queued unlock callbacks, column-family lifetime, the two clocks) andthis branch has three of its own. Merged:
main's become 23/24/25, this branch's become26/27/28. Every cross-reference on both sides was re-pointed, including the two that were
already stale on this branch, and
unregisterColumnFamily→retireColumnFamilyin the commentsthis branch owns.
Verification of the merge
pnpm check(type-check + lint + fmt) — clean.pnpm test:native— 234/234.pnpm test— 1045 passed, 10 skipped, 0 failed.destroy/drop/drop-deferred-reclamation/lifecycle/concurrent-teardownrun 3× —94/94 each time.
review), plus a delta round over the three fixes it produced. Round 1 recorded
independent=truewith Gemini + Cursor (Grok) + the Harper domain adjudicator; the Codex gradedleg failed with
Selected model is at capacity, so Claude ran as a same-family advisorypass that does not count as independent coverage. Findings ruled on below.
fork-destroy-open.mtssix-way concurrent, 30 runs — 0 failures.fork-open-attach-destroy.mtsfour-way concurrent, 12 runs — 0 failures.Second rebase: onto main's WriteBufferManager stall watchdog (#824)
mainmoved again (7213b98a, "Make a WriteBufferManager write stall observable") and this branch wasCONFLICTING. Four conflicts, all resolved keeping both sides:binding.cppShutdown— main's watchdogbegin…Shutdown()/join…Watchdog()bracket now wraps this PR'stry/catchGlobalEvents::Shutdown()afterDBRegistry::Shutdown()(a quarantining close emitsdatabase:closeFailed, which needs its listeners to still exist —test/fixtures/fork-shutdown-failure.mtsasserts the event), and the watchdog join must run before the throw, or a failedshutdown()leaves the 1 Hz thread alive. Same combination in the last-env cleanup hook.backup.cpp— both hunks were pure PR-side additions (BackupInFlightClaim, thenapi_cancelledin-flight decrement) that git could not place.db_registry.h— this PR'sCloseResult CloseDB(...)return type alongside main'sCollectWriteBufferManagerInventory.AGENTS.md— invariant renumbering. Main's WBM-stall invariant becomes 21; this PR's three become 22/23/24, with the two internal cross-references (see invariant 21→ 22,see invariant 22→ 23) updated. Main'sdroppedColumns.clear()infinishClose()and its two-argumentColumnFamilyDescriptorconstructor were carried into this PR's rewrittenfinishClose(bool destroying)/OpenDB().Fixed by the rebase's own independent pre-push review
Two full rounds against the rebased head (codex + gemini + harper-domain; Cursor legs are disabled on any diff that edits
AGENTS.md). Round 1 found five real items, all fixed; round 2 verified each against HEAD and produced no new actionable findings.ownerThreadIdguard that makes a foreign close napi-safe also means a cross-envdestroy()/shutdown()can no longer clearlogRefs, so a reopened handle handeduseLog()back aTransactionLogwhoseTransactionLogHandle::storeweak_ptrpointed at the unregistered store of the closed lifecycle. OnlyaddEntryre-resolves; every read accessor reported an empty log.DBHandle::open()now releases the cache on the owning thread beforeDBRegistry::OpenDB.test/fixtures/fork-foreign-close-log-cache.mtsholds the staleTransactionLogalive across the foreignshutdown()(the cache entry is a weaknapi_ref, so letting it be collected would mask the bug) and asserts log size, queried entry count, and object identity — it reports0 bytes, expected 31without the fix, 3/3 clean with it.transactionDetails—RegistryStatusDBdeclares it non-optional, so a monitor readingentry.transactionDetails.lengththrew on exactly the entry shape that only appears when a physical destroy failed, i.e. when the diagnostic is needed.::getenvoff the JS thread —ROCKSDB_JS_BACKUP_DELAY_MS(libuv worker),ROCKSDB_JS_DESTROY_DELAY_MSandROCKSDB_JS_CLOSE_RETRY_DELAY_MS(any teardown thread), all added by this PR, now snapshotted ininitializeTestSeams()like the fault flags beside them. AGENTS.md's entries for the three say so.fork-shutdown-retry.mtsraced the retry claim — the worker posts before callingshutdown(), so the parent's open could reach the still quarantined entry and fail with "previous close failed" instead of measuring the wait; and a handle opened while the process-wideshutdown()loop is still scanning can be force-closed before the data assertion reads it. Both barriers the two-descriptor fixture already had.fork-compact-cancel-async.mtscould drop the destroy result —await outcomeyields to the event loop with no'message'listener attached, so a{ destroyed: true }posted in that window was delivered to a zero-listener emitter and dropped, hanging the fixture to its runner timeout. The promise is now claimed before the await. The sync sibling is not affected (compactSync()blocks the loop, so the queued message is only delivered after the next listener attaches).Red CI caught a real path-identity defect
The first rebase push turned macOS CI red (Bun + Deno) with four quarantine tests failing on
/private/var/...vs/var/.... Root cause, not a test problem: a destroy-cleanup tombstone inregistryStatus()and everydatabase:closeFailedevent reported the registry key's resolved identity, so a caller matching either against the path it opened did not recognize it. That is exactly what AGENTS.md invariant 19 forbids ("returning only the resolved identity breaks callers that match paths against the spelling they supplied") and whatregistryStatus()already did correctly for a live descriptor.emitCloseFailures()reportsdescriptor->path.DBRegistryEntry::reportedPathremembers the opening caller's spelling, because a tombstone has no descriptor left to ask.DestroyDBcaptures it from the first claimed descriptor underdatabasesMutex, and a retry of a failed destroy — which finds only the tombstone — falls back to the spelling that entry already remembered (the round-3 review finding).Not caused by the rebase: the pre-rebase head carried identical code and was green on the older
macos-26-arm64runner image (20260728→20260831), whoseTMPDIRmade the latent mismatch reachable.test/destroy.test.tsnow covers it on every platform with an explicit symlink rather than depending on macOS's/varlink, and asserts both the initial failure and the retry. Each report site was reverted individually to confirm the test fails — withexpected undefined to be true, the same assertion macOS CI produced, andexpected [ …(2) ] to match object [ …(2) ]for the retry.A CI "flake" that was a real use-after-free
The first two rebase pushes each had exactly one job die with
SIGSEGVinfork-destroy-open.mts— a different runtime each time (Bun/ubuntu, then Deno/ubuntu), which reads like flake. It is not. Running that fixture six-way concurrently reproduced it locally at 2/60, and gdb named the frame:databasesMutexcovers the registry map, not a descriptor's own maps.registryStatus()walkeddescriptor->columns— guarded bycolumnsMutex— while a cross-envdestroy()'sfinishClose()cleared that map from the worker thread, soname.c_str()pointed into a freed map node andnapi_set_named_property()strlen()ed it.locks.size()was read the same way (a count, so a torn read rather than a fault). The column summary is now snapshotted undercolumnsMutex(plus the per-CFuserSharedBuffersMutexfor its buffer count) with the N-API values built after releasing it — holding it across those calls would risk a finalizer re-entering the same non-recursive mutex on this thread, which is whytransactionsalready had this shape undertxnsMutex.locks.size()is read underlocksMutex.No lock-order inversion:
OpenDBandCollectWriteBufferManagerInventoryalready establishdatabasesMutex → columnsMutex,getUserSharedBufferis a leaf, and nolocksMutexregion reaches the registry. Pre-existing onmain(itsregistryStatus()walkscolumnsunguarded too); this PR is what makes it reachable, becausedestroy()now force-closes every descriptor and the PR's own fixtures pollregistryStatus()across that window.test/fixtures/fork-registry-status-column-race.mtsmakes it deterministic instead of leaving it to CI luck: a worker churnsdropSync()against the shared descriptor while the main thread pollsregistryStatus(), with the newROCKSDB_JS_REGISTRY_STATUS_COLUMNS_DELAY_MSseam parking the walk per column family so an erase lands inside it. The column names run past libstdc++'s 15-char small-string buffer so the erase frees a separate heap allocation — an SSO name usually survives the free intact and hides the bug, which is why the first attempt at this fixture (driving the close-timecolumns.clear()instead) reproduced 0/7 and was deleted rather than shipped unreferenced. 5/5 abort without the fix, 3/3 clean with it, and the concurrentfork-destroy-open.mtsstress went 2/60 → 0/132. Reverting the snapshot fails the churn fixture, so the close-time path is covered by the same net.Review-thread adjudication
All three previously-unresolved threads were re-checked against the current code; none is left unanswered, and the one new thread is ruled on below.
shutdown()can abort process exit and skip later cleanup (new,binding.cpp) — claim verified, prescribed fix overruled. Details under## For the human reviewer.db_registry.cpp) — unchanged ruling: real, pre-existing (registry identity has always been the raw path string), and canonicalizing changes a user-visible identity contract. Wants its own PR.transaction.cpp) — unchangedruling: the window is real but is not specific to quarantine, and the suggested fix (reject admission when
descriptor->isClosing()) is what AGENTS.md invariant 18 exists to prevent. The fix that closes it — running the closables sweep before the flush infinishClose()— reorders the most delicate path here and wants its own change.Third rebase: onto main's tsdown bump
mainmoved again (dependabot's tsdown bump, merged as531af655). No conflicts — the only delta between the previous head's merge-base andorigin/mainwaspackage.json/pnpm-lock.yaml, andgit rebase origin/mainreplayed all 11 commits clean. Build,pnpm test:native(194/194), and the full Vitest suite (935 passed / 9 skipped) all pass post-rebase;pnpm fmt:check/lint/type-checkare clean.Independent pre-push review (round 25, full — a force-push is never an ancestor of the prior review): codex (graded) and Gemini both ran; the Harper-domain adjudicator hit its time budget and was killed (
SIGKILL, timeout) before it could rank/filter, so I adjudicated the raw output myself against the code:benchmark/setup.ts's teardownresolve()-then-throwsilently reports success: it conflates the serialization gate (activeBenchmark/promise, used only to sequence between benchmarks) with the current benchmark's own result (theasync setup()call's returned promise, which the trailingthrowgenuinely rejects —throws: trueis set exactly so vitest surfaces it). Not applied.db_registry.cpp:1029holdsdatabasesMutexacrossnapi_create_string_utf8calls inRegistryStatus(), and if V8 GC runs aDBHandlefinalizer synchronously mid-allocation, that finalizer'sDBRegistry::CloseDBwould re-lock the same non-recursive mutex on the same thread — checks out as a real hazard class (it's exactly what invariant 6'scolumnsMutex/txnsMutex/locksMutexnarrow-scoping exists to avoid for the other locks in this same function), but the outerdatabasesMutexhold is byte-for-byte unchanged fromorigin/main(confirmed viagit show origin/main:src/binding/database/db_registry.cpp) — it predates this PR and every prior round. Left alone as out-of-scope for a rebase; flagged separately for its own issue.For the human reviewer
Ruled on from the merge round's independent review
ReleaseLogRefsByEnvmisses a descriptor a foreigndestroy()alreadyerased. True as stated (the walk is over
instance->databases), and the crash it predicts isnot reachable. The last owner of that
DBHandleis thenapi_wrapfinalizer'sshared_ptr(
database.cpp);closablesholds only a weak reference, and that finalizer runs on the owning env's own threadwhile the env is still valid, so
~DBHandle→close()→releaseLogRefsLocked()deletes therefs against a live env exactly as in the non-destroyed case. The recycled-thread-id window
invariant 18 describes needs the handle to outlive its env; nothing here makes it. Not changed.
AsyncCatchUpStatewhen queueing fails. Factually wrong.Database::CatchUpWithPrimarydoes not callqueueAsyncWorkOrReject; it calls::napi_queue_async_workdirectly and, on failure, callsowned->deleteAsyncWork()and returnswith the
unique_ptrstill owning the state. One delete.destroy()while it holds thepath gate. Real, and it is this branch's known trade rather than a merge regression: invariant
16 records that close can still wedge on a stall and that fixing it must not flush into the
teardown race, and invariant 17 records why the drain cannot simply be bounded (a timed-out
drain reaches
db.reset()under a live flush). This is the bounded-vs-unbounded decisionalready listed for you below, now with a concrete reproduction sequence attached to it.
destroy(). Thequarantine text says "call
destroy()", andDatabase::Destroyrejects a read-only handle.Verified narrow rather than dismissed:
DBDescriptor::flush()returns OK immediately for aread-only descriptor, so the flush-failure quarantine the text was written for cannot happen
there; the only remaining source is a
WaitForCompactfailure on a database that runs nobackground compaction, and
shutdown()still retries the close. Flagged, not changed —conditionalizing that message is a product decision, not a merge fix.
Transaction::GetCountdereferenced a member reference aforeign teardown can null.
txnDbHandlewas bound as a reference toTransactionHandle::dbHandle, whichTransactionHandle::close()resets from whichever threaddrives a forced
destroy()/shutdown(), with no owner-thread check — so a reset landing betweenthe null check and the
->descriptorread is a null dereference. Fixed: a localshared_ptrcopy, which also pins the handle for the
OperationGuard's lifetime. The delta round thenpointed out, correctly, that copying a
shared_ptrthat another thread mayreset()isitself a race. That exposure is class-wide, not local:
dbHandleis read bare at about fifteensites across
transaction.cppandtransaction_handle.cpp, and the two candidate fixes —an owner-thread check on the reset, or a mutex around every read — collide with invariant 18
(
TransactionHandle::close()is deliberately napi-free and thread-agnostic precisely because arecycled
std::thread::idwas the Linux corrupting write). The copy removes thecheck-then-dereference window at the one site that was flagged; closing the class needs its own
change.
DBHandle::open()readdescriptor->identityPathafter the registrylock dropped, contradicting invariant 28. Fixed:
OpenDB()publishesidentityPathwith therest of the handle's descriptor-backed fields, and
Database::CatchUpWithPrimary's comment nolonger describes the async-work drain as bounded.
finishClose(destroying=true)still callsWaitForCompact(). It skips the flush and the close-time compaction when destroying but thenwaits out the whole pending compaction backlog, on the JS thread, holding the path gate — so a
destroy()after a bulk load can time out every concurrentopen(). The suggested fix(
CancelAllBackgroundWork(db, true)instead, ordered beforeTransactionLogStoreRegistry::Unregister) is plausible and the adjudicator rates its ownconfidence as moderate; it changes ordering on the path invariant 16 warns about, so it is
flagged, not changed in a merge commit. Worth its own PR.
the per-row iterator mutex, the comment-narration nit, the sticky-background-error quarantine,
and
shutdown()throwing).shutdown()now throws — the review claim is correct, and I kept the behavior. Your call to overrule me. The bot's facts check out, and I verified the Node semantics directly: a throw from aprocess.on('exit')listener skips everyexitlistener registered after it (unconditionally), and flips the exit code to 1 unless anuncaughtExceptionhandler is installed. Harper core calls it exactly that way —resources/RocksTransactionLogStore.ts:19,process.on('exit', () => shutdown()), bare, registered at module load — so on a close-time flush failure at exit it would skip laterexitlisteners, including ones registered by application components. Onmain,shutdown()genuinely never throws (descriptor->close()isvoidand the close-time flush status is discarded), so this is a behavior change, not a clarification.Three reasons I did not adopt "keep
shutdown()non-throwing":db.close()already throws the identical error (database.cpp:211, with the "Call shutdown() to retry close, or destroy() to delete the database" suffix). Makingshutdown()silent would have the two close entry points disagree about whether a failed close is an error.DBRegistry::Shutdown()'s four throw sites are lifecycle timeouts, not closefailures:
shutdownMutexcontention, the per-descriptor drain, and the destroy-in-flight wait.database:closeFailedcannot carry those — no descriptor failed. A blanket non-throwingshutdown()would return normally while databases are still open or files still being deleted, which is a worse silent failure than the one being avoided.database:closeFailed(no listener anywhere in the repo), so today the throw is the only channel that reaches it. Event-only reporting would make a close-time flush failure fully invisible there.What I did instead is make the contract explicit where a caller will see it: the README already documented the throw and the
try/catchexit-listener pattern, and the exportedshutdownnow carries the same JSDoc. The harper-side guard is still missing and is a one-line change atresources/RocksTransactionLogStore.ts:19; it needs to land with or before this, and it is outside this repo. If you would rather not ask that of consumers, the alternative is a return value plus a separate throwing API, and I'd want your direction before building it.Carried from the original PR body, still open (unchanged by either rebase):
../symlink alias can still bypass the destroy/open gate (narrower than it reads — RocksDB's ownLOCKfile blocks a second read-write open through an alias; a read-only alias during destroy is the live hazard). Wants its own PR — canonicalizing changes a user-visible identity contract (registryStatus().path, lock/backup file paths,TransactionLogStoreRegistrykeys).finishClose()) reorders the most delicate path here and wants its own change.--moduleRefCount == 0; a fresh env loading the module concurrently isn't coordinated against a concurrentShutdown()/Teardown(). Pre-existing.DBIteratorHandle::Next()'s unconditional per-rowiteratorMutexlock/unlock on the hottest read path. Estimated low single digits percent overhead against a several-hundred-ns per-row N-API cost — real, but the obvious fix (a relaxed-load gate) is insufficient (a closer could still free the iterator between the load and the mutex), and a correct reader/closer handshake plus apnpm benchrange-scan comparison is more than a rebase should take on.Declined this round, with reasons:
isClosing()as amemory_order_relaxedload (raised as a major by Gemini, downgraded to a nit by the adjudicator in both rounds, and independently scored a nit by Codex). Its own estimate is that theLDARis dwarfed byiterator->Next(), andgetKeysCount()is not a request hot path. The flag is also read under mutexes in the registry paths, so changing the default ordering of a widely-used accessor is not a local change.db_registry.cpp,db_descriptor.{h,cpp},async.h,db_handle.cpp,closable.h,binding.cppthat narrate implementation history or address the reviewer). The technical content is accurate; churning ten comment blocks across a 69-file lifecycle diff adds review surface for no behavior change. Better as a follow-up sweep over the whole file set at once.RegistryStatusDB.userSharedBuffersis declared non-optional but no branch ofregistryStatus()sets a top-level property of that name (it is per-column-family, which is what the README documents) — soentry.userSharedBuffers > 0silently readsundefined > 0. Verified pre-existing onmain; same forcolumnFamiliesbeing typedstring[]against an object. Wants its own change.TransactionLogretained across a foreign close and never reopened still reports an empty log to its holder. Unchanged frommain(which deleted the weak ref on close, leaving any retained object equally stale), and the open-time invalidation above fixes the reopen path, which is the reachable one.Remaining CI failure, not from this PR
test/txn-close-commit-uaf.test.tsaborts on Deno/macOS only. That is the pre-existing worker-env teardown abort the test itself names and already mitigates —const retry = process.versions.deno && process.platform === 'darwin' ? 1 : 0with a comment pointing at #746 — and it exhausted that single retry. It appeared on the same job before the path-spelling and use-after-free commits, and nothing in this PR touches that repro's path.Verification
pnpm check(type-check + lint + fmt) — clean.pnpm test:native— 194/194 passed.pnpm test(Vitest) — 933 passed, 9 skipped (expected Deno/GC-related skips), 0 failed.test/destroy.test.tsalone — 31/31, including the new foreign-close log-cache fixture.test/destroy.test.tsalone — 32/32, run 5× to check the new symlink test for flakiness.--full, forced by the rebase — the prior review is no longer an ancestor) found the five items above. Round 2 (delta) confirmed all five fixed against HEAD and surfaced no new actionable finding; its three surviving items are the two declined nits plus the pre-existinguserSharedBufferstype mismatch, and it recorded a correction to round 1 (which had claimed the tombstone branch omitteduserSharedBufferstoo — never a top-level field). Round 3, on the path-spelling fix, found the tombstone-retry gap, fixed above. Round 4 converged: the graded leg reported zero findings.RegistryStatusDBfields the tombstone branch fills (reworded); round 6 converged, with every item previously adjudicated.checkpoint.cppleakingoperationsInFlightwhen admission/queueing fails (CheckpointInFlightClaimis the RAII equivalent ofBackupInFlightClaim, andhandedOffis set only after both succeed), andbenchmark/setup.tsmaking concurrent workers share one database path (they already did before this PR — the declaration moved within the same scope chain, and sharing one database across workers is the point of those benchmarks);checkpoint.cpp'sCheckpointInFlightClaimagain;ReleaseLogRefsByEnv"permanently leaking"napi_refs for an erased descriptor (they are env-owned and reclaimed at env teardown, and once the entry is erased no foreignclose()can reach the handle throughclosables, which is the hazard invariant 18 exists for); and, for the third time, a per-call__cxa_guard_acquireon the seam flags (astd::atomic<int>with aconstexprconstructor is constant-initialized, so no guard is emitted — confirmed by inspecting the generated assembly in an earlier round).Refs #787
🤖 Generated with Claude Code
Final maintenance pass — 2026-09-09
This note supersedes earlier rebase, review-coverage, and verification statements above. The branch is rebased cleanly onto
17fceeef; PR headac31bdeacontains the final review-feedback fix.registryStatus()thread was valid: carrying ashared_ptr<DBDescriptor>beyonddatabasesMutexcould make a racing last-handle close see an extra owner and skip its purge with no release-side retry.RegistryStatusEntrynow contains only copied values, captured under the registry and owning child locks, and all N-API construction happens after unlocking. The audited registry-to-child ordering is documented beside each mutex, and invariant 6 records the ownership-pin failure mode.isClosing()violates invariant 18 and the safer closables-before-flush reorder needs a separate change; andshutdown()error reporting is a consumer-compatibility choice rather than a mechanical review fix.Changes
DBSettings::Configvalidates the new lifecycle wait bound stored in db_settings.h; database.ts documents close-failure events and carries read-only intent into destroy, while store.ts distinguishes closing from closed sync operations. db_options.h only updates the shifted invariant reference.Verification
pnpm check— clean (type-check, lint, formatting);git diff --check— clean.pnpm test:native— 194/194 passed.pnpm test— 956 passed, 9 skipped, 0 failed.87842947with a retained registry entry (followed by a teardownSIGSEGV) and passes on this head; the strengthened overlap assertion also passes.Framing-Verdict: chosen-approach-sound(Gemini). Final full review produced no new actionablefinding: Gemini repeated an adjudicated benchmark false positive and the existing comment nit; Claude failed at startup, Cursor was policy-pruned for the
AGENTS.mdedit, and the domain leg timed out. These are review-infrastructure coverage gaps, not test failures.Complexity: complicated
Origin — the dispatch brief this PR was written from
Serialize database destruction with concurrent opens
LIVE CONVERSATION about #787.
You are answering a person, in a thread, one turn at a time. Every turn:
each of your previous turns is in it. Read the PR/issue and the code as needed.
they are talking to you, and a status template is not an answer.
Each turn arrives as ASK (answer it, change nothing) or PERFORM (do it, then say what you did) —
the person chose which when they sent it, and the run's own prompt tells you which one this is.
Never infer it from the wording: an unrequested commit in the middle of a discussion and a polite
description of work that was supposed to happen are the two failures this exists to prevent.
Never mark a PR ready and never merge from this conversation.
Dispatch: task
chat-pr-rocksdb-js-787-kriszyp· queued by unknown · ran by claude/opus/low · worker kzyp-xps-1Review-Coverage: authored=claude; ran=codex,gemini; adjudicated=domain; blocked=cursor-grok(failed); declined=cursor-composer; rounds=27; full=1 @ 1947744
Human-Review-Need: 4 (decisions: quarantine-on-close-failure, unbounded-drains, backup-stream-on-libuv-pool, close-and-shutdown-throw, destroy-waits-for-copies, txn-dbhandle-reset-policy, lifecycle-wait-global-setting, pr-scope) @ 1947744