Skip to content

Serialize database destruction with concurrent opens - #787

Merged
kriszyp merged 20 commits into
mainfrom
kris/serialize-destroy-open
Sep 21, 2026
Merged

kriszyp merged 20 commits into
mainfrom
kris/serialize-destroy-open

Conversation

@kriszyp

@kriszyp kriszyp commented Aug 15, 2026 •

Copy link
Copy Markdown
Member

Summary

Rebases this PR onto latest main. While this was in flight, main independently landed an equivalent-but-different fix for the same destroy-vs-open race (2c2066aa, 84157bf9): a resolved-identity, per-entry gate (DBKey{path, readOnly, secondaryPath}) instead of this PR's original path-only destroyingPaths set. Per Kris's direction, this rebase adopts main's gate rather than porting the PR's original design: the first commit skips the PR's now-duplicate commit and adapts every dependent commit onto main's per-entry/cross-key conditions, and the second carries forward the PR's quarantine/retry/compaction-cancellation work and its test fixtures on top of that. destroyingPaths itself is re-derived (not ported) as a separate, lock-free-compatible gate for the window between "registry entries erased" and "physical files deleted" — physical deletion runs without databasesMutex held (it's I/O-bound), so main's design alone left that window ungated; this was found by test failures against the carried-forward fixtures, not assumed from the PR.

This is the rocksdb-js root-cause fix for Harper PR #2169: it moves a lifecycle invariant Harper worked around in JavaScript (locking root opens) into the native registry, makes teardown failures observable/recoverable via quarantine + database:closeFailed, and adds cancellation tokens so a manual compaction cannot block teardown indefinitely.

Merged onto main through #866, consolidated with #850 — 2026-09-21

This note supersedes the rebase and verification statements below. main moved from fb4ba092
to dcca4ed3; the load-bearing change in that range is
#850, Defer physical column-family drops behind admitted commits,
which rewrote the same lifecycle files this PR does. The branch is brought up to date with a
merge, not a rebase: its history already contains one merge from main, main restructured
db_registry.cpp by 430 lines, and replaying 16 commits over that would have meant resolving the
same conflict sixteen times in the most delicate file here.

Nine files conflicted (20 hunks). Every resolution kept both sides' intent, and where the two
branches had grown the same mechanism twice, one of them is now gone.

Kept one mechanism instead of two

  • Commit setup unwind. main's
    PendingTransactionCommitState
    guard already does everything this PR's admission/queue-failure paths did by hand — mark the
    state completed, delete the async work, drop the refs, and (which the hand-rolled version did
    not do) put the handle back to Pending. Both failure paths in Transaction::Commit now
    reject and return,
    letting the guard unwind, rather than calling admitAsyncWorkOrReject/queueAsyncWorkOrReject,
    whose delete state cannot coexist with the guard's ownership. Those two helpers remain for the
    eight database.cpp/backup.cpp/checkpoint.cpp call sites that have no such guard.
    finishCommitCompletion() became completion->finish() in b7a23da2 and is spelled that way now.
  • registryStatus() locking. Defer physical column-family drops behind admitted commits #850 added per-map locking inline around the N-API construction;
    this PR's value-only snapshot, taken under each owning mutex and built into JS values after
    every unlock, already covers the same race and additionally avoids the descriptor pin that can
    strand a last-handle purge. The snapshot wins; TxnSummary::id is widened to the 64-bit
    transaction id from fix(transaction): widen transaction ids to 64 bits so log writes survive past 2^31 #853.
  • ROCKSDB_JS_COMMIT_EXECUTE_DELAY_MS. Two copies existed — main's at the top of the execute
    callback (which also sets the …DelayActive flag its fixtures poll) and this PR's snapshot-read
    copy inside the else arm. One remains, at main's call site, reading this PR's
    initializeTestSeams() snapshot instead of ::getenv on a worker thread.
  • DBHandleParams is deleted. Invariant 28 exists because returning it and attaching in
    Database::Open() left a teardown gap; the struct had been left behind unused.

Reconciled, both sides kept

Dropped, because #850 fixed the underlying bug

drop.test.ts's pessimistic-cross-family case no longer asserts a poisoned environment, a failing
close(), and a destroy() recovery. #850 refuses the commit at admission with
ERR_COLUMN_FAMILY_DROPPED, so #726 no
longer latches a background error and main's "still writable afterwards" expectations are the
correct ones.

Invariant renumbering

main grew three invariants (queued unlock callbacks, column-family lifetime, the two clocks) and
this branch has three of its own. Merged: main's become 23/24/25, this branch's become
26/27/28. Every cross-reference on both sides was re-pointed, including the two that were
already stale on this branch, and unregisterColumnFamily → retireColumnFamily in the comments
this branch owns.

Verification of the merge

  • pnpm check (type-check + lint + fmt) — clean.
  • pnpm test:native — 234/234.
  • pnpm test — 1045 passed, 10 skipped, 0 failed.
  • destroy / drop / drop-deferred-reclamation / lifecycle / concurrent-teardown run 3× —
    94/94 each time.
  • Independent pre-push review, full round on the merge (a merge is not an ancestor of the prior
    review), plus a delta round over the three fixes it produced. Round 1 recorded
    independent=true with Gemini + Cursor (Grok) + the Harper domain adjudicator; the Codex graded
    leg failed with Selected model is at capacity, so Claude ran as a same-family advisory
    pass that does not count as independent coverage. Findings ruled on below.
  • fork-destroy-open.mts six-way concurrent, 30 runs — 0 failures.
    fork-open-attach-destroy.mts four-way concurrent, 12 runs — 0 failures.

Second rebase: onto main's WriteBufferManager stall watchdog (#824)

main moved again (7213b98a, "Make a WriteBufferManager write stall observable") and this branch was CONFLICTING. Four conflicts, all resolved keeping both sides:

  • binding.cpp Shutdown — main's watchdog begin…Shutdown()/join…Watchdog() bracket now wraps this PR's try/catch
    • throw. Two ordering constraints had to be reconciled: this PR moved GlobalEvents::Shutdown() after DBRegistry::Shutdown() (a quarantining close emits database:closeFailed, which needs its listeners to still exist — test/fixtures/fork-shutdown-failure.mts asserts the event), and the watchdog join must run before the throw, or a failed shutdown() leaves the 1 Hz thread alive. Same combination in the last-env cleanup hook.
  • backup.cpp — both hunks were pure PR-side additions (BackupInFlightClaim, the napi_cancelled in-flight decrement) that git could not place.
  • db_registry.h — this PR's CloseResult CloseDB(...) return type alongside main's CollectWriteBufferManagerInventory.
  • AGENTS.md — invariant renumbering. Main's WBM-stall invariant becomes 21; this PR's three become 22/23/24, with the two internal cross-references (see invariant 21 → 22, see invariant 22 → 23) updated. Main's droppedColumns.clear() in finishClose() and its two-argument ColumnFamilyDescriptor constructor were carried into this PR's rewritten finishClose(bool destroying) / OpenDB().

Fixed by the rebase's own independent pre-push review

Two full rounds against the rebased head (codex + gemini + harper-domain; Cursor legs are disabled on any diff that edits AGENTS.md). Round 1 found five real items, all fixed; round 2 verified each against HEAD and produced no new actionable findings.

  • Stale transaction-log cache survives a foreign close into the reopen — a regression from this PR's own invariant-18 fix. The ownerThreadId guard that makes a foreign close napi-safe also means a cross-env destroy()/shutdown() can no longer clear logRefs, so a reopened handle handed useLog() back a TransactionLog whose TransactionLogHandle::store weak_ptr pointed at the unregistered store of the closed lifecycle. Only addEntry re-resolves; every read accessor reported an empty log. DBHandle::open() now releases the cache on the owning thread before DBRegistry::OpenDB. test/fixtures/fork-foreign-close-log-cache.mts holds the stale TransactionLog alive across the foreign shutdown() (the cache entry is a weak napi_ref, so letting it be collected would mask the bug) and asserts log size, queried entry count, and object identity — it reports 0 bytes, expected 31 without the fix, 3/3 clean with it.
  • A destroy-cleanup tombstone omitted transactionDetails — RegistryStatusDB declares it non-optional, so a monitor reading entry.transactionDetails.length threw on exactly the entry shape that only appears when a physical destroy failed, i.e. when the diagnostic is needed.
  • Three test-delay seams still called ::getenv off the JS thread — ROCKSDB_JS_BACKUP_DELAY_MS (libuv worker), ROCKSDB_JS_DESTROY_DELAY_MS and ROCKSDB_JS_CLOSE_RETRY_DELAY_MS (any teardown thread), all added by this PR, now snapshotted in initializeTestSeams() like the fault flags beside them. AGENTS.md's entries for the three say so.
  • fork-shutdown-retry.mts raced the retry claim — the worker posts before calling shutdown(), so the parent's open could reach the still quarantined entry and fail with "previous close failed" instead of measuring the wait; and a handle opened while the process-wide shutdown() loop is still scanning can be force-closed before the data assertion reads it. Both barriers the two-descriptor fixture already had.
  • fork-compact-cancel-async.mts could drop the destroy result — await outcome yields to the event loop with no 'message' listener attached, so a { destroyed: true } posted in that window was delivered to a zero-listener emitter and dropped, hanging the fixture to its runner timeout. The promise is now claimed before the await. The sync sibling is not affected (compactSync() blocks the loop, so the queued message is only delivered after the next listener attaches).

Red CI caught a real path-identity defect

The first rebase push turned macOS CI red (Bun + Deno) with four quarantine tests failing on /private/var/... vs /var/.... Root cause, not a test problem: a destroy-cleanup tombstone in registryStatus() and every database:closeFailed event reported the registry key's resolved identity, so a caller matching either against the path it opened did not recognize it. That is exactly what AGENTS.md invariant 19 forbids ("returning only the resolved identity breaks callers that match paths against the spelling they supplied") and what registryStatus() already did correctly for a live descriptor.

  • emitCloseFailures() reports descriptor->path.
  • DBRegistryEntry::reportedPath remembers the opening caller's spelling, because a tombstone has no descriptor left to ask. DestroyDB captures it from the first claimed descriptor under databasesMutex, and a retry of a failed destroy — which finds only the tombstone — falls back to the spelling that entry already remembered (the round-3 review finding).

Not caused by the rebase: the pre-rebase head carried identical code and was green on the older macos-26-arm64 runner image (20260728 → 20260831), whose TMPDIR made the latent mismatch reachable. test/destroy.test.ts now covers it on every platform with an explicit symlink rather than depending on macOS's /var link, and asserts both the initial failure and the retry. Each report site was reverted individually to confirm the test fails — with expected undefined to be true, the same assertion macOS CI produced, and expected [ …(2) ] to match object [ …(2) ] for the retry.

A CI "flake" that was a real use-after-free

The first two rebase pushes each had exactly one job die with SIGSEGV in fork-destroy-open.mts — a different runtime each time (Bun/ubuntu, then Deno/ubuntu), which reads like flake. It is not. Running that fixture six-way concurrently reproduced it locally at 2/60, and gdb named the frame:

__strlen_avx2
v8::String::NewFromUtf8
napi_set_named_property
rocksdb_js::DBRegistry::RegistryStatus

databasesMutex covers the registry map, not a descriptor's own maps. registryStatus() walked descriptor->columns — guarded by columnsMutex — while a cross-env destroy()'s finishClose() cleared that map from the worker thread, so name.c_str() pointed into a freed map node and napi_set_named_property() strlen()ed it. locks.size() was read the same way (a count, so a torn read rather than a fault). The column summary is now snapshotted under columnsMutex (plus the per-CF userSharedBuffersMutex for its buffer count) with the N-API values built after releasing it — holding it across those calls would risk a finalizer re-entering the same non-recursive mutex on this thread, which is why transactions already had this shape under txnsMutex. locks.size() is read under locksMutex.

No lock-order inversion: OpenDB and CollectWriteBufferManagerInventory already establish databasesMutex → columnsMutex, getUserSharedBuffer is a leaf, and no locksMutex region reaches the registry. Pre-existing on main (its registryStatus() walks columns unguarded too); this PR is what makes it reachable, because destroy() now force-closes every descriptor and the PR's own fixtures poll registryStatus() across that window.

test/fixtures/fork-registry-status-column-race.mts makes it deterministic instead of leaving it to CI luck: a worker churns dropSync() against the shared descriptor while the main thread polls registryStatus(), with the new ROCKSDB_JS_REGISTRY_STATUS_COLUMNS_DELAY_MS seam parking the walk per column family so an erase lands inside it. The column names run past libstdc++'s 15-char small-string buffer so the erase frees a separate heap allocation — an SSO name usually survives the free intact and hides the bug, which is why the first attempt at this fixture (driving the close-time columns.clear() instead) reproduced 0/7 and was deleted rather than shipped unreferenced. 5/5 abort without the fix, 3/3 clean with it, and the concurrent fork-destroy-open.mts stress went 2/60 → 0/132. Reverting the snapshot fails the churn fixture, so the close-time path is covered by the same net.

Review-thread adjudication

All three previously-unresolved threads were re-checked against the current code; none is left unanswered, and the one new thread is ruled on below.

  • shutdown() can abort process exit and skip later cleanup (new, binding.cpp) — claim verified, prescribed fix overruled. Details under ## For the human reviewer.
  • Raw path aliases bypass destruction ownership (db_registry.cpp) — unchanged ruling: real, pre-existing (registry identity has always been the raw path string), and canonicalizing changes a user-visible identity contract. Wants its own PR.
  • Quarantined descriptors still admit transaction commits (transaction.cpp) — unchanged
    ruling: the window is real but is not specific to quarantine, and the suggested fix (reject admission when descriptor->isClosing()) is what AGENTS.md invariant 18 exists to prevent. The fix that closes it — running the closables sweep before the flush in finishClose() — reorders the most delicate path here and wants its own change.

Third rebase: onto main's tsdown bump

main moved again (dependabot's tsdown bump, merged as 531af655). No conflicts — the only delta between the previous head's merge-base and origin/main was package.json/pnpm-lock.yaml, and git rebase origin/main replayed all 11 commits clean. Build, pnpm test:native (194/194), and the full Vitest suite (935 passed / 9 skipped) all pass post-rebase; pnpm fmt:check/lint/ type-check are clean.

Independent pre-push review (round 25, full — a force-push is never an ancestor of the prior review): codex (graded) and Gemini both ran; the Harper-domain adjudicator hit its time budget and was killed (SIGKILL, timeout) before it could rank/filter, so I adjudicated the raw output myself against the code:

  • Codex's three "surviving findings" (iterator mutex cost, an atomic load's default memory order, a comment-style nit) are unchanged carries-forward from the ~24 prior rounds already reflected in this description — not new, not acted on.
  • Gemini reprised the shared-benchmark-db-path claim already dismissed above, plus a new minor claim that benchmark/setup.ts's teardown resolve()-then-throw silently reports success: it conflates the serialization gate (activeBenchmark/promise, used only to sequence between benchmarks) with the current benchmark's own result (the async setup() call's returned promise, which the trailing throw genuinely rejects — throws: true is set exactly so vitest surfaces it). Not applied.
  • Gemini's one new claim outside prior rounds — db_registry.cpp:1029 holds databasesMutex across napi_create_string_utf8 calls in RegistryStatus(), and if V8 GC runs a DBHandle finalizer synchronously mid-allocation, that finalizer's DBRegistry::CloseDB would re-lock the same non-recursive mutex on the same thread — checks out as a real hazard class (it's exactly what invariant 6's columnsMutex/txnsMutex/locksMutex narrow-scoping exists to avoid for the other locks in this same function), but the outer databasesMutex hold is byte-for-byte unchanged from origin/main (confirmed via git show origin/main:src/binding/database/db_registry.cpp) — it predates this PR and every prior round. Left alone as out-of-scope for a rebase; flagged separately for its own issue.

For the human reviewer

Ruled on from the merge round's independent review

  • Gemini, blocker — ReleaseLogRefsByEnv misses a descriptor a foreign destroy() already
    erased.
    True as stated (the walk is over instance->databases), and the crash it predicts is
    not reachable. The last owner of that DBHandle is the napi_wrap finalizer's shared_ptr
    (database.cpp);
    closables holds only a weak reference, and that finalizer runs on the owning env's own thread
    while the env is still valid, so ~DBHandle → close() → releaseLogRefsLocked() deletes the
    refs against a live env exactly as in the non-destroyed case. The recycled-thread-id window
    invariant 18 describes needs the handle to outlive its env; nothing here makes it. Not changed.
  • Gemini, major — double free of AsyncCatchUpState when queueing fails. Factually wrong.
    Database::CatchUpWithPrimary does not call queueAsyncWorkOrReject; it calls
    ::napi_queue_async_work directly and, on failure, calls owned->deleteAsyncWork() and returns
    with the unique_ptr still owning the state. One delete.
  • Cursor/Grok, blocker — a WriteBufferManager stall can wedge destroy() while it holds the
    path gate.
    Real, and it is this branch's known trade rather than a merge regression: invariant
    16 records that close can still wedge on a stall and that fixing it must not flush into the
    teardown race, and invariant 17 records why the drain cannot simply be bounded (a timed-out
    drain reaches db.reset() under a live flush). This is the bounded-vs-unbounded decision
    already listed for you below, now with a concrete reproduction sequence attached to it.
  • Cursor/Grok, significant — a quarantined read-only path has no in-process destroy(). The
    quarantine text says "call destroy()", and Database::Destroy rejects a read-only handle.
    Verified narrow rather than dismissed: DBDescriptor::flush() returns OK immediately for a
    read-only descriptor, so the flush-failure quarantine the text was written for cannot happen
    there; the only remaining source is a WaitForCompact failure on a database that runs no
    background compaction, and shutdown() still retries the close. Flagged, not changed —
    conditionalizing that message is a product decision, not a merge fix.
  • Harper domain adjudicator, kept — Transaction::GetCount dereferenced a member reference a
    foreign teardown can null.
    txnDbHandle was bound as a reference to
    TransactionHandle::dbHandle, which TransactionHandle::close() resets from whichever thread
    drives a forced destroy()/shutdown(), with no owner-thread check — so a reset landing between
    the null check and the ->descriptor read is a null dereference. Fixed: a local shared_ptr
    copy, which also pins the handle for the OperationGuard's lifetime. The delta round then
    pointed out, correctly, that copying a shared_ptr that another thread may reset() is
    itself a race. That exposure is class-wide, not local: dbHandle is read bare at about fifteen
    sites across transaction.cpp and transaction_handle.cpp, and the two candidate fixes —
    an owner-thread check on the reset, or a mutex around every read — collide with invariant 18
    (TransactionHandle::close() is deliberately napi-free and thread-agnostic precisely because a
    recycled std::thread::id was the Linux corrupting write). The copy removes the
    check-then-dereference window at the one site that was flagged; closing the class needs its own
    change.
  • Claude (advisory leg) — DBHandle::open() read descriptor->identityPath after the registry
    lock dropped
    , contradicting invariant 28. Fixed: OpenDB() publishes identityPath with the
    rest of the handle's descriptor-backed fields, and Database::CatchUpWithPrimary's comment no
    longer describes the async-work drain as bounded.
  • Harper domain adjudicator, kept — finishClose(destroying=true) still calls
    WaitForCompact().
    It skips the flush and the close-time compaction when destroying but then
    waits out the whole pending compaction backlog, on the JS thread, holding the path gate — so a
    destroy() after a bulk load can time out every concurrent open(). The suggested fix
    (CancelAllBackgroundWork(db, true) instead, ordered before
    TransactionLogStoreRegistry::Unregister) is plausible and the adjudicator rates its own
    confidence as moderate; it changes ordering on the path invariant 16 warns about, so it is
    flagged, not changed in a merge commit. Worth its own PR.
  • Everything else the round produced repeats items already adjudicated above (the untimed drain,
    the per-row iterator mutex, the comment-narration nit, the sticky-background-error quarantine,
    and shutdown() throwing).

shutdown() now throws — the review claim is correct, and I kept the behavior. Your call to overrule me. The bot's facts check out, and I verified the Node semantics directly: a throw from a process.on('exit') listener skips every exit listener registered after it (unconditionally), and flips the exit code to 1 unless an uncaughtException handler is installed. Harper core calls it exactly that way — resources/RocksTransactionLogStore.ts:19, process.on('exit', () => shutdown()), bare, registered at module load — so on a close-time flush failure at exit it would skip later exit listeners, including ones registered by application components. On main, shutdown() genuinely never throws (descriptor->close() is void and the close-time flush status is discarded), so this is a behavior change, not a clarification.

Three reasons I did not adopt "keep shutdown() non-throwing":

  1. db.close() already throws the identical error (database.cpp:211, with the "Call shutdown() to retry close, or destroy() to delete the database" suffix). Making shutdown() silent would have the two close entry points disagree about whether a failed close is an error.
  2. Three of DBRegistry::Shutdown()'s four throw sites are lifecycle timeouts, not close
    failures: shutdownMutex contention, the per-descriptor drain, and the destroy-in-flight wait. database:closeFailed cannot carry those — no descriptor failed. A blanket non-throwing shutdown() would return normally while databases are still open or files still being deleted, which is a worse silent failure than the one being avoided.
  3. Harper does not listen for database:closeFailed (no listener anywhere in the repo), so today the throw is the only channel that reaches it. Event-only reporting would make a close-time flush failure fully invisible there.

What I did instead is make the contract explicit where a caller will see it: the README already documented the throw and the try/catch exit-listener pattern, and the exported shutdown now carries the same JSDoc. The harper-side guard is still missing and is a one-line change at resources/RocksTransactionLogStore.ts:19; it needs to land with or before this, and it is outside this repo. If you would rather not ask that of consumers, the alternative is a return value plus a separate throwing API, and I'd want your direction before building it.

Carried from the original PR body, still open (unchanged by either rebase):

  • Registry identity compares raw path strings, so a ../symlink alias can still bypass the destroy/open gate (narrower than it reads — RocksDB's own LOCK file blocks a second read-write open through an alias; a read-only alias during destroy is the live hazard). Wants its own PR — canonicalizing changes a user-visible identity contract (registryStatus().path, lock/backup file paths, TransactionLogStoreRegistry keys).
  • A transaction can still be admitted against a quarantined descriptor via the legacy libuv commit path, since the closables sweep never ran for that descriptor. The fix (running the closables sweep before the flush in finishClose()) reorders the most delicate path here and wants its own change.
  • Last-env module cleanup still keys off an unsynchronized --moduleRefCount == 0; a fresh env loading the module concurrently isn't coordinated against a concurrent Shutdown()/Teardown(). Pre-existing.
  • DBIteratorHandle::Next()'s unconditional per-row iteratorMutex lock/unlock on the hottest read path. Estimated low single digits percent overhead against a several-hundred-ns per-row N-API cost — real, but the obvious fix (a relaxed-load gate) is insufficient (a closer could still free the iterator between the load and the mutex), and a correct reader/closer handshake plus a pnpm bench range-scan comparison is more than a rebase should take on.

Declined this round, with reasons:

  • isClosing() as a memory_order_relaxed load (raised as a major by Gemini, downgraded to a nit by the adjudicator in both rounds, and independently scored a nit by Codex). Its own estimate is that the LDAR is dwarfed by iterator->Next(), and getKeysCount() is not a request hot path. The flag is also read under mutexes in the registry paths, so changing the default ordering of a widely-used accessor is not a local change.
  • The aggregate comment-narration nit (a handful of blocks in db_registry.cpp, db_descriptor.{h,cpp}, async.h, db_handle.cpp, closable.h, binding.cpp that narrate implementation history or address the reviewer). The technical content is accurate; churning ten comment blocks across a 69-file lifecycle diff adds review surface for no behavior change. Better as a follow-up sweep over the whole file set at once.
  • RegistryStatusDB.userSharedBuffers is declared non-optional but no branch of registryStatus() sets a top-level property of that name (it is per-column-family, which is what the README documents) — so entry.userSharedBuffers > 0 silently reads undefined > 0. Verified pre-existing on main; same for columnFamilies being typed string[] against an object. Wants its own change.
  • A cached TransactionLog retained across a foreign close and never reopened still reports an empty log to its holder. Unchanged from main (which deleted the weak ref on close, leaving any retained object equally stale), and the open-time invalidation above fixes the reopen path, which is the reachable one.

Remaining CI failure, not from this PR

test/txn-close-commit-uaf.test.ts aborts on Deno/macOS only. That is the pre-existing worker-env teardown abort the test itself names and already mitigates — const retry = process.versions.deno && process.platform === 'darwin' ? 1 : 0 with a comment pointing at #746 — and it exhausted that single retry. It appeared on the same job before the path-spelling and use-after-free commits, and nothing in this PR touches that repro's path.

Verification

  • pnpm check (type-check + lint + fmt) — clean.
  • pnpm test:native — 194/194 passed.
  • pnpm test (Vitest) — 933 passed, 9 skipped (expected Deno/GC-related skips), 0 failed.
  • test/destroy.test.ts alone — 31/31, including the new foreign-close log-cache fixture.
  • test/destroy.test.ts alone — 32/32, run 5× to check the new symlink test for flakiness.
  • Independent pre-push review, six rounds on the rebased head. Round 1 (--full, forced by the rebase — the prior review is no longer an ancestor) found the five items above. Round 2 (delta) confirmed all five fixed against HEAD and surfaced no new actionable finding; its three surviving items are the two declined nits plus the pre-existing userSharedBuffers type mismatch, and it recorded a correction to round 1 (which had claimed the tombstone branch omitted userSharedBuffers too — never a top-level field). Round 3, on the path-spelling fix, found the tombstone-retry gap, fixed above. Round 4 converged: the graded leg reported zero findings.
  • Rounds 5 and 6 covered the use-after-free fix above: round 5 found the unreferenced fixture (removed) and a comment of mine that overclaimed which RegistryStatusDB fields the tombstone branch fills (reworded); round 6 converged, with every item previously adjudicated.
  • Gemini claims checked and rejected rather than applied, across the rounds: checkpoint.cpp leaking operationsInFlight when admission/queueing fails (CheckpointInFlightClaim is the RAII equivalent of BackupInFlightClaim, and handedOff is set only after both succeed), and benchmark/setup.ts making concurrent workers share one database path (they already did before this PR — the declaration moved within the same scope chain, and sharing one database across workers is the point of those benchmarks); checkpoint.cpp's CheckpointInFlightClaim again; ReleaseLogRefsByEnv "permanently leaking" napi_refs for an erased descriptor (they are env-owned and reclaimed at env teardown, and once the entry is erased no foreign close() can reach the handle through closables, which is the hazard invariant 18 exists for); and, for the third time, a per-call __cxa_guard_acquire on the seam flags (a std::atomic<int> with a constexpr constructor is constant-initialized, so no guard is emitted — confirmed by inspecting the generated assembly in an earlier round).
  • Earlier rounds on the pre-rebase head (four, converged) are unchanged and described in the commit history.

Refs #787

🤖 Generated with Claude Code

Final maintenance pass — 2026-09-09

This note supersedes earlier rebase, review-coverage, and verification statements above. The branch is rebased cleanly onto 17fceeef; PR head ac31bdea contains the final review-feedback fix.

  • The new registryStatus() thread was valid: carrying a shared_ptr<DBDescriptor> beyond databasesMutex could make a racing last-handle close see an extra owner and skip its purge with no release-side retry. RegistryStatusEntry now contains only copied values, captured under the registry and owning child locks, and all N-API construction happens after unlocking. The audited registry-to-child ordering is documented beside each mutex, and invariant 6 records the ownership-pin failure mode.
  • The new worker fixture uses a dedicated status worker, asserts that the close actually overlaps the delayed registry walk, and then verifies the entry was purged. It is wired into the destroy lifecycle suite. The earlier column-map race fixture now reports its own timeout before Vitest's outer deadline.
  • The three older open threads remain human decisions: canonical path identity changes user-visible behavior; rejecting transaction admission on isClosing() violates invariant 18 and the safer closables-before-flush reorder needs a separate change; and shutdown() error reporting is a consumer-compatibility choice rather than a mechanical review fix.

Changes

Verification

Superseded for the current head by "Verification of the merge" in the 2026-09-21 note above;
the figures below are from the pre-merge head ac31bdea.

  • pnpm check — clean (type-check, lint, formatting); git diff --check — clean.
  • pnpm test:native — 194/194 passed.
  • pnpm test — 956 passed, 9 skipped, 0 failed.
  • The regression failed on 87842947 with a retained registry entry (followed by a teardown SIGSEGV) and passes on this head; the strengthened overlap assertion also passes.
  • Planning fallback after the default Claude leg hit quota: Framing-Verdict: chosen-approach-sound (Gemini). Final full review produced no new actionable
    finding: Gemini repeated an adjudicated benchmark false positive and the existing comment nit; Claude failed at startup, Cursor was policy-pruned for the AGENTS.md edit, and the domain leg timed out. These are review-infrastructure coverage gaps, not test failures.
  • Current-head CI has reported: validation and security checks pass; platform, native, benchmark, and stress jobs are running with no reported failure.

Complexity: complicated

Origin — the dispatch brief this PR was written from

Serialize database destruction with concurrent opens

LIVE CONVERSATION about #787.

You are answering a person, in a thread, one turn at a time. Every turn:

  1. Read the whole thread in this dispatch file's # Log — it is the conversation so far, and
    each of your previous turns is in it. Read the PR/issue and the code as needed.
  2. Answer the LAST message. Append your answer to # Log as your turn. Prose, not a report:
    they are talking to you, and a status template is not an answer.
  3. Set status: needs-input and stop. The thread stays open; their next message resumes it.

Each turn arrives as ASK (answer it, change nothing) or PERFORM (do it, then say what you did) —
the person chose which when they sent it, and the run's own prompt tells you which one this is.
Never infer it from the wording: an unrequested commit in the middle of a discussion and a polite
description of work that was supposed to happen are the two failures this exists to prevent.

Never mark a PR ready and never merge from this conversation.

Dispatch: task chat-pr-rocksdb-js-787-kriszyp · queued by unknown · ran by claude/opus/low · worker kzyp-xps-1

Review-Coverage: authored=claude; ran=codex,gemini; adjudicated=domain; blocked=cursor-grok(failed); declined=cursor-composer; rounds=27; full=1 @ 1947744

Human-Review-Need: 4 (decisions: quarantine-on-close-failure, unbounded-drains, backup-stream-on-libuv-pool, close-and-shutdown-throw, destroy-waits-for-copies, txn-dbhandle-reset-policy, lifecycle-wait-global-setting, pr-scope) @ 1947744

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces robust database lifecycle management for RocksDB JS bindings. It implements a timed-wait mechanism (lifecycleWaitSeconds) for open, destroy, and shutdown operations to prevent concurrent lifecycle conflicts. It also introduces a "quarantine" state for database paths when a native close, flush, compaction, or physical directory cleanup fails, preventing subsequent opens until the cleanup is retried via destroy() or shutdown(). Additionally, it ensures that in-flight operations (like backups and checkpoints) are safely awaited before destruction, and that thread-affine N-API references are cleaned up safely. There are no review comments, so I have no feedback to provide.

@github-actions

github-actions Bot commented Aug 15, 2026 •

Copy link
Copy Markdown
Contributor

📊 Benchmark Results

get-sync.bench.ts

getSync() > random keys - small key size (100 records)

Implementation Rank Operations/sec Mean (ms) Min (ms) Max (ms) RME (%) Samples
🥇 lmdb 1 24.49K ops/sec 40.83 39.36 582.73 0.113 122,446
🥈 rocksdb 2 10.71K ops/sec 93.39 90.02 24,804.734 1.01 53,542

getSync() > sequential keys - small key size (100 records)

Implementation Rank Operations/sec Mean (ms) Min (ms) Max (ms) RME (%) Samples
🥇 lmdb 1 28.75K ops/sec 34.78 33.72 508.735 0.098 143,752
🥈 rocksdb 2 11.53K ops/sec 86.73 84.25 574.605 0.050 57,653

ranges.bench.ts

getRange() > small range (100 records, 50 range)

Implementation Rank Operations/sec Mean (ms) Min (ms) Max (ms) RME (%) Samples
🥇 lmdb 1 24.82K ops/sec 40.30 35.83 2,264.097 0.297 124,079
🥈 rocksdb 2 15.75K ops/sec 63.49 55.76 1,077.087 0.119 78,747

realistic-load.bench.ts

Realistic write load with workers > write variable records with transaction log

Implementation Rank Operations/sec Mean (ms) Min (ms) Max (ms) RME (%) Samples
🥇 rocksdb 1 362.00 ops/sec 2,762.432 110.333 79,096.801 17.43 724
🥈 lmdb 2 26.10 ops/sec 38,313.582 438.982 1,201,507.25 136.57 64.00

transaction-log.bench.ts

Transaction log > read 100 iterators while write log with 100 byte records

Implementation Rank Operations/sec Mean (ms) Min (ms) Max (ms) RME (%) Samples
🥇 rocksdb 1 38.42K ops/sec 26.03 11.36 7,708.659 0.387 192,099
🥈 lmdb 2 437.89 ops/sec 2,283.658 162.247 32,224.301 1.68 2,190

Transaction log > read one entry from random position from log with 1000 100 byte records

Implementation Rank Operations/sec Mean (ms) Min (ms) Max (ms) RME (%) Samples
🥇 rocksdb 1 730.81K ops/sec 1.37 1.19 4,861.915 0.205 3,654,026
🥈 lmdb 2 444.95K ops/sec 2.25 1.15 7,942.485 0.588 2,224,737

worker-put-sync.bench.ts

putSync() > random keys - small key size (100 records, 10 workers)

Implementation Rank Operations/sec Mean (ms) Min (ms) Max (ms) RME (%) Samples
🥇 rocksdb 1 834.49 ops/sec 1,198.33 1,028.429 1,917.183 0.341 1,669
🥈 lmdb 2 1.15 ops/sec 871,886.072 804,555.666 963,136.741 4.29 10.00

worker-transaction-log.bench.ts

Transaction log with workers > write log with 100 byte records

Implementation Rank Operations/sec Mean (ms) Min (ms) Max (ms) RME (%) Samples
🥇 rocksdb 1 22.82K ops/sec 43.83 29.45 20,444.944 2.08 45,636
🥈 lmdb 2 817.13 ops/sec 1,223.799 282.414 10,488.217 5.17 1,637

Results from commit 3842f2f

@kriszyp
kriszyp marked this pull request as ready for review August 15, 2026 11:52
Comment thread src/binding/database/db_registry.cpp Outdated
Comment thread src/binding/binding.cpp Outdated
Comment thread src/binding/database/db_registry.cpp Outdated
Comment thread src/binding/core/test_seam.h Outdated
Comment thread src/binding/database/db_handle.cpp
Comment thread src/binding/database/db_registry.cpp Outdated
Comment thread src/binding/database/db_handle.cpp
Comment thread src/binding/database/db_registry.cpp Outdated
Comment thread src/binding/database/database.cpp
Comment thread src/binding/iterator/db_iterator.cpp Outdated
Comment thread AGENTS.md Outdated
Comment thread benchmark/setup.ts
Comment thread test/destroy.test.ts
Comment thread src/binding/iterator/db_iterator.cpp Outdated
Comment thread src/binding/database/db_registry.cpp Outdated
Comment thread src/binding/database/database.cpp
@cb1kenobi

Copy link
Copy Markdown
Member

Reviewed f24ef7a5 — no issues found. This PR looks good, nice job!

Re-review of the one new commit since b35ad3f7 ("Address remaining lifecycle review threads"). Both previously-open findings are confirmed fixed in the code, not just marked resolved:

  • db_iterator.cpp:289 (Medium, per-row getenv) — fixed. The lookup is hoisted into initializeTestSeams(), which runs as the first statement of NAPI_MODULE_INIT, and Next() now does a relaxed atomic load. Verified at the object-code level: ROCKSDB_JS_ITERATOR_NEXT_DELAY_MS no longer appears in db_iterator.o (only in binding.o), and the compiled DBIterator::Next contains zero getenv calls. The new ROCKSDB_JS_COUNT_DELAY_MS seam got the same treatment up front.
  • db_registry.cpp:960/961 (Medium, closeError asymmetry) — fixed. The four copies of the finishClose() → erase-or-quarantine → notify → emit tail are collapsed into closeClaimedDescriptors(), and the policy that had drifted is now one named, documented option (failOnCompletedWithError). The gate reduces exactly to the old DestroyDB behavior when false and the old unconditional behavior when true, so the refactor is behavior-preserving while making the remaining asymmetry deliberate rather than accidental.
  • The earlier getCount Low is also addressed: countRemaining() polls isClosing() per row and reports the abort instead of a partial count, on both the database and transaction paths, and the inaccurate "compaction is the only unbounded in-flight op" comment is corrected.

Also checked and cleared: the dropped if (condition) null guard in PurgeIfUnreferenced is safe (both DBRegistryEntry constructors make_shared the condition; the old guard only mattered because the notify used to sit outside the if (descriptor) block); the newly-added closeRetrying = false on the PurgeAll/PurgeIfUnreferenced quarantine paths is a no-op, since only DestroyDB/Shutdown ever latch it and beginClose() is single-shot.

Verification: full suite 55 files, 773 passed / 1 skipped / 0 failed; targeted destroy + ranges 64/64 including the new aborts an in-flight getCount() when a foreign destroy begins fixture. CI green on the head (Windows jobs still pending at review time).

One merge-ordering note, not a defect in this PR: finishClose() still holds txnsMutex across cancelForDB(), which takes writerMutex_, and this PR makes that a routine path because destroy() now force-closes every descriptor. If #744 lands with PurgeIfUnreferenced still on the wake callback path, its lock-order inversion becomes materially more likely — worth sequencing #744's fix before or with this.

—
Generated by Barber AI

kriszyp added a commit that referenced this pull request Aug 25, 2026
…GetCount against concurrent close

finishClose() cleared compactCancelRequested right after the operationsInFlight
drain, but an async compact() releases its OperationGuard at setup handoff and
is not awaited until the closables sweep — so it can still be running after
the drain returns, and clearing the token there left it able to stall
teardown (and every concurrent open on the path) indefinitely. Keep the token
armed for finishClose()'s whole duration instead, and have the close-time
compact-on-close pass opt out via a new compactRange() `cancellable` param
rather than relying on the shared flag being cleared.

Transaction::GetCount now takes an OperationGuard and checks isClosing()
before scanning: without it, finishClose()'s drain can return immediately and
the closables sweep can roll back the transaction while the count scan is
parked between rows, reading freed memory.

Carries the in-progress PR #787 lifecycle repair plan describing the fuller
atomic-admission fix these two changes are a first slice of.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013CvCKcGKG7vyhRts1Mv4Wh
@kriszyp
kriszyp force-pushed the kris/serialize-destroy-open branch from f24ef7a to 5b459e4 Compare August 25, 2026 13:31
kriszyp added a commit that referenced this pull request Aug 25, 2026
…GetCount against concurrent close

finishClose() cleared compactCancelRequested right after the operationsInFlight
drain, but an async compact() releases its OperationGuard at setup handoff and
is not awaited until the closables sweep — so it can still be running after
the drain returns, and clearing the token there left it able to stall
teardown (and every concurrent open on the path) indefinitely. Keep the token
armed for finishClose()'s whole duration instead, and have the close-time
compact-on-close pass opt out via a new compactRange() `cancellable` param
rather than relying on the shared flag being cleared.

Transaction::GetCount now takes an OperationGuard and checks isClosing()
before scanning: without it, finishClose()'s drain can return immediately and
the closables sweep can roll back the transaction while the count scan is
parked between rows, reading freed memory.

Carries the in-progress PR #787 lifecycle repair plan describing the fuller
atomic-admission fix these two changes are a first slice of.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013CvCKcGKG7vyhRts1Mv4Wh
@kriszyp
kriszyp force-pushed the kris/serialize-destroy-open branch from 5b459e4 to 542b058 Compare August 25, 2026 14:59
@cb1kenobi

Copy link
Copy Markdown
Member

Reviewed 542b058a — no issues found. This PR looks good, nice job!

Re-review of the one new commit since f24ef7a5 (the branch was rebased onto main after rocksdb-js#744 merged; verified via git range-diff that all 35 prior commits carried forward unchanged modulo rebase context, with commit 542b058a new at the tip).

542b058a fixes a real race: compactCancelRequested now stays armed for finishClose()'s whole duration (an async compact-on-close pass opts out via a new cancellable param instead), and Transaction::GetCount now takes an OperationGuard + isClosing() check so finishClose()'s drain can't return early and let the closables sweep roll back the transaction mid-scan. Both changes are consistent with the existing OperationGuard/ACQUIRE_OPERATIONS_LOCK pattern elsewhere in the codebase.

Also re-verified at object-code level:

  • DBIterator::Next() has zero getenv calls in its compiled disassembly (ROCKSDB_JS_ITERATOR_NEXT_DELAY_MS is absent from db_iterator.o's string table); the seam is now a relaxed atomic load, set once in initializeTestSeams().
  • closeClaimedDescriptors() remains the single teardown tail for all four callers (CloseDB, DestroyDB, PurgeAll, Shutdown), with the completed-but-errored policy as the named ClaimedCloseOptions.failOnCompletedWithError option — false only for destroy(), true (fatal) everywhere else.
  • finishClose() still takes txnsMutex and holds it across cancelForDB(), which itself takes VT's writerMutex_ — the txnsMutex → writerMutex_ ordering is unchanged by rocksdb-js#744 merging.

pnpm test (destroy.test.ts lifecycle suite: 19/19), pnpm test:native (148/148, 3 expected macOS skips), and pnpm check all pass clean at this head.

—
Generated by Barber AI

kriszyp added a commit that referenced this pull request Aug 25, 2026
…GetCount against concurrent close

finishClose() cleared compactCancelRequested right after the operationsInFlight
drain, but an async compact() releases its OperationGuard at setup handoff and
is not awaited until the closables sweep — so it can still be running after
the drain returns, and clearing the token there left it able to stall
teardown (and every concurrent open on the path) indefinitely. Keep the token
armed for finishClose()'s whole duration instead, and have the close-time
compact-on-close pass opt out via a new compactRange() `cancellable` param
rather than relying on the shared flag being cleared.

Transaction::GetCount now takes an OperationGuard and checks isClosing()
before scanning: without it, finishClose()'s drain can return immediately and
the closables sweep can roll back the transaction while the count scan is
parked between rows, reading freed memory.

Carries the in-progress PR #787 lifecycle repair plan describing the fuller
atomic-admission fix these two changes are a first slice of.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013CvCKcGKG7vyhRts1Mv4Wh
@kriszyp
kriszyp force-pushed the kris/serialize-destroy-open branch 2 times, most recently from 542b058 to 3cdd9f9 Compare August 25, 2026 17:20
Comment thread src/binding/database/db_registry.cpp
kriszyp added a commit that referenced this pull request Aug 26, 2026
…GetCount against concurrent close

finishClose() cleared compactCancelRequested right after the operationsInFlight
drain, but an async compact() releases its OperationGuard at setup handoff and
is not awaited until the closables sweep — so it can still be running after
the drain returns, and clearing the token there left it able to stall
teardown (and every concurrent open on the path) indefinitely. Keep the token
armed for finishClose()'s whole duration instead, and have the close-time
compact-on-close pass opt out via a new compactRange() `cancellable` param
rather than relying on the shared flag being cleared.

Transaction::GetCount now takes an OperationGuard and checks isClosing()
before scanning: without it, finishClose()'s drain can return immediately and
the closables sweep can roll back the transaction while the count scan is
parked between rows, reading freed memory.

Carries the in-progress PR #787 lifecycle repair plan describing the fuller
atomic-admission fix these two changes are a first slice of.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013CvCKcGKG7vyhRts1Mv4Wh
@kriszyp
kriszyp force-pushed the kris/serialize-destroy-open branch from 107b216 to ab5cd96 Compare August 26, 2026 15:02
Comment thread AGENTS.md Outdated
Comment thread src/binding/database/db_descriptor.h

@cb1kenobi cb1kenobi left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Please rebase with main and resolve the merge conflicts.

@cb1kenobi cb1kenobi left a comment •

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed ba0e1c0 and found no blocking issues. No new blocking defects were confirmed on changed lines. Existing findings were not repeated.

—
Generated by Barber AI

kriszyp added a commit that referenced this pull request Sep 2, 2026
…GetCount against concurrent close

finishClose() cleared compactCancelRequested right after the operationsInFlight
drain, but an async compact() releases its OperationGuard at setup handoff and
is not awaited until the closables sweep — so it can still be running after
the drain returns, and clearing the token there left it able to stall
teardown (and every concurrent open on the path) indefinitely. Keep the token
armed for finishClose()'s whole duration instead, and have the close-time
compact-on-close pass opt out via a new compactRange() `cancellable` param
rather than relying on the shared flag being cleared.

Transaction::GetCount now takes an OperationGuard and checks isClosing()
before scanning: without it, finishClose()'s drain can return immediately and
the closables sweep can roll back the transaction while the count scan is
parked between rows, reading freed memory.

Carries the in-progress PR #787 lifecycle repair plan describing the fuller
atomic-admission fix these two changes are a first slice of.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013CvCKcGKG7vyhRts1Mv4Wh
kriszyp and others added 12 commits September 9, 2026 14:28
…ng, destroy flush waste

- CloseDB and the iterator finalizer detached from `closables` before
  close()/Reset(), leaving a concurrent destroy()/shutdown() sweep unable to
  see (and wait for) a handle/iterator still draining async work or mid-Next()
  -- a real use-after-free window on the shared rocksdb::DB. Detach after
  close() returns instead; closeMutex/iteratorMutex already serialize a
  foreign close arriving in that window.
- DBHandle::close() released `logRefs` napi_refs gated on a std::thread::id
  equality check, the same recycled-pthread-id hazard invariant 18 already
  fixed for TransactionHandle. Add DBRegistry::ReleaseLogRefsByEnv, wired into
  the env cleanup hook like CloseTransactionsByEnv, so a dying env's logRefs
  are emptied while it is still alive -- before its thread id could ever be
  reused against the stale guard.
- OpenDB's "still closing"/"retry in progress" waits parked on one matching
  descriptor's condition variable but re-scanned the whole path in their
  predicate, so two descriptors closing on one path (e.g. a writable and a
  secondary) could leave an opener asleep for the full lifecycleWaitSeconds
  even after the path was free. Track the specific selected entry instead.
- finishClose(destroying=true) still ran a full flush + close-time compaction
  before the files are unlinked -- wasted I/O, and with allow_write_stall's
  default a destroy() of a stalled database could hang indefinitely while
  holding the destroyingPaths gate. Skip both when destroying.
- Add a fixture covering a lifecycleWaitSeconds timeout actually firing and
  the path recovering afterward (previously only config validation was
  tested).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01122SfiNfCtiLgWvQTD6wZ7
…eal two-descriptor regression test

- fork-lifecycle-timeout.mts raced the shutdown retry claim against the main
  thread's open() with no barrier, so under load the open could win and throw
  the wrong error instead of timing out -- a flake, not a proof. Expose
  `closeRetrying` on registryStatus() and poll for it before racing the open.
- fork-compact-cancel-destroy.mts lost its discriminating power once destroy()
  skips compactOnClose (previous commit): with nothing left to block on
  compactMutex, an early vs. late cancellation arm became timing-indistinguishable.
  Drive it through shutdown() instead, which still runs compactOnClose.
  Verified: reverting the early arm now makes the fixture fail again (7.7s vs
  the ~500ms bound), confirming this restores the fixture's purpose.
- Added a genuine two-descriptor regression test for the OpenDB
  condition-variable/predicate fix (db_registry.cpp:603/:655): a writable and
  a read-only descriptor quarantine, then retry concurrently under one
  shutdown() call, with an opener racing in. Verified against the pre-fix
  predicate: 4/4 runs stalled to the full 8s deadline instead of the expected
  ~3s, confirming this catches the regression the prior test explicitly could
  not.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01122SfiNfCtiLgWvQTD6wZ7
shutdown() is process-wide, not scoped to one path, and its worker call could
still be re-scanning for anything left to close after the racing open() call
returned but before it posted shutdownResult -- a handle opened in that
window is not safe to read from or hold onto (shutdown() could sweep and
force-close it too). Keep the timing-only open (the actual regression proof)
but defer the data-preservation check to a fresh open, strictly after the
worker's shutdown() call and the worker itself are both confirmed done.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01122SfiNfCtiLgWvQTD6wZ7
Review follow-up on the `shutdown()`-throws thread: the behavior is kept
(db.close() throws the identical quarantine error, and three of the four
throw sites are lifecycle timeouts that `database:closeFailed` cannot
carry), but the only place it was written down was the README. The
exported symbol now carries the same contract, including why a
`process.on('exit')` listener has to wrap the call: a throw from an exit
listener skips every listener registered after it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…eams

Independent pre-push review round 1 on the rebased head (codex + gemini +
harper-domain) found five real items:

- `DBHandle::open()` now drops the previous lifecycle's `logRefs`. The
  owner-thread guard that makes a foreign close napi-safe (AGENTS.md
  invariant 18) also means a cross-env `destroy()`/`shutdown()` leaves the
  cache populated, so a reopened handle handed `useLog()` back a
  `TransactionLog` whose store `weak_ptr` pointed at the unregistered store
  of the closed lifecycle. Only `addEntry` re-resolves, so every read
  accessor reported an empty log — `getLogFileSize()` returned 0 instead of
  31 in the new fixture, which fails 1/1 without the fix.
- `registryStatus()`'s destroy-cleanup tombstone branch omitted
  `transactionDetails`, which `RegistryStatusDB` declares non-optional; a
  monitor reading `entry.transactionDetails.length` threw on exactly the
  entry shape that only appears when a physical destroy failed.
- The three test-delay seams this PR added (`ROCKSDB_JS_BACKUP_DELAY_MS`,
  `ROCKSDB_JS_DESTROY_DELAY_MS`, `ROCKSDB_JS_CLOSE_RETRY_DELAY_MS`) still
  called `::getenv` from a libuv worker / arbitrary teardown thread. They
  are snapshotted in `initializeTestSeams()` now, like the fault flags
  beside them.
- `fork-shutdown-retry.mts` raced the worker's retry claim: the worker posts
  before calling `shutdown()`, so the parent's open could hit the still
  quarantined entry. It now polls `closeRetrying` first and re-reads data
  from a fresh handle after the worker is done — the two barriers the
  two-descriptor fixture already had.
- `fork-compact-cancel-async.mts` had no `'message'` listener attached
  across `await outcome`, so a `{destroyed:true}` posted in that window was
  dropped and the fixture would hang to its timeout.

Declined, with reasons in the PR body: making `isClosing()` a relaxed load
(the reviewer's own estimate is that it is dwarfed by `iterator->Next()`,
and the flag is read under mutexes elsewhere), and the aggregate
comment-narration nit.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
CI turned red on macOS (Bun + Deno) with four quarantine tests failing on
`/private/var/...` vs `/var/...`: a tombstoned `registryStatus()` entry and
every `database:closeFailed` event reported the registry key's resolved
identity, so a caller matching either against the path it opened did not
recognize it. That is exactly what AGENTS.md invariant 19 forbids
("returning only the resolved identity breaks callers that match paths
against the spelling they supplied") and what `registryStatus()` already
does correctly for a live descriptor.

- `emitCloseFailures()` reports `descriptor->path`.
- `DBRegistryEntry::reportedPath` remembers the opening caller's spelling so
  a destroy-cleanup tombstone — which has no descriptor left to ask — can
  still report it, in `registryStatus()` and in its own emit.

Not caused by the rebase (the pre-rebase head carried identical code and was
green on the older macos-26-arm64 runner image); the new image's TMPDIR made
the latent mismatch reachable. `test/destroy.test.ts` now covers it on every
platform with an explicit symlink instead of depending on macOS's `/var`
link — verified to fail with `expected undefined to be true`, the same
assertion macOS CI produced, when either report site is reverted.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…tombstone

Round 3 of the independent pre-push review (codex) found the one gap in the
previous commit: a second `destroy()` after a failed physical cleanup starts
with `reportedPath = identityPath` and its scan skipped the existing
tombstone, because that entry has no descriptor. So `registryStatus().path`
kept reporting the opened spelling while the retry's `database:closeFailed`
reverted to the resolved identity — the two disagreeing is the same defect
one step later. The scan now falls back to a matching entry's remembered
spelling, still preferring a live descriptor's.

The symlink test asserts the event path on the retry as well, and fails with
`expected [ …(2) ] to match object [ …(2) ]` without the fallback.

Also dropped from that round, verified rather than assumed: Gemini's blocker
on `checkpoint.cpp` leaking `operationsInFlight` when admission or queueing
fails — `CheckpointInFlightClaim` is the RAII equivalent of
`BackupInFlightClaim` and `handedOff` is only set after both succeed
(`src/binding/database/checkpoint.cpp:69-76,114-115,206`).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
CI's remaining SIGSEGV (`fork-destroy-open.mts`, one Deno/Bun job per run
since the rebase) is a real native use-after-free on the JS thread, not a
flake. Reproduced locally at 2/60 by running that fixture 6-way concurrently,
and gdb put it exactly here:

  __strlen_avx2
  v8::String::NewFromUtf8
  napi_set_named_property
  rocksdb_js::DBRegistry::RegistryStatus

`registryStatus()` holds `databasesMutex`, which covers the registry map but
not a descriptor's own maps. It walked `descriptor->columns` — guarded by
`columnsMutex` — while a cross-env `destroy()`'s `finishClose()` cleared that
map from the worker thread, so `name.c_str()` pointed into a freed map node
and `napi_set_named_property()` `strlen()`ed it. `locks.size()` was read
unguarded the same way (a count, so a torn read rather than a fault).

The column summary is now snapshotted under `columnsMutex` (plus the per-CF
`userSharedBuffersMutex` for its buffer count) and the JS values are built
after releasing it — holding it across the N-API calls would risk a finalizer
re-entering the same non-recursive mutex on this thread, which is why
`transactions` already had this shape under `txnsMutex`. `locks.size()` is
read under `locksMutex`. Neither adds a lock-order inversion: `OpenDB` and
`CollectWriteBufferManagerInventory` already establish
`databasesMutex → columnsMutex`, `getUserSharedBuffer` is a leaf, and no
`locksMutex` region reaches the registry.

Pre-existing on `main` (its `registryStatus()` walks `columns` unguarded too);
this PR is what makes it reachable, because `destroy()` now force-closes every
descriptor and its own fixtures poll `registryStatus()` across that window.

`test/fixtures/fork-registry-status-column-race.mts` makes it deterministic
rather than leaving it to CI luck: a worker churns `dropSync()` against the
shared descriptor while the main thread polls `registryStatus()`, with the new
`ROCKSDB_JS_REGISTRY_STATUS_COLUMNS_DELAY_MS` seam parking the walk per column
family so an erase lands inside it. Column names run past libstdc++'s 15-char
small-string buffer so the erase frees a separate heap allocation (an SSO name
usually survives the free intact and hides the bug). 5/5 abort without the
fix, 3/3 clean with it; the concurrent `fork-destroy-open.mts` stress went
2/60 → 0/132.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TMokk8DsJyGwpHz4Lmts85
…iming

Round 5 of the independent pre-push review (codex, seconded by the domain
adjudicator) caught `test/fixtures/fork-registry-status-destroy-race.mts`
shipping unreferenced: it was the first attempt at the `registryStatus()`
column-walk regression test and does not reproduce (0/7 against the unfixed
walk), because a single close-time `columns.clear()` rarely lands inside a
walk and a small-string column name usually survives the free intact. The
drop-churn fixture that replaced it does reproduce deterministically (5/5),
and it covers the same thing — the fix is the snapshot in the walk, not
anything per-mutator, so reverting that snapshot fails the churn fixture too.

Also reworded the tombstone branch's comment, which claimed it fills "every
non-optional field of RegistryStatusDB": `userSharedBuffers` is declared
non-optional and set by no branch, and `columnFamilies` is typed `string[]`
against an object. Both predate this PR and are noted rather than fixed here.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TMokk8DsJyGwpHz4Lmts85
Co-Authored-By: GPT-5 Codex <noreply@openai.com>
Co-Authored-By: GPT-5 Codex <noreply@openai.com>
Co-Authored-By: GPT-5 Codex <noreply@openai.com>
@kriszyp
kriszyp force-pushed the kris/serialize-destroy-open branch from 5e8803d to 8784294 Compare September 9, 2026 20:44
Comment thread src/binding/database/db_registry.cpp
kriszyp and others added 3 commits September 9, 2026 15:52
Co-Authored-By: GPT-5 Codex <noreply@openai.com>
Co-Authored-By: GPT-5 Codex <noreply@openai.com>

@cb1kenobi cb1kenobi left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This was a big one, but it looks great! I love the OperationGuard!

kriszyp and others added 3 commits September 21, 2026 07:55
Brings the branch up to `main` (through #866) and reconciles it with
#850, which landed the deferred column-family reclamation work in the
same lifecycle files. Nine files conflicted; the resolutions keep both
sides' intent and, where the two branches had grown the same mechanism
twice, keep one:

- `DBRegistry::OpenDB`: #850's outer reclaim-retry loop now wraps this
  branch's destroy/quarantine/close-retry wait loop, and the retiring-
  generation reclaim runs inside the column block. The handle is still
  adopted, cancellation-cleared and attached under `databasesMutex`
  (invariant 28), so `DBHandleParams` stays unused and is removed.
- `Transaction::Commit`: main's `PendingTransactionCommitState` guard
  subsumes this branch's hand-rolled unwind, so the async-work refusal
  and queue-failure paths reject and let the guard release the claim,
  delete the async work and restore the handle to `Pending` — instead of
  `admitAsyncWorkOrReject`/`queueAsyncWorkOrReject`, whose `delete state`
  cannot coexist with the guard's ownership. `finishCommitCompletion()`
  became `completion->finish()` in b7a23da.
- `DBDescriptor::finishClose`: keeps this branch's `cancelBlockingWork()`
  sweep first and its `closeWorkersStopped` retry gate, with main's newer
  per-env `CommitCompletion::release()` pass inside it.
- `registryStatus()`: keeps this branch's value-only snapshot (which
  already fixes the `columns`/`transactions` races #850 addressed inline)
  and widens `TxnSummary::id` to the 64-bit transaction id from #853.
- `ROCKSDB_JS_COMMIT_EXECUTE_DELAY_MS`: one seam, at main's call site
  (which also sets the `…DelayActive` flag), read from this branch's
  `initializeTestSeams()` snapshot rather than `::getenv`.
- `drop.test.ts`: drops this branch's poisoned-environment assertions —
  #850 refuses the commit at admission, so #726 no longer poisons the
  database and main's "still writable" expectations are the correct ones.
- `db_handle.cpp`: `commitCompletion.reset()` moves into `open()` beside
  `releaseLogRefs()`, both owner-thread-only pre-adoption resets.
- AGENTS.md: main's three new invariants become 23/24/25 and this
  branch's three become 26/27/28; stale cross-references on both sides
  (including `unregisterColumnFamily` → `retireColumnFamily`) updated.

Verification: `pnpm check` clean, `pnpm test:native` 234/234,
`pnpm test` 1045 passed / 10 skipped.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Pre-push review (Claude advisory leg) flagged that `DBHandle::open()`
read `this->descriptor->identityPath` after `DBRegistry::OpenDB()`
returned, which contradicts invariant 28's "no handle shared-pointer
field may be read after the registry lock drops". OpenDB already
publishes every other descriptor-backed field under `databasesMutex`;
`identityPath` now joins them.

Not a live defect — a foreign close resets `descriptor` only on the
owning thread, which is the thread inside `open()` — but the code and
the invariant it cites have to agree, and this is the cheaper side to
change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Pre-push review (harper-domain, kept from cursor-grok's pass):
`Transaction::GetCount` bound `txnDbHandle` as a reference to
`TransactionHandle::dbHandle`, null-checked it, and then dereferenced
`->descriptor` to build its OperationGuard. `TransactionHandle::close()`
resets that member from whichever thread drives a forced teardown, with
no owner-thread check, so a foreign destroy()/shutdown() landing between
the two is a null dereference. A local shared_ptr copy closes the window
and pins the handle for the guard's lifetime.

Also drops a comment in `Database::CatchUpWithPrimary` that still
described `DBHandle::close()`'s async-work drain as bounded with its
failure ignored; this branch made that drain untimed (invariant 17).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

@cb1kenobi cb1kenobi left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No new blocking defects were confirmed on changed lines. Existing findings at the same code paths were not repeated.

—
Reviewed 1947744

@kriszyp
kriszyp merged commit 5df2c3e into main Sep 21, 2026
44 of 45 checks passed
@kriszyp
kriszyp deleted the kris/serialize-destroy-open branch September 21, 2026 19:37
kriszyp added a commit that referenced this pull request Sep 21, 2026
…ebase

main's own #787 landed invariants 26-28 in the same slot this branch's
purge invariant occupied (25), so the rebase conflict resolution moved
purge to 29 and its sibling databaseFlushed invariant (which collided
with main's new 26) to 30. Fixed the three stale in-code "invariant 25"
cross-references this displaced (AGENTS.md, transaction-log-reader.ts,
transaction_log_store.cpp), following the whole-tree re-grep lesson
already recorded in this PR's Findings.

Dispatch-Task: pr-maint-68bd037170be49b2a01e194607ecae40
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
kriszyp added a commit that referenced this pull request Sep 29, 2026
…g to read it (#820)

* Release a purged segment's mapping instead of pinning it

A purge unlinks the segment, which removes one link to an inode whose bytes a
retired segment never changes again; a reader's MemoryMap is the other link, so
the entries it mapped are still exactly the committed history and stay readable.
The bug in HarperFast/harper#2337 was never that those bytes were served — it was
that the mapping was never released, so the purge reclaimed no space: 16 MiB of a
deleted .txnlog resident until restart.

TransactionLog._currentLogBuffer, the fast path over the already-weak
_logBuffers cache, held a strong reference and is only refreshed by query(), so
a reader that calls query() once and next() forever (harper's audit
subscription) froze it on whatever segment was current then. It is now a
WeakRef: the mapping goes at the next GC once the iterator holding it moves on,
with no purge-time invalidation and no cross-handle signalling.

Also here, because they are the same reclaim path:

- nextReadableLogBuffer() skips a run retention deleted when an iterator
  advances. Stopping at the hole stopped the iterator permanently, since every
  later poll stopped in the same place. _findPosition(0) names the oldest
  survivor, so a purged prefix costs one native call rather than a probe per
  segment, and only a run the store no longer has is skipped: a segment it still
  knows is merely unmappable for now, so iteration stops and retries.
- readableExtent() bounds a read of a purged segment by its mapping, since the
  store reports no size for a segment it has forgotten. The 0 it reports dropped
  every entry the reader had not reached yet — including entries appended after
  it last polled, which the writer's overlay extension made visible in that same
  mapping.
- removeFile() uses the non-throwing std::filesystem::remove overloads on both
  platforms; a Windows sharing violation used to unwind a C++ exception through
  the N-API purge boundary.
- A segment that vanished between the purge's scan and its unlink is forgotten
  from sequenceFiles the way the scan forgets an already-missing one, and a
  segment that could not be deleted for a real reason is reported once per purge
  run via log.warn instead of silently stalling retention.

Refs HarperFast/harper#2337

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* Gate the txn.state rewrite on the file's existence, not the stream's is_open()

databaseFlushed() keeps the flushed-state stream open across flushes, and a
stream describes a descriptor, not a pathname: once txn.state (or the whole
store directory) is unlinked, every write lands in the orphaned inode while
getLastFlushedPosition(), which reads by path, returns the {0,0} sentinel and
retention never advances.

The pathname is now checked before the unchanged-position shortcut, since a
flush resolving to the already-recorded position must still restore a missing
file. The directory is recreated the way getLogFile() does, isClosing is
re-checked under flushedStateMutex so a concurrent destroy cannot be
resurrected, and the reopen is in-place (in | out) rather than truncating: the
8-byte record is overwritten whole, and a truncating reopen after a failed write
would erase the last durable position before a retry that can fail again. The
creating open is taken only after the file is verified absent.

This runs on RocksDB's flush thread, where an escaping exception ends the
process, so the whole rewrite sits behind a catch-all and every failure is
reported once via log.warn, leaves lastWrittenFlushedPosition untouched, and is
retried on the next flush.

Hardening rather than a live bug: only purgeLogs({ destroy: true }) removes the
directory today, and Harper does not call it in production.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* Fix AGENTS formatting

Co-Authored-By: GPT-5 Codex <noreply@openai.com>

* Correct the oxfmt scope claim that misdiagnosed this branch's CI failure

AGENTS.md said oxfmt formats "TS/JS/JSON only" and does "not touch C++ or
Markdown". It does format Markdown, including AGENTS.md — that file is what
the failing `Check` job named, and renumbering its ordered invariant list is
what oxfmt objected to. Believing the doc is why the failure was first read as
a hand-fixable numbering slip.

Also record the trap the doc hid: a `pull_request` build formats the merge
commit, so a branch whose own `fmt:check` is green fails CI whenever it and
main have each appended an invariant and the numbers collide.

Verified: oxfmt scans .md and skips .cpp entirely, renumbers `19, 19, 20` to
`19, 20, 21`, and leaves lazy `1.` numbering alone.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GePs22dNhggr8XKDc1DThG

* Resolve the next readable segment by successor, and stop a read recreating one

Two independent defects in the purge-advance path, both found by the pre-push
review of this branch.

nextReadableLogBuffer() asked _findPosition(0) for "what comes after N". That
walks backward from the current sequence and stops at the first gap, so it
names the bottom of the contiguous run ending at the current segment. That is
the oldest survivor only when the deletions form a single prefix. With a
survivor between two holes - a purge({all}) that continued past a segment it
could not unlink, or segments deleted out of band and registered that way at
load - it lands past the survivor, and that segment's committed entries are
never yielded to the reader. Silently: the skip has no signal.

_nextLogId() (TransactionLogStore::nextSequenceAfter, sequenceFiles.upper_bound)
answers the question actually being asked, in one O(log n) lookup. The probe is
a bounded loop rather than one hop, because a registered successor can be absent
too - unlinked out of band, or by another process's retention, before this
process's purge run forgets it - and stopping at the first one that will not map
is the same permanent wedge. It terminates because _nextLogId() strictly
increases and is capped at the latest sequence.

Separately, resolving a registered-but-closed segment opened it, and
TransactionLogFile::open() creates (O_RDWR | O_CREAT, and OPEN_ALWAYS on
Windows). Probing a segment that discovery registered but that is no longer on
disk therefore recreated it as a header-only ghost that the next startup
registers again. openIfPresent() skips a definite absence, using the same
reasoning ensureExtent() already documents - only a definite absence skips,
since a stat that errors leaves the extent unresolved - and is meaningful
because dataSetsMutex is held across the check and the open. All three paths
that resolve a segment go through it: getLogFileSize(), getMemoryMap(), and
findPositionByTimestamp()'s backward walk, which skips to the previous sequence
instead of opening. The walk's outcome is unchanged, since an absent file
yielded position 0 and continued anyway; only the ghost goes away.

An already-open segment is unaffected by any of this: its handle still describes
the file, and on POSIX its unlinked inode is still exactly the committed history
this branch exists to keep serving.

The regression covers the successor lookup against a two-hole registry, a
registered successor that is itself absent, and the absence of resurrection
across all three resolve paths. It deliberately does not drive an iterator end
to end: findPositionByTimestamp() keeps the backward-walk shape, so after a
restart with holes no entry point positions a reader below one - measured
_findPosition(0) = 5 and startFromLastFlushed = 5 on a 1/3/5 layout - and that
shape is only reachable for an iterator already live when the holes appear.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GePs22dNhggr8XKDc1DThG

* Bound the resync scan by the written extent, and descend holes when positioning

Two follow-ups approved after review, both regressions this branch introduced or
left behind.

The resync bound at src/transaction-log-reader.ts:133 fell back to the raw
mapping length when the store reports no extent for a purged segment. Invariant
11 forbids exactly that, and prior behaviour was dataEnd = 0 (no scan), so this
branch introduced the failure rather than inheriting it: findResyncPosition
tries every start offset, so a mapped-capacity bound byte-scans the whole
pre-extended map on the JS thread, finds nothing (zeros never satisfy
frameFits), and returns undefined - reporting a recoverable mid-log break as a
torn tail, the harper#2016 amputation the invariant exists to prevent. It now
uses the cached extent or readableExtent(), which walks the frames to the
end-of-entries marker when the store has forgotten the segment.

That does not cover every case, and the gap is worth naming: endOfEntries()
deliberately returns the whole mapping when framing is broken, so that it never
becomes the thing that decides a corrupt frame ends the log. For a purged
segment whose framing is broken - the case that reaches corruptFrame in the
first place - the bound is therefore still the mapped capacity. Every other
purged-segment case is fixed and none is made worse, but choosing a bound when
no authoritative extent exists (the store has forgotten the segment and the
frame walk is defeated by the break) is a design question left open.

findPositionByTimestamp() carried the same defect just fixed in
nextReadableLogBuffer(), and a worse consequence. It stepped with
sequenceFiles.find(--sequenceNumber), so the walk ended at the first missing
sequence: after out-of-band deletion a reader asking for timestamp 0 silently
received only the newest contiguous run while every older survivor sat
registered, on disk, with a valid extent. It now descends by map order, and
tracks the next registered sequence above the entry being examined so the two
"the timestamp belongs further up" exits name a segment that exists rather than
sequenceNumber + 1, which a hole may have removed. Only a segment that actually
opened becomes that tracked sequence: both exits hand it back as a position to
read from, so a registered-but-absent one would send the reader to a file that
is not on disk.

Fixing positioning is what makes the end-to-end case testable at all: before it,
no entry point could put a reader below a hole, which is why the earlier
regression asserted the successor primitive instead. The test now covers a
three-wide gap - 2 and 4 never registered, 3 registered but absent - asserting
query({start: 0}) yields [1, 5] where the old lookup started at 5 and yielded
[5], plus a timestamp past every segment so the other exit is exercised rather
than only the position-zero path. Its read order is load-bearing and says so:
reading a segment's extent opens it, and an open handle keeps reporting the real
size after an unlink (invariant 20), so nothing may touch a segment before it is
meant to be gone.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GePs22dNhggr8XKDc1DThG

* fix: preserve txnlog purge read boundaries

Carry the append-owned readable extent with retained memory maps and invalidate stale flushed-state correlations when destructive purge empties a live store.

Co-Authored-By: GPT-5 Codex <noreply@openai.com>

* fix: harden txnlog purge safeguards

Co-Authored-By: GPT-5 Codex <noreply@openai.com>

* test: close txnlog review gaps

Co-Authored-By: GPT-5 Codex <noreply@openai.com>

* fix: close txnlog pathname races

Use atomic no-create opens for read probes and verify flushed-state writes through their pathname so purge cannot leave ghost segments or stale retention state.

Co-Authored-By: GPT-5 Codex <noreply@openai.com>

* test: assert only the portable purge-vs-mapping contract

The Windows-only "converge on a purge refused by a live mapping" test asserted
that a live reader mapping refuses the unlink there. Windows CI showed both
outcomes for that test across this branch's heads (refused at 9a8606a, removed
at c8e790a) with no change to the mapping's lifetime in between, so the outcome
of any single purge run is not something to assert on that platform.

Replace it with a cross-platform test of the contract the code actually
implements: the reader keeps every entry it mapped, and retention reclaims the
segment once nothing maps it — immediately where the unlink lands, on the next
run where it did not. It reads the mapping reference after the purge so V8
cannot collect it first, which is what made the old test's premise unverifiable.
The three POSIX-only tests that assert a first-run deletion now say why they are
skipped on Windows.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CVAGTbmYzkKxvuH4HQjXKd

* test: force the refused-unlink path, and lock the error it reports

The purge's refused-unlink branch (segment stays registered, one warn line per
run, next run reclaims it) had no test that reached it: POSIX always unlinks and
the Windows outcome is not predictable. An unwritable store directory fails the
unlink with EACCES, which is the same branch a sharing violation takes, so the
contract — including end-to-end delivery of the `log.warn` line — is now covered
deterministically where permissions apply.

`lastRemoveError` is a plain std::error_code written under fileMutex; the purge
read it unlocked, which can tear its value/category pair against a concurrent
retirement of the same file. Read it through a locked accessor.

Also trims the `nextReadableLogBuffer` header to the two constraints the code
cannot state itself; the rest restated invariant 22 verbatim.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CVAGTbmYzkKxvuH4HQjXKd

* docs: renumber duplicate AGENTS.md invariant 23 to 24

The rebase onto main moved this branch's purge invariant from 22 to 23
(main appended its own 22, "queued unlock callback"), which collided
with this branch's existing invariant 23 ("databaseFlushed persists and
verifies txn.state by pathname"). Bump the latter to 24.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs: point purge-invariant cross-references at 23, not 22

Three comments still cited "invariant 22" for the purge/read-coherence
rule after the rebase renumbered it to 23 (main's own new invariant 22,
queued unlock callback, now occupies that number). Caught by the
independent pre-push review on the rebased head.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix: preserve committed txnlog bounds across rotation

Co-Authored-By: GPT-5 Codex <noreply@openai.com>

* docs: renumber AGENTS.md invariants 23-24 to 24-25 after rebase

main's own rebase-picked invariant 23 (deferred column-family drops,
#850) now occupies the number this branch's purge invariant held.
Move purge to 24 and databaseFlushed to 25, and fix the two source
comments and the one AGENTS.md self-reference that cited the old
number 23 for the purge invariant.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Dispatch-Task: pr-maint-68bd037170be49b2a01e194607ecae40

* fix: close the last two creating-open ghost-recreate paths, and drop a hot-path refcount

Pre-push review (round 30, codex+cursor-grok+gemini+harper-domain) found
two more places this PR's own non-creating-open fix missed:

- TransactionLogStore::load() opened both the surviving current segment
  and each older segment scanned for a recovery boundary with the
  creating open(). A file deleted out of band between the directory
  scan and that open (this same load()) would silently come back as a
  header-only ghost, exactly the resurrection this PR closed on the
  read path (openIfPresent/openExisting). Both call sites now use
  openExisting() and treat a genuine absence like any other
  non-fatal open failure already handled there.

- TransactionLogFile::publishReadableExtentLocked() runs on every
  append (the commit hot path) and copied a shared_ptr<MemoryMap> to
  read one field, paying two atomic refcount ops for no ownership
  need. Now takes memoryMap.get() and only falls back to
  frozenMapCache.lock() (which has no raw-pointer equivalent) when
  there is no live memoryMap.

The refused-unlink-drops-current-segment-mapping finding from the same
round is real (three independent legs converged on it) but reachable
only through purgeLogs({ destroy: true }), which Harper does not run
in production, and touches the destroy path this PR has deliberately
left alone since the September 2 scope correction — recorded in the PR
body for the human reviewer instead of fixed here.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Dispatch-Task: pr-maint-68bd037170be49b2a01e194607ecae40

* docs: correct frozen-map ownership in README, trim review-narration comments

README claimed native code holds every file's map until purge/close;
frozen (rotated) files are only weakly cached (frozenMapCache) and
survive solely through JS Buffer references, independent of purge or
close. Also corrects stats.memory.activeMaps: it counts only the
current file's strongly-held map, not a frozen one kept alive by JS.

Trims three comments added by the last commit that narrated this PR's
own history ("this PR closed on the read path") instead of stating the
invariant itself.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Dispatch-Task: pr-maint-68bd037170be49b2a01e194607ecae40

* docs: renumber duplicate AGENTS.md invariant 25, fix stale invariant-24 refs

Rebasing onto main (which landed its own new invariant 24, the two
process-wide clocks note) collided with this branch's purge invariant,
also numbered 24. Resolved by keeping main's 24 and bumping the purge
invariant to 25 — but that collided with this branch's own
databaseFlushed invariant, already numbered 25 from an earlier rebase's
renumbering. Bumped databaseFlushed to 26 and fixed the three
"invariant 24" cross-references (AGENTS.md, transaction-log-reader.ts,
transaction_log_store.cpp) that meant the purge invariant.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Dispatch-Task: pr-maint-68bd037170be49b2a01e194607ecae40

* docs: qualify frozen-map weak-ownership claim as POSIX-only

The README's memory-map section said a frozen (rotated/purged) file's
map is weakly held and excluded from stats.memory.activeMaps
unconditionally. That's only true on POSIX. On Windows,
TransactionLogFile::getMemoryMapLocked() never applies the weak-for-
frozen optimization -- every frozen read re-creates the mapping and
re-pins it strongly for the file's life (by design, per the comment
above that function), and getStats() counts memoryMap into activeMaps
regardless of platform. Surfaced by this rebase's pre-push review.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Dispatch-Task: pr-maint-68bd037170be49b2a01e194607ecae40

* docs: renumber purge invariant 25->29, databaseFlushed 26->30 after rebase

main's own #787 landed invariants 26-28 in the same slot this branch's
purge invariant occupied (25), so the rebase conflict resolution moved
purge to 29 and its sibling databaseFlushed invariant (which collided
with main's new 26) to 30. Fixed the three stale in-code "invariant 25"
cross-references this displaced (AGENTS.md, transaction-log-reader.ts,
transaction_log_store.cpp), following the whole-tree re-grep lesson
already recorded in this PR's Findings.

Dispatch-Task: pr-maint-68bd037170be49b2a01e194607ecae40
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs: fix a stale invariant number and a wrong function name in comments

Pre-push review round 1 (codex) caught two factually wrong references
that predate this rebase and were never corrected by earlier renumbering
passes:

- test/transaction-log.test.ts:3112 cited "invariant 20" (now the
  transaction-timestamp invariant) for a property that invariant 29
  (the purge/mapping invariant) actually documents.
- src/transaction-log-reader.ts's corruptFrame() docstring named
  `getLogFileSize` as the store-mutex-taking call being avoided, but
  the caller resolves the extent via `readableExtent()` (a lock-free
  atomic accessor), not `getLogFileSize`.

Dispatch-Task: pr-maint-68bd037170be49b2a01e194607ecae40
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs: correct attribution in corruptFrame docstring

Pre-push delta review caught that my prior fix (404b19ef) misattributed
the readableExtent() resolution to "the caller" -- corruptFrame() itself
calls it (conditionally, on a logBuffer.size cache miss) at line 134 for
the readUncommitted branch. Describe the actual cache-then-native-fallback
mechanism instead.

Dispatch-Task: pr-maint-68bd037170be49b2a01e194607ecae40
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs: renumber purge invariant 30->30, databaseFlushed 30->31 after second rebase

main advanced again mid-review (PR #868, "Keep user shared buffers for
the life of the column family") and landed its own new invariant 29,
colliding with this branch's purge invariant. Kept both (main's 29
stays, purge moves to 30, databaseFlushed to 31) and re-grepped the
whole tree for stale "invariant 29" cross-references left behind by
the collision, per the lesson already recorded in this PR's Findings.

Dispatch-Task: pr-maint-68bd037170be49b2a01e194607ecae40
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs: trim redundant corruptFrame docstring

Round 3 (domain) flagged the docstring as restating the cache-then-
native-fallback pattern already visible in the code two lines below
(logBuffer.size ?? readableExtent(logBuffer)) plus the existing inline
comment above it. Cut it down to the one-line purpose statement.

Dispatch-Task: pr-maint-68bd037170be49b2a01e194607ecae40
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: GPT-5 Codex <noreply@openai.com>
Co-authored-by: Chris Barber <chris@harperdb.io>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants