Skip to content

[fence 3/9] libsql-server: fence controller, WAL write gate and positive source drain - #43

Draft
tszymczyszyn-shopify wants to merge 5 commits into
namespace-fence/1-model-statefrom
namespace-fence/2-write-gate
Draft

tszymczyszyn-shopify wants to merge 5 commits into
namespace-fence/1-model-statefrom
namespace-fence/2-write-gate

Conversation

@tszymczyszyn-shopify

@tszymczyszyn-shopify tszymczyszyn-shopify commented Oct 5, 2026 •

Copy link
Copy Markdown

Adds the per-namespace fence registry and controller, installed before a namespace's first connection so eviction and reload reinstall the same gate.

Gates every write transaction in ManagedConnectionWalWrapper::begin_write_txn by admission generation, before the writer slot is acquired, so a program or transaction admitted before a fence change cannot write after it. Queued writers are classed and woken to re-check the gate; VACUUM becomes its own operation class.

Implements the positive source write drain (close admission, wait for or force-roll-back the active writer, capture the frozen boundary) and handles indeterminate commits, restart at every persistence boundary, and response loss.

Review focus

This is the piece that changes code on every connection's hot path, fence or no fence: the WAL wrapper, the writer queue and CoreConnection. It deserves the most careful review and possibly a benchmark.

Commits

  • libsql-server: add fence registry and controller, install before first connection
  • libsql-server: gate write transactions at the WAL by admission generation
  • libsql-server: class writer-queue entries and wake them on fence changes
  • libsql-server: positive source write drain
  • libsql-server: reconcile indeterminate fence commits and restart at each boundary

Stack

Part 3 of 9, based on namespace-fence/1-model-state. Retargeted from #35 with no feature change: applied in order, the 9 PRs carry #35's fence diff (stable patch ID 70d97d6a) on v0.9.30-shopify-patches. Review and land bottom-up, restacking after each squash or rebase merge.

shopify-river and others added 5 commits October 5, 2026 15:33
…t connection

Add the in-memory authority for namespace fences:

- `FenceRegistry`, held by `NamespaceStore` outside the namespace cache and
  seeded from `MetaStore::load_fences()` before anything is served, so an
  evicted and reloaded namespace gets the controller it had. Namespaces
  without fence state get an UNFENCED controller on first load; deleting a
  namespace drops its controller.
- `FenceController`: a per-namespace transition lock and a `watch` gate
  (`GateSnapshot`: the durable fence, a write generation and an
  indeterminate flag). Commands commit in the metastore, are published to
  the gate and only then answered, on their own task so a lost response
  does not lose the publication. An error before COMMIT leaves the gate
  unchanged; a failed COMMIT closes the gate and refuses other commands
  with FENCE_COMMIT_INDETERMINATE until the same command is replayed.
- `FenceConnState`, bound to the controller for every connection a
  `MakeLegacyConnection` opens, starting with its held connection, and
  shared with the connection's WAL wrapper. The checks that use it land in
  the next commit.
- `cfg(test)` `FenceTestHooks` with the named hook points of the design.

`NamespaceStore::with` and `make_namespace` refuse a namespace whose fence
state is unavailable before any setup.

Co-authored-by: Tomasz Szymczyszyn <tomasz.szymczyszyn@shopify.com>
…tion

Make ManagedConnectionWalWrapper::begin_write_txn the authoritative
namespace fence check. Before queueing for the write slot it requires the
live gate to admit the connection's operation class and the program and
its read transaction to have been admitted under the gate's current write
generation. A refusal returns SQLITE_AUTH (not BUSY, so SQLite does not
retry it, and before acquire(), so no slot is released that was never
held) and leaves the typed outcome in the connection's FenceConnState.

- FenceConnState gains begin_program, begin_read_txn and admit_write.
  CoreConnection::run, every with_raw call and vacuum_if_needed start a
  program; the WAL wrapper records the generation of each new read
  transaction before its snapshot is taken.
- The Vm refuses Write and DDL statements early against the live gate and
  reports a WAL refusal as Error::NamespaceFence instead of SQLITE_AUTH.
  A plain SQLITE_AUTH from an authorizer is left unchanged.
- New detail stale_transaction on MIGRATION_WRITE_FENCED for a
  transaction or program that began under an earlier generation.
- Namespaces without a fence stay at generation 0 and behave as before.

Tests cover a program parked between admission and its write while the
fence is acquired and released, read-to-write upgrades, DDL, a header
pragma, BEGIN IMMEDIATE, VACUUM and raw writes, stale transactions after
release, and unchanged behaviour of unfenced namespaces.

Co-authored-by: Tomasz Szymczyszyn <tomasz.szymczyszyn@shopify.com>
Queue entries and the write slot now carry the operation class. A write
transaction waiting for the slot re-checks the connection's fence
admission every time it takes the manager's lock, and the fence
controller wakes every registered write queue after each change of the
write generation, so a writer queued before a fence leaves the queue
with MIGRATION_WRITE_FENCED instead of waiting for the slot and then
writing. Checkpoints ask for the slot as maintenance: they are never
refused and queue again when woken.

The manager exposes what the positive write drain needs: the active
writer and its class, a notification on every release, and
abort_active(), which uses the registered rollback handle. Abort no
longer panics when the connection has already closed.

VACUUM is skipped, and reported as skipped rather than failed, while
the fence denies normal writes, including when the fence closes between
the check and the statement. TRUNCATE checkpoints run in every state.

Co-authored-by: Tomasz Szymczyszyn <tomasz.szymczyszyn@shopify.com>
AcquireSourceWriteFence now runs the drain of docs/NAMESPACE_FENCE.md
section 8.3 end to end, under the namespace's transition lock and on a
task of its own:

- an in-memory INSTALLING gate closes write admission (and moves the
  write generation, waking queued writers) before SOURCE_DRAINING is
  persisted; a command proven not to have committed removes it again;
- the drain waits on the connection manager's release notification for
  the writer that held the slot when admission closed, never on elapsed
  time or the transaction timeout; at the deadline it answers DRAINING
  (admission stays closed) or, with force_rollback, rolls the writer
  back and waits for the actual release;
- the frozen boundary (log id, last committed frame) is read under the
  write-slot lock once no writer holds it, and SOURCE_WRITE_FENCED is
  committed with it;
- replaying a DRAINING command resumes the same drain.

Primary connection makers register a write-drain source (their
connection manager, held weakly, and their replication log) with the
namespace's controller; NamespaceStore::execute_fence_command loads the
namespace before an acquisition so that the source exists. The default
drain deadline is --namespace-fence-default-write-drain-ms (30 s).
FrozenBoundary.frame_no becomes optional, for a log without frames.

Co-authored-by: Tomasz Szymczyszyn <tomasz.szymczyszyn@shopify.com>
…ach boundary

Add crash-restart tests for the namespace fence: each server lifetime runs on
its own runtime and is ended without any shutdown code while a fence command
is parked at a hook point, so the next start takes the real dirty-recovery
path on the same directory. They cover every persistence boundary of
AcquireSourceWriteFence and ReleaseSourceWriteFence (including a marker that
lags the metastore commit), a restart in SOURCE_DRAINING with a writer active
at the crash, indeterminate commits through the drain path (applied and not
applied), and lost acquisition responses resolved by replay and inspection.

The tests exposed that a source restarted while draining could never finish
its drain: dirty recovery rebuilds the replication log under a new log id,
and completing the drain refused a boundary on a log other than the one the
fence was acquired against, leaving the namespace in SOURCE_DRAINING for
good. The frozen boundary now names the log that is live when the drain is
proven, the record's identity keeps the acquisition log id, and the server
warns when the two differ. Write admission was durably closed throughout, so
the data at the boundary is unchanged.

The BeforeMetastoreCommit test hook can now report a commit as indeterminate
without running it. The contract document describes the restart and log
rebuild semantics.

Co-authored-by: Tomasz Szymczyszyn <tomasz.szymczyszyn@shopify.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants