Warning
This is not authoritative documentation. It describes a plan of work that is not yet implemented and will change as it lands.
Task 7: DROP reclamation, fencing behavior, and operational surface¶
Repo: https://opendev.org/drizzle/drizzle. Floats after task 5.
Compaction-filter DROP¶
Replace task 3’s foreground prefix delete for large tables:
doDropTablewrites meta deletions plus the0x00 'd' <table-id>marker in one transaction and returns.The shim gains the only planned additions to its surface after task 1: a prefix-tombstone
CompactionFilterSupplierregistered at open throughDbBuilder::with_compaction_filter_supplier, plusslatedb_dropped_ids_add(db, id, out_err)andslatedb_dropped_ids_remove(db, id, out_err).The API, checked against 0.14.1 rather than assumed (
compaction_filter.rs):CompactionFilterSupplier:: create_compaction_filter(&CompactionJobContext)is called per compaction job; the returnedCompactionFiltersees every entry throughfilter(&RowEntry)and answersKeep,Drop, orModify(ValueDeletable::Tombstone);on_compaction_endis where the per-job tombstone count goes into engine metrics. We answerModify(Tombstone)and neverDrop— upstream documents thatDropremoves an entry without shadowing it, which can resurrect older versions from lower runs. The API is behind the non-defaultcompaction_filterscargo feature (slatedb/Cargo.toml); task 1 enables it. If that pin was missed, flipping it is the first commit of this task.CompactionFilter’s snapshot-consistency warning is acceptable here by construction: dropped tables have no live readers — the server serialized the DROP against open cursors already, and the ids we filter appear in no committed metadata.
The dropped-id set is live¶
A set read once at open reclaims nothing dropped since the server started, which in a long-running server is every table anyone actually dropped. The set is shared, mutable, and advanced only by committed facts:
The supplier owns an
Arc<RwLock<HashSet<u32>>>, seeded at open from a scan of the0x00 'd'prefix.create_compaction_filtertakes the read lock once and copies the set into the job-local filter, so the per-entryfilter()call touches no lock. Per-job filter instances are exactly what makes a per-job snapshot the right shape; a job that started before an id was added misses that table this round, and the next compaction takes it.slatedb_dropped_ids_addis called only after the DROP transaction’s commit returns OK. Not before it, not onCONFLICT, not on rollback, not from a destructor, not on any path that is not a returned-OK commit. An id in the filter set whose DROP did not commit tombstones a live table’s rows.slatedb_dropped_ids_removeis called only after the marker-deletion transaction commits, for the mirror-image reason. Removing first and then failing to commit merely leaves a marker nothing acts on until the next restart — not a correctness bug, but it is the same rule, so it is the same rule.Crash safety needs no extra machinery: the
0x00 'd'prefix is the authority, the in-memory set is a cache rebuilt from it at every open, and both operations are idempotent. A crash between commit and set update costs a delayed reclamation, never a row.Marker retirement: a periodic sweep (or the next open) probes each dropped prefix; empty → delete the marker, then remove the id.
Keep the foreground path for
DELETE FROM t(delete_all_rows) and decide with perf-suite data whether small-table DROP keeps it too; record the threshold or its absence.
Fencing behavior¶
Engine handling of FENCED from any shim call: log at the highest
severity with the operational explanation (“another server has opened
this SlateDB path and now owns it”), fail the current statement with
the mapped error, and poison the engine into refuse-all-writes until
restart. A test in drizzle-test’s framework limits permitting a second
server process (or a spike-binary stand-in opening the same path)
asserts the behavior. Deployment documentation leads with the
inverted-failover semantics per the spec.
Operational surface¶
slatedb_flush()UDF (mirrorswiredtiger_checkpoint()’s shape and registration).A
DATA_DICTIONARY.SLATEDB_STATUStable function: database path, format version, last flush, fenced flag, per-table id map — read from meta keys and engine state; modeled on the existing dictionary plugins.User documentation (
docs/index.rst): the durability triangle with the task-1 measured numbers, the single-writer/fencing model, checkpoint-based backup via the upstreamslatedbCLI, and the refusals table from the spec.
Recorded follow-ons (not in this series)¶
BLOB support; AUTO_INCREMENT via reserved counter blocks; SSI as a
session option; hidden-PK tables; DbReader-backed read-only
replica mode (its own spec); exposing SlateDB checkpoints as a UDF;
approximate row counters via SlateDB’s merge operator plus
unmark_write (only if measurement ever shows the capped probe is
not good enough — and with the caveat from the spec that an unmarked
key is no longer transactional with the rows it counts).
Verification¶
drizzle-test green; DROP of a large seeded table returns fast and space is reclaimed after induced compaction (assert via SlateDB admin/manifest inspection through the CLI in the test harness).
The reclaimed table is dropped while the server keeps running, with no restart before the induced compaction. This is the regression test for the open-time-snapshot bug: a restart-based test passes on the broken design, so it must not be the only coverage.
A DROP whose transaction is made to fail (rollback or induced
CONFLICT) leaves the table’s rows intact after forced compaction, and the id never enters the filter set.Fencing test green.
Docs build green under the Sphinx
-Wgate.