Warning

This is not authoritative documentation. It describes a plan of work that is not yet implemented and will change as it lands.

Task 7: DROP reclamation, fencing behavior, and operational surface

Repo: https://opendev.org/drizzle/drizzle. Floats after task 5.

Compaction-filter DROP

Replace task 3’s foreground prefix delete for large tables:

  • doDropTable writes meta deletions plus the 0x00 'd' <table-id> marker in one transaction and returns.

  • The shim gains the only planned additions to its surface after task 1: a prefix-tombstone CompactionFilterSupplier registered at open through DbBuilder::with_compaction_filter_supplier, plus slatedb_dropped_ids_add(db, id, out_err) and slatedb_dropped_ids_remove(db, id, out_err).

  • The API, checked against 0.14.1 rather than assumed (compaction_filter.rs): CompactionFilterSupplier:: create_compaction_filter(&CompactionJobContext) is called per compaction job; the returned CompactionFilter sees every entry through filter(&RowEntry) and answers Keep, Drop, or Modify(ValueDeletable::Tombstone); on_compaction_end is where the per-job tombstone count goes into engine metrics. We answer Modify(Tombstone) and never Drop — upstream documents that Drop removes an entry without shadowing it, which can resurrect older versions from lower runs. The API is behind the non-default compaction_filters cargo feature (slatedb/Cargo.toml); task 1 enables it. If that pin was missed, flipping it is the first commit of this task.

  • CompactionFilter’s snapshot-consistency warning is acceptable here by construction: dropped tables have no live readers — the server serialized the DROP against open cursors already, and the ids we filter appear in no committed metadata.

The dropped-id set is live

A set read once at open reclaims nothing dropped since the server started, which in a long-running server is every table anyone actually dropped. The set is shared, mutable, and advanced only by committed facts:

  • The supplier owns an Arc<RwLock<HashSet<u32>>>, seeded at open from a scan of the 0x00 'd' prefix.

  • create_compaction_filter takes the read lock once and copies the set into the job-local filter, so the per-entry filter() call touches no lock. Per-job filter instances are exactly what makes a per-job snapshot the right shape; a job that started before an id was added misses that table this round, and the next compaction takes it.

  • slatedb_dropped_ids_add is called only after the DROP transaction’s commit returns OK. Not before it, not on CONFLICT, not on rollback, not from a destructor, not on any path that is not a returned-OK commit. An id in the filter set whose DROP did not commit tombstones a live table’s rows.

  • slatedb_dropped_ids_remove is called only after the marker-deletion transaction commits, for the mirror-image reason. Removing first and then failing to commit merely leaves a marker nothing acts on until the next restart — not a correctness bug, but it is the same rule, so it is the same rule.

  • Crash safety needs no extra machinery: the 0x00 'd' prefix is the authority, the in-memory set is a cache rebuilt from it at every open, and both operations are idempotent. A crash between commit and set update costs a delayed reclamation, never a row.

  • Marker retirement: a periodic sweep (or the next open) probes each dropped prefix; empty → delete the marker, then remove the id.

  • Keep the foreground path for DELETE FROM t (delete_all_rows) and decide with perf-suite data whether small-table DROP keeps it too; record the threshold or its absence.

Fencing behavior

Engine handling of FENCED from any shim call: log at the highest severity with the operational explanation (“another server has opened this SlateDB path and now owns it”), fail the current statement with the mapped error, and poison the engine into refuse-all-writes until restart. A test in drizzle-test’s framework limits permitting a second server process (or a spike-binary stand-in opening the same path) asserts the behavior. Deployment documentation leads with the inverted-failover semantics per the spec.

Operational surface

  • slatedb_flush() UDF (mirrors wiredtiger_checkpoint()’s shape and registration).

  • A DATA_DICTIONARY.SLATEDB_STATUS table function: database path, format version, last flush, fenced flag, per-table id map — read from meta keys and engine state; modeled on the existing dictionary plugins.

  • User documentation (docs/index.rst): the durability triangle with the task-1 measured numbers, the single-writer/fencing model, checkpoint-based backup via the upstream slatedb CLI, and the refusals table from the spec.

Recorded follow-ons (not in this series)

BLOB support; AUTO_INCREMENT via reserved counter blocks; SSI as a session option; hidden-PK tables; DbReader-backed read-only replica mode (its own spec); exposing SlateDB checkpoints as a UDF; approximate row counters via SlateDB’s merge operator plus unmark_write (only if measurement ever shows the capped probe is not good enough — and with the caveat from the spec that an unmarked key is no longer transactional with the rows it counts).

Verification

  • drizzle-test green; DROP of a large seeded table returns fast and space is reclaimed after induced compaction (assert via SlateDB admin/manifest inspection through the CLI in the test harness).

  • The reclaimed table is dropped while the server keeps running, with no restart before the induced compaction. This is the regression test for the open-time-snapshot bug: a restart-based test passes on the broken design, so it must not be the only coverage.

  • A DROP whose transaction is made to fail (rollback or induced CONFLICT) leaves the table’s rows intact after forced compaction, and the id never enters the filter set.

  • Fencing test green.

  • Docs build green under the Sphinx -W gate.