.. warning:: This is not authoritative documentation. It describes a plan of work that is not yet implemented and will change as it lands. Task 7: DROP reclamation, fencing behavior, and operational surface =================================================================== Repo: https://opendev.org/drizzle/drizzle. Floats after task 5. Compaction-filter DROP ---------------------- Replace task 3's foreground prefix delete for large tables: - ``doDropTable`` writes meta deletions plus the ``0x00 'd' `` marker in one transaction and returns. - The shim gains the only planned additions to its surface after task 1: a prefix-tombstone ``CompactionFilterSupplier`` registered at open through ``DbBuilder::with_compaction_filter_supplier``, plus ``slatedb_dropped_ids_add(db, id, out_err)`` and ``slatedb_dropped_ids_remove(db, id, out_err)``. - **The API, checked against 0.14.1 rather than assumed** (``compaction_filter.rs``): ``CompactionFilterSupplier:: create_compaction_filter(&CompactionJobContext)`` is called **per compaction job**; the returned ``CompactionFilter`` sees every entry through ``filter(&RowEntry)`` and answers ``Keep``, ``Drop``, or ``Modify(ValueDeletable::Tombstone)``; ``on_compaction_end`` is where the per-job tombstone count goes into engine metrics. We answer ``Modify(Tombstone)`` and never ``Drop`` — upstream documents that ``Drop`` removes an entry without shadowing it, which can resurrect older versions from lower runs. The API is behind the **non-default** ``compaction_filters`` cargo feature (``slatedb/Cargo.toml``); task 1 enables it. If that pin was missed, flipping it is the first commit of this task. - ``CompactionFilter``'s snapshot-consistency warning is acceptable here by construction: dropped tables have no live readers — the server serialized the DROP against open cursors already, and the ids we filter appear in no committed metadata. The dropped-id set is live ~~~~~~~~~~~~~~~~~~~~~~~~~~ A set read once at open reclaims nothing dropped since the server started, which in a long-running server is every table anyone actually dropped. The set is shared, mutable, and advanced **only by committed facts**: - The supplier owns an ``Arc>>``, seeded at open from a scan of the ``0x00 'd'`` prefix. - ``create_compaction_filter`` takes the read lock once and copies the set into the job-local filter, so the per-entry ``filter()`` call touches no lock. Per-job filter instances are exactly what makes a per-job snapshot the right shape; a job that started before an id was added misses that table this round, and the next compaction takes it. - ``slatedb_dropped_ids_add`` is called **only after the DROP transaction's commit returns OK**. Not before it, not on ``CONFLICT``, not on rollback, not from a destructor, not on any path that is not a returned-OK commit. An id in the filter set whose DROP did not commit tombstones a live table's rows. - ``slatedb_dropped_ids_remove`` is called **only after the marker-deletion transaction commits**, for the mirror-image reason. Removing first and then failing to commit merely leaves a marker nothing acts on until the next restart — not a correctness bug, but it is the same rule, so it is the same rule. - Crash safety needs no extra machinery: the ``0x00 'd'`` prefix is the authority, the in-memory set is a cache rebuilt from it at every open, and both operations are idempotent. A crash between commit and set update costs a delayed reclamation, never a row. - Marker retirement: a periodic sweep (or the next open) probes each dropped prefix; empty → delete the marker, then remove the id. - Keep the foreground path for ``DELETE FROM t`` (``delete_all_rows``) and decide with perf-suite data whether small-table DROP keeps it too; record the threshold or its absence. Fencing behavior ---------------- Engine handling of ``FENCED`` from any shim call: log at the highest severity with the operational explanation ("another server has opened this SlateDB path and now owns it"), fail the current statement with the mapped error, and poison the engine into refuse-all-writes until restart. A test in drizzle-test's framework limits permitting a second server process (or a spike-binary stand-in opening the same path) asserts the behavior. Deployment documentation leads with the inverted-failover semantics per the spec. Operational surface ------------------- - ``slatedb_flush()`` UDF (mirrors ``wiredtiger_checkpoint()``'s shape and registration). - A ``DATA_DICTIONARY.SLATEDB_STATUS`` table function: database path, format version, last flush, fenced flag, per-table id map — read from meta keys and engine state; modeled on the existing dictionary plugins. - User documentation (``docs/index.rst``): the durability triangle with the task-1 measured numbers, the single-writer/fencing model, checkpoint-based backup via the upstream ``slatedb`` CLI, and the refusals table from the spec. Recorded follow-ons (not in this series) ---------------------------------------- BLOB support; AUTO_INCREMENT via reserved counter blocks; SSI as a session option; hidden-PK tables; ``DbReader``-backed read-only replica mode (its own spec); exposing SlateDB checkpoints as a UDF; approximate row counters via SlateDB's merge operator plus ``unmark_write`` (only if measurement ever shows the capped probe is not good enough — and with the caveat from the spec that an unmarked key is no longer transactional with the rows it counts). Verification ------------ - drizzle-test green; DROP of a large seeded table returns fast and space is reclaimed after induced compaction (assert via SlateDB admin/manifest inspection through the CLI in the test harness). - **The reclaimed table is dropped while the server keeps running**, with no restart before the induced compaction. This is the regression test for the open-time-snapshot bug: a restart-based test passes on the broken design, so it must not be the only coverage. - A DROP whose transaction is made to fail (rollback or induced ``CONFLICT``) leaves the table's rows intact after forced compaction, and the id never enters the filter set. - Fencing test green. - Docs build green under the Sphinx ``-W`` gate.