SlateDB Storage Engine: Design

Scope and thesis

A Drizzle storage engine plugin (plugin/slatedb) that stores OLTP tables in object storage (S3, GCS, ABS, MinIO) through SlateDB — a Rust embedded LSM engine whose SSTs, WAL, and manifest all live in the object store. The use case: a Drizzle server whose durable state is a bucket. No local data directory to lose, no EBS volume to snapshot, no replication to configure — durability and capacity are the object store’s problem.

The relationship to the Iceberg engine is complementary, not competitive: Iceberg is the open-format analytical/cold tier that other engines can read; SlateDB is a private-format transactional tier that happens to live in the same kind of bucket. Nothing but this engine reads a SlateDB bucket. The two engines cover the two halves of “Drizzle on object storage.”

The subtractive frame, as always: this engine is a deliberate subset. One SlateDB database per server. Snapshot isolation only. No savepoints. Refused features are refused loudly with a specific error, naming what and why — never approximated. The WiredTiger PLAN.md Tier-0 audit is treated as a checklist of silent-wrong-results bugs this engine must be structurally unable to have: every one of them (dup-key overwrite, find-flag ignorance, fabricated records(), no-op statement rollback, secondary-index write corruption) has a named answer in this design.

Why SlateDB specifically

Evaluated against the actual 0.14.1 tree, not the marketing page:

  • Real transactions. Db::begin(IsolationLevel) → DbTransaction with buffered writes, read-your-own-writes inside the transaction (the write batch backs reads — verified in db_transaction.rs), write-write conflict detection at commit, and both Snapshot and SerializableSnapshot isolation. This maps onto Drizzle’s TransactionalStorageEngine contract almost one-to-one.

  • Ordered scans both directions. scan/scan_prefix over byte ranges with IterationOrder::Ascending/Descending and seek — everything index_next/index_prev/ records_in_range need, provided our keys are memcomparable (below).

  • Writer fencing. Manifest-epoch fencing (fence.rs): a second writer opening the same path fences the first, which starts failing with Fenced. Split-brain is structurally impossible; the operational consequence is documented below.

  • Read replicas for free (future). DbReader opens the same bucket read-only, following the latest state or pinned to a checkpoint. A read-only Drizzle replica is a config change away once the engine exists. Designed-for, not built.

  • Explicit conflict-set control. DbTransaction::mark_read adds keys to the read set even under Snapshot — whose get and scan otherwise track nothing — and unmark_write removes keys from the write set. Both verified in db_transaction.rs. The design below leans hard on knowing exactly which set a call lands in, so this is listed as a feature, not trivia.

  • Compaction filters. CompactionFilter / CompactionFilterSupplier (compaction_filter.rs): a filter sees every entry during compaction and returns Keep, Drop, or Modify(ValueDeletable::Tombstone) — the lazy-reclamation mechanism for DROP TABLE (phase 2 of drop, below). Caveat recorded once here and honoured in the build: the API sits behind the non-default compaction_filters cargo feature in 0.14.1 (slatedb/Cargo.toml), which the shim crate enables from its first commit.

  • Tunable durability. WriteOptions::await_durable and Settings::flush_interval expose the exact latency/cost/durability triangle object storage forces; we surface it as engine options instead of hiding it.

Library landscape (as of July 2026)

  • slatedb 0.14.1 (Apache-2.0, Commonhaus Foundation). Workspace pins Rust 1.91.1 via rust-toolchain.toml. Pre-1.0: API churn expected; pin the exact version in the shim’s Cargo.toml and Cargo.lock, both living in one place (the shim crate).

  • object_store 0.14 (Rust crate, Apache-2.0) is SlateDB’s storage substrate: S3, GCS, Azure, local filesystem, in-memory. Db::resolve_object_store(url) builds one from a URL plus environment credentials — s3://bucket/path for production, file:///path for CI without MinIO, MinIO for the S3-semantics tests. We expose the URL directly as the engine’s storage option and invent no abstraction over it.

  • No C ABI upstream. SlateDB’s foreign bindings (Go, Java, Python, Node) all ride uniffi, which has no C++ target. We write our own shim (next section) — narrower and better-fitted than anything generated.

License note, following the posture established by the Iceberg series (Arrow/iceberg-cpp are likewise Apache-2.0): SlateDB is Apache-2.0, consumed as a separately-built shared library through a C ABI. Task 1’s README must include the transitive dependency list with licenses (cargo license) and confirm nothing GPLv2-incompatible beyond the already-accepted Apache-2.0 posture rides along.

The FFI shim: libslatedb_capi

The plugin never sees Rust. A small Rust crate, slatedb-capi (living in plugin/slatedb/capi/ in the drizzle tree), builds a cdylib — libslatedb_capi.so — plus one hand-written, versioned C header. The Drizzle plugin is ordinary C++ that links -lslatedb_capi, exactly as the WiredTiger plugin links -lwiredtiger.

Design rules for the shim:

  • Narrow. Only what the engine calls: database open/close/flush, transaction begin/commit/rollback, transactional get/put/delete, scan open/seek/next/close (with direction), and error-message retrieval. No settings struct mirroring — configuration crosses the boundary as one string (SlateDB’s own Settings file format, which figment parses), so new SlateDB knobs cost zero shim changes.

  • Sync bridge. The shim owns a multi-threaded tokio runtime inside the database handle; every call is runtime.block_on. Drizzle sessions are threads that already block on disk I/O; blocking them on object-store I/O is the same shape with bigger constants. No async leaks into C++.

  • C ABI, not CXX. This boundary is bytes-in/bytes-out with opaque handles — the degenerate case where CXX buys nothing. Plain extern "C", opaque pointers, (ptr, len) byte slices, and integer status codes paired with owned error objects (next bullet). (The transaction-replication Rust work remains the planned CXX proving-ground; this shim neither depends on nor advances that.)

  • Owned errors everywhere; no per-handle ``last_error``. Every shim function returns a status code and takes a trailing slatedb_error_t **out_err; on any non-OK status it writes an owned error object there, which the caller releases with slatedb_error_free (slatedb_error_message borrows the string until then). Uniformly — no function is exempt, because a rule with exceptions is a rule nobody remembers at the call site. The rejected alternative is the conventional per-handle last_error string, and the reason it is rejected is concrete: the handle can already be gone when you want the message. DbTransaction::commit and rollback take self by value in 0.14.1 (db_transaction.rs), so a failing commit has no transaction left to hang a message on; slatedb_scan_close has the identical hazard; and hoisting the string to the parent Db handle would make it shared mutable state raced by every session thread.

  • Consuming calls are stated, not inferred. slatedb_txn_commit(txn, out_err) and slatedb_txn_rollback(txn, out_err) always consume the transaction handle — on success and on failure alike. After either call the pointer is dangling, must never be passed to another shim function, and needs no separate free: there is no slatedb_txn_free, and a second call on the same pointer is a use-after-free, not a leak. The shim moves the DbTransaction out of its box and drops the box, mirroring the Rust signature exactly rather than hiding it behind an Option<DbTransaction> that would let C++ keep using a dead handle and call it safe. The engine consequently clears its session slot before it looks at the returned status.

  • Error taxonomy at the boundary. The header defines the small set the engine dispatches on: OK, NOT_FOUND, CONFLICT (transaction commit conflict), FENCED, INVALID, IO. Everything else is IO plus the message string. The C++ side maps these to HA_ERR_* in one function (slatedb_to_drizzle_error.cc), mirroring wt_to_drizzle_error.cc.

  • Ownership across the boundary is the registry pattern. Handles are created and destroyed by paired shim functions; no Rust allocation is ever freed by C++ or vice versa — except for the consuming pair above, which is the whole reason that pair is written down. Scan results are valid until the next next/close on that scan — the C++ side copies into the Drizzle record buffer immediately, so no lifetime-extension API is needed.

Build and dependency policy: container-first, and the composition note

The binding policy from the Iceberg series applies verbatim: the only build command is podman build fronted by just targets; every pin lives in a Containerfile (or, here, Cargo.lock + rust-toolchain.toml inside the image build); no dependency detection anywhere; plugin inclusion is an explicit build switch (--with-slatedb, default off) that the containerized build always passes. Cargo is upstream’s private build system, invoked by a RUN line, exactly as CMake is for WiredTiger — it never leaks into autotools, bindep, or Zuul configuration. This deliberately does not advance the “cargo integrated into the drizzle build” prerequisite from the broader Rust plan; it sidesteps it, and that is a feature: the shim is a leaf artifact with its own toolchain, not the first tenant of a shared Rust build.

The recorded open puzzle from the Iceberg spec is now due. The WiredTiger image landed; slatedb-capi is the second source-built engine dependency, which was the named trigger for deciding how N per-engine dependencies compose into one builder-image lineage. The decision, made here so it is a decision and not an accident:

Per-dependency build-stage images, composed by ``COPY –from=``. Each engine dependency gets its own image built FROM quay.io/drizzle/libdrizzle (WiredTiger already does this; slatedb-capi follows), installing into /usr/local through a DESTDIR so the payload is an exact file set. The drizzle builder image composes the enabled set with one COPY --from=quay.io/drizzle/<dep> /install/usr/local/ /usr/local/ line per dependency, then ldconfig. Discipline required and checked in CI: payloads must be file-disjoint (a trivial find | sort | uniq -d check across payload manifests in the builder build) and sonames unique. Rejected alternatives: one fat deps image (accretes, single cache-bust point, couples release cadences) and per-plugin builder images (multiplies the expensive drizzle build). This paragraph is the “short design note” the Iceberg spec called for; if composition grows real complexity later it graduates to its own spec.

Data model: one database, prefixed keyspaces

One SlateDB ``Db`` per server, opened at plugin init from slatedb.url + slatedb.path, closed at shutdown. Not one per table: each Db carries a tokio runtime, WAL flush loop, manifest poller, and compactor — per-table instances would multiply all of that and turn CREATE TABLE into an object-store fencing ceremony. Tables are key prefixes.

Key layout (all keys in one keyspace, first byte discriminates):

0x00 'p' <table-path>                       → serialized message::Table proto
0x00 'i' <table-path>                       → table-id (u32 BE)
0x00 'c'                                    → next-table-id counter
0x00 'd' <table-id BE>                      → dropped-table marker (phase-2 drop)
0x01 <table-id u32 BE> <index-id u8> <encoded-key>   → row / index entry

index-id 0 is the primary key; 1..N are secondary indexes in proto order. Every table or index scan is a SlateDB prefix scan on the 6-byte 0x01 || table-id || index-id prefix; records_in_range is a bounded range scan. Table protos live in the meta prefix exactly as the WiredTiger engine stores them in its table:table_definitions WT table — doGetTableDefinition reads the proto key (EEXIST/ENOENT per convention), doGetTableIdentifiers scans the 0x00 'p' prefix, and no table-definition files are ever written.

DDL is transactional by construction: CREATE TABLE writes the proto, the id mapping, and the bumped counter in one transaction commit; RENAME rewrites the two meta keys (data keys carry only the id, so rename never touches rows); DROP deletes the meta keys and writes the 0x00 'd' marker in one commit (reclamation below).

Key and value encoding

SlateDB compares keys as raw bytes, so the key codec must be memcomparable — encoded order equals SQL order. This is the one substantial piece of new C++ in the engine and it is written as a standalone, exhaustively unit-tested codec library (plugin/slatedb/codec/) before any cursor code exists. Prior art is MyRocks’ key encoding; we take the ideas, not the code.

Per-type key encoding:

  • Nullability prefix: nullable key parts get one byte — 0x00 NULL, 0x01 present. NULLs sort first, matching the server’s ordering.

  • Signed integers (LONG, LONGLONG, DATE, TIME, ENUM as its underlying int): big-endian with the sign bit flipped.

  • Unsigned / TIMESTAMP / MICROTIME: big-endian.

  • DOUBLE: IEEE-754 total-order trick — positive: flip sign bit; negative: flip all bits.

  • BOOLEAN, UUID, IPV6: fixed-width big-endian byte images.

  • DECIMAL: Drizzle’s my_decimal binary format is already memcomparable for fixed (precision, scale) — used as-is, which is also what the server’s own key format relies on.

  • VARCHAR/CHAR (collated text): the column collation’s strnxfrm weight string, then byte-escaped (each 0x00 → 0x00 0xFF) and terminated with 0x00 0x00 so shorter strings sort before their extensions and multi-part keys cannot alias.

  • BLOB in keys: refused in v1 (HTON_NO_BLOBS — matching the WiredTiger engine’s current posture; blob columns and blob keys arrive together in a later task).

Keys are never decoded. strnxfrm is one-way, so the design commits to it globally rather than special-casing: the row value stores all columns including primary-key columns, and every read materializes from the value. A secondary index entry carries the encoded PK, which is consumed verbatim as bytes (it is already the primary key’s encoded form) to point-read the primary entry — but where it carries it depends on whether the index is unique, and the reason is concurrency control, not space:

  • Non-unique index: key = index-cols-encoded || pk-encoded, value = empty. The PK suffix is what keeps equal-valued rows distinct; a lookup slices it off at the encoded-PK boundary.

  • Unique index, no NULL in the index columns: key = index-cols-encoded — no PK suffix — value = pk-encoded. Two sessions inserting the same unique value with different primary keys then write the same physical key, which is the only thing SlateDB’s conflict detector looks at (below); exactly one of them commits. A lookup reads the PK out of the value.

  • Unique index, any NULL in the index columns: the non-unique shape, PK suffix and empty value. SQL says NULL never equals NULL, so multiple NULL-bearing rows must coexist in a UNIQUE index — giving them distinct physical keys is that rule, expressed in the layout. The alternative (one shape plus a “skip the conflict when NULL” rule at commit time) is not implementable: conflict detection is a set intersection over key bytes inside SlateDB, with no hook to consult and no schema to consult it with. Push the semantics into the bytes and there is nothing left to get wrong.

Key structure stays parseable even though key values are not decodable — nullability prefix bytes, fixed widths, and the 0x00 0x00 string terminator make part boundaries recoverable. That is what lets a reader find the PK-suffix boundary (already true before this refinement) and what lets the encoder decide, locally and cheaply, whether a unique key has a NULL part and therefore which shape to write. No decode anywhere, no HA_KEYREAD_ONLY, and index-only-scan optimization is explicitly refused rather than half-supported.

The consequence worth writing down twice: key bytes embed collation weight tables, so a collation-table change is an on-disk compatibility event. The charset/collation codegen work must treat shipped weight tables as frozen-per-version; the engine stamps a format version into the meta prefix at database creation and refuses to open a future incompatible format rather than misreading it.

Value encoding: version byte, null bitmap, then each non-null field in proto column order — fixed-width types as fixed-width little-endian images, var-width as varint-length-prefixed bytes. Not memcomparable, not clever; versioned so it can evolve. position()/rnd_pos use the encoded PK bytes length-prefixed in ref, exactly the WiredTiger cursor’s pattern.

Transactions and isolation

The engine is a TransactionalStorageEngine. Per-session state lives in the Session::getEngineData slot (the innobase/WiredTiger pattern): the shim transaction handle plus a small statement-state struct.

  • Begin on first use, via doStartStatement — adopting the WiredTiger engine’s fix verbatim (startTransaction stays a no-op so non-SlateDB transactions don’t inflate Handler_commit). First touch calls slatedb_txn_begin and registers the engine as a transaction resource.

  • Isolation: snapshot only. begin(IsolationLevel::Snapshot) always. The server’s READ COMMITTED / READ UNCOMMITTED requests are refused with an error, not silently upgraded — the same collapse the WiredTiger 11 port made, for the same reason, with the same user-visible honesty. SlateDB’s SerializableSnapshot is real and tested upstream; exposing it as a per-session option is a cheap later task, listed, not done.

  • Commit: slatedb_txn_commit. A CONFLICT return (write-write conflict with a concurrently committed transaction) maps to HA_ERR_LOCK_DEADLOCK — the one error code every SQL application already retries on. This is first-committer-wins optimistic concurrency and the docs say so plainly: hot-row workloads will see deadlock-shaped retries instead of lock waits.

  • Rollback: whole-transaction rollback is free (DbTransaction::rollback drops the buffered batch). Statement rollback and savepoints are refused: SlateDB’s transaction batch cannot be truncated to a mark, so doSetSavepoint returns not-supported and a failed statement inside a multi-statement transaction rolls back the whole transaction and reports exactly that — the NDB-style answer the WiredTiger plan chose over the silent-no-op it audited (Tier 0.7). Never claim success for a rollback that didn’t happen.

  • Read-your-own-writes works — the transaction’s write batch backs its own reads and scans, so multi-statement INSERT-then-SELECT behaves as SQL requires. (This is the capability that separates this engine from the Iceberg engine’s deliberately-blind commit buffer.)

  • DDL vs. fencing: unlike WiredTiger there is no EBUSY-style DDL/checkpoint interaction — DDL is ordinary transactional writes to meta keys. The pre-DDL-checkpoint fixture pattern does not carry over because the problem it fixes does not exist here.

Duplicate keys are detected, not overwritten (WiredTiger Tier 0.2/0.3 learned the hard way): doInsertRecord probes the primary key with a transactional get before put (a memtable/cache hit in the common case), and probes each non-NULL-bearing UNIQUE secondary with a point get on its unique-columns-only key, returning HA_ERR_FOUND_DUPP_KEY with the correct key number; the concurrent case falls through to the commit-time conflict described below. INSERT … ON DUPLICATE KEY UPDATE therefore works through the standard kernel path with no engine-side replace shortcut. A key-changing UPDATE is delete-old + insert-new with dup probe, decided by comparing old/new encoded keys (Tier 0.4). Writes during a secondary-index-driven scan go through primary-key point writes, never through the scan’s iterator (Tier 0.5) — which is natural here since scans and writes are separate shim objects by construction.

What the conflict detector actually sees

Uniqueness enforcement rests on this, so it is verified and written down rather than assumed. All of it read out of 0.14.1’s db_transaction.rs and transaction_manager.rs:

  • Under IsolationLevel::Snapshot only the write set is tracked. get/scan add nothing: get_key_value_with_options guards its track_read_keys call on isolation_level == SerializableSnapshot, and check_has_conflict says in as many words that Snapshot “only checks write-write conflicts”.

  • The commit-time check is has_write_write_conflict: a set intersection of this transaction’s write keys against the write keys of every transaction that committed after this one started. So two transactions that put the same key do conflict, and the loser gets SlateDBError::TransactionConflict. Exact bytes, exact keys — no ranges, no predicates, no schema.

  • mark_read(keys) opts named keys into read-write detection even under Snapshot, and unmark_write(keys) opts named keys out of write-write detection. Both are real; neither takes a range, which is why neither rescues a predicate probe.

The consequence, and it is the load-bearing sentence of this section: a probe is not a lock. Reading a key and finding it absent buys nothing under Snapshot; only writing a key does. Every uniqueness guarantee in this engine must therefore reduce to “the competing transactions write the same physical key” — which is precisely why the unique-index layout above drops the PK suffix. With the suffix, two inserters of the same unique value write two different keys, both probe-miss, and both commit: a silently violated constraint, the WiredTiger Tier 0.3 failure mode wearing a different hat.

The probe stays, as an optimization on top of the conflict rather than as the mechanism. The two outcomes, both correct, stated so that neither surprises anyone:

  • Uncontended — the duplicate was already committed when this statement started: the probe hits and the engine returns HA_ERR_FOUND_DUPP_KEY with the right key number. This is essentially all real duplicate-key traffic.

  • Contended — two concurrent inserts of the same primary key, or of the same unique value: both probe-miss, both write the same physical key, and the second committer gets CONFLICT → HA_ERR_LOCK_DEADLOCK rather than HA_ERR_FOUND_DUPP_KEY. This is acceptable and it is not worked around. The loser’s transaction is rolled back either way; deadlock is the one error every SQL client already retries; and on retry the winner’s row is committed, so the probe hits and the duplicate is reported properly. The only way to return the “right” error here would be to serialize inserts, which is the thing optimistic concurrency exists to avoid. The user documentation states it next to the hot-row retry note.

Hot shared keys, audited

First-committer-wins turns any key that many concurrent transactions write into a serialization point delivered as retries. The design therefore owns a list of every shared key it creates, and says why each is acceptable — “fine” as a judgement, not as an oversight:

  • 0x00 'c' next-table-id counter — fine. Only CREATE TABLE read-modify-writes it, DDL is rare and human-paced, and a conflict costs one retry of one cheap transaction.

  • 0x00 'd' <id> dropped-table markers — fine. Written once by the DROP that creates one, deleted once by the sweep that retires it; no two transactions contend for the same marker.

  • Per-table meta keys (0x00 'p', 0x00 'i') — fine. DDL-only and per-table, and concurrent DDL on one table is already serialized above the engine.

  • Unique-index keys — contended by design. That contention is the enforcement, and it only occurs for exactly the duplicate the application asked us to reject.

  • A maintained per-table row-count key — rejected; see “Statistics honesty” below. It would have made every writer to a table conflict with every other writer to that table, on one key.

  • AUTO_INCREMENT counters — refused in v1, and this is why the recorded plan is block reservation rather than a shared counter key touched once per insert.

Durability, latency, and cost

Object storage makes the write-latency triangle explicit; the engine exposes SlateDB’s own knobs and adds none:

  • slatedb.url (required) and slatedb.path — the object store URL and database path.

  • slatedb.settings-file — optional path handed through the shim verbatim to SlateDB’s figment-based Settings loader (flush_interval, caches, compactor tuning, all of it). We do not mirror individual SlateDB settings as engine options; one passthrough, zero drift.

  • slatedb.await-durable (default on): commits block until the WAL batch is durable in object storage — correctness-first, with commit latency ≈ flush_interval + one PUT round trip, and the documentation states those numbers rather than apologizing for them. Off: commit returns at memtable acceptance; a crash can lose the tail up to the last flush, stated equally plainly. This is per-server policy, not per-table.

  • Read latency is governed by SlateDB’s in-memory block cache and optional local-disk object cache (cached_object_store), configured through the settings file; the perf-tracking suite measures hot/cold read split explicitly.

flush() on the engine maps to Db::flush; a slatedb_flush() UDF mirrors wiredtiger_checkpoint() for operational use.

Statistics honesty (Tier 0.6): no ``HTON_HAS_RECORDS``, no HTON_STATS_RECORDS_IS_EXACT. info() reports estimates and says so. Nothing in this engine claims an exact row count, and nothing in it maintains one.

  • Row count, v1: a capped prefix probe. Scan the table’s 0x01 || table-id || 0x00 prefix up to a fixed row cap (a constant, tuned against the perf suite); exhausted under the cap → the exact count; hit the cap → a scaled estimate from key-space sampling. Cheap, honest, and wrong only in the direction optimizers already assume estimates are wrong.

  • Cardinality and records_in_range: bounded range probes under the same cap, and honest unknown-cardinality defaults where a probe is not worth its I/O. Never a fabricated floor — that is the Tier 0.6 anti-pattern itself.

Rejected, explicitly: a per-table row-count key maintained transactionally with writes. It is the obvious way to get an exact count and it is a design bug in this engine. Under first-committer-wins every concurrent writer to a table would write that one key and therefore write-write conflict with every other writer to the same table: a single hot key serializing all writes to the table it is counting, with the serialization delivered as HA_ERR_LOCK_DEADLOCK storms rather than as waits. Exactness is not worth converting a concurrent engine into a serial one, and this engine does not need it, because it does not claim it.

For the record, since it will be asked: a conflict-free counter is constructible on 0.14.1. There is a merge operator (merge_operator.rs, DbTransaction::merge), and unmark_write lifts a key out of write-write detection, so a per-statement delta merged into an unmarked counter key would not conflict. It is listed as a possible later task, not v1 work, and the caveat travels with the listing: merge operands still land in the write batch (batch.rs: keys() returns every op’s key, merges included), so the counter is conflict-free only if the unmark is applied without exception — and an unmarked key is by definition no longer transactional with the rows it counts. That tradeoff earns its own evaluation with measurements, not a paragraph here.

DROP TABLE reclamation

Two phases, both landing eventually, first one first:

  1. v1 — foreground delete: DROP scans the table’s data prefix and issues batched deletes, then removes meta keys. Correct, simple, and slow for huge tables over S3; the docs say so.

  2. Compaction-filter reclamation: DROP writes only the meta deletion + 0x00 'd' <id> marker; a compaction filter tombstones dead-prefix entries as compaction naturally rewrites them; a background sweep retires the marker once the prefix is empty. This is the steady-state design; v1’s foreground path remains as DROP TABLE’s behavior under a size threshold or explicit option if the split proves worth keeping — decide with data from the perf suite, not in advance.

The API as it actually exists in 0.14.1 (compaction_filter.rs, read, not assumed): a CompactionFilterSupplier is registered once at open through DbBuilder::with_compaction_filter_supplier, and its create_compaction_filter(&CompactionJobContext) is called per compaction job; the returned CompactionFilter sees every entry through filter(&RowEntry) and answers Keep, Drop, or Modify(ValueDeletable::Tombstone), with on_compaction_end as the hook for per-job counters. We answer Modify(Tombstone), never Drop: upstream documents that Drop removes an entry without shadowing it, which can resurrect older versions from lower runs. The feature is off by default (compaction_filters in slatedb/Cargo.toml) and the shim crate turns it on. Upstream’s snapshot-consistency warning is acceptable here by construction — the ids we filter belong to tables that appear in no committed metadata and have no live readers, because the server serialized the DROP against open cursors already.

The dropped-id set is live, not a snapshot taken at open. A set read once at open would reclaim nothing dropped since the server started — which, in a long-running server, is every table anyone actually dropped. The shim owns it:

  • The filter supplier holds an Arc<RwLock<HashSet<u32>>>, seeded at open from a scan of the 0x00 'd' prefix.

  • create_compaction_filter takes the read lock once and copies the set into the job-local filter, so the per-entry filter() call touches no lock at all. Per-job filter instances exist precisely so that per-job snapshots are the intended shape; a job that began before an id was added simply misses that table this round and the next compaction takes it.

  • Two shim functions, slatedb_dropped_ids_add(db, id, out_err) and slatedb_dropped_ids_remove(db, id, out_err), mutate the set. They are the only additions to the shim surface after task 1.

  • The ordering is the entire correctness argument. doDropTable calls add only after the DROP transaction’s commit returns OK — not before it, not on CONFLICT, not on rollback, not from a destructor, not on any path that is not a returned-OK commit. An id in the filter set whose DROP did not commit tombstones a live table’s rows.

  • Symmetrically the retirement sweep calls remove only after the marker-deletion transaction commits. Removing first and then failing to commit merely leaves a marker nothing acts on until the next restart — but it is the same rule, so it is the same rule.

  • Crash safety needs no extra machinery: the 0x00 'd' prefix in the database is the authority, the in-memory set is a cache rebuilt from it at every open, and both operations are idempotent. A crash between commit and set update costs a delayed reclamation, never a row.

Operational model

  • Exactly one writing server per (url, path). SlateDB’s epoch fencing makes a second writer win: the new server fences the old, whose SlateDB calls start failing FENCED. The engine treats FENCED as fatal — log loudly and refuse further writes (HA_ERR_* mapped to a read-only/IO error), because the honest description of a fenced writer is “another server owns this bucket now.” This inverts the usual failover ergonomics (starting the replacement kills the incumbent) and the deployment docs must lead with it. It is also exactly what makes lease-less failover work.

  • Backup = checkpoint. SlateDB checkpoints (create_checkpoint, Admin::create_detached_checkpoint) are consistent named references in the same bucket; the slatedb CLI manages them. A slatedb_checkpoint() UDF is a natural follow-on once wanted.

  • Read replicas (future, designed-for): DbReader against the same path, following latest or pinned to a checkpoint. The engine split (writer engine vs. reader engine mode) is a later spec; the key/value codec and the one-Db model are the parts that must not foreclose it, and they don’t.

Refused in v1 (all loud, all specific)

Refusal

Why / when it could return

READ COMMITTED / READ UNCOMMITTED

Snapshot-only, per the WiredTiger precedent; never returns

Savepoints / statement rollback

SlateDB txn batch is not truncatable; whole-txn rollback reported honestly

SERIALIZABLE (SSI)

Cheap later task — SlateDB supports it

BLOB columns and blob keys

Later task (value codec is ready; cursor blob buffer management is the work)

AUTO_INCREMENT

Later task: transactional counter key with block reservation; refuse rather than ship a racy version

Index-only scans (HA_KEYREAD_ONLY)

Keys are strnxfrm-encoded, never decoded; by design, likely permanent

ALTER, TEMPORARY tables

HTON_ALTER_NOT_SUPPORTED and HTON_TEMPORARY_NOT_SUPPORTED, as WiredTiger

Multiple SlateDB databases per server

One Db; revisit only with a concrete need

Known limits, stated not hidden

  • Transaction write sets are buffered in memory (SlateDB’s WriteBatch); a multi-gigabyte transaction is a multi-gigabyte allocation. Document; no artificial cap in v1.

  • First-committer-wins conflicts replace lock waits; hot-row contention becomes retry loops. Honest tradeoff of the optimistic model.

  • Two concurrent inserts of the same duplicate key — primary or unique secondary — surface as HA_ERR_LOCK_DEADLOCK on the loser, not HA_ERR_FOUND_DUPP_KEY. Correct, retryable, and different from what a lock-based engine returns; a retry reports the duplicate normally.

  • Commit latency has an object-store floor when await-durable is on. That floor is the product’s identity, not a bug.

Implementation plan

Seven task handoffs in doc/source/specs/implementation-plans/slatedb-engine/; critical chain 1 → 2 → 3 → 4 → 5, then 6; 7 floats after 5. Tasks 1–4 deliver a readable/writable single-index engine under autocommit; task 5 delivers real transactions; task 6 delivers secondary and unique indexes; task 7 is reclamation and operational polish.