SlateDB Storage Engine: Design ============================== Scope and thesis ---------------- A Drizzle storage engine plugin (``plugin/slatedb``) that stores OLTP tables in object storage (S3, GCS, ABS, MinIO) through `SlateDB `__ — a Rust embedded LSM engine whose SSTs, WAL, and manifest all live in the object store. The use case: a Drizzle server whose durable state is a bucket. No local data directory to lose, no EBS volume to snapshot, no replication to configure — durability and capacity are the object store's problem. The relationship to the Iceberg engine is complementary, not competitive: Iceberg is the open-format analytical/cold tier that other engines can read; SlateDB is a *private-format transactional* tier that happens to live in the same kind of bucket. Nothing but this engine reads a SlateDB bucket. The two engines cover the two halves of "Drizzle on object storage." The subtractive frame, as always: this engine is a deliberate subset. One SlateDB database per server. Snapshot isolation only. No savepoints. Refused features are refused loudly with a specific error, naming what and why — never approximated. The WiredTiger PLAN.md Tier-0 audit is treated as a checklist of silent-wrong-results bugs this engine must be structurally unable to have: every one of them (dup-key overwrite, find-flag ignorance, fabricated ``records()``, no-op statement rollback, secondary-index write corruption) has a named answer in this design. Why SlateDB specifically ------------------------ Evaluated against the actual 0.14.1 tree, not the marketing page: - **Real transactions.** ``Db::begin(IsolationLevel) → DbTransaction`` with buffered writes, read-your-own-writes inside the transaction (the write batch backs reads — verified in ``db_transaction.rs``), write-write conflict detection at commit, and both ``Snapshot`` and ``SerializableSnapshot`` isolation. This maps onto Drizzle's ``TransactionalStorageEngine`` contract almost one-to-one. - **Ordered scans both directions.** ``scan``/``scan_prefix`` over byte ranges with ``IterationOrder::Ascending``/``Descending`` and ``seek`` — everything ``index_next``/``index_prev``/ ``records_in_range`` need, provided our keys are memcomparable (below). - **Writer fencing.** Manifest-epoch fencing (``fence.rs``): a second writer opening the same path fences the first, which starts failing with ``Fenced``. Split-brain is structurally impossible; the operational consequence is documented below. - **Read replicas for free (future).** ``DbReader`` opens the same bucket read-only, following the latest state or pinned to a checkpoint. A read-only Drizzle replica is a config change away once the engine exists. Designed-for, not built. - **Explicit conflict-set control.** ``DbTransaction::mark_read`` adds keys to the read set even under ``Snapshot`` — whose ``get`` and ``scan`` otherwise track nothing — and ``unmark_write`` removes keys from the write set. Both verified in ``db_transaction.rs``. The design below leans hard on knowing exactly which set a call lands in, so this is listed as a feature, not trivia. - **Compaction filters.** ``CompactionFilter`` / ``CompactionFilterSupplier`` (``compaction_filter.rs``): a filter sees every entry during compaction and returns ``Keep``, ``Drop``, or ``Modify(ValueDeletable::Tombstone)`` — the lazy-reclamation mechanism for ``DROP TABLE`` (phase 2 of drop, below). Caveat recorded once here and honoured in the build: the API sits behind the **non-default** ``compaction_filters`` cargo feature in 0.14.1 (``slatedb/Cargo.toml``), which the shim crate enables from its first commit. - **Tunable durability.** ``WriteOptions::await_durable`` and ``Settings::flush_interval`` expose the exact latency/cost/durability triangle object storage forces; we surface it as engine options instead of hiding it. Library landscape (as of July 2026) ----------------------------------- - **slatedb 0.14.1** (Apache-2.0, Commonhaus Foundation). Workspace pins Rust 1.91.1 via ``rust-toolchain.toml``. Pre-1.0: API churn expected; **pin the exact version** in the shim's ``Cargo.toml`` and ``Cargo.lock``, both living in one place (the shim crate). - **object_store 0.14** (Rust crate, Apache-2.0) is SlateDB's storage substrate: S3, GCS, Azure, local filesystem, in-memory. ``Db::resolve_object_store(url)`` builds one from a URL plus environment credentials — ``s3://bucket/path`` for production, ``file:///path`` for CI without MinIO, MinIO for the S3-semantics tests. We expose the URL directly as the engine's storage option and invent no abstraction over it. - **No C ABI upstream.** SlateDB's foreign bindings (Go, Java, Python, Node) all ride uniffi, which has no C++ target. We write our own shim (next section) — narrower and better-fitted than anything generated. License note, following the posture established by the Iceberg series (Arrow/iceberg-cpp are likewise Apache-2.0): SlateDB is Apache-2.0, consumed as a separately-built shared library through a C ABI. Task 1's README must include the transitive dependency list with licenses (``cargo license``) and confirm nothing GPLv2-incompatible beyond the already-accepted Apache-2.0 posture rides along. The FFI shim: ``libslatedb_capi`` --------------------------------- The plugin never sees Rust. A small Rust crate, ``slatedb-capi`` (living in ``plugin/slatedb/capi/`` in the drizzle tree), builds a ``cdylib`` — ``libslatedb_capi.so`` — plus one hand-written, versioned C header. The Drizzle plugin is ordinary C++ that links ``-lslatedb_capi``, exactly as the WiredTiger plugin links ``-lwiredtiger``. Design rules for the shim: - **Narrow.** Only what the engine calls: database open/close/flush, transaction begin/commit/rollback, transactional get/put/delete, scan open/seek/next/close (with direction), and error-message retrieval. No settings struct mirroring — configuration crosses the boundary as one string (SlateDB's own ``Settings`` file format, which ``figment`` parses), so new SlateDB knobs cost zero shim changes. - **Sync bridge.** The shim owns a multi-threaded tokio runtime inside the database handle; every call is ``runtime.block_on``. Drizzle sessions are threads that already block on disk I/O; blocking them on object-store I/O is the same shape with bigger constants. No async leaks into C++. - **C ABI, not CXX.** This boundary is bytes-in/bytes-out with opaque handles — the degenerate case where CXX buys nothing. Plain ``extern "C"``, opaque pointers, ``(ptr, len)`` byte slices, and integer status codes paired with owned error objects (next bullet). (The transaction-replication Rust work remains the planned CXX proving-ground; this shim neither depends on nor advances that.) - **Owned errors everywhere; no per-handle ``last_error``.** Every shim function returns a status code and takes a trailing ``slatedb_error_t **out_err``; on any non-``OK`` status it writes an owned error object there, which the caller releases with ``slatedb_error_free`` (``slatedb_error_message`` borrows the string until then). Uniformly — no function is exempt, because a rule with exceptions is a rule nobody remembers at the call site. The rejected alternative is the conventional per-handle ``last_error`` string, and the reason it is rejected is concrete: the handle can already be gone when you want the message. ``DbTransaction::commit`` and ``rollback`` take ``self`` **by value** in 0.14.1 (``db_transaction.rs``), so a failing commit has no transaction left to hang a message on; ``slatedb_scan_close`` has the identical hazard; and hoisting the string to the parent ``Db`` handle would make it shared mutable state raced by every session thread. - **Consuming calls are stated, not inferred.** ``slatedb_txn_commit(txn, out_err)`` and ``slatedb_txn_rollback(txn, out_err)`` **always consume the transaction handle — on success and on failure alike.** After either call the pointer is dangling, must never be passed to another shim function, and **needs no separate free**: there is no ``slatedb_txn_free``, and a second call on the same pointer is a use-after-free, not a leak. The shim moves the ``DbTransaction`` out of its box and drops the box, mirroring the Rust signature exactly rather than hiding it behind an ``Option`` that would let C++ keep using a dead handle and call it safe. The engine consequently clears its session slot *before* it looks at the returned status. - **Error taxonomy at the boundary.** The header defines the small set the engine dispatches on: ``OK``, ``NOT_FOUND``, ``CONFLICT`` (transaction commit conflict), ``FENCED``, ``INVALID``, ``IO``. Everything else is ``IO`` plus the message string. The C++ side maps these to ``HA_ERR_*`` in one function (``slatedb_to_drizzle_error.cc``), mirroring ``wt_to_drizzle_error.cc``. - **Ownership across the boundary is the registry pattern.** Handles are created and destroyed by paired shim functions; no Rust allocation is ever freed by C++ or vice versa — except for the consuming pair above, which is the whole reason that pair is written down. Scan results are valid until the next ``next``/``close`` on that scan — the C++ side copies into the Drizzle record buffer immediately, so no lifetime-extension API is needed. Build and dependency policy: container-first, and the composition note ---------------------------------------------------------------------- The binding policy from the Iceberg series applies verbatim: the only build command is ``podman build`` fronted by ``just`` targets; every pin lives in a Containerfile (or, here, ``Cargo.lock`` + ``rust-toolchain.toml`` inside the image build); **no dependency detection anywhere**; plugin inclusion is an explicit build switch (``--with-slatedb``, default off) that the containerized build always passes. Cargo is upstream's private build system, invoked by a ``RUN`` line, exactly as CMake is for WiredTiger — it never leaks into autotools, bindep, or Zuul configuration. This deliberately does *not* advance the "cargo integrated into the drizzle build" prerequisite from the broader Rust plan; it sidesteps it, and that is a feature: the shim is a leaf artifact with its own toolchain, not the first tenant of a shared Rust build. **The recorded open puzzle from the Iceberg spec is now due.** The WiredTiger image landed; ``slatedb-capi`` is the second source-built engine dependency, which was the named trigger for deciding how N per-engine dependencies compose into one builder-image lineage. The decision, made here so it is a decision and not an accident: **Per-dependency build-stage images, composed by ``COPY --from=``.** Each engine dependency gets its own image built ``FROM quay.io/drizzle/libdrizzle`` (WiredTiger already does this; ``slatedb-capi`` follows), installing into ``/usr/local`` through a ``DESTDIR`` so the payload is an exact file set. The drizzle builder image composes the enabled set with one ``COPY --from=quay.io/drizzle/ /install/usr/local/ /usr/local/`` line per dependency, then ``ldconfig``. Discipline required and checked in CI: payloads must be file-disjoint (a trivial ``find | sort | uniq -d`` check across payload manifests in the builder build) and sonames unique. Rejected alternatives: one fat ``deps`` image (accretes, single cache-bust point, couples release cadences) and per-plugin builder images (multiplies the expensive drizzle build). This paragraph is the "short design note" the Iceberg spec called for; if composition grows real complexity later it graduates to its own spec. Data model: one database, prefixed keyspaces -------------------------------------------- **One SlateDB ``Db`` per server**, opened at plugin init from ``slatedb.url`` + ``slatedb.path``, closed at shutdown. Not one per table: each ``Db`` carries a tokio runtime, WAL flush loop, manifest poller, and compactor — per-table instances would multiply all of that and turn ``CREATE TABLE`` into an object-store fencing ceremony. Tables are key prefixes. Key layout (all keys in one keyspace, first byte discriminates): :: 0x00 'p' → serialized message::Table proto 0x00 'i' → table-id (u32 BE) 0x00 'c' → next-table-id counter 0x00 'd' → dropped-table marker (phase-2 drop) 0x01 → row / index entry ``index-id`` 0 is the primary key; 1..N are secondary indexes in proto order. Every table or index scan is a SlateDB prefix scan on the 6-byte ``0x01 || table-id || index-id`` prefix; ``records_in_range`` is a bounded range scan. Table protos live in the meta prefix exactly as the WiredTiger engine stores them in its ``table:table_definitions`` WT table — ``doGetTableDefinition`` reads the proto key (``EEXIST``/``ENOENT`` per convention), ``doGetTableIdentifiers`` scans the ``0x00 'p'`` prefix, and no table-definition files are ever written. DDL is transactional by construction: ``CREATE TABLE`` writes the proto, the id mapping, and the bumped counter in one transaction commit; ``RENAME`` rewrites the two meta keys (data keys carry only the id, so rename never touches rows); ``DROP`` deletes the meta keys and writes the ``0x00 'd'`` marker in one commit (reclamation below). Key and value encoding ---------------------- SlateDB compares keys as raw bytes, so **the key codec must be memcomparable** — encoded order equals SQL order. This is the one substantial piece of new C++ in the engine and it is written as a standalone, exhaustively unit-tested codec library (``plugin/slatedb/codec/``) before any cursor code exists. Prior art is MyRocks' key encoding; we take the ideas, not the code. Per-type key encoding: - **Nullability prefix**: nullable key parts get one byte — ``0x00`` NULL, ``0x01`` present. NULLs sort first, matching the server's ordering. - **Signed integers** (LONG, LONGLONG, DATE, TIME, ENUM as its underlying int): big-endian with the sign bit flipped. - **Unsigned / TIMESTAMP / MICROTIME**: big-endian. - **DOUBLE**: IEEE-754 total-order trick — positive: flip sign bit; negative: flip all bits. - **BOOLEAN, UUID, IPV6**: fixed-width big-endian byte images. - **DECIMAL**: Drizzle's ``my_decimal`` binary format is already memcomparable for fixed (precision, scale) — used as-is, which is also what the server's own key format relies on. - **VARCHAR/CHAR (collated text)**: the column collation's ``strnxfrm`` weight string, then byte-escaped (each ``0x00`` → ``0x00 0xFF``) and terminated with ``0x00 0x00`` so shorter strings sort before their extensions and multi-part keys cannot alias. - **BLOB in keys**: refused in v1 (``HTON_NO_BLOBS`` — matching the WiredTiger engine's current posture; blob columns and blob keys arrive together in a later task). **Keys are never decoded.** ``strnxfrm`` is one-way, so the design commits to it globally rather than special-casing: the row value stores *all* columns including primary-key columns, and every read materializes from the value. A secondary index entry carries the encoded PK, which is consumed *verbatim as bytes* (it is already the primary key's encoded form) to point-read the primary entry — but *where* it carries it depends on whether the index is unique, and the reason is concurrency control, not space: - **Non-unique index**: ``key = index-cols-encoded || pk-encoded``, ``value = empty``. The PK suffix is what keeps equal-valued rows distinct; a lookup slices it off at the encoded-PK boundary. - **Unique index, no NULL in the index columns**: ``key = index-cols-encoded`` — **no PK suffix** — ``value = pk-encoded``. Two sessions inserting the same unique value with different primary keys then write the *same physical key*, which is the only thing SlateDB's conflict detector looks at (below); exactly one of them commits. A lookup reads the PK out of the value. - **Unique index, any NULL in the index columns**: the non-unique shape, PK suffix and empty value. SQL says NULL never equals NULL, so multiple NULL-bearing rows must coexist in a UNIQUE index — giving them distinct physical keys *is* that rule, expressed in the layout. The alternative (one shape plus a "skip the conflict when NULL" rule at commit time) is not implementable: conflict detection is a set intersection over key bytes inside SlateDB, with no hook to consult and no schema to consult it with. Push the semantics into the bytes and there is nothing left to get wrong. Key *structure* stays parseable even though key *values* are not decodable — nullability prefix bytes, fixed widths, and the ``0x00 0x00`` string terminator make part boundaries recoverable. That is what lets a reader find the PK-suffix boundary (already true before this refinement) and what lets the encoder decide, locally and cheaply, whether a unique key has a NULL part and therefore which shape to write. No decode anywhere, no ``HA_KEYREAD_ONLY``, and index-only-scan optimization is explicitly refused rather than half-supported. The consequence worth writing down twice: **key bytes embed collation weight tables**, so a collation-table change is an on-disk compatibility event. The charset/collation codegen work must treat shipped weight tables as frozen-per-version; the engine stamps a format version into the meta prefix at database creation and refuses to open a future incompatible format rather than misreading it. Value encoding: version byte, null bitmap, then each non-null field in proto column order — fixed-width types as fixed-width little-endian images, var-width as varint-length-prefixed bytes. Not memcomparable, not clever; versioned so it can evolve. ``position()``/``rnd_pos`` use the encoded PK bytes length-prefixed in ``ref``, exactly the WiredTiger cursor's pattern. Transactions and isolation -------------------------- The engine is a ``TransactionalStorageEngine``. Per-session state lives in the ``Session::getEngineData`` slot (the innobase/WiredTiger pattern): the shim transaction handle plus a small statement-state struct. - **Begin on first use**, via ``doStartStatement`` — adopting the WiredTiger engine's fix verbatim (``startTransaction`` stays a no-op so non-SlateDB transactions don't inflate ``Handler_commit``). First touch calls ``slatedb_txn_begin`` and registers the engine as a transaction resource. - **Isolation: snapshot only.** ``begin(IsolationLevel::Snapshot)`` always. The server's READ COMMITTED / READ UNCOMMITTED requests are *refused with an error*, not silently upgraded — the same collapse the WiredTiger 11 port made, for the same reason, with the same user-visible honesty. SlateDB's ``SerializableSnapshot`` is real and tested upstream; exposing it as a per-session option is a cheap later task, listed, not done. - **Commit**: ``slatedb_txn_commit``. A ``CONFLICT`` return (write-write conflict with a concurrently committed transaction) maps to ``HA_ERR_LOCK_DEADLOCK`` — the one error code every SQL application already retries on. This is first-committer-wins optimistic concurrency and the docs say so plainly: hot-row workloads will see deadlock-shaped retries instead of lock waits. - **Rollback**: whole-transaction rollback is free (``DbTransaction::rollback`` drops the buffered batch). **Statement rollback and savepoints are refused**: SlateDB's transaction batch cannot be truncated to a mark, so ``doSetSavepoint`` returns not-supported and a failed statement inside a multi-statement transaction rolls back the whole transaction and reports exactly that — the NDB-style answer the WiredTiger plan chose over the silent-no-op it audited (Tier 0.7). Never claim success for a rollback that didn't happen. - **Read-your-own-writes works** — the transaction's write batch backs its own reads and scans, so multi-statement INSERT-then-SELECT behaves as SQL requires. (This is the capability that separates this engine from the Iceberg engine's deliberately-blind commit buffer.) - **DDL vs. fencing**: unlike WiredTiger there is no ``EBUSY``-style DDL/checkpoint interaction — DDL is ordinary transactional writes to meta keys. The pre-DDL-checkpoint fixture pattern does not carry over because the problem it fixes does not exist here. Duplicate keys are detected, not overwritten (WiredTiger Tier 0.2/0.3 learned the hard way): ``doInsertRecord`` probes the primary key with a transactional ``get`` before ``put`` (a memtable/cache hit in the common case), and probes each non-NULL-bearing UNIQUE secondary with a point ``get`` on its unique-columns-only key, returning ``HA_ERR_FOUND_DUPP_KEY`` with the correct key number; the concurrent case falls through to the commit-time conflict described below. INSERT … ON DUPLICATE KEY UPDATE therefore works through the standard kernel path with no engine-side replace shortcut. A key-changing UPDATE is delete-old + insert-new with dup probe, decided by comparing old/new encoded keys (Tier 0.4). Writes during a secondary-index-driven scan go through primary-key point writes, never through the scan's iterator (Tier 0.5) — which is natural here since scans and writes are separate shim objects by construction. What the conflict detector actually sees ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Uniqueness enforcement rests on this, so it is verified and written down rather than assumed. All of it read out of 0.14.1's ``db_transaction.rs`` and ``transaction_manager.rs``: - Under ``IsolationLevel::Snapshot`` **only the write set is tracked**. ``get``/``scan`` add nothing: ``get_key_value_with_options`` guards its ``track_read_keys`` call on ``isolation_level == SerializableSnapshot``, and ``check_has_conflict`` says in as many words that Snapshot "only checks write-write conflicts". - The commit-time check is ``has_write_write_conflict``: a set intersection of this transaction's write keys against the write keys of every transaction that committed after this one started. So two transactions that ``put`` the same key **do** conflict, and the loser gets ``SlateDBError::TransactionConflict``. Exact bytes, exact keys — no ranges, no predicates, no schema. - ``mark_read(keys)`` opts named keys into read-write detection even under Snapshot, and ``unmark_write(keys)`` opts named keys out of write-write detection. Both are real; neither takes a range, which is why neither rescues a predicate probe. The consequence, and it is the load-bearing sentence of this section: **a probe is not a lock.** Reading a key and finding it absent buys nothing under Snapshot; only *writing* a key does. Every uniqueness guarantee in this engine must therefore reduce to "the competing transactions write the same physical key" — which is precisely why the unique-index layout above drops the PK suffix. With the suffix, two inserters of the same unique value write two different keys, both probe-miss, and both commit: a silently violated constraint, the WiredTiger Tier 0.3 failure mode wearing a different hat. The probe stays, as an optimization on top of the conflict rather than as the mechanism. The two outcomes, both correct, stated so that neither surprises anyone: - **Uncontended** — the duplicate was already committed when this statement started: the probe hits and the engine returns ``HA_ERR_FOUND_DUPP_KEY`` with the right key number. This is essentially all real duplicate-key traffic. - **Contended** — two concurrent inserts of the same primary key, or of the same unique value: both probe-miss, both write the same physical key, and the second committer gets ``CONFLICT`` → ``HA_ERR_LOCK_DEADLOCK`` rather than ``HA_ERR_FOUND_DUPP_KEY``. **This is acceptable and it is not worked around.** The loser's transaction is rolled back either way; deadlock is the one error every SQL client already retries; and on retry the winner's row is committed, so the probe hits and the duplicate is reported properly. The only way to return the "right" error here would be to serialize inserts, which is the thing optimistic concurrency exists to avoid. The user documentation states it next to the hot-row retry note. Hot shared keys, audited ~~~~~~~~~~~~~~~~~~~~~~~~ First-committer-wins turns any key that many concurrent transactions write into a serialization point delivered as retries. The design therefore owns a list of every shared key it creates, and says why each is acceptable — "fine" as a judgement, not as an oversight: - ``0x00 'c'`` **next-table-id counter** — fine. Only ``CREATE TABLE`` read-modify-writes it, DDL is rare and human-paced, and a conflict costs one retry of one cheap transaction. - ``0x00 'd' `` **dropped-table markers** — fine. Written once by the DROP that creates one, deleted once by the sweep that retires it; no two transactions contend for the same marker. - **Per-table meta keys** (``0x00 'p'``, ``0x00 'i'``) — fine. DDL-only and per-table, and concurrent DDL on one table is already serialized above the engine. - **Unique-index keys** — contended *by design*. That contention is the enforcement, and it only occurs for exactly the duplicate the application asked us to reject. - **A maintained per-table row-count key** — rejected; see "Statistics honesty" below. It would have made every writer to a table conflict with every other writer to that table, on one key. - ``AUTO_INCREMENT`` **counters** — refused in v1, and this is why the recorded plan is block reservation rather than a shared counter key touched once per insert. Durability, latency, and cost ----------------------------- Object storage makes the write-latency triangle explicit; the engine exposes SlateDB's own knobs and adds none: - ``slatedb.url`` (required) and ``slatedb.path`` — the object store URL and database path. - ``slatedb.settings-file`` — optional path handed through the shim verbatim to SlateDB's ``figment``-based ``Settings`` loader (``flush_interval``, caches, compactor tuning, all of it). We do not mirror individual SlateDB settings as engine options; one passthrough, zero drift. - ``slatedb.await-durable`` (default **on**): commits block until the WAL batch is durable in object storage — correctness-first, with commit latency ≈ ``flush_interval`` + one PUT round trip, and the documentation states those numbers rather than apologizing for them. Off: commit returns at memtable acceptance; a crash can lose the tail up to the last flush, stated equally plainly. This is per-server policy, not per-table. - Read latency is governed by SlateDB's in-memory block cache and optional local-disk object cache (``cached_object_store``), configured through the settings file; the perf-tracking suite measures hot/cold read split explicitly. ``flush()`` on the engine maps to ``Db::flush``; a ``slatedb_flush()`` UDF mirrors ``wiredtiger_checkpoint()`` for operational use. Statistics honesty (Tier 0.6): **no ``HTON_HAS_RECORDS``**, no ``HTON_STATS_RECORDS_IS_EXACT``. ``info()`` reports estimates and says so. Nothing in this engine claims an exact row count, and nothing in it maintains one. - Row count, v1: a **capped prefix probe**. Scan the table's ``0x01 || table-id || 0x00`` prefix up to a fixed row cap (a constant, tuned against the perf suite); exhausted under the cap → the exact count; hit the cap → a scaled estimate from key-space sampling. Cheap, honest, and wrong only in the direction optimizers already assume estimates are wrong. - Cardinality and ``records_in_range``: bounded range probes under the same cap, and honest unknown-cardinality defaults where a probe is not worth its I/O. Never a fabricated floor — that is the Tier 0.6 anti-pattern itself. **Rejected, explicitly: a per-table row-count key maintained transactionally with writes.** It is the obvious way to get an exact count and it is a design bug in this engine. Under first-committer-wins every concurrent writer to a table would write that one key and therefore write-write conflict with every other writer to the same table: a single hot key serializing all writes to the table it is counting, with the serialization delivered as ``HA_ERR_LOCK_DEADLOCK`` storms rather than as waits. Exactness is not worth converting a concurrent engine into a serial one, and this engine does not need it, because it does not claim it. For the record, since it will be asked: a conflict-free counter *is* constructible on 0.14.1. There is a merge operator (``merge_operator.rs``, ``DbTransaction::merge``), and ``unmark_write`` lifts a key out of write-write detection, so a per-statement delta merged into an unmarked counter key would not conflict. It is listed as a **possible later task, not v1 work**, and the caveat travels with the listing: merge operands still land in the write batch (``batch.rs``: ``keys()`` returns every op's key, merges included), so the counter is conflict-free only if the unmark is applied without exception — and an unmarked key is by definition no longer transactional with the rows it counts. That tradeoff earns its own evaluation with measurements, not a paragraph here. DROP TABLE reclamation ---------------------- Two phases, both landing eventually, first one first: 1. **v1 — foreground delete**: DROP scans the table's data prefix and issues batched deletes, then removes meta keys. Correct, simple, and slow for huge tables over S3; the docs say so. 2. **Compaction-filter reclamation**: DROP writes only the meta deletion + ``0x00 'd' `` marker; a compaction filter tombstones dead-prefix entries as compaction naturally rewrites them; a background sweep retires the marker once the prefix is empty. This is the steady-state design; v1's foreground path remains as ``DROP TABLE``'s behavior under a size threshold or explicit option if the split proves worth keeping — decide with data from the perf suite, not in advance. The API as it actually exists in 0.14.1 (``compaction_filter.rs``, read, not assumed): a ``CompactionFilterSupplier`` is registered once at open through ``DbBuilder::with_compaction_filter_supplier``, and its ``create_compaction_filter(&CompactionJobContext)`` is called **per compaction job**; the returned ``CompactionFilter`` sees every entry through ``filter(&RowEntry)`` and answers ``Keep``, ``Drop``, or ``Modify(ValueDeletable::Tombstone)``, with ``on_compaction_end`` as the hook for per-job counters. We answer ``Modify(Tombstone)``, never ``Drop``: upstream documents that ``Drop`` removes an entry without shadowing it, which can resurrect older versions from lower runs. The feature is off by default (``compaction_filters`` in ``slatedb/Cargo.toml``) and the shim crate turns it on. Upstream's snapshot-consistency warning is acceptable here by construction — the ids we filter belong to tables that appear in no committed metadata and have no live readers, because the server serialized the DROP against open cursors already. **The dropped-id set is live, not a snapshot taken at open.** A set read once at open would reclaim nothing dropped since the server started — which, in a long-running server, is every table anyone actually dropped. The shim owns it: - The filter supplier holds an ``Arc>>``, seeded at open from a scan of the ``0x00 'd'`` prefix. - ``create_compaction_filter`` takes the read lock once and copies the set into the job-local filter, so the per-entry ``filter()`` call touches no lock at all. Per-job filter instances exist precisely so that per-job snapshots are the intended shape; a job that began before an id was added simply misses that table this round and the next compaction takes it. - Two shim functions, ``slatedb_dropped_ids_add(db, id, out_err)`` and ``slatedb_dropped_ids_remove(db, id, out_err)``, mutate the set. They are the only additions to the shim surface after task 1. - **The ordering is the entire correctness argument.** ``doDropTable`` calls ``add`` **only after the DROP transaction's commit returns OK** — not before it, not on ``CONFLICT``, not on rollback, not from a destructor, not on any path that is not a returned-OK commit. An id in the filter set whose DROP did not commit tombstones a live table's rows. - Symmetrically the retirement sweep calls ``remove`` **only after the marker-deletion transaction commits**. Removing first and then failing to commit merely leaves a marker nothing acts on until the next restart — but it is the same rule, so it is the same rule. - Crash safety needs no extra machinery: the ``0x00 'd'`` prefix in the database is the authority, the in-memory set is a cache rebuilt from it at every open, and both operations are idempotent. A crash between commit and set update costs a delayed reclamation, never a row. Operational model ----------------- - **Exactly one writing server per (url, path).** SlateDB's epoch fencing makes a second writer *win*: the new server fences the old, whose SlateDB calls start failing ``FENCED``. The engine treats ``FENCED`` as fatal — log loudly and refuse further writes (``HA_ERR_*`` mapped to a read-only/IO error), because the honest description of a fenced writer is "another server owns this bucket now." This inverts the usual failover ergonomics (starting the replacement kills the incumbent) and the deployment docs must lead with it. It is also exactly what makes lease-less failover work. - **Backup = checkpoint.** SlateDB checkpoints (``create_checkpoint``, ``Admin::create_detached_checkpoint``) are consistent named references in the same bucket; the ``slatedb`` CLI manages them. A ``slatedb_checkpoint()`` UDF is a natural follow-on once wanted. - **Read replicas (future, designed-for)**: ``DbReader`` against the same path, following latest or pinned to a checkpoint. The engine split (writer engine vs. reader engine mode) is a later spec; the key/value codec and the one-Db model are the parts that must not foreclose it, and they don't. Refused in v1 (all loud, all specific) -------------------------------------- +---------------------------------------+-----------------------------------------------+ | Refusal | Why / when it could return | +=======================================+===============================================+ | READ COMMITTED / READ UNCOMMITTED | Snapshot-only, per the WiredTiger precedent; | | | never returns | +---------------------------------------+-----------------------------------------------+ | Savepoints / statement rollback | SlateDB txn batch is not truncatable; | | | whole-txn rollback reported honestly | +---------------------------------------+-----------------------------------------------+ | SERIALIZABLE (SSI) | Cheap later task — SlateDB supports it | +---------------------------------------+-----------------------------------------------+ | BLOB columns and blob keys | Later task (value codec is ready; cursor blob | | | buffer management is the work) | +---------------------------------------+-----------------------------------------------+ | AUTO_INCREMENT | Later task: transactional counter key with | | | block reservation; refuse rather than ship a | | | racy version | +---------------------------------------+-----------------------------------------------+ | Index-only scans (HA_KEYREAD_ONLY) | Keys are strnxfrm-encoded, never decoded; by | | | design, likely permanent | +---------------------------------------+-----------------------------------------------+ | ALTER, TEMPORARY tables | ``HTON_ALTER_NOT_SUPPORTED`` and | | | ``HTON_TEMPORARY_NOT_SUPPORTED``, as | | | WiredTiger | +---------------------------------------+-----------------------------------------------+ | Multiple SlateDB databases per server | One ``Db``; revisit only with a concrete need | +---------------------------------------+-----------------------------------------------+ Known limits, stated not hidden ------------------------------- - Transaction write sets are buffered in memory (SlateDB's ``WriteBatch``); a multi-gigabyte transaction is a multi-gigabyte allocation. Document; no artificial cap in v1. - First-committer-wins conflicts replace lock waits; hot-row contention becomes retry loops. Honest tradeoff of the optimistic model. - Two *concurrent* inserts of the same duplicate key — primary or unique secondary — surface as ``HA_ERR_LOCK_DEADLOCK`` on the loser, not ``HA_ERR_FOUND_DUPP_KEY``. Correct, retryable, and different from what a lock-based engine returns; a retry reports the duplicate normally. - Commit latency has an object-store floor when ``await-durable`` is on. That floor is the product's identity, not a bug. Implementation plan ------------------- Seven task handoffs in ``doc/source/specs/implementation-plans/slatedb-engine/``; critical chain 1 → 2 → 3 → 4 → 5, then 6; 7 floats after 5. Tasks 1–4 deliver a readable/writable single-index engine under autocommit; task 5 delivers real transactions; task 6 delivers secondary and unique indexes; task 7 is reclamation and operational polish.