SlateDB Storage Engine: Design¶
Scope and thesis¶
A Drizzle storage engine plugin (plugin/slatedb) that stores OLTP
tables in object storage (S3, GCS, ABS, MinIO) through SlateDB — a Rust embedded LSM engine
whose SSTs, WAL, and manifest all live in the object store. The use
case: a Drizzle server whose durable state is a bucket. No local data
directory to lose, no EBS volume to snapshot, no replication to
configure — durability and capacity are the object store’s problem.
The relationship to the Iceberg engine is complementary, not competitive: Iceberg is the open-format analytical/cold tier that other engines can read; SlateDB is a private-format transactional tier that happens to live in the same kind of bucket. Nothing but this engine reads a SlateDB bucket. The two engines cover the two halves of “Drizzle on object storage.”
The subtractive frame, as always: this engine is a deliberate subset.
One SlateDB database per server. Snapshot isolation only. No
savepoints. Refused features are refused loudly with a specific error,
naming what and why — never approximated. The WiredTiger PLAN.md
Tier-0 audit is treated as a checklist of silent-wrong-results bugs
this engine must be structurally unable to have: every one of them
(dup-key overwrite, find-flag ignorance, fabricated records(),
no-op statement rollback, secondary-index write corruption) has a named
answer in this design.
Why SlateDB specifically¶
Evaluated against the actual 0.14.1 tree, not the marketing page:
Real transactions.
Db::begin(IsolationLevel) → DbTransactionwith buffered writes, read-your-own-writes inside the transaction (the write batch backs reads — verified indb_transaction.rs), write-write conflict detection at commit, and bothSnapshotandSerializableSnapshotisolation. This maps onto Drizzle’sTransactionalStorageEnginecontract almost one-to-one.Ordered scans both directions.
scan/scan_prefixover byte ranges withIterationOrder::Ascending/Descendingandseek— everythingindex_next/index_prev/records_in_rangeneed, provided our keys are memcomparable (below).Writer fencing. Manifest-epoch fencing (
fence.rs): a second writer opening the same path fences the first, which starts failing withFenced. Split-brain is structurally impossible; the operational consequence is documented below.Read replicas for free (future).
DbReaderopens the same bucket read-only, following the latest state or pinned to a checkpoint. A read-only Drizzle replica is a config change away once the engine exists. Designed-for, not built.Explicit conflict-set control.
DbTransaction::mark_readadds keys to the read set even underSnapshot— whosegetandscanotherwise track nothing — andunmark_writeremoves keys from the write set. Both verified indb_transaction.rs. The design below leans hard on knowing exactly which set a call lands in, so this is listed as a feature, not trivia.Compaction filters.
CompactionFilter/CompactionFilterSupplier(compaction_filter.rs): a filter sees every entry during compaction and returnsKeep,Drop, orModify(ValueDeletable::Tombstone)— the lazy-reclamation mechanism forDROP TABLE(phase 2 of drop, below). Caveat recorded once here and honoured in the build: the API sits behind the non-defaultcompaction_filterscargo feature in 0.14.1 (slatedb/Cargo.toml), which the shim crate enables from its first commit.Tunable durability.
WriteOptions::await_durableandSettings::flush_intervalexpose the exact latency/cost/durability triangle object storage forces; we surface it as engine options instead of hiding it.
Library landscape (as of July 2026)¶
slatedb 0.14.1 (Apache-2.0, Commonhaus Foundation). Workspace pins Rust 1.91.1 via
rust-toolchain.toml. Pre-1.0: API churn expected; pin the exact version in the shim’sCargo.tomlandCargo.lock, both living in one place (the shim crate).object_store 0.14 (Rust crate, Apache-2.0) is SlateDB’s storage substrate: S3, GCS, Azure, local filesystem, in-memory.
Db::resolve_object_store(url)builds one from a URL plus environment credentials —s3://bucket/pathfor production,file:///pathfor CI without MinIO, MinIO for the S3-semantics tests. We expose the URL directly as the engine’s storage option and invent no abstraction over it.No C ABI upstream. SlateDB’s foreign bindings (Go, Java, Python, Node) all ride uniffi, which has no C++ target. We write our own shim (next section) — narrower and better-fitted than anything generated.
License note, following the posture established by the Iceberg series
(Arrow/iceberg-cpp are likewise Apache-2.0): SlateDB is Apache-2.0,
consumed as a separately-built shared library through a C ABI. Task 1’s
README must include the transitive dependency list with licenses
(cargo license) and confirm nothing GPLv2-incompatible beyond the
already-accepted Apache-2.0 posture rides along.
The FFI shim: libslatedb_capi¶
The plugin never sees Rust. A small Rust crate, slatedb-capi
(living in plugin/slatedb/capi/ in the drizzle tree), builds a
cdylib — libslatedb_capi.so — plus one hand-written, versioned
C header. The Drizzle plugin is ordinary C++ that links
-lslatedb_capi, exactly as the WiredTiger plugin links
-lwiredtiger.
Design rules for the shim:
Narrow. Only what the engine calls: database open/close/flush, transaction begin/commit/rollback, transactional get/put/delete, scan open/seek/next/close (with direction), and error-message retrieval. No settings struct mirroring — configuration crosses the boundary as one string (SlateDB’s own
Settingsfile format, whichfigmentparses), so new SlateDB knobs cost zero shim changes.Sync bridge. The shim owns a multi-threaded tokio runtime inside the database handle; every call is
runtime.block_on. Drizzle sessions are threads that already block on disk I/O; blocking them on object-store I/O is the same shape with bigger constants. No async leaks into C++.C ABI, not CXX. This boundary is bytes-in/bytes-out with opaque handles — the degenerate case where CXX buys nothing. Plain
extern "C", opaque pointers,(ptr, len)byte slices, and integer status codes paired with owned error objects (next bullet). (The transaction-replication Rust work remains the planned CXX proving-ground; this shim neither depends on nor advances that.)Owned errors everywhere; no per-handle ``last_error``. Every shim function returns a status code and takes a trailing
slatedb_error_t **out_err; on any non-OKstatus it writes an owned error object there, which the caller releases withslatedb_error_free(slatedb_error_messageborrows the string until then). Uniformly — no function is exempt, because a rule with exceptions is a rule nobody remembers at the call site. The rejected alternative is the conventional per-handlelast_errorstring, and the reason it is rejected is concrete: the handle can already be gone when you want the message.DbTransaction::commitandrollbacktakeselfby value in 0.14.1 (db_transaction.rs), so a failing commit has no transaction left to hang a message on;slatedb_scan_closehas the identical hazard; and hoisting the string to the parentDbhandle would make it shared mutable state raced by every session thread.Consuming calls are stated, not inferred.
slatedb_txn_commit(txn, out_err)andslatedb_txn_rollback(txn, out_err)always consume the transaction handle — on success and on failure alike. After either call the pointer is dangling, must never be passed to another shim function, and needs no separate free: there is noslatedb_txn_free, and a second call on the same pointer is a use-after-free, not a leak. The shim moves theDbTransactionout of its box and drops the box, mirroring the Rust signature exactly rather than hiding it behind anOption<DbTransaction>that would let C++ keep using a dead handle and call it safe. The engine consequently clears its session slot before it looks at the returned status.Error taxonomy at the boundary. The header defines the small set the engine dispatches on:
OK,NOT_FOUND,CONFLICT(transaction commit conflict),FENCED,INVALID,IO. Everything else isIOplus the message string. The C++ side maps these toHA_ERR_*in one function (slatedb_to_drizzle_error.cc), mirroringwt_to_drizzle_error.cc.Ownership across the boundary is the registry pattern. Handles are created and destroyed by paired shim functions; no Rust allocation is ever freed by C++ or vice versa — except for the consuming pair above, which is the whole reason that pair is written down. Scan results are valid until the next
next/closeon that scan — the C++ side copies into the Drizzle record buffer immediately, so no lifetime-extension API is needed.
Build and dependency policy: container-first, and the composition note¶
The binding policy from the Iceberg series applies verbatim: the only
build command is podman build fronted by just targets; every
pin lives in a Containerfile (or, here, Cargo.lock +
rust-toolchain.toml inside the image build); no dependency
detection anywhere; plugin inclusion is an explicit build switch
(--with-slatedb, default off) that the containerized build always
passes. Cargo is upstream’s private build system, invoked by a RUN
line, exactly as CMake is for WiredTiger — it never leaks into
autotools, bindep, or Zuul configuration. This deliberately does not
advance the “cargo integrated into the drizzle build” prerequisite from
the broader Rust plan; it sidesteps it, and that is a feature: the shim
is a leaf artifact with its own toolchain, not the first tenant of a
shared Rust build.
The recorded open puzzle from the Iceberg spec is now due. The
WiredTiger image landed; slatedb-capi is the second source-built
engine dependency, which was the named trigger for deciding how N
per-engine dependencies compose into one builder-image lineage. The
decision, made here so it is a decision and not an accident:
Per-dependency build-stage images, composed by ``COPY –from=``. Each engine dependency gets its own image built
FROM quay.io/drizzle/libdrizzle(WiredTiger already does this;slatedb-capifollows), installing into/usr/localthrough aDESTDIRso the payload is an exact file set. The drizzle builder image composes the enabled set with oneCOPY --from=quay.io/drizzle/<dep> /install/usr/local/ /usr/local/line per dependency, thenldconfig. Discipline required and checked in CI: payloads must be file-disjoint (a trivialfind | sort | uniq -dcheck across payload manifests in the builder build) and sonames unique. Rejected alternatives: one fatdepsimage (accretes, single cache-bust point, couples release cadences) and per-plugin builder images (multiplies the expensive drizzle build). This paragraph is the “short design note” the Iceberg spec called for; if composition grows real complexity later it graduates to its own spec.
Data model: one database, prefixed keyspaces¶
One SlateDB ``Db`` per server, opened at plugin init from
slatedb.url + slatedb.path, closed at shutdown. Not one per
table: each Db carries a tokio runtime, WAL flush loop, manifest
poller, and compactor — per-table instances would multiply all of that
and turn CREATE TABLE into an object-store fencing ceremony. Tables
are key prefixes.
Key layout (all keys in one keyspace, first byte discriminates):
0x00 'p' <table-path> → serialized message::Table proto
0x00 'i' <table-path> → table-id (u32 BE)
0x00 'c' → next-table-id counter
0x00 'd' <table-id BE> → dropped-table marker (phase-2 drop)
0x01 <table-id u32 BE> <index-id u8> <encoded-key> → row / index entry
index-id 0 is the primary key; 1..N are secondary indexes in proto
order. Every table or index scan is a SlateDB prefix scan on the 6-byte
0x01 || table-id || index-id prefix; records_in_range is a
bounded range scan. Table protos live in the meta prefix exactly as the
WiredTiger engine stores them in its table:table_definitions WT
table — doGetTableDefinition reads the proto key
(EEXIST/ENOENT per convention), doGetTableIdentifiers
scans the 0x00 'p' prefix, and no table-definition files are ever
written.
DDL is transactional by construction: CREATE TABLE writes the
proto, the id mapping, and the bumped counter in one transaction
commit; RENAME rewrites the two meta keys (data keys carry only the
id, so rename never touches rows); DROP deletes the meta keys and
writes the 0x00 'd' marker in one commit (reclamation below).
Key and value encoding¶
SlateDB compares keys as raw bytes, so the key codec must be
memcomparable — encoded order equals SQL order. This is the one
substantial piece of new C++ in the engine and it is written as a
standalone, exhaustively unit-tested codec library
(plugin/slatedb/codec/) before any cursor code exists. Prior art is
MyRocks’ key encoding; we take the ideas, not the code.
Per-type key encoding:
Nullability prefix: nullable key parts get one byte —
0x00NULL,0x01present. NULLs sort first, matching the server’s ordering.Signed integers (LONG, LONGLONG, DATE, TIME, ENUM as its underlying int): big-endian with the sign bit flipped.
Unsigned / TIMESTAMP / MICROTIME: big-endian.
DOUBLE: IEEE-754 total-order trick — positive: flip sign bit; negative: flip all bits.
BOOLEAN, UUID, IPV6: fixed-width big-endian byte images.
DECIMAL: Drizzle’s
my_decimalbinary format is already memcomparable for fixed (precision, scale) — used as-is, which is also what the server’s own key format relies on.VARCHAR/CHAR (collated text): the column collation’s
strnxfrmweight string, then byte-escaped (each0x00→0x00 0xFF) and terminated with0x00 0x00so shorter strings sort before their extensions and multi-part keys cannot alias.BLOB in keys: refused in v1 (
HTON_NO_BLOBS— matching the WiredTiger engine’s current posture; blob columns and blob keys arrive together in a later task).
Keys are never decoded. strnxfrm is one-way, so the design
commits to it globally rather than special-casing: the row value stores
all columns including primary-key columns, and every read
materializes from the value. A secondary index entry carries the
encoded PK, which is consumed verbatim as bytes (it is already the
primary key’s encoded form) to point-read the primary entry — but
where it carries it depends on whether the index is unique, and the
reason is concurrency control, not space:
Non-unique index:
key = index-cols-encoded || pk-encoded,value = empty. The PK suffix is what keeps equal-valued rows distinct; a lookup slices it off at the encoded-PK boundary.Unique index, no NULL in the index columns:
key = index-cols-encoded— no PK suffix —value = pk-encoded. Two sessions inserting the same unique value with different primary keys then write the same physical key, which is the only thing SlateDB’s conflict detector looks at (below); exactly one of them commits. A lookup reads the PK out of the value.Unique index, any NULL in the index columns: the non-unique shape, PK suffix and empty value. SQL says NULL never equals NULL, so multiple NULL-bearing rows must coexist in a UNIQUE index — giving them distinct physical keys is that rule, expressed in the layout. The alternative (one shape plus a “skip the conflict when NULL” rule at commit time) is not implementable: conflict detection is a set intersection over key bytes inside SlateDB, with no hook to consult and no schema to consult it with. Push the semantics into the bytes and there is nothing left to get wrong.
Key structure stays parseable even though key values are not
decodable — nullability prefix bytes, fixed widths, and the
0x00 0x00 string terminator make part boundaries recoverable. That
is what lets a reader find the PK-suffix boundary (already true before
this refinement) and what lets the encoder decide, locally and cheaply,
whether a unique key has a NULL part and therefore which shape to
write. No decode anywhere, no HA_KEYREAD_ONLY, and index-only-scan
optimization is explicitly refused rather than half-supported.
The consequence worth writing down twice: key bytes embed collation weight tables, so a collation-table change is an on-disk compatibility event. The charset/collation codegen work must treat shipped weight tables as frozen-per-version; the engine stamps a format version into the meta prefix at database creation and refuses to open a future incompatible format rather than misreading it.
Value encoding: version byte, null bitmap, then each non-null field in
proto column order — fixed-width types as fixed-width little-endian
images, var-width as varint-length-prefixed bytes. Not memcomparable,
not clever; versioned so it can evolve. position()/rnd_pos use
the encoded PK bytes length-prefixed in ref, exactly the WiredTiger
cursor’s pattern.
Transactions and isolation¶
The engine is a TransactionalStorageEngine. Per-session state lives
in the Session::getEngineData slot (the innobase/WiredTiger
pattern): the shim transaction handle plus a small statement-state
struct.
Begin on first use, via
doStartStatement— adopting the WiredTiger engine’s fix verbatim (startTransactionstays a no-op so non-SlateDB transactions don’t inflateHandler_commit). First touch callsslatedb_txn_beginand registers the engine as a transaction resource.Isolation: snapshot only.
begin(IsolationLevel::Snapshot)always. The server’s READ COMMITTED / READ UNCOMMITTED requests are refused with an error, not silently upgraded — the same collapse the WiredTiger 11 port made, for the same reason, with the same user-visible honesty. SlateDB’sSerializableSnapshotis real and tested upstream; exposing it as a per-session option is a cheap later task, listed, not done.Commit:
slatedb_txn_commit. ACONFLICTreturn (write-write conflict with a concurrently committed transaction) maps toHA_ERR_LOCK_DEADLOCK— the one error code every SQL application already retries on. This is first-committer-wins optimistic concurrency and the docs say so plainly: hot-row workloads will see deadlock-shaped retries instead of lock waits.Rollback: whole-transaction rollback is free (
DbTransaction::rollbackdrops the buffered batch). Statement rollback and savepoints are refused: SlateDB’s transaction batch cannot be truncated to a mark, sodoSetSavepointreturns not-supported and a failed statement inside a multi-statement transaction rolls back the whole transaction and reports exactly that — the NDB-style answer the WiredTiger plan chose over the silent-no-op it audited (Tier 0.7). Never claim success for a rollback that didn’t happen.Read-your-own-writes works — the transaction’s write batch backs its own reads and scans, so multi-statement INSERT-then-SELECT behaves as SQL requires. (This is the capability that separates this engine from the Iceberg engine’s deliberately-blind commit buffer.)
DDL vs. fencing: unlike WiredTiger there is no
EBUSY-style DDL/checkpoint interaction — DDL is ordinary transactional writes to meta keys. The pre-DDL-checkpoint fixture pattern does not carry over because the problem it fixes does not exist here.
Duplicate keys are detected, not overwritten (WiredTiger Tier 0.2/0.3
learned the hard way): doInsertRecord probes the primary key with a
transactional get before put (a memtable/cache hit in the
common case), and probes each non-NULL-bearing UNIQUE secondary with a
point get on its unique-columns-only key, returning
HA_ERR_FOUND_DUPP_KEY with the correct key number; the concurrent
case falls through to the commit-time conflict described below. INSERT
… ON DUPLICATE KEY UPDATE therefore works through the standard kernel
path with no engine-side replace shortcut. A key-changing UPDATE is
delete-old + insert-new with dup probe, decided by comparing old/new
encoded keys (Tier 0.4). Writes during a secondary-index-driven scan go
through primary-key point writes, never through the scan’s iterator
(Tier 0.5) — which is natural here since scans and writes are separate
shim objects by construction.
What the conflict detector actually sees¶
Uniqueness enforcement rests on this, so it is verified and written
down rather than assumed. All of it read out of 0.14.1’s
db_transaction.rs and transaction_manager.rs:
Under
IsolationLevel::Snapshotonly the write set is tracked.get/scanadd nothing:get_key_value_with_optionsguards itstrack_read_keyscall onisolation_level == SerializableSnapshot, andcheck_has_conflictsays in as many words that Snapshot “only checks write-write conflicts”.The commit-time check is
has_write_write_conflict: a set intersection of this transaction’s write keys against the write keys of every transaction that committed after this one started. So two transactions thatputthe same key do conflict, and the loser getsSlateDBError::TransactionConflict. Exact bytes, exact keys — no ranges, no predicates, no schema.mark_read(keys)opts named keys into read-write detection even under Snapshot, andunmark_write(keys)opts named keys out of write-write detection. Both are real; neither takes a range, which is why neither rescues a predicate probe.
The consequence, and it is the load-bearing sentence of this section: a probe is not a lock. Reading a key and finding it absent buys nothing under Snapshot; only writing a key does. Every uniqueness guarantee in this engine must therefore reduce to “the competing transactions write the same physical key” — which is precisely why the unique-index layout above drops the PK suffix. With the suffix, two inserters of the same unique value write two different keys, both probe-miss, and both commit: a silently violated constraint, the WiredTiger Tier 0.3 failure mode wearing a different hat.
The probe stays, as an optimization on top of the conflict rather than as the mechanism. The two outcomes, both correct, stated so that neither surprises anyone:
Uncontended — the duplicate was already committed when this statement started: the probe hits and the engine returns
HA_ERR_FOUND_DUPP_KEYwith the right key number. This is essentially all real duplicate-key traffic.Contended — two concurrent inserts of the same primary key, or of the same unique value: both probe-miss, both write the same physical key, and the second committer gets
CONFLICT→HA_ERR_LOCK_DEADLOCKrather thanHA_ERR_FOUND_DUPP_KEY. This is acceptable and it is not worked around. The loser’s transaction is rolled back either way; deadlock is the one error every SQL client already retries; and on retry the winner’s row is committed, so the probe hits and the duplicate is reported properly. The only way to return the “right” error here would be to serialize inserts, which is the thing optimistic concurrency exists to avoid. The user documentation states it next to the hot-row retry note.
Durability, latency, and cost¶
Object storage makes the write-latency triangle explicit; the engine exposes SlateDB’s own knobs and adds none:
slatedb.url(required) andslatedb.path— the object store URL and database path.slatedb.settings-file— optional path handed through the shim verbatim to SlateDB’sfigment-basedSettingsloader (flush_interval, caches, compactor tuning, all of it). We do not mirror individual SlateDB settings as engine options; one passthrough, zero drift.slatedb.await-durable(default on): commits block until the WAL batch is durable in object storage — correctness-first, with commit latency ≈flush_interval+ one PUT round trip, and the documentation states those numbers rather than apologizing for them. Off: commit returns at memtable acceptance; a crash can lose the tail up to the last flush, stated equally plainly. This is per-server policy, not per-table.Read latency is governed by SlateDB’s in-memory block cache and optional local-disk object cache (
cached_object_store), configured through the settings file; the perf-tracking suite measures hot/cold read split explicitly.
flush() on the engine maps to Db::flush; a
slatedb_flush() UDF mirrors wiredtiger_checkpoint() for
operational use.
Statistics honesty (Tier 0.6): no ``HTON_HAS_RECORDS``, no
HTON_STATS_RECORDS_IS_EXACT. info() reports estimates and says
so. Nothing in this engine claims an exact row count, and nothing in it
maintains one.
Row count, v1: a capped prefix probe. Scan the table’s
0x01 || table-id || 0x00prefix up to a fixed row cap (a constant, tuned against the perf suite); exhausted under the cap → the exact count; hit the cap → a scaled estimate from key-space sampling. Cheap, honest, and wrong only in the direction optimizers already assume estimates are wrong.Cardinality and
records_in_range: bounded range probes under the same cap, and honest unknown-cardinality defaults where a probe is not worth its I/O. Never a fabricated floor — that is the Tier 0.6 anti-pattern itself.
Rejected, explicitly: a per-table row-count key maintained
transactionally with writes. It is the obvious way to get an exact
count and it is a design bug in this engine. Under first-committer-wins
every concurrent writer to a table would write that one key and
therefore write-write conflict with every other writer to the same
table: a single hot key serializing all writes to the table it is
counting, with the serialization delivered as HA_ERR_LOCK_DEADLOCK
storms rather than as waits. Exactness is not worth converting a
concurrent engine into a serial one, and this engine does not need it,
because it does not claim it.
For the record, since it will be asked: a conflict-free counter is
constructible on 0.14.1. There is a merge operator
(merge_operator.rs, DbTransaction::merge), and unmark_write
lifts a key out of write-write detection, so a per-statement delta
merged into an unmarked counter key would not conflict. It is listed as
a possible later task, not v1 work, and the caveat travels with the
listing: merge operands still land in the write batch (batch.rs:
keys() returns every op’s key, merges included), so the counter is
conflict-free only if the unmark is applied without exception — and an
unmarked key is by definition no longer transactional with the rows it
counts. That tradeoff earns its own evaluation with measurements, not a
paragraph here.
DROP TABLE reclamation¶
Two phases, both landing eventually, first one first:
v1 — foreground delete: DROP scans the table’s data prefix and issues batched deletes, then removes meta keys. Correct, simple, and slow for huge tables over S3; the docs say so.
Compaction-filter reclamation: DROP writes only the meta deletion +
0x00 'd' <id>marker; a compaction filter tombstones dead-prefix entries as compaction naturally rewrites them; a background sweep retires the marker once the prefix is empty. This is the steady-state design; v1’s foreground path remains asDROP TABLE’s behavior under a size threshold or explicit option if the split proves worth keeping — decide with data from the perf suite, not in advance.
The API as it actually exists in 0.14.1 (compaction_filter.rs,
read, not assumed): a CompactionFilterSupplier is registered once
at open through DbBuilder::with_compaction_filter_supplier, and its
create_compaction_filter(&CompactionJobContext) is called per
compaction job; the returned CompactionFilter sees every entry
through filter(&RowEntry) and answers Keep, Drop, or
Modify(ValueDeletable::Tombstone), with on_compaction_end as
the hook for per-job counters. We answer Modify(Tombstone), never
Drop: upstream documents that Drop removes an entry without
shadowing it, which can resurrect older versions from lower runs. The
feature is off by default (compaction_filters in
slatedb/Cargo.toml) and the shim crate turns it on. Upstream’s
snapshot-consistency warning is acceptable here by construction — the
ids we filter belong to tables that appear in no committed metadata and
have no live readers, because the server serialized the DROP against
open cursors already.
The dropped-id set is live, not a snapshot taken at open. A set read once at open would reclaim nothing dropped since the server started — which, in a long-running server, is every table anyone actually dropped. The shim owns it:
The filter supplier holds an
Arc<RwLock<HashSet<u32>>>, seeded at open from a scan of the0x00 'd'prefix.create_compaction_filtertakes the read lock once and copies the set into the job-local filter, so the per-entryfilter()call touches no lock at all. Per-job filter instances exist precisely so that per-job snapshots are the intended shape; a job that began before an id was added simply misses that table this round and the next compaction takes it.Two shim functions,
slatedb_dropped_ids_add(db, id, out_err)andslatedb_dropped_ids_remove(db, id, out_err), mutate the set. They are the only additions to the shim surface after task 1.The ordering is the entire correctness argument.
doDropTablecallsaddonly after the DROP transaction’s commit returns OK — not before it, not onCONFLICT, not on rollback, not from a destructor, not on any path that is not a returned-OK commit. An id in the filter set whose DROP did not commit tombstones a live table’s rows.Symmetrically the retirement sweep calls
removeonly after the marker-deletion transaction commits. Removing first and then failing to commit merely leaves a marker nothing acts on until the next restart — but it is the same rule, so it is the same rule.Crash safety needs no extra machinery: the
0x00 'd'prefix in the database is the authority, the in-memory set is a cache rebuilt from it at every open, and both operations are idempotent. A crash between commit and set update costs a delayed reclamation, never a row.
Operational model¶
Exactly one writing server per (url, path). SlateDB’s epoch fencing makes a second writer win: the new server fences the old, whose SlateDB calls start failing
FENCED. The engine treatsFENCEDas fatal — log loudly and refuse further writes (HA_ERR_*mapped to a read-only/IO error), because the honest description of a fenced writer is “another server owns this bucket now.” This inverts the usual failover ergonomics (starting the replacement kills the incumbent) and the deployment docs must lead with it. It is also exactly what makes lease-less failover work.Backup = checkpoint. SlateDB checkpoints (
create_checkpoint,Admin::create_detached_checkpoint) are consistent named references in the same bucket; theslatedbCLI manages them. Aslatedb_checkpoint()UDF is a natural follow-on once wanted.Read replicas (future, designed-for):
DbReaderagainst the same path, following latest or pinned to a checkpoint. The engine split (writer engine vs. reader engine mode) is a later spec; the key/value codec and the one-Db model are the parts that must not foreclose it, and they don’t.
Refused in v1 (all loud, all specific)¶
Refusal |
Why / when it could return |
|---|---|
READ COMMITTED / READ UNCOMMITTED |
Snapshot-only, per the WiredTiger precedent; never returns |
Savepoints / statement rollback |
SlateDB txn batch is not truncatable; whole-txn rollback reported honestly |
SERIALIZABLE (SSI) |
Cheap later task — SlateDB supports it |
BLOB columns and blob keys |
Later task (value codec is ready; cursor blob buffer management is the work) |
AUTO_INCREMENT |
Later task: transactional counter key with block reservation; refuse rather than ship a racy version |
Index-only scans (HA_KEYREAD_ONLY) |
Keys are strnxfrm-encoded, never decoded; by design, likely permanent |
ALTER, TEMPORARY tables |
|
Multiple SlateDB databases per server |
One |
Implementation plan¶
Seven task handoffs in
doc/source/specs/implementation-plans/slatedb-engine/; critical
chain 1 → 2 → 3 → 4 → 5, then 6; 7 floats after 5. Tasks 1–4 deliver a
readable/writable single-index engine under autocommit; task 5 delivers
real transactions; task 6 delivers secondary and unique indexes; task 7
is reclamation and operational polish.