Warning

This is not authoritative documentation. It describes a plan of work that is not yet implemented and will change as it lands.

Task 5: multi-statement transactions

Repo: https://opendev.org/drizzle/drizzle. Depends on task 4. Delivers BEGIN/COMMIT/ROLLBACK with snapshot isolation and honest refusals.

Session state and lifecycle

Per-session struct in the Session::getEngineData slot (the WiredTiger/innobase pattern): the shim txn handle plus flags. Lifecycle adopts the WiredTiger engine’s begin-on-first-use shape verbatim:

  • startTransaction: no-op (with the same explanatory comment — the kernel notifies every transactional engine on BEGIN; registering eagerly double-counts commits for transactions that never touch this engine).

  • doStartStatement: if no txn handle, slatedb_txn_begin (Snapshot) and register as a transaction resource.

  • doCommit(all=true): slatedb_txn_commit; CONFLICT → HA_ERR_LOCK_DEADLOCK; clear the slot. await-durable governs whether commit blocks to object-store durability.

  • doRollback(all=true): slatedb_txn_rollback; clear the slot.

  • doRollback(all=false) (statement rollback): roll back the whole transaction and return the error state that says so — the NDB-shaped answer chosen over WiredTiger’s audited silent no-op (Tier 0.7). Verify precisely which error propagation the kernel expects here by reading the callers in drizzled/transaction_services.cc before implementing; record findings in the commit message.

  • doSetSavepoint / doRollbackToSavepoint / doReleaseSavepoint: return not-supported. Result files assert the user-visible error.

  • Isolation requests other than the server default snapshot behavior (READ COMMITTED / READ UNCOMMITTED / SERIALIZABLE): refuse with the named error, matching the WiredTiger 11 collapse. SSI is recorded as a follow-on, not smuggled in.

  • closeConnection: roll back any live txn, free the slot.

Read-your-own-writes needs no engine code — the shim transaction’s scans and gets already see its buffered writes (verified in the task-1 spike, step 2) — but gets its own drizzle-test case anyway, because it is the semantic line between this engine and the Iceberg engine’s buffer-blind design.

Statistics

No row-count key, and this task must not invent one. The spec’s “Statistics honesty” section rejects a per-table counter maintained transactionally with writes, and the reason is structural: under first-committer-wins every concurrent writer to a table would write that one key and conflict with every other writer to the table, serializing all of them and reporting it as deadlocks. info() keeps task 4’s capped-prefix-probe estimate, which transactions do not change.

Commit boundary

Two commits: transaction lifecycle (begin-on-first-use, commit, rollback, CONFLICT mapping) with tests; refusals (savepoints, statement rollback, isolation levels) with result files asserting the user-visible errors.

Verification

  • drizzle-test: BEGIN/COMMIT visibility across sessions, ROLLBACK, read-your-own-writes, snapshot repeatable-read (session A reads, session B commits a change, A re-reads same value, A re-reads new value after its own commit), two-session write-write conflict returning the deadlock-mapped error with the loser able to retry, savepoint refusal, statement-rollback whole-txn semantics asserted explicitly in a result file (this is a documented behavior, not a bug to hide).

  • Handler_commit counting: a transaction touching only another engine must not increment SlateDB commits (the regression the begin-on-first-use shape exists to prevent).

  • No engine-manufactured hot keys: N sessions concurrently INSERT/UPDATE distinct rows of one table and every commit succeeds. This is the standing regression test for the deleted row-count key and for anything like it — a conflict here means the engine invented a shared key behind the user’s back.

  • Concurrency smoke under MinIO in the periodic pipeline: N sessions mixed workload, assertions on invariant sums.