Task 5: Transactional buffered append¶
Warning
This is not authoritative documentation. It describes a plan of work that is not yet implemented and will change as it lands. Once the work is complete this document should be removed; the finished system is described by the Iceberg storage engine spec, the plugin’s user documentation, and the repos themselves.
Repo: https://opendev.org/drizzle/drizzle. Depends on task 4. Two commits. Goal: INSERT works, buffered until commit, one Iceberg append commit per SQL commit. After this task, use case 2’s write half (cold-tier roll-off) works against tables created elsewhere.
Binding constraints: doInsertRecord never touches object storage;
rollback is buffer-drop; UPDATE and row-level DELETE remain refused
(HA_ERR_WRONG_COMMAND) with tests asserting so.
Commit 1: engine becomes transactional; write path¶
Reparent¶
IcebergEngine : StorageEngine → : TransactionalStorageEngine
(drizzled/plugin/transactional_storage_engine.h). Registration moves
to the transactional addPlugin chain automatically via
Registry::add’s type dispatch — verify, don’t assume, that the
engine now appears in the transactional engine vector and participates
in SQL transactions (participatesInSqlTransaction true, XA false).
session_state grows write buffers¶
Per (session, table): an Arrow builder set (one per column, by
field-id), row count, and savepoint marks.
start_bulk_insert(ha_rows) pre-sizes the builders;
end_bulk_insert is a no-op (commit is the flush point,
unconditionally).
Cursor write hooks¶
doInsertRecord(buf): unpack the record buffer via the write map (the inverse of task 4’s read conversion — same file,executor’s conversion unit or a sibling; every type round-trips) into the session’s builders. Auto-increment: not supported; the synthesized proto never declares one.doUpdateRecord/doDeleteRecord:HA_ERR_WRONG_COMMAND.
Transaction hooks (iceberg_engine.cc)¶
doStartTransaction: create the slot entry; nothing external.doCommit(session, normal_transaction): for each buffered table with rows — finalize Arrow table → write Parquet data file(s) through iceberg-cpp’s writer at Iceberg-normal target sizes (library defaults; do not invent tuning options yet) → single append commit via the transaction API. Catalog CAS conflict → library-level retry of the metadata commit (data files remain valid), bounded attempts, then return commit failure to the server so the SQL layer reports it. Partitioned tables: rows are fanned to partitions by the table’s spec via the library’s partitioned writer — verify the pinned iceberg-cpp version supports partitioned appends (the spike README should already say; if it cannot, appends to partitioned tables are refused with a specific error and that refusal is tested and documented — do not write unpartitioned files into a partitioned table).doRollback: drop buffers. Also handles statement rollback semantics per thenormal_transactionflag exactly as the base contract requires — statement-level rollback truncates buffers to the statement-start mark.Savepoints (
doSetSavepoint/doRollbackToSavepoint/doReleaseSavepoint— pure virtuals, must be implemented): row-count marks per table buffer; rollback-to truncates builders to the mark.Pin-clearing moves here from task 4’s statement-end interim for transactional sessions; resolve the marker task 4 left. Autocommit still clears at statement end.
Commit 2: test suites¶
Single INSERT autocommit → readable by pyiceberg (interop assert), one new snapshot.
Multi-statement transaction: INSERT, INSERT, ROLLBACK → no snapshot; INSERT, SAVEPOINT, INSERT, ROLLBACK TO, COMMIT → exactly the first rows, one snapshot.
The roll-off shape end to end: populate an innobase hot table,
BEGIN; INSERT INTO iceberg_t SELECT ... WHERE aged; COMMIT;then verify counts in Iceberg via a second SELECT and via pyiceberg; thenDELETE FROM hot. Document in the suite (and user docs) that the two commits are not atomic with each other and the safe ordering is insert-verify-delete.No read-your-own-writes: INSERT then SELECT in one transaction does not see the buffered rows; test asserts the documented behavior.
Commit conflict: background pyiceberg append racing a Drizzle commit; the Drizzle commit succeeds via retry (or fails cleanly if retries exhaust — either way, no torn state: re-scan shows a consistent snapshot).
UPDATE/DELETE refused with the specific error.
Verification¶
Build green in both switch states (
--with-icebergon/off) per commit; suites green.Valgrind + massif on a large bulk INSERT…SELECT: builder memory scales with buffered rows and is fully released post-commit; no leaks.
Confirm in MinIO that a failed/rolled-back transaction left no committed snapshot (orphaned data files from a failed CAS after exhausted retries are acceptable and noted — Iceberg maintenance territory, not ours).