Task 5: Transactional buffered append

Warning

This is not authoritative documentation. It describes a plan of work that is not yet implemented and will change as it lands. Once the work is complete this document should be removed; the finished system is described by the Iceberg storage engine spec, the plugin’s user documentation, and the repos themselves.

Repo: https://opendev.org/drizzle/drizzle. Depends on task 4. Two commits. Goal: INSERT works, buffered until commit, one Iceberg append commit per SQL commit. After this task, use case 2’s write half (cold-tier roll-off) works against tables created elsewhere.

Binding constraints: doInsertRecord never touches object storage; rollback is buffer-drop; UPDATE and row-level DELETE remain refused (HA_ERR_WRONG_COMMAND) with tests asserting so.

Commit 1: engine becomes transactional; write path

Reparent

IcebergEngine : StorageEngine → : TransactionalStorageEngine (drizzled/plugin/transactional_storage_engine.h). Registration moves to the transactional addPlugin chain automatically via Registry::add’s type dispatch — verify, don’t assume, that the engine now appears in the transactional engine vector and participates in SQL transactions (participatesInSqlTransaction true, XA false).

session_state grows write buffers

Per (session, table): an Arrow builder set (one per column, by field-id), row count, and savepoint marks. start_bulk_insert(ha_rows) pre-sizes the builders; end_bulk_insert is a no-op (commit is the flush point, unconditionally).

Cursor write hooks

  • doInsertRecord(buf): unpack the record buffer via the write map (the inverse of task 4’s read conversion — same file, executor’s conversion unit or a sibling; every type round-trips) into the session’s builders. Auto-increment: not supported; the synthesized proto never declares one.

  • doUpdateRecord / doDeleteRecord: HA_ERR_WRONG_COMMAND.

Transaction hooks (iceberg_engine.cc)

  • doStartTransaction: create the slot entry; nothing external.

  • doCommit(session, normal_transaction): for each buffered table with rows — finalize Arrow table → write Parquet data file(s) through iceberg-cpp’s writer at Iceberg-normal target sizes (library defaults; do not invent tuning options yet) → single append commit via the transaction API. Catalog CAS conflict → library-level retry of the metadata commit (data files remain valid), bounded attempts, then return commit failure to the server so the SQL layer reports it. Partitioned tables: rows are fanned to partitions by the table’s spec via the library’s partitioned writer — verify the pinned iceberg-cpp version supports partitioned appends (the spike README should already say; if it cannot, appends to partitioned tables are refused with a specific error and that refusal is tested and documented — do not write unpartitioned files into a partitioned table).

  • doRollback: drop buffers. Also handles statement rollback semantics per the normal_transaction flag exactly as the base contract requires — statement-level rollback truncates buffers to the statement-start mark.

  • Savepoints (doSetSavepoint/doRollbackToSavepoint/doReleaseSavepoint — pure virtuals, must be implemented): row-count marks per table buffer; rollback-to truncates builders to the mark.

  • Pin-clearing moves here from task 4’s statement-end interim for transactional sessions; resolve the marker task 4 left. Autocommit still clears at statement end.

Commit 2: test suites

  • Single INSERT autocommit → readable by pyiceberg (interop assert), one new snapshot.

  • Multi-statement transaction: INSERT, INSERT, ROLLBACK → no snapshot; INSERT, SAVEPOINT, INSERT, ROLLBACK TO, COMMIT → exactly the first rows, one snapshot.

  • The roll-off shape end to end: populate an innobase hot table, BEGIN; INSERT INTO iceberg_t SELECT ... WHERE aged; COMMIT; then verify counts in Iceberg via a second SELECT and via pyiceberg; then DELETE FROM hot. Document in the suite (and user docs) that the two commits are not atomic with each other and the safe ordering is insert-verify-delete.

  • No read-your-own-writes: INSERT then SELECT in one transaction does not see the buffered rows; test asserts the documented behavior.

  • Commit conflict: background pyiceberg append racing a Drizzle commit; the Drizzle commit succeeds via retry (or fails cleanly if retries exhaust — either way, no torn state: re-scan shows a consistent snapshot).

  • UPDATE/DELETE refused with the specific error.

Verification

  • Build green in both switch states (--with-iceberg on/off) per commit; suites green.

  • Valgrind + massif on a large bulk INSERT…SELECT: builder memory scales with buffered rows and is fully released post-commit; no leaks.

  • Confirm in MinIO that a failed/rolled-back transaction left no committed snapshot (orphaned data files from a failed CAS after exhausted retries are acceptable and noted — Iceberg maintenance territory, not ours).