Warning

This is not authoritative documentation. It describes a plan of work that is not yet implemented and will change as it lands. Once the work is complete this document should be removed; the finished system is described by the Iceberg storage engine spec, the plugin’s user documentation, and the repos themselves.

Task 6: DDL — CREATE with transforms, DROP without purge, RENAME

Repo: https://opendev.org/drizzle/drizzle. Depends on task 5. Two commits. Goal: self-service table lifecycle. Zero parser changes — verified: the grammar already accepts arbitrary engine options (sql_yacc.yy:1098-1105, ident_or_text '=' engine_option_value → parser::buildEngineOption → Engine::Option pairs, message/engine.proto:12-21).

Commit 1: transforms mini-language + CREATE

transforms.{h,cc} — parse and validate, isolated and unit-tested

Input: the PARTITION_BY option string, a comma-separated list of Iceberg’s own transform spellings — identity(col) or bare col, year(col), month(col), day(col) (accept days etc. aliases exactly as Iceberg SQL does), hour(col), bucket(N, col), truncate(W, col). And SORT_ORDER: comma-separated col [asc|desc] [nulls first|nulls last].

Validation, each with a specific error: unknown transform; unknown column; transform/type incompatibility (e.g. day() on a non-date/timestamp column — mirror Iceberg’s own compatibility matrix, don’t invent one); duplicate source column where Iceberg forbids it; malformed syntax with position. Output: iceberg-cpp partition-spec / sort-order builder inputs.

doCreateTable (iceberg_engine.cc)

  • Flip doCanCreateTable to true (delete the attach-only comment; the commit message notes v1’s attach-only period ends here).

  • Translate the message::Table field list → Iceberg schema — the inverse of task 3’s schema_map; it lives there, and every type round-trips (create-then-describe is identity). Drizzle types with no Iceberg inverse (ENUM, IPV6) are refused at create with a specific error.

  • Read PARTITION_BY / SORT_ORDER from the proto’s engine options (absent → unpartitioned/unsorted); parse via transforms; create through the catalog. Unknown engine options are an error, not ignored (catches typos like PARITION_BY).

  • Table properties passthrough: any engine option of the form ICEBERG.<prop>='v' becomes an Iceberg table property verbatim — the escape hatch that avoids growing bespoke options for every knob.

DROP and RENAME

  • doDropTable: catalog drop without purge — deregister only, never delete data files. The doxygen comment and user docs carry the rationale (attached tables are the ecosystem’s data; Drizzle tidying its view must not destroy it). Purge is unsupported.

  • doRenameTable: catalog rename within a namespace; cross-namespace → specific refusal.

  • ALTER remains refused via HTON_ALTER_NOT_SUPPORTED (schema evolution belongs to a future task once iceberg-cpp’s update-schema surface is proven; do not partially wire it).

Commit 2: test suites

  • CREATE TABLE ... ENGINE=iceberg PARTITION_BY='day(created_at), bucket(16, user_id)' SORT_ORDER='created_at desc' → pyiceberg reads back the identical spec (interop assert on partition spec JSON, not just success).

  • Create-describe round-trip identity for every supported type.

  • Every transforms validation error has a result-file test.

  • INSERT into a self-created partitioned table → pyiceberg confirms rows landed in correct partitions (exercises task 5’s partitioned writer against our own spec).

  • DROP → table gone from catalog; data files still present in MinIO (asserted); re-CREATE with the same name works.

  • RENAME within namespace; cross-namespace refusal.

Verification

  • Build green in both switch states (--with-iceberg on/off) per commit; suites green.

  • transforms unit coverage: full transform matrix × valid/invalid.

  • Docs updated: the DDL page with the option mini-language, the property passthrough, and the drop-never-purges guarantee in its own box.