Warning
This is not authoritative documentation. It describes a plan of work that is not yet implemented and will change as it lands. Once the work is complete this document should be removed; the finished system is described by the Iceberg storage engine spec, the plugin’s user documentation, and the repos themselves.
Task 6: DDL — CREATE with transforms, DROP without purge, RENAME¶
Repo: https://opendev.org/drizzle/drizzle. Depends on task 5. Two
commits. Goal: self-service table lifecycle. Zero parser changes —
verified: the grammar already accepts arbitrary engine options
(sql_yacc.yy:1098-1105, ident_or_text '=' engine_option_value →
parser::buildEngineOption → Engine::Option pairs,
message/engine.proto:12-21).
Commit 1: transforms mini-language + CREATE¶
transforms.{h,cc} — parse and validate, isolated and unit-tested¶
Input: the PARTITION_BY option string, a comma-separated list of
Iceberg’s own transform spellings — identity(col) or bare col,
year(col), month(col), day(col) (accept days etc.
aliases exactly as Iceberg SQL does), hour(col), bucket(N, col),
truncate(W, col). And SORT_ORDER: comma-separated
col [asc|desc] [nulls first|nulls last].
Validation, each with a specific error: unknown transform; unknown
column; transform/type incompatibility (e.g. day() on a
non-date/timestamp column — mirror Iceberg’s own compatibility matrix,
don’t invent one); duplicate source column where Iceberg forbids it;
malformed syntax with position. Output: iceberg-cpp partition-spec /
sort-order builder inputs.
doCreateTable (iceberg_engine.cc)¶
Flip
doCanCreateTableto true (delete the attach-only comment; the commit message notes v1’s attach-only period ends here).Translate the
message::Tablefield list → Iceberg schema — the inverse of task 3’sschema_map; it lives there, and every type round-trips (create-then-describe is identity). Drizzle types with no Iceberg inverse (ENUM, IPV6) are refused at create with a specific error.Read
PARTITION_BY/SORT_ORDERfrom the proto’s engine options (absent → unpartitioned/unsorted); parse viatransforms; create through the catalog. Unknown engine options are an error, not ignored (catches typos likePARITION_BY).Table properties passthrough: any engine option of the form
ICEBERG.<prop>='v'becomes an Iceberg table property verbatim — the escape hatch that avoids growing bespoke options for every knob.
DROP and RENAME¶
doDropTable: catalog drop without purge — deregister only, never delete data files. The doxygen comment and user docs carry the rationale (attached tables are the ecosystem’s data; Drizzle tidying its view must not destroy it). Purge is unsupported.doRenameTable: catalog rename within a namespace; cross-namespace → specific refusal.ALTER remains refused via
HTON_ALTER_NOT_SUPPORTED(schema evolution belongs to a future task once iceberg-cpp’s update-schema surface is proven; do not partially wire it).
Commit 2: test suites¶
CREATE TABLE ... ENGINE=iceberg PARTITION_BY='day(created_at), bucket(16, user_id)' SORT_ORDER='created_at desc'→ pyiceberg reads back the identical spec (interop assert on partition spec JSON, not just success).Create-describe round-trip identity for every supported type.
Every transforms validation error has a result-file test.
INSERT into a self-created partitioned table → pyiceberg confirms rows landed in correct partitions (exercises task 5’s partitioned writer against our own spec).
DROP → table gone from catalog; data files still present in MinIO (asserted); re-CREATE with the same name works.
RENAME within namespace; cross-namespace refusal.
Verification¶
Build green in both switch states (
--with-icebergon/off) per commit; suites green.transformsunit coverage: full transform matrix × valid/invalid.Docs updated: the DDL page with the option mini-language, the property passthrough, and the drop-never-purges guarantee in its own box.