Task 3: Catalog-backed metadata

Warning

This is not authoritative documentation. It describes a plan of work that is not yet implemented and will change as it lands. Once the work is complete this document should be removed; the finished system is described by the Iceberg storage engine spec, the plugin’s user documentation, and the repos themselves.

Repo: https://opendev.org/drizzle/drizzle. Depends on task 2. Two commits. Goal: Iceberg namespaces appear as Drizzle schemas and Iceberg tables as tables — listable, DESCRIBE-able, SHOW CREATE-able — with the catalog as the sole source of truth. No data access yet (SELECT still fails via the task-2 stub cursor). The model is plugin/function_engine (function.cc:54-67): synthesize protos, own no files.

Commit 1: namespace/table enumeration and schema translation

Engine hooks (iceberg_engine.cc)

  • doGetSchemaIdentifiers(identifier::schema::vector&): list namespaces from the catalog. Map Iceberg’s multi-level namespaces by joining levels with a separator only if the fixture catalogs actually nest; otherwise support single-level and refuse nested namespaces explicitly (subtractive — decide from what pyiceberg+Polaris actually produce, note the decision in the commit message).

  • doGetSchemaDefinition(const identifier::Schema&): synthesize a message::Schema (utf8_general_ci, is_replicated=false — same settings function_engine uses) for namespaces that exist; empty shared_ptr otherwise.

  • doGetTableIdentifiers(...): catalog table list for the namespace; ignore the CachedDirectory parameter (as function_engine does).

  • doDoesTableExist: direct catalog check (cheaper than definition fetch).

  • doGetTableDefinition(...): load the table, translate schema (below), return EEXIST; ENOENT for absent; a specific, named error for tables containing unsupported types.

Name-collision rule: if an Iceberg namespace collides with a local schema name, local wins and the engine logs; do not invent merge semantics.

schema_map.{h,cc} — the translation, isolated and unit-testable

Iceberg schema → message::Table, per the design-doc type table: boolean→ BOOLEAN, int→INTEGER, long→BIGINT, float→DOUBLE, double→DOUBLE, decimal(p,s)→DECIMAL (p over Drizzle’s max → refuse), date→DATE, time→TIME, timestamp→DATETIME, timestamptz→EPOCH (UTC), string→VARCHAR, binary/fixed→BLOB, uuid→UUID. Nullability from Iceberg required/optional. Iceberg field-ids must be retained in the mapping structure (task 4 needs them for projection; task 5 for writes) — carry a side table drizzle field index ↔ iceberg field id, owned by the engine’s per-table share, not smuggled through the proto.

Refusals (error names column and type): struct, list, map, variant, geometry/geography, timestamp_ns, unknown-to-us type ids (fail closed on anything a newer library version introduces). One table = fully representable or not attached; no column hiding.

Set message::Table::TableOptions comment from the Iceberg table’s comment property if present. Engine name in the proto: iceberg.

Commit 2: drizzle-test suites

Against the task-1 fixtures: SHOW SCHEMAS includes the fixture namespace; SHOW TABLES; DESCRIBE each fixture table with expected column types (including the uuid and decimal columns); SHOW CREATE TABLE renders; the struct-column fixture table errors with the specific refusal message (result-file asserts the message, not just failure); a table created by pyiceberg mid-suite appears after FLUSH TABLES (documents the cache behavior rather than hiding it).

Flip the Zuul iceberg job voting for these suites.

Verification

  • Build green in both switch states (--with-iceberg on/off); suites green.

  • schema_map unit-style coverage for every mapped type and every refusal.

  • No file writes anywhere under the datadir attributable to the plugin (bas_ext empty; nothing creates .dfes).