Warning
This is not authoritative documentation. It describes a plan of work that is not yet implemented and will change as it lands. Once the work is complete this document should be removed; the finished system is described by the Iceberg storage engine spec, the plugin’s user documentation, and the repos themselves.
Task 1: iceberg-cpp spike and CI fixtures — container-first¶
Repos: https://opendev.org/drizzle/drizzle-test (images/fixtures) plus a spike directory (not merged into drizzle). Goal: de-risk the dependency and stand up the test infrastructure every later task needs, before any engine code exists.
Build philosophy (binding for this task and everything downstream)¶
Container-first. The only build command a human or Zuul ever runs is
podman build (wrapped in just targets for convenience). All
dependencies are pinned exactly — image bases by digest, source deps by
tag/sha, Python deps by version — in Containerfiles, which are the
single pin location. There is no dependency detection anywhere: no
CMake, no configure probes, no “found libfoo: yes”. We know exactly what
is in the image because we put it there; if a pin is wrong the build
fails loudly at podman build time, which is the correct failure
mode. Rootless podman throughout, matching the
system-config/perf-tracking infrastructure.
One necessary asterisk: iceberg-cpp and (if built from source) Arrow use
CMake as their internal build systems. That is upstream’s private
detail, invoked by a RUN line inside podman build of the deps
image, and it never leaks: nothing of ours is configured, detected, or
built by CMake, and no CMake exists in the runtime layers or in any of
our own build steps.
Part 1: images¶
Debian trixie bases, pinned by digest. Layered so the expensive parts cache:
``iceberg-deps``: Arrow C++/Parquet — from the Apache Arrow apt repository at an exact pinned version if trixie-compatible packages exist, else built from a pinned source tag; decide once, record why in the Containerfile comment. iceberg-cpp built from its pinned release tag (0.3.0) and installed to a known prefix. Also: g++, make, valgrind, gdb — the toolchain the spike (and later the plugin CI) compiles and debugs with. This image is the ancestor of the drizzle builder image the plugin work uses from task 2 on; the version pins live here and nowhere else.
``iceberg-spike``:
FROM iceberg-deps, copies the spike source, runsmake. The Makefile is plain make with explicit-I/-L/-lflags against the known prefix — no pkg-config, nothing conditional. Under ten lines of Makefile is the right size; if it wants to grow logic, the logic belongs in the Containerfile instead.Fixture images, pinned by digest: the REST catalog (lightest faithful implementation — Polaris standalone or the reference
iceberg-rest-fixture; pin and note why), MinIO, and a seeder image (python + pinned pyiceberg) whose entrypoint creates the fixture tables: one simple, one partitioned (day(...)+bucket(...)), one with position deletes, one containing a struct column for the refusal tests, several snapshots deep.
Orchestration: a kube yaml for the catalog+MinIO+seeder stack, run via
podman kube play — the same shape as the quadlet .kube units in
production infrastructure, so the fixture definition is directly
reusable if a long-lived test environment ever wants it. A Justfile
fronts everything: just images, just up, just seed,
just spike, just valgrind, just down. Zuul runs the
identical targets — gate/dev parity by construction, not by copied
package lists.
Part 2: the spike program¶
A standalone C++ program, built by the plain Makefile inside
iceberg-spike, run via podman run on the fixture network.
Against the seeded stack it must:
Connect to the catalog from URI + credentials; list namespaces and tables.
Load a table created and populated by pyiceberg (interop is the thing under test, not self-consistency); print schema, current snapshot id, partition spec.
Plan a scan of the current snapshot, iterate FileScanTasks, read Arrow record batches, print per-task row counts and a column checksum. Include the position-deletes table and confirm deleted rows are absent.
Append rows: build an Arrow table, write a Parquet data file, commit an append through the transaction API. Re-scan and confirm old + new rows; then read back with pyiceberg (via the seeder image) and confirm it sees the appended rows — bidirectional interop.
Exercise the commit-conflict path: two spike invocations race an append; confirm the loser retries (or fails cleanly) as the library documents.
README deliverables: the exact pins and why; which iceberg-cpp APIs served each step (these become the engine’s integration points); API gaps and sharp edges; whether partitioned appends are supported at the pin (task 5 depends on this answer); the transitive dependency list with licenses (Arrow is Apache-2.0; confirm nothing GPLv2-incompatible rides along).
Gate: if step 3 or 4 cannot be completed against the pinned release, stop and report — the plan re-sequences around what the library can actually do, and that is cheaper to learn here than in task 4.
Part 3: the Zuul job¶
A job in the vouched pipeline (non-voting until task 4 lands) that runs
the same just targets: build images, play the fixture stack, seed,
run the spike, tear down. Later tasks point real drizzle-test suites at
the same fixture stack and flip it voting. Image builds publish through
the existing promote pipeline pattern so gate runs pull, not rebuild,
where caching allows.
Verification¶
From a clean host with only podman and just:
just images && just up && just seed && just spikepasses end to end. No host-installed toolchain, no host-installed libraries.just valgrind(the spike under valgrind, inside the container) clean — the library’s leak behavior is worth knowing before it lives in a server.Fixture bring-up reproducible from scratch;
just down && just upleaves no state between runs.README findings reviewed before task 2 starts — it is the input to the builder-image and linkage decisions, and to task 5’s partitioned-write question.