Warning

This is not authoritative documentation. It describes a plan of work that is not yet implemented and will change as it lands.

Task 2: the memcomparable key codec and row value codec

Repo: https://opendev.org/drizzle/drizzle, plugin/slatedb/codec/. Pure C++, no SlateDB dependency, no cursor code — the codec is a standalone library with its own exhaustive tests, landed and reviewed before anything consumes it. Independent of task 1; depends on nothing.

Why first-class: this is the one component where a bug is silent wrong results (mis-sorted scans, aliased multi-part keys), the exact failure family the WiredTiger audit spent its Tier 0 on. It gets the test budget accordingly.

Deliverables

key_codec.{h,cc}:

  • encodeKeyPart(Field&, uchar* out) and encodeKey over a KeyInfo, implementing the spec’s per-type rules: null prefix byte, sign-flipped big-endian integers, IEEE-754 total-order doubles, fixed-width images (BOOLEAN/UUID/IPV6), my_decimal binary passthrough, and collation strnxfrm weights with 0x00 → 0x00 0xFF escaping and 0x00 0x00 termination for text.

  • encodeKeyFromIndexBuf — the server-side key-buffer form used by index_read (the path WiredTiger’s Tier 0.9 NULL bug lived in; NULL handling here is tested per type, not FIXME’d).

  • Prefix-key support: encoding a leading subset of key parts for HA_READ_KEY_OR_NEXT-family positioning, with the invariant that a prefix encoding is a byte-prefix of every full encoding it matches.

  • keyHasNullPart over an encoded key or a KeyInfo + record — the predicate the unique-index layout selects on (task 6): a unique key with a NULL part gets the PK-suffixed shape, one without gets the suffix-free shape.

  • encodedKeyLength over a KeyInfo — part-boundary recovery from key bytes alone. Structure is parseable (nullability prefix bytes, fixed widths, the 0x00 0x00 terminator) even though values are not; this is what slices a PK suffix off an index entry, and it is tested as its own property.

  • No decode functions exist. Their absence is the design; recovering boundaries is not recovering values.

value_codec.{h,cc}:

  • encodeRow(Table&, const uchar* record) / decodeRow — version byte (0x01), null bitmap, proto-order fields, varint-length var-width values. All columns present including PK columns.

keyspace.{h,cc}: the layout constants and builders from the spec (meta prefixes, 0x01 || table-id || index-id data prefix, prefixUpperBound for scan ranges).

Tests (the point of the task)

Property-style table-driven tests in plugin/slatedb/codec/tests/ runnable in the plain build (no server): for every supported type and for representative multi-part keys with and without NULLs, generate value sets, encode, and assert byte order equals server comparison order (compare against Field::cmp / key_cmp ground truth), including: INT_MIN/−1/0/1/INT_MAX, ±0.0, ±inf, NaN policy (documented: NaN refused at the Field layer before the codec — verify the server already guarantees this and record where), empty strings, strings differing only in trailing space per collation, embedded NUL bytes, and utf8 multi-byte boundaries under the general and binary collations. Round-trip tests for the value codec over randomized rows. Boundary tests for encodedKeyLength: for every multi-part key in the corpus, concatenating an index-cols encoding with a PK encoding and recovering the split must reproduce both halves exactly, including keys whose text parts contain escaped NUL bytes. keyHasNullPart is tested per type and per part position, not just “all NULL”.

Commit boundary

Two commits: key codec + tests; value codec + keyspace + tests.

Verification

  • Build and codec tests green in the standard builder image (no slatedb-capi needed).

  • A short docs/ note freezing the format: version byte meanings, the collation-weight compatibility caveat (cross-referenced from the charset codegen work), and the rule that format changes bump the version and never reinterpret 0x01.