.. warning:: This is not authoritative documentation. It describes a plan of work that is not yet implemented and will change as it lands. Task 2: the memcomparable key codec and row value codec ======================================================= Repo: https://opendev.org/drizzle/drizzle, ``plugin/slatedb/codec/``. Pure C++, no SlateDB dependency, no cursor code — the codec is a standalone library with its own exhaustive tests, landed and reviewed before anything consumes it. Independent of task 1; depends on nothing. Why first-class: this is the one component where a bug is silent wrong results (mis-sorted scans, aliased multi-part keys), the exact failure family the WiredTiger audit spent its Tier 0 on. It gets the test budget accordingly. Deliverables ------------ ``key_codec.{h,cc}``: - ``encodeKeyPart(Field&, uchar* out)`` and ``encodeKey`` over a ``KeyInfo``, implementing the spec's per-type rules: null prefix byte, sign-flipped big-endian integers, IEEE-754 total-order doubles, fixed-width images (BOOLEAN/UUID/IPV6), ``my_decimal`` binary passthrough, and collation ``strnxfrm`` weights with ``0x00 → 0x00 0xFF`` escaping and ``0x00 0x00`` termination for text. - ``encodeKeyFromIndexBuf`` — the server-side key-buffer form used by ``index_read`` (the path WiredTiger's Tier 0.9 NULL bug lived in; NULL handling here is tested per type, not FIXME'd). - Prefix-key support: encoding a leading subset of key parts for ``HA_READ_KEY_OR_NEXT``-family positioning, with the invariant that a prefix encoding is a byte-prefix of every full encoding it matches. - ``keyHasNullPart`` over an encoded key or a ``KeyInfo`` + record — the predicate the unique-index layout selects on (task 6): a unique key with a NULL part gets the PK-suffixed shape, one without gets the suffix-free shape. - ``encodedKeyLength`` over a ``KeyInfo`` — part-boundary recovery from key bytes alone. Structure is parseable (nullability prefix bytes, fixed widths, the ``0x00 0x00`` terminator) even though values are not; this is what slices a PK suffix off an index entry, and it is tested as its own property. - No decode functions exist. Their absence is the design; recovering *boundaries* is not recovering *values*. ``value_codec.{h,cc}``: - ``encodeRow(Table&, const uchar* record)`` / ``decodeRow`` — version byte (``0x01``), null bitmap, proto-order fields, varint-length var-width values. All columns present including PK columns. ``keyspace.{h,cc}``: the layout constants and builders from the spec (meta prefixes, ``0x01 || table-id || index-id`` data prefix, ``prefixUpperBound`` for scan ranges). Tests (the point of the task) ----------------------------- Property-style table-driven tests in ``plugin/slatedb/codec/tests/`` runnable in the plain build (no server): for every supported type and for representative multi-part keys with and without NULLs, generate value sets, encode, and assert **byte order equals server comparison order** (compare against ``Field::cmp`` / ``key_cmp`` ground truth), including: INT_MIN/−1/0/1/INT_MAX, ±0.0, ±inf, NaN policy (documented: NaN refused at the Field layer before the codec — verify the server already guarantees this and record where), empty strings, strings differing only in trailing space per collation, embedded NUL bytes, and utf8 multi-byte boundaries under the general and binary collations. Round-trip tests for the value codec over randomized rows. Boundary tests for ``encodedKeyLength``: for every multi-part key in the corpus, concatenating an index-cols encoding with a PK encoding and recovering the split must reproduce both halves exactly, including keys whose text parts contain escaped NUL bytes. ``keyHasNullPart`` is tested per type and per part position, not just "all NULL". Commit boundary --------------- Two commits: key codec + tests; value codec + keyspace + tests. Verification ------------ - Build and codec tests green in the standard builder image (no slatedb-capi needed). - A short ``docs/`` note freezing the format: version byte meanings, the collation-weight compatibility caveat (cross-referenced from the charset codegen work), and the rule that format changes bump the version and never reinterpret ``0x01``.