Performance tracking — implementation plan

Warning

This is not authoritative documentation. It describes a plan of work that is not yet implemented and will change as it lands. Once the work is complete this document should be removed; the finished system is described by the Performance tracking spec and the repos themselves.

Tactical companion to the Performance tracking spec. The spec covers the design and the why; this is the delivery plan — a reviewable patch series, one repo per patch, each with a Zuul story and explicit cross-repo Depends-On.

This plan assumes an AGENTS.md is present in each repo. The prompts below do not restate OpenDev/Gerrit/Zuul mechanics, commit-splitting norms, the no-pkg-config rule, no-gravestone-comments, rootless-podman/bindep conventions, or the vouched/gate/promote pipeline shape — all of that lives in AGENTS.md and the agent is expected to have read it. Prompts carry only what is specific to the patch and not derivable from AGENTS.md plus the repo tree. The first patch to touch a repo that lacks AGENTS.md adds it (from the drafted per-repo files).

Gerrit project paths: drizzle/drizzle-test, drizzle/drizzle, drizzle/perf-tracking.

Reference facts the prompts rely on

Collected here so each prompt can stay terse.

Distro / images

Base is Debian trixie (default), Ubuntu 26.04 (resolute) secondary. Image tags: trixie / resolute / latest. The 12.04–24.04 numbers are historical backfill tags only, never a live distro.

Server image

quay.io/drizzle/drizzle is the running server. It speaks the MySQL wire protocol on the standard port 3306 (the mysql_protocol plugin’s default; the image does not override it) and the native Drizzle protocol on 4427.

Load generator

sysbench (modern akopytov/sysbench, Debian sysbench package — bundles LuaJIT and a MySQL client, no compile, no Perl) reaches Drizzle with --db-driver=mysql --mysql-host=<server> --mysql-port=3306. Its OLTP tests are Lua scripts speaking the MySQL protocol. Drizzle honors ENGINE=InnoDB and MySQL-compatible AUTO_INCREMENT, so the bundled scripts are expected to work unmodified. Escalation if a quirk surfaces: sysbench’s OLTP logic is all Lua (bundled scripts require("oltp_common")), so ship a drop-in drizzle-oltp.lua that shadows the one offending function and point sysbench’s testname at it — no binary fork, no server change. The vendored Perl sql-bench tree is no longer used.

Workload constants (define the baseline — do not change casually)

sysbench oltp_read_write, --threads=1, a fixed --events=N, --table-size=N, --tables=N; single-threaded fixed-event for callgrind determinism. Pick concrete N values when authoring and treat them as the baseline contract.

Metric parsing (port from drizzle/drizzle’s tools/perf-report.pl)

callgrind PROGRAM TOTALS → ir; estimated_cycles = ir + 10·L1 + 100·LLC + 10·branch_mispredict; L1/LLC/branch sums; the Ge column → global_bus_events (present because the callgrind entrypoint passes --collect-bus=yes); massif peak heap; size text → drizzled / plugins_total / plugin_count. The original perf-report.pl predates --collect-bus, so the Ge column is a new field to add to the parser’s PROGRAM TOTALS handling.

Existing prior art

drizzle/drizzle’s .zuul.yaml has the image-build pattern (opendev-build-container-image, container_images:, provides/requires, buildset-registry) and a promote pipeline. drizzle/drizzle-test’s .zuul.yaml has the image-consumer pattern (opendev-buildset-registry-consumer, requires: an image) and reusable playbooks/ + a podman role.

Dependency graph

P1 server images ─┐
P2 sysbench image ─┴─► P3 runner + job ──► P6 gating + upload ──► P7 drizzle triggers
                                                ▲
P5 perf-tracking service ───────────────────────┘
P4 drizzle removes old harness  (Depends-On P3; lands late)

P1, P2 and P5 start independently. P3 waits on P1+P2; P6 on P3+P5; P4 and P7 land last.

Patch 1 — drizzle/drizzle-test: measured-server images + build job

Scope:

Two measured-server images — one running drizzled under callgrind, one under massif — and an unpublished build job. They differ only in their ENTRYPOINT.

Zuul story:

New job drizzle-perf-server-image (opendev-build-container-image, provides: drizzle-perf-callgrind-image and drizzle-perf-massif-image, requires: drizzle-server-container-image, required-projects: drizzle/drizzle), builds both into the buildset registry. No upload/promote jobs — but the build job runs in both vouched and gate (in gate it occupies the slot the upload job would normally hold), so the images exist in-registry for consumers in both pipelines.

Depends-On:

None (builds FROM the published server image).

Prompt

Full plan: see doc/source/specs/implementation-plans/perf-testing.rst in the drizzle/drizzle repo.

Add the measured-server images for the perf harness under a new perf/ directory, without touching the existing pydtr package or its jobs.

callgrind and massif are separate valgrind tools that can’t run in one pass, so the workload runs twice — once under each. Express that as two derivative images built FROM quay.io/drizzle/drizzle, identical except for the ENTRYPOINT, each with valgrind and nothing else (no Perl, driver, or sysbench — this is the process being measured, keep it clean). Baking the tool into the ENTRYPOINT keeps every call site free of valgrind boilerplate.

  • callgrind image: ENTRYPOINT runs valgrind --tool=callgrind --cache-sim=yes --branch-sim=yes --collect-bus=yes --callgrind-out-file=/out/callgrind.out wrapping drizzled. (--collect-bus=yes counts atomic instructions — the Ge global-bus-event — which matters because Drizzle leans on atomics; the collector reads it, below.)

  • massif image: ENTRYPOINT runs valgrind --tool=massif --trace-children=no --massif-out-file=/out/massif.out wrapping drizzled.

Both write to /out, a volume the runner mounts and reads. Pass through args so each stays a drop-in for the stock server. Since the only difference is one line, a shared base stage with two final stages (or two tiny Containerfiles) is fine — whichever reads cleaner. Note the callgrind output is post-processed with callgrind_annotate, so ensure valgrind’s full tool set is present.

Add an unpublished build job drizzle-perf-server-image. Model it on drizzle-build-server-image in drizzle/drizzle’s .zuul.yaml but drop the upload/promote jobs — these images are never published; they only need to exist in the buildset registry for the perf job to consume. It must provides both drizzle-perf-callgrind-image and drizzle-perf-massif-image, requires: drizzle-server-container-image (so they track the change-under-test’s server), and list drizzle/drizzle in required-projects. Put the build job in both vouched and gate (in gate it takes the slot the upload job would normally hold), so the images are available to consumers in both.

Patch 2 — drizzle/drizzle-test: sysbench load-gen image + build job

Scope:

The load-generator image (sysbench from the distribution) and an unpublished build job. Placeholder entrypoint until P3 adds the runner.

Zuul story:

New job drizzle-perf-sysbench-image (opendev-build-container-image, provides: drizzle-perf-sysbench-image, not published). Build job runs in both vouched and gate (gate slot the upload job would normally hold), so the image is available to consumers in both.

Depends-On:

None.

Prompt

Full plan: see doc/source/specs/implementation-plans/perf-testing.rst in the drizzle/drizzle repo.

Add the sysbench load-generator image under perf/, building on patch 1’s layout. This is just a package install — no compiler, no driver, no Perl.

Base it on Debian trixie and apt-get install sysbench (the modern akopytov/sysbench; the package bundles LuaJIT and a MySQL client). sysbench reaches Drizzle over the MySQL protocol with --db-driver=mysql --mysql-host=<server> --mysql-port=3306 — 3306 is the Drizzle server image’s default mysql-protocol port; don’t override it. The workload is the bundled oltp_read_write script.

Drizzle honors ENGINE=InnoDB and MySQL-compatible AUTO_INCREMENT, so the bundled scripts are expected to run unmodified — try them as-is first. If oltp_common’s prepare trips on a Drizzle idiosyncrasy, do not fork sysbench or change the server: ship a perf/drizzle-oltp.lua that requires oltp_common and shadows only the offending function (e.g. create_table), and point sysbench’s testname at it. As a build-time smoke, run sysbench oltp_read_write ... --tables=1 --table-size=1000 prepare then a short --events=100 --threads=1 run against a quay.io/drizzle/drizzle container on 3306, and record whether the stock scripts worked or a Lua shim was needed.

A placeholder entrypoint is fine; the runner becomes the real one in patch 3. Add an unpublished build job drizzle-perf-sysbench-image mirroring patch 1’s server image job (provides: drizzle-perf-sysbench-image); put it in both vouched and gate (gate takes the upload slot), so the image is available to consumers in both.

Patch 3 — drizzle/drizzle-test: runner + compose + perf job (artifact only)

Scope:

The runner script (single script, not a package) with unit tests; two compose files (callgrind and massif) wiring the measured-server images to the load generator; the sysbench image entrypoint set to the runner; the perf job, reporting only.

Zuul story:

New drizzle-perf (consumes the three perf images — callgrind, massif, sysbench — publishes metrics.json as an artifact, non-voting) + drizzle-perf-unit (ruff + unit tests, no containers, voting). Vouched+gate.

Depends-On:

Patches 1 and 2.

Prompt

Full plan: see doc/source/specs/implementation-plans/perf-testing.rst in the drizzle/drizzle repo.

Under perf/, write the runner that drives the perf images and emits perf metrics as a Zuul artifact. No DB and no gating yet — this milestone just makes every change show its numbers.

The runner is a single script, deliberately not a package (no discovery or fixtures to justify scaffolding; see the spec). It’s Python because patch 6 adds an HTTP history fetch and a stddev verdict that want a real language and unit tests — keep it that simple for now. Port the metric parsing and the fixed workload from the reference facts (originals: drizzle/drizzle’s tools/perf-report.pl and tools/perf.sh — this patch is their replacement).

The workload runs twice because callgrind and massif are separate passes. Write two compose files — perf/compose.callgrind.yaml and perf/compose.massif.yaml — each with a server service (the patch-1 callgrind or massif image respectively; its entrypoint already wraps drizzled in the right tool writing to /out, so no command override — just mount /out) and a sysbench service (the patch-2 image, depends_on server, runs the oltp_read_write workload over port 3306). The two files differ only in the server image. The runner does two podman-compose up cycles, reads callgrind.out then massif.out off /out between them, runs size, writes metrics.json, and exits 0 regardless (emit JSON even on a failed run). metrics.json carries the run metadata the service’s schema needs alongside the metrics — at minimum arch, distro, branch and commit_sha (read from the environment / Zuul vars and the measured image) — because the gating and upload in patch 6 key the trailing window on (workload, branch, arch, distro, metric_key); capturing them now means patch 6 adds no new collection. callgrind output is post-processed with callgrind_annotate. Set the sysbench image’s entrypoint to the runner. Heads-up: rootless-podman volume permissions under the pasta/netavark setup the podman role configures are the most likely first-run snag on the shared /out volume.

Unit-test the parsing under perf/tests/. Add two jobs: drizzle-perf (opendev-buildset-registry-consumer, requires the three perf images, runs a playbooks/perf.yaml that calls the runner and publishes metrics.json as an artifact, non-voting) and drizzle-perf-unit (ruff + the unit tests, no containers, voting). Both vouched+gate. Footers: Depends-On patches 1 and 2.

Patch 4 — drizzle/drizzle: remove the stranded perf harness

Scope:

Delete tools/perf.sh, tools/perf-report.pl, the [perf] bindep profile, and future-zuul.d/. Keep perf/*.json; rewrite perf/README.rst for the new flow.

Zuul story:

No new jobs; existing jobs stay green (pure removal).

Depends-On:

Patch 3 — the runner that replaces perf.sh/perf-report.pl must land first. This is why removal follows the runner rather than preceding it.

Prompt

Full plan: see doc/source/specs/implementation-plans/perf-testing.rst in the drizzle/drizzle repo.

Remove the old perf harness now that it’s reimplemented in drizzle/drizzle-test (patch 3 is its replacement).

Delete tools/perf.sh, tools/perf-report.pl, and the whole future-zuul.d/ directory (vestigial intent-only config; .zuul.yaml is authoritative). Remove the three [perf] profile lines (valgrind, libdbi-perl, zlib1g-dev) and their comment from bindep.txt, leaving other profiles untouched — the server repo shouldn’t carry test/perf deps once the harness lives elsewhere.

Keep every perf/*.json (the committed baseline + per-release series is still the long-horizon reference). Rewrite perf/README.rst: it currently describes a perf Containerfile stage and a CPAN ADD that no longer exist. Replace that with the real flow — performance is measured in drizzle/drizzle-test by running drizzled under valgrind (callgrind for instructions, massif for heap; two minimal server images) driven by a sysbench load generator over the MySQL protocol, orchestrated by a runner script — and perf/baseline.json + perf/<release>.json remain the committed series. Pure removal/docs; don’t change image-build behavior beyond dropping the unused deps. Footer: Depends-On: <Gerrit URL of patch 3>.

Patch 5 — drizzle/perf-tracking: results service + dogfood integration job

Scope:

The whole service in the currently-bare repo: Flask + flask-restx + SQLAlchemy over Drizzle, the two-table schema and v1 API, two-token auth, a container image, a collection-host compose.yaml. Adds AGENTS.md/CLAUDE.md.

Zuul story:

New voting perf-service-integration (service against a real Drizzle container) + a ruff job, replacing the noop .zuul.yaml. Vouched+gate.

Depends-On:

None.

Prompt

Full plan: see doc/source/specs/implementation-plans/perf-testing.rst in the drizzle/drizzle repo.

Build the results service that stores perf runs. It dogfoods — Drizzle is its own store. The repo is currently just .gitreview + a noop .zuul.yaml; add an AGENTS.md (Python service variant; mirror the drizzle-test agent file’s structure and norms, retargeted to a Flask service) and a CLAUDE.md pointing at it.

A uv-managed Flask + flask-restx + SQLAlchemy service, synchronous WSGI (no asyncio), talking to Drizzle via a MySQL/MariaDB driver. Implement the generic schema and v1 API from the spec (runs + metrics tables; POST /v1/runs with a write token, idempotent on the full series plus commit and tag — (commit_sha, workload, branch, arch, distro, tag) — enforced as a DB unique constraint; GET /v1/history filtered by the full series key — workload, branch, arch, distro, metric_key — with a limit, read token; GET /v1/baseline; GET /v1/runs/{id}). Note the spec’s metric key uses a plain dim_key string column ('' for the scalar v1 metrics), not a JSON column, and tag is a non-null column defaulted to '' so it can sit in that unique constraint — follow both exactly so the schema is portable on Drizzle. Two bearer tokens via config/env — a broad read token and a separately-scoped write token. The service makes no regression decisions; it stores and returns.

Ship a Containerfile and a compose.yaml wiring {perf-service, drizzle} for the collection host. Replace the noop .zuul.yaml with a voting perf-service-integration job (stands the service + a quay.io/drizzle/drizzle container via compose; a pytest suite POSTs a synthetic run, GETs history, asserts the round-trip and that the read token cannot write) plus a ruff job, in vouched+gate. Model job structure on drizzle/drizzle-test’s .zuul.yaml.

Patch 6 — drizzle/drizzle-test: close the loop (gating + upload)

Scope:

Extend the runner with the history fetch + stddev verdict and non-zero exit on regression; add the promote-pipeline upload job (write token, post-merge only) and a one-shot backfill of perf/*.json.

Zuul story:

drizzle-perf goes voting in gate; new drizzle-perf-promote in the promote pipeline uploads with the write token; backfill is one-shot.

Depends-On:

Patch 5.

Prompt

Full plan: see doc/source/specs/implementation-plans/perf-testing.rst in the drizzle/drizzle repo.

Extend the patch-3 runner from “report only” to “gate + record,” now that the service (patch 5) exists.

Add a history client that GETs /v1/history (read token) for the candidate’s own series — (workload=perf, branch, arch, distro, metric_key) — and computes mean and stddev over last-5 and last-20 windows. The arch and distro filters are not optional: callgrind counts differ by ISA, so an amd64 run must never be compared against arm64 history, and a distro change is a rebaseline (fresh series), not a comparison across the boundary. Fail the run (non-zero exit) if a metric is worse than the mean by more than max(k·σ, floor), with k=3 and floors 3% for callgrind.ir/callgrind.estimated_cycles, 5% for callgrind.global_bus_events (atomic counts swing more with thread scheduling), 1% for massif.peak_heap_bytes, 0% for size.*; “worse” is higher for all v1 metrics. Always print the full two-sided delta table (improvements included); only regressions fail. Put k and the floors in a committed, reviewable config (e.g. perf/thresholds.toml). Unit-test the verdict with synthetic windows.

Add drizzle-perf-promote in the promote pipeline only: download the gate build’s metrics.json artifact and POST it to /v1/runs with the write token (a Zuul secret scoped to promote only), pipeline='promote', commit_sha=<merge sha>. The vouched/gate drizzle-perf job must never hold the write token. Flip drizzle-perf to voting in gate. Finally, add a one-shot backfill task that reads drizzle/drizzle’s perf/*.json and POSTs each as a tagged run (tag='baseline', tag='14.04' … tag='24.04', change/patchset null). These historical files predate the arch/distro split, so set their arch/distro to whatever they actually were (the legacy amd64 series) explicitly — they form their own tagged series and won’t be compared against live trixie/resolute runs. Footer: Depends-On: <Gerrit URL of patch 5>.

Patch 7 — drizzle/drizzle: trigger perf on server changes

Scope:

List the drizzle-perf job (defined in drizzle-test) in the server repo’s vouched+gate, with the docs irrelevant-files skip.

Zuul story:

No new job definitions — references the existing job cross-repo.

Depends-On:

Patch 6 (job should be gating-capable first).

Prompt

Full plan: see doc/source/specs/implementation-plans/perf-testing.rst in the drizzle/drizzle repo.

Make a server change run the perf job. In .zuul.yaml, add the drizzle-perf job — defined in drizzle/drizzle-test — to the vouched and gate pipelines, referencing it cross-repo the same way drizzle-test-dtr is referenced (it requires the three perf images and the server image). Apply the same docs irrelevant-files list that drizzle-build-server-image uses, so docs-only changes skip perf. Nothing else changes. Footer: Depends-On: <Gerrit URL of patch 6>.

Open items (none blocking)

  • Confirm sysbench’s oltp_common prepare + oltp_read_write run cleanly against Drizzle on port 3306 (surfaces in patch 2’s build-time smoke). Stock scripts are expected to work since Drizzle honors ENGINE=InnoDB + AUTO_INCREMENT; if a quirk appears, the fix is a drop-in drizzle-oltp.lua shadowing the offending function — no server change. Record which was needed.

  • quay.io/drizzle/drizzle-ccache-data availability — orthogonal, but the server image the measured image builds on inherits that build chain.

  • Rootless-podman shared-volume permissions for the callgrind /out mount (surfaces in patch 3; worth a local run before relying on the gate).