Performance tracking — implementation plan¶
Warning
This is not authoritative documentation. It describes a plan of work that is not yet implemented and will change as it lands. Once the work is complete this document should be removed; the finished system is described by the Performance tracking spec and the repos themselves.
Tactical companion to the Performance tracking spec. The spec covers
the design and the why; this is the delivery plan — a reviewable patch
series, one repo per patch, each with a Zuul story and explicit
cross-repo Depends-On.
This plan assumes an AGENTS.md is present in each repo. The
prompts below do not restate OpenDev/Gerrit/Zuul mechanics,
commit-splitting norms, the no-pkg-config rule, no-gravestone-comments,
rootless-podman/bindep conventions, or the vouched/gate/promote pipeline
shape — all of that lives in AGENTS.md and the agent is expected to
have read it. Prompts carry only what is specific to the patch and not
derivable from AGENTS.md plus the repo tree. The first patch to
touch a repo that lacks AGENTS.md adds it (from the drafted per-repo
files).
Gerrit project paths: drizzle/drizzle-test, drizzle/drizzle,
drizzle/perf-tracking.
Reference facts the prompts rely on¶
Collected here so each prompt can stay terse.
- Distro / images
Base is Debian
trixie(default), Ubuntu26.04(resolute) secondary. Image tags:trixie/resolute/latest. The12.04–24.04numbers are historical backfill tags only, never a live distro.- Server image
quay.io/drizzle/drizzleis the running server. It speaks the MySQL wire protocol on the standard port 3306 (the mysql_protocol plugin’s default; the image does not override it) and the native Drizzle protocol on 4427.- Load generator
sysbench (modern akopytov/sysbench, Debian
sysbenchpackage — bundles LuaJIT and a MySQL client, no compile, no Perl) reaches Drizzle with--db-driver=mysql --mysql-host=<server> --mysql-port=3306. Its OLTP tests are Lua scripts speaking the MySQL protocol. Drizzle honorsENGINE=InnoDBand MySQL-compatibleAUTO_INCREMENT, so the bundled scripts are expected to work unmodified. Escalation if a quirk surfaces: sysbench’s OLTP logic is all Lua (bundled scriptsrequire("oltp_common")), so ship a drop-indrizzle-oltp.luathat shadows the one offending function and point sysbench’s testname at it — no binary fork, no server change. The vendored Perl sql-bench tree is no longer used.- Workload constants (define the baseline — do not change casually)
sysbench
oltp_read_write,--threads=1, a fixed--events=N,--table-size=N,--tables=N; single-threaded fixed-event for callgrind determinism. Pick concrete N values when authoring and treat them as the baseline contract.- Metric parsing (port from
drizzle/drizzle’stools/perf-report.pl) callgrind
PROGRAM TOTALS→ir;estimated_cycles = ir + 10·L1 + 100·LLC + 10·branch_mispredict; L1/LLC/branch sums; theGecolumn →global_bus_events(present because the callgrind entrypoint passes--collect-bus=yes); massif peak heap;sizetext → drizzled / plugins_total / plugin_count. The originalperf-report.plpredates--collect-bus, so theGecolumn is a new field to add to the parser’sPROGRAM TOTALShandling.- Existing prior art
drizzle/drizzle’s.zuul.yamlhas the image-build pattern (opendev-build-container-image,container_images:,provides/requires, buildset-registry) and apromotepipeline.drizzle/drizzle-test’s.zuul.yamlhas the image-consumer pattern (opendev-buildset-registry-consumer,requires:an image) and reusableplaybooks/+ apodmanrole.
Dependency graph¶
P1 server images ─┐
P2 sysbench image ─┴─► P3 runner + job ──► P6 gating + upload ──► P7 drizzle triggers
▲
P5 perf-tracking service ───────────────────────┘
P4 drizzle removes old harness (Depends-On P3; lands late)
P1, P2 and P5 start independently. P3 waits on P1+P2; P6 on P3+P5; P4 and P7 land last.
Patch 1 — drizzle/drizzle-test: measured-server images + build job¶
- Scope:
Two measured-server images — one running drizzled under callgrind, one under massif — and an unpublished build job. They differ only in their ENTRYPOINT.
- Zuul story:
New job
drizzle-perf-server-image(opendev-build-container-image,provides: drizzle-perf-callgrind-imageanddrizzle-perf-massif-image,requires: drizzle-server-container-image,required-projects: drizzle/drizzle), builds both into the buildset registry. No upload/promote jobs — but the build job runs in both vouched and gate (in gate it occupies the slot the upload job would normally hold), so the images exist in-registry for consumers in both pipelines.- Depends-On:
None (builds
FROMthe published server image).
Prompt
Full plan: see doc/source/specs/implementation-plans/perf-testing.rst in the drizzle/drizzle repo.
Add the measured-server images for the perf harness under a new
perf/ directory, without touching the existing pydtr package
or its jobs.
callgrind and massif are separate valgrind tools that can’t run in
one pass, so the workload runs twice — once under each. Express that
as two derivative images built FROM quay.io/drizzle/drizzle,
identical except for the ENTRYPOINT, each with valgrind and
nothing else (no Perl, driver, or sysbench — this is the process
being measured, keep it clean). Baking the tool into the ENTRYPOINT
keeps every call site free of valgrind boilerplate.
callgrind image: ENTRYPOINT runs
valgrind --tool=callgrind --cache-sim=yes --branch-sim=yes --collect-bus=yes --callgrind-out-file=/out/callgrind.outwrapping drizzled. (--collect-bus=yescounts atomic instructions — theGeglobal-bus-event — which matters because Drizzle leans on atomics; the collector reads it, below.)massif image: ENTRYPOINT runs
valgrind --tool=massif --trace-children=no --massif-out-file=/out/massif.outwrapping drizzled.
Both write to /out, a volume the runner mounts and reads. Pass
through args so each stays a drop-in for the stock server. Since the
only difference is one line, a shared base stage with two final
stages (or two tiny Containerfiles) is fine — whichever reads
cleaner. Note the callgrind output is post-processed with
callgrind_annotate, so ensure valgrind’s full tool set is present.
Add an unpublished build job drizzle-perf-server-image. Model it
on drizzle-build-server-image in drizzle/drizzle’s
.zuul.yaml but drop the upload/promote jobs — these images are
never published; they only need to exist in the buildset registry for
the perf job to consume. It must provides both
drizzle-perf-callgrind-image and drizzle-perf-massif-image,
requires: drizzle-server-container-image (so they track the
change-under-test’s server), and list drizzle/drizzle in
required-projects. Put the build job in both vouched and
gate (in gate it takes the slot the upload job would normally hold),
so the images are available to consumers in both.
Patch 2 — drizzle/drizzle-test: sysbench load-gen image + build job¶
- Scope:
The load-generator image (sysbench from the distribution) and an unpublished build job. Placeholder entrypoint until P3 adds the runner.
- Zuul story:
New job
drizzle-perf-sysbench-image(opendev-build-container-image,provides: drizzle-perf-sysbench-image, not published). Build job runs in both vouched and gate (gate slot the upload job would normally hold), so the image is available to consumers in both.- Depends-On:
None.
Prompt
Full plan: see doc/source/specs/implementation-plans/perf-testing.rst in the drizzle/drizzle repo.
Add the sysbench load-generator image under perf/, building on
patch 1’s layout. This is just a package install — no compiler, no
driver, no Perl.
Base it on Debian trixie and apt-get install sysbench (the modern
akopytov/sysbench; the package bundles LuaJIT and a MySQL client).
sysbench reaches Drizzle over the MySQL protocol with
--db-driver=mysql --mysql-host=<server> --mysql-port=3306 — 3306
is the Drizzle server image’s default mysql-protocol port; don’t
override it. The workload is the bundled oltp_read_write script.
Drizzle honors ENGINE=InnoDB and MySQL-compatible
AUTO_INCREMENT, so the bundled scripts are expected to run
unmodified — try them as-is first. If oltp_common’s prepare
trips on a Drizzle idiosyncrasy, do not fork sysbench or change
the server: ship a perf/drizzle-oltp.lua that requires
oltp_common and shadows only the offending function (e.g.
create_table), and point sysbench’s testname at it. As a
build-time smoke, run sysbench oltp_read_write ... --tables=1
--table-size=1000 prepare then a short --events=100 --threads=1
run against a quay.io/drizzle/drizzle container on 3306, and
record whether the stock scripts worked or a Lua shim was needed.
A placeholder entrypoint is fine; the runner becomes the real one in
patch 3. Add an unpublished build job drizzle-perf-sysbench-image
mirroring patch 1’s server image job (provides:
drizzle-perf-sysbench-image); put it in both vouched and gate
(gate takes the upload slot), so the image is available to consumers
in both.
Patch 3 — drizzle/drizzle-test: runner + compose + perf job (artifact only)¶
- Scope:
The runner script (single script, not a package) with unit tests; two compose files (callgrind and massif) wiring the measured-server images to the load generator; the sysbench image entrypoint set to the runner; the perf job, reporting only.
- Zuul story:
New
drizzle-perf(consumes the three perf images — callgrind, massif, sysbench — publishesmetrics.jsonas an artifact, non-voting) +drizzle-perf-unit(ruff + unit tests, no containers, voting). Vouched+gate.- Depends-On:
Patches 1 and 2.
Prompt
Full plan: see doc/source/specs/implementation-plans/perf-testing.rst in the drizzle/drizzle repo.
Under perf/, write the runner that drives the perf images and
emits perf metrics as a Zuul artifact. No DB and no gating yet —
this milestone just makes every change show its numbers.
The runner is a single script, deliberately not a package (no
discovery or fixtures to justify scaffolding; see the spec). It’s
Python because patch 6 adds an HTTP history fetch and a stddev verdict
that want a real language and unit tests — keep it that simple for
now. Port the metric parsing and the fixed workload from the
reference facts (originals: drizzle/drizzle’s
tools/perf-report.pl and tools/perf.sh — this patch is their
replacement).
The workload runs twice because callgrind and massif are separate
passes. Write two compose files — perf/compose.callgrind.yaml and
perf/compose.massif.yaml — each with a server service (the
patch-1 callgrind or massif image respectively; its entrypoint
already wraps drizzled in the right tool writing to /out, so no
command override — just mount /out) and a sysbench service
(the patch-2 image, depends_on server, runs the oltp_read_write
workload over port 3306). The two files differ only in the server
image. The runner does two podman-compose up cycles, reads
callgrind.out then massif.out off /out between them, runs
size, writes metrics.json, and exits 0 regardless (emit JSON
even on a failed run). metrics.json carries the run metadata the
service’s schema needs alongside the metrics — at minimum arch,
distro, branch and commit_sha (read from the environment /
Zuul vars and the measured image) — because the gating and upload in
patch 6 key the trailing window on (workload, branch, arch, distro,
metric_key); capturing them now means patch 6 adds no new
collection. callgrind output is post-processed with
callgrind_annotate. Set the sysbench image’s entrypoint to the
runner. Heads-up: rootless-podman volume permissions under the
pasta/netavark setup the podman role configures are the most
likely first-run snag on the shared /out volume.
Unit-test the parsing under perf/tests/. Add two jobs:
drizzle-perf (opendev-buildset-registry-consumer, requires
the three perf images, runs a playbooks/perf.yaml that calls the
runner and publishes metrics.json as an artifact, non-voting)
and drizzle-perf-unit (ruff + the unit tests, no containers,
voting). Both vouched+gate. Footers: Depends-On patches 1 and 2.
Patch 4 — drizzle/drizzle: remove the stranded perf harness¶
- Scope:
Delete
tools/perf.sh,tools/perf-report.pl, the[perf]bindep profile, andfuture-zuul.d/. Keepperf/*.json; rewriteperf/README.rstfor the new flow.- Zuul story:
No new jobs; existing jobs stay green (pure removal).
- Depends-On:
Patch 3 — the runner that replaces
perf.sh/perf-report.plmust land first. This is why removal follows the runner rather than preceding it.
Prompt
Full plan: see doc/source/specs/implementation-plans/perf-testing.rst in the drizzle/drizzle repo.
Remove the old perf harness now that it’s reimplemented in
drizzle/drizzle-test (patch 3 is its replacement).
Delete tools/perf.sh, tools/perf-report.pl, and the whole
future-zuul.d/ directory (vestigial intent-only config;
.zuul.yaml is authoritative). Remove the three [perf] profile
lines (valgrind, libdbi-perl, zlib1g-dev) and their
comment from bindep.txt, leaving other profiles untouched — the
server repo shouldn’t carry test/perf deps once the harness lives
elsewhere.
Keep every perf/*.json (the committed baseline + per-release
series is still the long-horizon reference). Rewrite
perf/README.rst: it currently describes a perf Containerfile
stage and a CPAN ADD that no longer exist. Replace that with the
real flow — performance is measured in drizzle/drizzle-test by
running drizzled under valgrind (callgrind for instructions, massif
for heap; two minimal server images) driven by a sysbench load
generator over the MySQL protocol, orchestrated by a runner script —
and perf/baseline.json + perf/<release>.json remain the
committed series. Pure removal/docs; don’t change image-build
behavior beyond dropping the unused deps. Footer: Depends-On:
<Gerrit URL of patch 3>.
Patch 5 — drizzle/perf-tracking: results service + dogfood integration job¶
- Scope:
The whole service in the currently-bare repo: Flask + flask-restx + SQLAlchemy over Drizzle, the two-table schema and v1 API, two-token auth, a container image, a collection-host
compose.yaml. AddsAGENTS.md/CLAUDE.md.- Zuul story:
New voting
perf-service-integration(service against a real Drizzle container) + a ruff job, replacing the noop.zuul.yaml. Vouched+gate.- Depends-On:
None.
Prompt
Full plan: see doc/source/specs/implementation-plans/perf-testing.rst in the drizzle/drizzle repo.
Build the results service that stores perf runs. It dogfoods —
Drizzle is its own store. The repo is currently just .gitreview +
a noop .zuul.yaml; add an AGENTS.md (Python service variant;
mirror the drizzle-test agent file’s structure and norms, retargeted
to a Flask service) and a CLAUDE.md pointing at it.
A uv-managed Flask + flask-restx + SQLAlchemy service, synchronous
WSGI (no asyncio), talking to Drizzle via a MySQL/MariaDB driver.
Implement the generic schema and v1 API from the spec (runs + metrics
tables; POST /v1/runs with a write token, idempotent on the full
series plus commit and tag — (commit_sha, workload, branch, arch,
distro, tag) — enforced as a DB unique constraint; GET /v1/history
filtered by the full series key — workload, branch, arch, distro,
metric_key — with a limit, read token; GET /v1/baseline;
GET /v1/runs/{id}). Note the spec’s metric key uses a plain
dim_key string column ('' for the scalar v1 metrics), not a
JSON column, and tag is a non-null column defaulted to '' so
it can sit in that unique constraint — follow both exactly so the
schema is portable on Drizzle. Two bearer tokens via config/env — a
broad read token and a separately-scoped write token. The service
makes no regression decisions; it stores and returns.
Ship a Containerfile and a compose.yaml wiring {perf-service,
drizzle} for the collection host. Replace the noop .zuul.yaml
with a voting perf-service-integration job (stands the service +
a quay.io/drizzle/drizzle container via compose; a pytest suite
POSTs a synthetic run, GETs history, asserts the round-trip and that
the read token cannot write) plus a ruff job, in vouched+gate. Model
job structure on drizzle/drizzle-test’s .zuul.yaml.
Patch 6 — drizzle/drizzle-test: close the loop (gating + upload)¶
- Scope:
Extend the runner with the history fetch + stddev verdict and non-zero exit on regression; add the promote-pipeline upload job (write token, post-merge only) and a one-shot backfill of
perf/*.json.- Zuul story:
drizzle-perfgoes voting in gate; newdrizzle-perf-promotein the promote pipeline uploads with the write token; backfill is one-shot.- Depends-On:
Patch 5.
Prompt
Full plan: see doc/source/specs/implementation-plans/perf-testing.rst in the drizzle/drizzle repo.
Extend the patch-3 runner from “report only” to “gate + record,” now that the service (patch 5) exists.
Add a history client that GETs /v1/history (read token) for
the candidate’s own series — (workload=perf, branch, arch, distro,
metric_key) — and computes mean and stddev over last-5 and last-20
windows. The arch and distro filters are not optional:
callgrind counts differ by ISA, so an amd64 run must never be compared
against arm64 history, and a distro change is a rebaseline (fresh
series), not a comparison across the boundary. Fail the run (non-zero
exit) if a metric is worse than the mean by more than max(k·σ,
floor), with k=3 and floors 3% for
callgrind.ir/callgrind.estimated_cycles, 5% for
callgrind.global_bus_events (atomic counts swing more with thread
scheduling), 1% for massif.peak_heap_bytes, 0% for size.*;
“worse” is higher for all v1 metrics. Always print the full two-sided
delta table (improvements included); only regressions fail. Put
k and the floors in a committed, reviewable config (e.g.
perf/thresholds.toml). Unit-test the verdict with synthetic
windows.
Add drizzle-perf-promote in the promote pipeline only:
download the gate build’s metrics.json artifact and POST it to
/v1/runs with the write token (a Zuul secret scoped to promote
only), pipeline='promote', commit_sha=<merge sha>. The
vouched/gate drizzle-perf job must never hold the write token.
Flip drizzle-perf to voting in gate. Finally, add a one-shot
backfill task that reads drizzle/drizzle’s perf/*.json and
POSTs each as a tagged run (tag='baseline', tag='14.04' …
tag='24.04', change/patchset null). These historical files
predate the arch/distro split, so set their arch/distro to
whatever they actually were (the legacy amd64 series) explicitly —
they form their own tagged series and won’t be compared against live
trixie/resolute runs. Footer: Depends-On: <Gerrit URL of
patch 5>.
Patch 7 — drizzle/drizzle: trigger perf on server changes¶
- Scope:
List the
drizzle-perfjob (defined indrizzle-test) in the server repo’s vouched+gate, with the docsirrelevant-filesskip.- Zuul story:
No new job definitions — references the existing job cross-repo.
- Depends-On:
Patch 6 (job should be gating-capable first).
Prompt
Full plan: see doc/source/specs/implementation-plans/perf-testing.rst in the drizzle/drizzle repo.
Make a server change run the perf job. In .zuul.yaml, add the
drizzle-perf job — defined in drizzle/drizzle-test — to the
vouched and gate pipelines, referencing it cross-repo the same way
drizzle-test-dtr is referenced (it requires the three perf images
and the server image). Apply the same docs irrelevant-files list
that drizzle-build-server-image uses, so docs-only changes skip
perf. Nothing else changes. Footer: Depends-On: <Gerrit URL of
patch 6>.
Open items (none blocking)¶
Confirm sysbench’s
oltp_commonprepare +oltp_read_writerun cleanly against Drizzle on port 3306 (surfaces in patch 2’s build-time smoke). Stock scripts are expected to work since Drizzle honorsENGINE=InnoDB+AUTO_INCREMENT; if a quirk appears, the fix is a drop-indrizzle-oltp.luashadowing the offending function — no server change. Record which was needed.quay.io/drizzle/drizzle-ccache-dataavailability — orthogonal, but the server image the measured image builds on inherits that build chain.Rootless-podman shared-volume permissions for the callgrind
/outmount (surfaces in patch 3; worth a local run before relying on the gate).