Performance tracking

Warning

This is not authoritative documentation. It describes a plan that is not yet implemented. Details may change significantly before completion, or may never be fully realized.

The following specification describes a plan to measure Drizzle’s performance on every change, store the results so they accumulate into a history, and gate changes that regress. It covers the measurement technique, the container images that produce the numbers, the service that stores them, and how regression decisions are made.

Introduction

Drizzle had a performance story once: an old drizzle-automation tree drove sql-bench, drizzleslap and sysbench against sandboxed servers and recorded the results in a database. That machinery is Python 2, BZR-era, and tied to MySQL-Sandbox; the orchestration is gone. What survives, and what is worth keeping, is the domain knowledge — which workloads matter, how their metrics are parsed, and the trailing-window method used to decide whether a run regressed.

In its place the revival left a smaller, sharper seed: a shell script (tools/perf.sh) and a Perl parser (tools/perf-report.pl) in the server repo that run a fixed sql-bench workload against drizzled under valgrind callgrind, count instructions rather than seconds, and diff the result against a committed baseline. That script is the right idea at the wrong scale. This spec generalizes it into a proper CI story: real Zuul jobs, a results service, and regression gating — while keeping the one insight that makes performance measurable in a noisy cloud at all.

The core problem and the core decision

CI runs on shared cloud hardware. Wall-clock throughput and latency on such hardware are dominated by noisy neighbours, thermal behaviour, and contention for cores and memory bandwidth — the variance swamps the signal we care about, which is whether our code got slower. No amount of eBPF or hardware perf-counter cleverness fixes this: those read the real CPU, so their numbers are host-specific by construction and reintroduce exactly the cross-node noise we are trying to escape.

There are only two ways to get a stable number: remove the real CPU from the measurement, or own dedicated bare metal. We choose the first. Valgrind’s callgrind runs the workload on a simulated CPU and counts instructions; those counts are reproducible to a tiny fraction of a percent and identical on any host. We therefore do not measure time at all. We run the workload under callgrind and measure instruction counts, simulated cache behaviour, and peak heap.

This has a counterintuitive but liberating consequence: the metric changes, not the workload type. sysbench drives the server with a real OLTP workload, but we read the server’s instruction count, not its transactions per second. The same sysbench workload extends to a future wall-clock tier — keep the load generator, change only what we measure and where we run it.

Trustworthy wall-clock numbers remain valuable, but they require dedicated, consistent hardware. The design leaves room for a future “real hardware” tier — a separate node type and a metric family flagged as wall-clock that only gates when run there — but until such machines exist, every cloud performance job is a callgrind job.

What we measure

For each run we record, all from the server process:

  • callgrind.ir — instructions retired. The headline metric.

  • callgrind.estimated_cycles — cachegrind’s documented cycle estimate, ir + 10·L1 + 100·LLC + 10·branch_mispredict, folding simulated cache and branch behaviour into one number.

  • callgrind.l1_misses, callgrind.llc_misses, callgrind.branch_mispredicts — the components.

  • callgrind.global_bus_events — the count of atomic operations (callgrind’s Ge event, enabled with --collect-bus=yes): lock-prefixed instructions on x86_64, store-conditional instructions on aarch64. Drizzle leans on atomics for cross-core coordination, so this is a direct measure of synchronization cost — a change that adds lock contention or atomic churn moves it even when ir barely shifts. The same event counts the equivalent primitive on both target architectures.

  • massif.peak_heap_bytes — peak heap from a separate massif pass.

  • size.drizzled_bytes, size.plugins_total_bytes, size.plugin_count — binary sizes from size(1).

The workload is a bundled sysbench OLTP script (oltp_read_write) run single-threaded with a fixed event count (--threads=1 --events=N), chosen so the run is deterministic under callgrind’s ~50x slowdown. Single-threaded fixed-event is deliberate: callgrind serializes execution anyway, and a fixed event count gives an identical instruction sequence every run. sysbench’s concurrency and timed modes are not used here — they belong to the wall-clock tier. The parameters (script, thread count, event count, table size and count) are part of the contract: changing any of them invalidates the committed baseline and must be done deliberately, as its own change.

A note on noise. Even instruction counts carry a small run-to-run floor — drizzled is multithreaded and callgrind sums every thread, so background-thread work varies between runs by a couple of percent. size is bit-exact and peak heap is near-stable. These properties set the regression thresholds below.

Architecture

Measurement is split between a measured server and a load generator, plus a results service. The split exists because valgrind must wrap the server — the database’s instruction count and heap are the metrics — while the load generator is merely the source of work and must not be counted. The measured server ships as two near-identical images (callgrind and massif; see below), but conceptually it is one role facing one load generator.

The measured-server images

Two minimal images built FROM the published Drizzle server image plus valgrind, and nothing else. Deliberately minimal: no compiler, no Perl, no sql-bench, because anything installed alongside the measured process is a distraction at best. callgrind and massif are separate valgrind tools and cannot share a pass, so the workload runs twice — once for instruction counts, once for peak heap. Rather than smuggle a tool-selection flag through a single image, we build two images that differ only in their entrypoint: one wraps drizzled in callgrind (with cache, branch, and global-bus-event collection enabled, the last for atomic-instruction counts), the other in massif. Each entrypoint is unconditional, so call sites carry no valgrind boilerplate and the compose wiring needs no command overrides.

These images are built by a Zuul job in the test repository, not by a stage in the server repository’s container build — the server image stays a pure server image. The build job tracks the change under test: it requires the server container image from the buildset registry, so performance is always measured against the Drizzle of the change being tested. They are built on the node and never published; nothing downstream consumes them.

The load-generator image

A plain image with sysbench installed from the distribution — nothing else of substance. No compiler, no Perl, no driver to build: the sysbench package bundles its own LuaJIT and a MySQL client, so the image is essentially apt-get install sysbench. It also carries the runner (see below) and is its entrypoint.

sysbench (the modern akopytov/sysbench) is the contemporary standard for database benchmarking, and its choice modernizes the workload. Its OLTP tests are Lua scripts driving an abstract connection; with --db-driver=mysql they speak the MySQL wire protocol directly, so sysbench reaches Drizzle over the server’s standard MySQL-protocol port with no Perl, no DBI, and no native driver. The workload is one of the bundled OLTP scripts (e.g. oltp_read_write), which exercises point selects, range scans, aggregates, inserts, updates and deletes — far more representative of real database work than the old sql-bench test-insert.

Compatibility with Drizzle is expected to need little or nothing. sysbench’s prepare step issues ordinary SQL — CREATE TABLE ... ENGINE=InnoDB, secondary-index creation, inserts — and Drizzle honors ENGINE=InnoDB and MySQL-compatible AUTO_INCREMENT. If a quirk does surface, the escalation ladder stays out of the server: sysbench’s OLTP logic is entirely in Lua (the bundled scripts simply require("oltp_common")), so a drop-in drizzle-oltp.lua in the image can shadow the one offending function — the same technique other forks use to adapt sysbench — with no binary fork and no server change. Charset is pinned via the MySQL client option the same way, since Drizzle is UTF-8 natively.

This route also unifies the two measurement tiers under one tool. sysbench’s real strength is concurrent, timed throughput — which is exactly the future wall-clock tier (see below), where it runs with real --threads and --time on dedicated hardware. Adopting it now for the callgrind tier means the same tool, schema and workload vocabulary carry forward to that tier; only the metric and the node type change.

The runner

A single script — not a package. It has no test discovery, fixtures or scenarios to justify scaffolding; it runs the workload against the measured server (twice — once under callgrind for instructions, once under massif for heap), parses the server’s output plus binary sizes into a metrics document, and (once gating is enabled) compares against recent history and exits non-zero on regression. It is written in Python rather than the original shell-plus-Perl because the comparison step wants a history fetch over HTTP and a standard-deviation verdict — logic that benefits from a real language and unit tests — but it stays a script.

The wiring is two compose files, one per pass, each defining the relevant measured-server image and the load generator that depends on it. Because the tool is baked into each server image’s entrypoint, neither file needs a command override; they differ only in which server image they name. The runner brings each up in turn; the server writes its callgrind or massif file to a shared volume and the runner reads it between cycles. The same compose files serve local “run it on my laptop” use. Rootless podman with podman-compose is already standard in the test repository’s CI, so this needs no new node setup.

The results service

A small HTTP service that stores runs and answers history queries, living in its own repository and backed by Drizzle itself — the performance store for Drizzle dogfoods Drizzle. It is a Flask plus flask-restx application over SQLAlchemy, synchronous, containerized, and deployed on a collection host as a compose of the service and a Drizzle instance.

The service is a store, not a judge. It records runs and returns history; the regression decision lives in the job, where it is reviewable as code. This keeps the service generic and the policy visible.

The schema is intentionally generic — one runs table and one metrics table — so new workloads and metrics need no schema change:

runs (
  run_id      BIGINT PRIMARY KEY AUTO_INCREMENT,
  workload    VARCHAR(40)   NOT NULL,   -- 'perf', later 'slap', 'sysbench'
  repo        VARCHAR(80)   NOT NULL,
  branch      VARCHAR(120)  NOT NULL,
  commit_sha  CHAR(40)      NOT NULL,
  change      VARCHAR(40)   NULL,       -- gerrit change; null for promote/backfill
  patchset    VARCHAR(20)   NULL,
  pipeline    VARCHAR(40)   NOT NULL,   -- 'promote' for authoritative rows
  distro      VARCHAR(40)   NOT NULL,   -- 'trixie', 'resolute'
  arch        VARCHAR(20)   NOT NULL,   -- 'amd64', 'arm64'
  host_class  VARCHAR(40)   NOT NULL,   -- 'callgrind'; a future wall-clock tier only
  tag         VARCHAR(80)   NOT NULL DEFAULT '',  -- 'baseline','14.04','release/x.y'; '' = none
  run_date    DATETIME      NOT NULL,
  succeeded   BOOLEAN       NOT NULL,
  INDEX series (workload, branch, arch, distro, run_date),
  UNIQUE KEY ident (commit_sha, workload, branch, arch, distro, tag)
);

metrics (
  run_id       BIGINT        NOT NULL,
  metric_key   VARCHAR(80)   NOT NULL,  -- 'callgrind.ir', 'massif.peak_heap_bytes'
  dim_key      VARCHAR(120)  NOT NULL DEFAULT '',  -- '' for scalar metrics
  value        DECIMAL(20,4) NOT NULL,
  wallclock    BOOLEAN       NOT NULL DEFAULT FALSE,
  PRIMARY KEY (run_id, metric_key, dim_key),
  FOREIGN KEY (run_id) REFERENCES runs(run_id)
);

A note on the metric key. Every v1 metric is scalar, so dim_key is the empty string throughout v1 and the primary key is effectively (run_id, metric_key). dim_key exists for a future where a metric is parameterized — concurrency, say — at which point it holds a canonical, deterministically-serialized string (e.g. threads=16), not structured JSON. Keeping it an ordinary defaulted VARCHAR in the primary key, rather than a JSON column with an expression index, keeps the schema portable and directly implementable on Drizzle; the producer is responsible for emitting the canonical form.

The v1 API is small:

  • POST /v1/runs — a metrics document plus run metadata and an optional tag. Write token. Idempotent on the full series plus commit and tag: (commit_sha, workload, branch, arch, distro, tag), enforced by the ident unique constraint above. This must match the series dimensions: now that amd64/trixie and arm64/trixie are distinct series, both upload a row for the same merge commit_sha and a narrower key would collapse them into one. tag is a non-null column defaulted to '' (the “no tag” sentinel) precisely so it can sit in that unique constraint — a nullable column would not enforce uniqueness, since NULL is distinct from NULL in SQL.

  • GET /v1/history — the trailing window the job compares against, filtered by the full series key: workload, branch, arch, distro and metric_key, with a limit. Read token. (See Regression gating for why arch and distro are part of the key and not just recorded.)

  • GET /v1/baseline — the committed reference series.

  • GET /v1/runs/{id} — a single run.

The service’s own CI stands it up against a real Drizzle container and exercises the round trip — the ultimate dogfood.

Authentication and the upload flow

Two tokens enforce a simple but important property: a change cannot write to the authoritative history before it merges.

A broad read token is available to the measurement job so it can fetch history for comparison. A tightly scoped write token exists only in the promote pipeline. The benchmark job that runs on proposed changes never holds the write token.

The flow follows from that:

  • In the vouched and gate pipelines the job runs the workload, produces a metrics document, publishes it as a Zuul artifact so reviewers can see the delta on the change, and — once gating is enabled — compares against history and fails on regression. It does not upload.

  • In the promote pipeline, after the change merges, a separate job downloads the metrics artifact from the gate build and POSTs it with the write token, recording an authoritative row.

The committed baseline series (the per-release perf/*.json files inherited from the server repo) is backfilled into the store as tagged runs — baseline, 14.04 through 24.04 — so the long historical series and live runs share one space, distinguished by tag. They are callgrind numbers like any other and need no special treatment.

Regression gating

Gating ports the old trailing-window method, modernized, and lives in the job.

For each metric the job fetches a trailing window of merged runs for the same series — the same workload, branch, arch and distro — a last-5 and a last-20 window, matching the historical model — and computes the mean and standard deviation over each. A candidate run fails if a metric is worse than the mean by more than max(k·σ, floor):

  • k is 3.

  • floor is per-metric: 3% for callgrind.ir and callgrind.estimated_cycles (the documented multithreaded noise floor), a wider 5% for callgrind.global_bus_events (atomic counts swing more with thread scheduling than instruction counts do), about 1% for massif.peak_heap_bytes, and 0% for size.* (bit-exact).

  • “Worse” means higher for every v1 metric. (Lower-is-worse only becomes relevant on a future wall-clock tier.)

The report is two-sided but the gate is one-sided: the job always prints the full delta table, improvements included, and fails only on regressions. The thresholds live in a committed config file so they are reviewed like any other code.

The series key is (workload, branch, arch, distro, metric_key). Each part earns its place. Callgrind counts are host-independent — the simulated CPU does not depend on which cloud node runs the job, so host_class is not a key part (it stays dormant in the schema for a future wall-clock tier). But host-independent is not ISA-independent: the same source compiles to different instruction sequences on amd64 and arm64, so the two architectures are separate series and must never share a trailing window — a candidate must be compared only against prior runs on its own arch. distro is in the key for the same reason at one remove: instruction counts shift with the toolchain and libc, so a distro change is a rebaseline event, starting a fresh series rather than comparing across the boundary. Omitting arch or distro from the key would let a migration or a mixed-arch history gate a candidate against the wrong baseline.

Where the work lives

The work spans three repositories, each owning its piece:

  • drizzle-test — the perf images, the runner, the compose wiring, and the measurement job; this is the harness’s home. The load generator installs sysbench from the distribution, so no separate driver or workload repository is involved.

  • perf-tracking — the results service and its dogfood integration job.

  • drizzle (the server) — loses the stranded perf.sh / perf-report.pl / [perf] bindep harness once the replacement runs, and gains a reference to the measurement job so a server change triggers it.

Each piece is developed and gated in its own repository first; the server repository only references the finished job.

See Performance tracking — implementation plan for work items and proposed agent prompts.

Follow-ups

These are out of scope for the initial work but the design accommodates them without change:

  • A dashboard. A static page on drizzle.org reading the service’s history API and plotting each metric over time, with release and distro tags marked. The API and the tag field are shaped for this already.

  • More workloads and the wall-clock tier. Additional sysbench OLTP scripts (read-only, write-only, point-select) under callgrind, each a new workload value needing no schema change. And the bare-metal wall-clock tier: the same sysbench, run with real --threads and --time on dedicated hardware, recording wallclock=true metrics that gate only there. randgen as a pass/fail stress job with no trend tracking.

Alternatives

A perf stage in the server’s container build. The original harness lived in the server repository and the obvious path is to add a perf build stage there. We reject this: the server repository should build a server image and nothing else, and coupling the perf toolchain into it makes the server build carry test concerns. Building the perf images in the test repository, consuming the server image via the buildset registry, keeps the separation clean and still guarantees perf runs against the exact change under test.

One combined perf image. A single image FROM the server image with valgrind and sysbench both installed, run as two containers, is simpler to build. We reject it because it puts the load-generation tooling inside the image whose instruction count we measure. Keeping the measured server (minimal, valgrind only) separate from the load generator keeps the measurement clean and makes swapping or adding sysbench workloads a matter of changing only the load generator.

The native Drizzle protocol. Measuring over Drizzle’s native protocol (port and wire format distinct from MySQL’s) would exercise that code path rather than the MySQL-compatibility one. Doing so would mean a native-protocol load generator — historically the dbd-drizzle Perl driver over libdrizzle, whose only real motivation was a BSD-licensed client, which is not a concern for us, and which would need modernization and a C/XS build. sysbench speaks the MySQL protocol out of the box with no such cost. We accept that this measures the MySQL-protocol path rather than the native one; for tracking server work that is fine, and it is arguably more representative of real load. Native-protocol measurement remains possible later if a specific need arises.

The vendored Perl sql-bench. The revival seed used the vendored sql-bench corpus driven by Perl DBI. We reject continuing it: sql-bench is a frozen relic, its test-insert workload is narrow, and driving it needs a Perl/DBI/driver stack in the load-generator image. sysbench is the modern standard, connects natively over the MySQL protocol, installs as a single package, offers more representative OLTP workloads, and — the decisive point — is the same tool we want for the future wall-clock tier.

Hardware perf counters or eBPF. Faster and “more accurate” in an absolute sense, but host-specific, which defeats the entire purpose on shared cloud hardware. They remain the right tools for diagnosing a regression once callgrind has detected one, but they cannot be the gating metric.

A Python package for the runner. The runner could be a proper package mirroring the integration-test driver. It has none of that driver’s structure — no discovery, no fixtures — so a package would be scaffolding without value. A single script is the honest size.