Benchmark parity
How the benchmark suite keeps overhead numbers honest - equivalent work on both sides, verified before anything is timed.
The benchmark suite has one hard rule:
Compare equivalent work, not just equivalent SQL count.
A wrapper looks slow if its rich result is timed against a raw query that does less, and looks fast if the raw side does more. Parity keeps both sides doing the same thing. The resulting numbers are on the benchmarks page; this page explains how they are produced.
What "API parity" means
Every published scenario is a pair of functions in benchmark/scenarios.ts: a raw Drizzle version and a better-drizzle version. For a pair to count as parity:
- Same result. The raw side returns the same value, deep-equal: the same rows, the same nested relation payloads, the same pagination metadata. Suites check this with
deepStrictEqualbefore timing. - Same work. The raw side runs the queries a careful developer would write to produce that result: batched relation loading instead of N+1 queries, a real
countfor offset totals, a real existence check for cursor navigation. - Same database. Both sides run against the same schema and seed data, on SQLite, so wrapper overhead is not hidden by network or disk I/O.
Scenarios that intentionally do less work on the raw side, such as a flat join instead of a nested payload or a cursor page without navigation flags, live in a separate manual drizzle reference group. They show what is possible without the wrapper, but they are never used as overhead claims.
Example: cursor hasPrevious
cursor() returns hasNext, hasPrevious, nextCursor, and previousCursor. An earlier raw scenario hardcoded hasPrevious: true, so it ran one query while better-drizzle resolved the flag for real. That made cursor() look about 2x slower than raw Drizzle, from work only one side did.
The current raw scenario computes the flag in the same query, with a correlated exists subquery, and falls back to a probe query for empty pages:
export const rawCursorPaginate = async (context: BenchmarkContext) => {
const data = await context.raw
.select({
...userColumns,
__hasPrevious:
sql`exists (select 1 from ${users} as prior where prior.id <= ${context.ids.cursorAfterId})`.mapWith(
Boolean,
),
})
.from(users)
.where(gt(users.id, context.ids.cursorAfterId))
.orderBy(asc(users.id))
.limit(26);
const hasPrevious = data.length
? data[0].__hasPrevious
: (await context.raw.select({ id: users.id }).from(users).limit(1))
.length > 0;
// ...strip __hasPrevious, slice to 25, build the same pagination object
};The data-only version, one query with no navigation flags, is kept as drizzle manual: cursor data only in the manual reference group. Re-check this kind of detail whenever a pagination scenario changes.
How the published numbers are measured
Absolute timings drift with machine load: the same operation can read 53 µs on an idle machine and 93 µs under load. The raw/better ratio is stable across runs, so that is what gets published. benchmark/report.ts is built around it:
Verify parity
Each read pair runs once on both sides and the results are compared with deepStrictEqual. A mismatch stops the run before anything is timed. Write and transaction pairs are skipped here, because they change state on each call; bench:full checks those.
Interleave samples
For each pair, both sides are sampled inside the same window. The leading side alternates on every sample, so a machine that is warming up or cooling down does not favour whichever side runs first. The sample count is BENCH_SAMPLES, 7 by default.
Delegate each measurement to mitata
Each sample is one mitata measure() call, and its p50 is recorded. Warmup, JIT settling, GC handling, and outlier trimming therefore match bun run bench. An earlier hand-rolled timing loop disagreed with mitata by about 30 points on point lookup, so timing is not done by hand.
Take the median
The report uses the median of each side's samples and publishes better / raw as the overhead percentage, alongside machine-specific absolute values.
Running the suites
| Command | What it does |
|---|---|
bun run bench | Latency per operation with mitata, grouped as api parity: reads, writes, transactions, and manual drizzle reference. Raw and better sides use separate temporary databases. |
bun run bench:verify | Deep result parity for every read, write, and transaction pair in benchmark/full.ts, with no timing. |
bun run bench:full | Runs the same verification, then times all of those pairs, including batch writes, relation writes, relation counts, and raw SQL. |
bun run bench:memory | Heap and RSS deltas across batches of reads, writes, and transactions. |
bun run bench:report | Verifies read parity, then prints the interleaved, median-of-samples Markdown tables used on the benchmarks page. |
bun run bench:all | bench, bench:full, and bench:memory in sequence. |
bun run bench:jsonb | PostgreSQL JSONB path filters with row parity and plan parity (both sides must use, or both skip, the expression index). Needs DATABASE_URL. |
bun run bench:arrays | PostgreSQL native array containment with raw Drizzle row parity. Needs DATABASE_URL and a GIN index. |
bun run bench:verify
BENCH_SAMPLES=11 bun run bench:reportWhen you change performance-sensitive code, run bench, bench:verify, bench:full, and bench:memory, and read regressions against the parity groups first.
Adding a fair scenario
Write both sides
Add a raw... and a better... function to benchmark/scenarios.ts. Each takes the BenchmarkContext and returns its result.
export const rawTopPosts = async (context: BenchmarkContext) =>
context.raw
.select()
.from(posts)
.where(eq(posts.published, true))
.orderBy(desc(posts.score), asc(posts.id))
.limit(10);
export const betterTopPosts = async (context: BenchmarkContext) =>
betterClient(context).posts.findMany({
where: { published: true },
orderBy: [{ score: 'desc' }, { id: 'asc' }],
take: 10,
});Give both sides the same total order, with a tiebreaker such as id. Rows that tie on the sort key can come back in a different order and fail the deep-equal check.
Match the work, not the call count
If the better side loads relations, the raw side loads the same relations with batched queries and assembles the same nesting. If it returns pagination metadata, the raw side computes every field. If a plugin rewrites the operation, the raw side performs the same filtering or rewrite.
Keep writes repeatable
Mutation scenarios run thousands of times. Use the deterministic ID and token counters in scenarios.ts, or restore state inside the scenario, so every iteration does the same work and the two sides never collide on unique keys.
Register the pair
Add it to the api parity group in benchmark/time.ts, to readPairs or writePairs in benchmark/full.ts, and, if it should be published, to READS or WRITES in benchmark/report.ts. A raw shortcut that does less work goes only in the manual drizzle reference group.
Verify before reading numbers
Run bun run bench:verify. Only once parity passes do the timings mean anything.
Reading the results
- A hand-tuned manual query can beat the wrapper. The question is whether the wrapper stayed small for the behavior it added.
- The wrapper can also win, when its default is better than the obvious raw code. Relation graph reads are the main case: the batched loader beats the equivalent raw code.
- Heap numbers matter as much as latency. Some batches, such as mixed reads and transactions, use more heap than raw Drizzle.