Repository navigation
Conversation
This was referenced Oct 7, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Part of #1.
What changes
src/tpch.rsis the TPC-H driver of spec/20 §20.7. The stepgenrunsdbgen3.0.1 and writes the 22 queries ofqgen -dtoDATA/queries. Q15 uses the approved variant A of the kit, with aWITHclause in place of the view. qgen prints the row count as--LIMIT n, and the driver turns it into aLIMITclause.loadcreates the 8 tables with the types, primary keys and foreign keys of the specification. DuckDB and PostgreSQL read the same.tblfiles. For a server, the driver sends the rows withCOPY ... FROM STDINthrough the harness client, adds the keys after the rows, and runsVACUUM ANALYZEandCHECKPOINT.runis insrc/suite.rs, which ClickBench will also use. Each query runs 3 times, and each run is a new session. The first run is cold: the driver drops the page cache, and for a server unit it also restarts the unit. The hot time is the smallest of the other runs. Each run has the counters of its cgroup. A failed query is left out of the time sums and counted as failed.--answerschecks the answers. An.outfile of the kit at SF1 matches with the rules of TPC-H clause 2.1.3.5. A.tsvfile from--save-answersof another run matches as a multiset of rows. A space at the start of a text in an.outfile is kept.src/pg.rshascopy_in.machines/install/tpch-tools.shbuildsdbgenandqgenfrom the pin withmakefile.suite, so a changed Makefile in the copy does not count.machines/install/postgresql.shgivespg_checkpointto the rolebench.tpch-sf0.1-postgresql, now works, and the driver callsdbgenandqgenwith absolute paths.reports/2026-10-07/68cad8fc-server3-*are the reports of the 5 smoke runs below. Their names and their first line mark them as smoke runs.Test
srvtestpassed on server3: clippy is clean and 50 tests pass. The tests cover the schema, the qgen output, the.outformat and the rules of clause 2.1.3.5.Smoke runs on server3, 7 October 2026. These are not baselines. server3 is a shared machine with 8 cores, and its load average was between 41 and 85 when I read it that day. PostgreSQL is REL_19_STABLE at
7d3d2db7(19beta4), and DuckDB is 1.5.6. The script ran these commands in/root/work/tpchwithRUPG_BENCH_COMMIT=68cad8fcand--smoke --report /root/work/reports --machine server3. The data of SF0.1 and SF1 was made bydbgenearlier that day.tpch --scale 0.1 --engine postgresql --unit postgresql@19-main --save-answers ans-pg-sf0.1instructions --set tpch --queries sf0.1/queries --engine postgresql --unit postgresql@19-maintpch --scale 0.1 --duckdb sf0.1.db --answers ans-pg-sf0.1instructions --set tpch --queries sf0.1/queries --duckdb sf0.1.dbtpch --scale 1 --duckdb sf1.db --answers dbgen/answers --save-answers ans-duckdb-sf1The PostgreSQL hot total is mostly Q20 (465.3 s) and Q17 (236.4 s). The schema has only the keys of the specification, so these correlated subqueries scan
lineitemagain for each outer row. At SF1 this is not practical on server3, so the PostgreSQL run is at SF0.1. The answer of PostgreSQL at SF0.1 has no reference in the kit, so it was saved and the DuckDB run checked against it.For the instruction counts, the median spread of the 3 runs of a query is 2.8 percent for PostgreSQL and 3.1 percent for DuckDB, and the largest spread is 51 percent (PostgreSQL Q3) and 15 percent (DuckDB Q2). The spread is (largest run minus smallest run) divided by the smallest run. On this shared machine it is above the 3 percent budget of spec/21 §21.14, so
ratchet.tomlstays empty.No line of #1 is complete with this PR. The TPC-H baseline line needs the
tpchmachine, and the instruction line also needs the YCSB and TPC-C sets.