Skip to content

Add the YCSB driver - #17

Merged
tamnd merged 2 commits into
mainfrom
ycsb
Oct 7, 2026
Merged

tamnd merged 2 commits into
mainfrom
ycsb

Conversation

@tamnd

@tamnd tamnd commented Oct 7, 2026

Copy link
Copy Markdown
Owner

Part of #1.

What changes

  • src/ycsb.rs is the YCSB driver of spec/20 §20.9. It follows the core workload of YCSB: the table usertable with ten text fields of 100 bytes, keys from the FNV hash of the record number, and the scrambled zipfian distribution with the constant 0.99. The workloads are A, B, C and F.
  • The step load makes the table and loads --records rows with COPY, then runs VACUUM ANALYZE and CHECKPOINT. The step run runs each workload for each row of --rows for --time seconds. A row is a client count such as 16, or a client count and a pipeline depth such as 16x64.
  • Each client has its own connection and thread and sends prepared statements with the extended protocol. src/pg.rs gets a pipeline: up to DEPTH statements wait for their results on one connection, and each has its own Sync, as in the pipeline mode of libpq.
  • The server cgroup is measured over each row, with the idle base first. Each row has the throughput, the server CPU for each operation, memory.peak, the p50, p95, p99 and p99.9 latency of each operation from a log linear histogram, and the update rate of the hottest key. --sync on|off sets synchronous_commit on each connection.
  • The checks: each read must return one row of its key, and each update must change one row. After each row the driver reads every field that the row updated, and the field must hold an update that no later acknowledged update replaced. A row that fails a check is a wrong answer, and the command fails.
  • reports/2026-10-07/c7a1474-server3-ycsb-postgresql-smoke.{md,json} is the report of the smoke run below. Its name and its first line mark it as a smoke run.
  • The README explains the command.

Test

srvtest passed on server3. The new unit tests cover the FNV keys, the zipfian skew, the workloads and rows, the histogram, the load rows and the rule for the allowed values of a field after a run. The pipeline ran against PostgreSQL in the smoke run below.

Smoke run on server3, 7 October 2026. It is not a baseline. server3 is a shared machine with 8 cores and a load average between 41 and 85 when I read it that day. The run used a second PostgreSQL cluster, 19/ycsb on port 5433 (REL_19_STABLE at 7d3d2db7, 19beta4), so that it did not touch the TPC-H data in 19/main. The command was rupg-bench ycsb --unit postgresql@19-ycsb --conn "host=/var/run/postgresql port=5433 user=bench dbname=bench" --records 100000 --time 10 --smoke with the default workloads and rows, and synchronous_commit was on from the server. The report has every setting.

Workload 1 client, ops/s 16 clients, ops/s 16 clients, depth 64, ops/s Fields checked, wrong
A 117 418 592 5,398, 0
B 138 859 2,158 1,618, 0
C 84 1,209 4,623 0, 0
F 85 377 542 4,984, 0

The load of 100,000 rows took 43.32 s. No operation failed in any row.

The check finds lost updates. At 13:43 UTC on server3, I added a trigger to usertable in 19/ycsb that returns OLD for 10 percent of the updates, so the update reports one row but does not store the value. Then rupg-bench ycsb --unit postgresql@19-ycsb --conn ... --steps run --workloads a --rows 1 --time 5 --smoke failed with wrong answers, the rows are not numbers: workload a with 1 clients: 3 fields do not hold their last acknowledged update. After I dropped the trigger, the same command passed with 385 fields checked and 0 wrong.

The cost of the sampler. The cgroup runner reads /proc/<pid>/smaps_rollup of each server process every 100 ms for the Pss. At 13:44 UTC on server3, /usr/bin/time rupg-bench measure --unit postgresql@19-ycsb --idle (14 server processes, 10 s) used 4.58 s and 4.71 s of system CPU in the harness in two tries, and took 74 and 72 samples, not 100. That CPU is in the harness and not in the server cgroup, so the server numbers do not include it, but on a loaded machine it takes CPU from the server. A later change can read the Pss less often.

A source that is missing. spec/20 §20.9 cites ../2140/bench/ycsb/04-the-driver.md. That file is not in the checkouts here, so the driver follows the YCSB core workload and the spec text.

No line of #1 is complete with this PR. The YCSB baseline line needs the oltp machine.

@tamnd tamnd added kind/feature It does not do something that it must do. area/ycsb YCSB workloads A to F, with and without pipelining. labels Oct 7, 2026
@tamnd tamnd self-assigned this Oct 7, 2026
@tamnd
tamnd merged commit 8246808 into main Oct 7, 2026
2 checks passed
@tamnd
tamnd deleted the ycsb branch October 7, 2026 13:55
@tamnd tamnd mentioned this pull request Oct 7, 2026
5 of 17 tasks
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area/ycsb YCSB workloads A to F, with and without pipelining. kind/feature It does not do something that it must do.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant