Chapter 24 · Benchmark report: -O0 vs your -O1 vs LLVM's -O2¶
Eight Pebble programs from labs/ch24-capstone/inputs/bench/ (each also an end-to-end test with its expected output), compiled by pebblec from a -DPEBBLE_USE_SOLUTION=all build with LLVM 23.1.2 on this course's x86-64 Linux container, and timed with the provided driver:
uv run python labs/ch24-capstone/provided/bench.py --pebblec build/<preset>/bin/pebblec \
--repeat 5 --markdown labs/ch24-capstone/inputs/bench/*.pbl
-O1 is the course pipeline: Chapter 12's function(pebble-strength) followed by pebble-o1, which runs the six stages of the reference designCoursePipeline (Lesson 24.1). -O2 is LLVM's buildPerModuleDefaultPipeline(O2). Instructions are LLVM IR instructions after the pipeline (counted syntactically in --emit=llvm output); object bytes are the .o size; compile time is one --emit=obj run (wall clock, ms); run time is the minimum of 5 executions (wall clock, ms). Timings on a shared machine move by ±10–20 % between runs (a re-run gave geometric-mean speedups of 3.4× and 3.7× for the same instruction counts and object sizes); compare ratios, not absolutes, and expect the 5–8 ms programs to swap places between -O1 and -O2.
| program | level | instructions | object bytes | compile ms | run ms (min of 5) |
|---|---|---|---|---|---|
| bench-checks.pbl | -O0 | 297 | 3368 | 16.3 | 28.7 |
| bench-checks.pbl | -O1 | 137 | 2528 | 26.4 | 13.0 |
| bench-checks.pbl | -O2 | 86 | 2480 | 35.7 | 11.5 |
| bench-collatz.pbl | -O0 | 182 | 2448 | 19.1 | 72.0 |
| bench-collatz.pbl | -O1 | 58 | 1552 | 23.3 | 18.8 |
| bench-collatz.pbl | -O2 | 41 | 1552 | 26.1 | 15.9 |
| bench-fib.pbl | -O0 | 78 | 1720 | 14.2 | 9.5 |
| bench-fib.pbl | -O1 | 36 | 1584 | 17.5 | 6.7 |
| bench-fib.pbl | -O2 | 19 | 1352 | 19.0 | 5.9 |
| bench-matmul.pbl | -O0 | 482 | 5136 | 16.3 | 62.4 |
| bench-matmul.pbl | -O1 | 145 | 2752 | 35.4 | 7.6 |
| bench-matmul.pbl | -O2 | 203 | 2344 | 47.2 | 6.6 |
| bench-nbody.pbl | -O0 | 426 | 4752 | 19.4 | 111.0 |
| bench-nbody.pbl | -O1 | 190 | 2608 | 28.3 | 29.5 |
| bench-nbody.pbl | -O2 | 122 | 2808 | 29.8 | 28.0 |
| bench-sieve.pbl | -O0 | 144 | 2176 | 13.7 | 52.6 |
| bench-sieve.pbl | -O1 | 60 | 1800 | 24.8 | 16.1 |
| bench-sieve.pbl | -O2 | 40 | 1536 | 30.8 | 14.0 |
| bench-sort.pbl | -O0 | 300 | 3352 | 18.5 | 28.7 |
| bench-sort.pbl | -O1 | 100 | 2120 | 23.7 | 5.8 |
| bench-sort.pbl | -O2 | 73 | 1752 | 26.0 | 5.3 |
| bench-structs.pbl | -O0 | 243 | 2656 | 15.7 | 46.2 |
| bench-structs.pbl | -O1 | 107 | 1808 | 25.0 | 9.4 |
| bench-structs.pbl | -O2 | 124 | 1776 | 31.3 | 8.8 |
Summary¶
| Metric (geometric mean over the 8 programs) | your -O1 |
LLVM -O2 |
|---|---|---|
run-time speedup over -O0 |
3.62× | 4.05× |
| run time relative to the other | 1.00 | 0.89 |
LLVM instructions relative to -O2 |
1.31× (0.71×–1.89×) | 1.00 |
compile time relative to -O0 |
1.22×–2.17× | 1.34×–2.90× |
What each ratio is made of¶
The trap calls left in each module (grep -c 'call void @pebble_trap' on --emit=llvm) say which checks each pipeline proved away:
| program | -O0 |
-O1 |
-O2 |
|---|---|---|---|
| bench-checks | 14 | 12 | 3 |
| bench-collatz | 7 | 3 | 2 |
| bench-fib | 3 | 3 | 1 |
| bench-matmul | 20 | 6 | 0 |
| bench-nbody | 8 | 0 | 0 |
| bench-sieve | 5 | 4 | 1 |
| bench-sort | 11 | 6 | 1 |
| bench-structs | 4 | 3 | 1 |
bench-matmul(8.2× at-O1, 20 → 6 traps). The-O0module reloadsa[i][k],b[k][j]andsfrom memory on every step and bounds-checks every index;pebble-mem2regputssand the induction variables in registers,pebble-inlineinlinesmulintomain,pebble-bceremoves the checks whose index is the loop counter bounded by the array length (the%i.addr.phi/%k.addr.philoops),pebble-licmhoists the row addressesa[i]andc[i](thegetelementptr … %i.addr.philines in the-O1output), andpebble-osrturns ther % 24and(r &* 5) % 24address arithmetic into the%.osrinduction variables.-O2is about 10–15 % faster with more instructions (203 vs 145): it unrolls thekloop and removes the last 6 checks withIndVarSimplify's range facts (at 5–8 ms per run the two levels are within the measurement noise of each other).bench-fib(1.42× at-O1, 3 → 3 traps). Calls dominate;-O1gains only from the promoted allocas and the branch simplification around the three checks.-O2(1.61×) folds two of the checks (n - 1andn - 2cannot overflow oncen >= 2is known on that path), marks the recursive callstail call fastccand keeps a singlellvm.sadd.with.overflowforfib(n-1) + fib(n-2).bench-checks(2.21×, 14 → 12 → 3).pebble-bceproves the two prefix-sum indicesprefix[i]andxs[i]from theforbounds;-O2'sIndVarSimplifyandLoopIdiom-era range analysis remove 11 of the 14. The remaining gap between 13.0 ms and 11.5 ms is the nine extra checks-O1executes per iteration.bench-nbody,bench-structs,bench-sort(3.8×–4.9×). Store-to-load forwarding (pebble-loadfwd) through&mutreferences and the inlining ofaccel,advanceandlcgdo most of the work;-O2adds SROA of theBodyfields and GVN across the unrolled inner loop, worth 2–7 %.bench-nbodyhas only float arithmetic and constant-bounded loops, so both pipelines remove every check.- Compile time.
pebble-o1runs itssimplifystage up to 3 rounds and itslatestage up to 2 (the trace box of Lesson 24.1); the second round rarely changes anything, and the measured cost is the pass work itself (hashing the module for change detection costs less than 1 ms on these programs).
Reproduce a single line with pebblec -O1 --emit=llvm labs/ch24-capstone/inputs/bench/bench-matmul.pbl -o - | grep -c 'call void @pebble_trap' (checks left) and pebblec --passes='pebble-o1<trace>' --emit=llvm … -o /dev/null (rounds).