Skip to content

Chapter 24 · Benchmark report: -O0 vs your -O1 vs LLVM's -O2

Eight Pebble programs from labs/ch24-capstone/inputs/bench/ (each also an end-to-end test with its expected output), compiled by pebblec from a -DPEBBLE_USE_SOLUTION=all build with LLVM 23.1.2 on this course's x86-64 Linux container, and timed with the provided driver:

uv run python labs/ch24-capstone/provided/bench.py --pebblec build/<preset>/bin/pebblec \
    --repeat 5 --markdown labs/ch24-capstone/inputs/bench/*.pbl

-O1 is the course pipeline: Chapter 12's function(pebble-strength) followed by pebble-o1, which runs the six stages of the reference designCoursePipeline (Lesson 24.1). -O2 is LLVM's buildPerModuleDefaultPipeline(O2). Instructions are LLVM IR instructions after the pipeline (counted syntactically in --emit=llvm output); object bytes are the .o size; compile time is one --emit=obj run (wall clock, ms); run time is the minimum of 5 executions (wall clock, ms). Timings on a shared machine move by ±10–20 % between runs (a re-run gave geometric-mean speedups of 3.4× and 3.7× for the same instruction counts and object sizes); compare ratios, not absolutes, and expect the 5–8 ms programs to swap places between -O1 and -O2.

program level instructions object bytes compile ms run ms (min of 5)
bench-checks.pbl -O0 297 3368 16.3 28.7
bench-checks.pbl -O1 137 2528 26.4 13.0
bench-checks.pbl -O2 86 2480 35.7 11.5
bench-collatz.pbl -O0 182 2448 19.1 72.0
bench-collatz.pbl -O1 58 1552 23.3 18.8
bench-collatz.pbl -O2 41 1552 26.1 15.9
bench-fib.pbl -O0 78 1720 14.2 9.5
bench-fib.pbl -O1 36 1584 17.5 6.7
bench-fib.pbl -O2 19 1352 19.0 5.9
bench-matmul.pbl -O0 482 5136 16.3 62.4
bench-matmul.pbl -O1 145 2752 35.4 7.6
bench-matmul.pbl -O2 203 2344 47.2 6.6
bench-nbody.pbl -O0 426 4752 19.4 111.0
bench-nbody.pbl -O1 190 2608 28.3 29.5
bench-nbody.pbl -O2 122 2808 29.8 28.0
bench-sieve.pbl -O0 144 2176 13.7 52.6
bench-sieve.pbl -O1 60 1800 24.8 16.1
bench-sieve.pbl -O2 40 1536 30.8 14.0
bench-sort.pbl -O0 300 3352 18.5 28.7
bench-sort.pbl -O1 100 2120 23.7 5.8
bench-sort.pbl -O2 73 1752 26.0 5.3
bench-structs.pbl -O0 243 2656 15.7 46.2
bench-structs.pbl -O1 107 1808 25.0 9.4
bench-structs.pbl -O2 124 1776 31.3 8.8

Summary

Metric (geometric mean over the 8 programs) your -O1 LLVM -O2
run-time speedup over -O0 3.62× 4.05×
run time relative to the other 1.00 0.89
LLVM instructions relative to -O2 1.31× (0.71×–1.89×) 1.00
compile time relative to -O0 1.22×–2.17× 1.34×–2.90×

What each ratio is made of

The trap calls left in each module (grep -c 'call void @pebble_trap' on --emit=llvm) say which checks each pipeline proved away:

program -O0 -O1 -O2
bench-checks 14 12 3
bench-collatz 7 3 2
bench-fib 3 3 1
bench-matmul 20 6 0
bench-nbody 8 0 0
bench-sieve 5 4 1
bench-sort 11 6 1
bench-structs 4 3 1
  • bench-matmul (8.2× at -O1, 20 → 6 traps). The -O0 module reloads a[i][k], b[k][j] and s from memory on every step and bounds-checks every index; pebble-mem2reg puts s and the induction variables in registers, pebble-inline inlines mul into main, pebble-bce removes the checks whose index is the loop counter bounded by the array length (the %i.addr.phi/%k.addr.phi loops), pebble-licm hoists the row addresses a[i] and c[i] (the getelementptr … %i.addr.phi lines in the -O1 output), and pebble-osr turns the r % 24 and (r &* 5) % 24 address arithmetic into the %.osr induction variables. -O2 is about 10–15 % faster with more instructions (203 vs 145): it unrolls the k loop and removes the last 6 checks with IndVarSimplify's range facts (at 5–8 ms per run the two levels are within the measurement noise of each other).
  • bench-fib (1.42× at -O1, 3 → 3 traps). Calls dominate; -O1 gains only from the promoted allocas and the branch simplification around the three checks. -O2 (1.61×) folds two of the checks (n - 1 and n - 2 cannot overflow once n >= 2 is known on that path), marks the recursive calls tail call fastcc and keeps a single llvm.sadd.with.overflow for fib(n-1) + fib(n-2).
  • bench-checks (2.21×, 14 → 12 → 3). pebble-bce proves the two prefix-sum indices prefix[i] and xs[i] from the for bounds; -O2's IndVarSimplify and LoopIdiom-era range analysis remove 11 of the 14. The remaining gap between 13.0 ms and 11.5 ms is the nine extra checks -O1 executes per iteration.
  • bench-nbody, bench-structs, bench-sort (3.8×–4.9×). Store-to-load forwarding (pebble-loadfwd) through &mut references and the inlining of accel, advance and lcg do most of the work; -O2 adds SROA of the Body fields and GVN across the unrolled inner loop, worth 2–7 %. bench-nbody has only float arithmetic and constant-bounded loops, so both pipelines remove every check.
  • Compile time. pebble-o1 runs its simplify stage up to 3 rounds and its late stage up to 2 (the trace box of Lesson 24.1); the second round rarely changes anything, and the measured cost is the pass work itself (hashing the module for change detection costs less than 1 ms on these programs).

Reproduce a single line with pebblec -O1 --emit=llvm labs/ch24-capstone/inputs/bench/bench-matmul.pbl -o - | grep -c 'call void @pebble_trap' (checks left) and pebblec --passes='pebble-o1<trace>' --emit=llvm … -o /dev/null (rounds).