Skip to content

Chapter 24 exercises

The capstone has three kinds of work. E1 is the one piece of pebblec this chapter adds: your -O1 pipeline, as data, run by the provided pebble-o1 module pass (pebble/lib/Passes/Pipeline/). E2 and lab L1 are the evidence that it is correct and worth having: a random program generator (labs/ch24-capstone/), the differential fuzzer and the benchmark driver that use it. The ★ parts are three independent projects: E3 synthetic debug info (pebble/lib/Passes/DebugInfo/), J1 an ORC JIT and REPL (labs/ch24-jit/), and a second front end in Rust (labs/ch24-rust-frontend/). Run the tests after every step:

./course test 24                                        # builds, then runs every test labelled ch24
ctest --preset linux -L '^ch24$' -R 'ch24.Pipeline'     # one suite while iterating (macos preset on a Mac)
build/linux/bin/pebblec --passes='pebble-o1<print>' --emit=llvm x.pbl -o /dev/null   # the design a build runs
build/linux/bin/pebblec --passes='pebble-o1<trace>' --emit=llvm x.pbl -o /dev/null   # ... with per-round counts

Before you start: ch24.Pipeline.* fail with no stages: designCoursePipeline returned an empty design, pebblec -O1 runs Chapter 12's function(pebble-strength) and then a pebble-o1 that does nothing (so every -O1 test still builds and produces correct code, only o-levels.pbl's instruction-count check fails), the generator and JIT tests stop with TODO(ch24), and opt -passes=pebble-debugify reports unknown pass name 'pebble-debugify'. That is expected.

Stuck? Each SPEC ends with hints, in order. The reference solutions are in solutions/pebble/lib/Passes/{Pipeline,DebugInfo}/ and solutions/labs/ch24-*/. Look after you have passed the tests, or after an honest hour.


E1: your -O1 pipeline

Contract: pebble::passes::designCoursePipeline in pebble/include/pebble/Passes/CoursePipeline.h; your code in pebble/lib/Passes/Pipeline/src/ (replace Stub.cpp) Tests: ch24.Pipeline.* (tests/ch24/unit/PipelineTest.cpp), tests/ch24/lit/{registry.test,pipeline-print.pbl,o-levels.pbl}, ch24.e2e-ch11 (all 66 Chapter 11 programs at -O1), ch24.e2e (the benchmark programs at -O0/-O1/-O2)

pebblec -O1 runs the course pipeline: every PEBBLE_COURSE_PIPELINE_STEP linked into the driver, by order key (Lesson 11.9, Algorithm 11.9.5's note "What -O1 is in your build"). Chapter 12 registered function(pebble-strength) at key 1210; this chapter registers the module pass pebble-o1 at key 2400, so -O1 is function(pebble-strength),pebble-o1 and pebble-o1 runs your design:

struct PipelineStage { std::string Name; std::string Pipeline; unsigned MaxRounds = 1; };
struct PipelineDesign { std::vector<PipelineStage> Stages; };
PipelineDesign designCoursePipeline(const std::map<std::string, std::string> &Available);

Available maps each course pass that this build registered to the kind of pipeline it fits in ("function", "module", "loop", "cgscc"); coursePassCandidates() lists, per chapter, every course pass the probe asks about. Pipeline is a module pipeline in opt -passes= syntax; LLVM's own passes (loop-simplify, lcssa, loop-rotate, verify) are always available. pebble-o1 runs each stage up to MaxRounds times and stops a stage early when a round changes nothing (it hashes the module; Lesson 24.1, Algorithm 24.1.7's Changed at module granularity).

Requirements (the unit test's constraints C1–C6, which are Lesson 24.1's Theorem 24.1.10 and Proposition 24.1.11 written down):

  • C1 SSA construction first: pebble-mem2reg precedes every register-level pass (sccp, gvn, adce, licm, osr, reassociate, lvn).
  • C2 loop canonicalization before loop passes: loop-simplify (and lcssa) precede pebble-licm, pebble-osr, pebble-unroll, pebble-bce, in the same stage or earlier (Definition 24.1.9).
  • C3 pebble-funcattrs after pebble-inline.
  • C4 the last stage deletes dead code (pebble-dce or pebble-adce).
  • C5 at least one stage has MaxRounds > 1 (the fixpoint block after inlining, Algorithm 24.1.5).
  • C6 the inliner runs no later than the stage that first runs pebble-gvn.
  • Every stage parses with the course PassBuilder; every transformation pass in Available is used at least once (pebble-bbcount, pebble-uninit and the printers are not transformations); nothing outside Available is named — a build with only Chapters 12–13 must still get a working design (test UsesOnlyAvailablePasses); the design is deterministic.
  • -O0, -O1, -O2 agree on every end-to-end program, and -O1 leaves fewer LLVM instructions than -O0 on o-levels.pbl.

What the tests check: the constraints above on the full candidate set and on a 4-pass subset; print<pebble-passes> lists ch24 module pass pebble-o1 after pebble-strength; pebble-o1<print> prints one stage <name> (max N round(s)) line per stage; all 66 Chapter 11 programs and the 8 benchmarks give identical stdout, stderr and exit status at the three levels.

Suggested order: write a one-stage design (function(pebble-mem2reg,pebble-simplifycfg,pebble-constfold,pebble-peephole,pebble-dce)) and get ch24.e2e-ch11 green; add the scalar and loop stages; add inlining and a fixpoint block; then run E2's benchmark and move passes until the numbers stop improving. Lesson 24.1 §2's trace box and §3 walk order.pbl through the reference design; §7 shows how PassBuilderPipelines.cpp orders the same passes.

E2: fuzz and benchmark it

Contract: none to write: labs/ch24-capstone/provided/fuzz.py and provided/bench.py are provided and need L1's generator (ch24-gen) and E1's pipeline Tests: tests/ch24/lit/{fuzz-run.test,bench-smoke.test}

B=build/linux/bin
uv run python labs/ch24-capstone/provided/fuzz.py --gen $B/ch24-gen --pebblec $B/pebblec \
    --pir-run $B/pir-run --pir-opt $B/pir-opt --opt /opt/llvm-23/bin/opt --seed 1 --count 500 --keep /tmp/fz
uv run python labs/ch24-capstone/provided/bench.py --pebblec $B/pebblec --markdown --json bench.json \
    labs/ch24-capstone/inputs/bench/*.pbl

The fuzzer's oracle is pir-run (the PIR interpreter of Chapter 8–9): every generated program has defined behavior (Pebble spec §11.2), so the interpreter's stdout, stderr and exit status are the answer, and every level's native executable must match. A difference is a miscompilation witness; the programs, PIR and IR at each level are kept in --keep for reduction (Chapter 12, ddmin) and bisection (pebblec --passes= with fewer stages). Traps count as agreement when both sides trap with the same message and status 101. The benchmark driver measures compile time, run time (minimum of --repeat runs), LLVM instruction count and object size per level, and refuses to report a program whose levels disagree.

Requirements: 500 seeds with traps and 100 without report 0 differences, 0 front-end failures; fill in the measurement table of SPEC §9 for the eight programs and write three sentences: which stage buys the most, on which program -O2 wins by the largest factor and why (Lesson 24.2's trap-count table is the place to look), and one thing you would change in E1 as a result. The chapter's own numbers are in benchmark.md.

What the tests check: 30 seeds with traps and 10 without agree at -O0/-O1/-O2 (a bounded, deterministic run, so CI catches a regression in any chapter's pass); bench.py --quick runs two programs and prints the table.

Optional, from Lesson 24.7: validate the intraprocedural stages of your pipeline with alive-tv on the kept programs (Algorithm 24.7.5): opt -load-pass-plugin=build/linux/lib/PebblePasses.so -passes='<stage>' on the -O0 IR, then alive-tv --func=<fn> --disable-undef-input before.ll after.ll. Stop before pebble-inline.

E3 ★: pebble-debugify

Contract: a module pass registered as pebble-debugify (parameter strip) in pebble/lib/Passes/DebugInfo/ (PEBBLE_MODULE_PASS("pebble-debugify", YourPass); or the parameterized form, as pebble-inline does) Tests: tests/ch24/lit/debugify.ll (labels ch24, star)

pebblec's code generator (Chapter 11) does not carry PIR's @line:col into LLVM IR. This pass gives a module synthetic debug info the way LLVM's debugify utility does, so that the chain of Lesson 24.4 can be exercised end to end: DIBuilder metadata → #dbg_declare/#dbg_value records → llc → DWARF line table and variable locations → gdb. Algorithm 24.4.6 is the specification; the tests check:

  • a DICompileUnit (DW_LANG_C, producer "pebblec") for the module's source_filename, the module flags Debug Info Version = 3 and Dwarf Version = 5 (Definition 24.4.1);
  • a distinct DISubprogram per defined function (name = linkageName = the function's name, a subroutine type from its LLVM signature), attached with !dbg;
  • a DILocation on every instruction: a fresh line per instruction, column 1, scope = the subprogram;
  • a #dbg_declare for every entry-block alloca whose name is a Pebble variable (count.addr → variable count; the temporaries _N are skipped), placed right after the alloca; a #dbg_value after every named scalar SSA definition (count.next, count.addr.phi3 → count), one DILocalVariable per (function, name);
  • basic types bool (i1/i8), int (other integers), float (double), a pointer type for ptr, arrays for [N x T];
  • the module verifies; llc -O0 -filetype=obj produces a line table with is_stmt prologue_end and a DW_TAG_variable with DW_AT_location (DW_OP_fbreg ...) and DW_AT_name ("count");
  • idempotence: running the pass on a module that already has a compile unit changes nothing (Proposition 24.4.9's practical form: the second run must not double the records); pebble-debugify<strip> removes every !dbg and #dbg_ (use llvm::StripDebugInfo).

Then try it: pebblec --emit=llvm -O0 bench-sieve.pbl, the pass, -O1's stages with opt, and gdb on the result: Lesson 24.4 §7 shows what info locals prints after mem2reg turns declares into values, and why an -O1 variable is <optimized out> where a #dbg_value was dropped.

J1 ★: the Pebble JIT and REPL

Contract: pebblejit::Session in labs/ch24-jit/include/pebblejit/JIT.h; your code in labs/ch24-jit/src/ (replace Stub.cpp); the driver pebble-jit is provided (tools/pebble-jit.cpp) Tests: ch24.JIT.* (tests/ch24/unit/JITTest.cpp), tests/ch24/lit/{jit.pbl,repl.test} (labels ch24, star)

struct Options { bool Lazy = false; bool Trace = false; };
class Session {
  static std::expected<std::unique_ptr<Session>, pebble::Error> create(const Options &);
  virtual std::expected<int64_t, pebble::Error>
  run(std::unique_ptr<llvm::Module>, std::unique_ptr<llvm::LLVMContext>, std::string_view Entry) = 0;
  virtual unsigned compiledFunctions() const = 0;
};

run adds a module and calls i64 Entry(); the driver compiles a .pbl in process (front end → PIR → lowerPIRToLLVM → optimizeModule at the requested level) and calls run with codegen::EntryPointSymbol (pebble_main). The REPL (--repl) keeps every declaration typed so far, wraps an expression line in print(...), compiles a fresh module per line — each defining pebble_main — and runs it in the same session. SPEC R1–R5:

  • R1 the runtime's symbols (pebble_print_int, pebble_trap, ..., and fmod) resolve in every module: publish them with absoluteSymbols (Lesson 24.3, Algorithm 24.3.3 step 2), do not rely on DynamicLibrarySearchGenerator finding them in the process, since pebble-jit is statically linked;
  • R2 each run uses a fresh JITDylib linked against the runtime dylib, so that many modules defining pebble_main coexist (ManyModulesInOneSession);
  • R3 eager (LLJIT) compiles every defined function of the module at run (compiledFunctions() == 4 on the test program); lazy (LLLazyJIT, addLazyIRModule) compiles only the functions that are called (== 2: pebble_main and used); count bodies in an IRTransformLayer transform, not symbols looked up;
  • R4 Trace prints pebble-jit: compiled <function> per body, in compilation order; the driver's --jit-stats prints pebble-jit: compiled N function(s);
  • R5 a trap inside JITed code behaves as in native code: the runtime's pebble_trap prints pebble: trap: ... at file:line:col and exits 101 (Inputs/trap.pbl); a missing entry symbol is an error, not a crash.

What the tests check: the four unit tests (eager count, lazy count, three modules in one session, missing entry); jit.pbl at -O0/-O1/-O2 with the eager and lazy counts; the REPL session in Inputs/session.txt (declarations, expressions, statements, an erroneous line that is reported and skipped, recursion, a struct).

Then try it: --lazy --jit-trace on bench-fib.pbl shows the order of first calls; Lesson 24.3's Algorithm 24.3.6 (ReOptimizeLayer) is the third tier, which the lab leaves to you.

★ A Pebble front end in Rust

Contract: the out-of-process front-end protocol of docs/architecture.md §4 (--emit=pir --pir-version=1.0 FILE; PIR on stdout; JSON-lines diagnostics on stderr; exit 0/½), as a cargo crate in labs/ch24-rust-frontend/ Tests: ch24.rust-e2e-ch11, ch24.rust-e2e-ch24 (labels ch24, star; configured only when cargo is on PATH)

This is the largest ★ project (a reference implementation of about 2 700 lines is in the lab; write your own in a copy of the crate, or replace its modules one at a time, lexer first). The definition of done is the conformance suite: every program under tests/ch11/e2e/ and tests/ch24/e2e/ gives the same PIR-level behavior (stdout, stderr, exit status through pebblec --frontend=<path>) as the C++ front end, and the SPEC's error programs are rejected with the same error codes, through the JSON-lines protocol. The SPEC lists which parts of the language spec each module implements.

What the tests check: pebble-e2e.py runs both suites with --param frontend=<your binary>.