Skip to content

References — Chapter 0 · Compiler Architectures & Your Toolchain

Every source this chapter cites, grouped by kind. Lessons cite entries inline as [KEY]; each entry says why and when to read it. Core reading marks the entries the chapter assumes you will open.

Foundational and research papers

  • [BBB+57] John W. Backus, R. J. Beeber, S. Best, R. Goldberg, L. M. Haibt, H. L. Herrick, R. A. Nelson, D. Sayre, P. B. Sheridan, H. Stern, I. Ziller, R. A. Hughes, and R. Nutt. The FORTRAN Automatic Coding System. Proc. Western Joint Computer Conference, pp. 188–198, 1957. doi:10.1145/1455567.1455599
    Why and when: The first optimizing compiler, organized as six sections run in sequence: the origin of the multi-pass pipeline and of AOT compilation (Lessons 0.1 and 0.3). Short and readable; skim the description of the sections.
    Cited in: 01-pipeline-shapes, 03-aot-jit-transpilers

  • [BCFR09] Carl Friedrich Bolz, Antonio Cuni, Maciej Fijałkowski, and Armin Rigo. Tracing the Meta-Level: PyPy's Tracing JIT Compiler. ICOOOLPS 2009, pp. 18–25, 2009. doi:10.1145/1565824.1565827
    Why and when: Core reading. Meta-tracing: trace the interpreter, not the program, with hints marking the dispatch loop — a practical, dynamic form of the first Futamura projection (Lesson 0.3).
    Cited in: 03-aot-jit-transpilers

  • [BDB00] Vasanth Bala, Evelyn Duesterwald, and Sanjeev Banerjia. Dynamo: A Transparent Dynamic Optimization System. PLDI 2000, pp. 1–12, 2000. doi:10.1145/349299.349303
    Why and when: Hot-path (trace) selection and optimization of running native code; the origin of the tracing idea that TraceMonkey and LuaJIT apply to language VMs (Lesson 0.3).
    Cited in: 03-aot-jit-transpilers

  • [Bel73] James R. Bell. Threaded Code. Communications of the ACM 16(6), pp. 370–372, 1973. doi:10.1145/362248.362270
    Why and when: Core reading. Three pages that introduce threaded code: the program as a list of routine addresses, each routine jumping to the next. The origin of Definition 0.2.8's direct threading.
    Cited in: 02-interpreters

  • [Bra61] Harvey Bratman. An Alternate Form of the "UNCOL Diagram". Communications of the ACM 4(3), p. 142, 1961. doi:10.1145/366199.366249
    Why and when: One page that introduces the T-shaped diagram for compilers (source, target, implementation language) — the notation of Lesson 0.5.
    Cited in: 05-bootstrapping-and-trust

  • [DS84] L. Peter Deutsch and Allan M. Schiffman. Efficient Implementation of the Smalltalk-80 System. POPL 1984, pp. 297–302, 1984. doi:10.1145/800017.800542
    Why and when: Dynamic translation of Smalltalk bytecode to native code on first use, and inline caches: the origin of the method JIT in Lesson 0.3.
    Cited in: 03-aot-jit-transpilers

  • [EG03] M. Anton Ertl and David Gregg. The Structure and Performance of Efficient Interpreters. Journal of Instruction-Level Parallelism 5, 2003. pdf
    Why and when: Core reading. Measures how indirect-branch mispredictions dominate interpreter time and why threaded code predicts better than switch dispatch — the model of Definition 0.2.9. Read after Lesson 0.2 §3's dispatch trace.
    Cited in: overview, 02-interpreters

  • [EG03b] M. Anton Ertl and David Gregg. Optimizing Indirect Branch Prediction Accuracy in Virtual Machine Interpreters. PLDI 2003, 2003. doi:10.1145/781131.781162
    Why and when: Replicated handlers and superinstructions to improve prediction (Lesson 0.2 §6). Read after the lab: try replicating the load handler and measure.
    Cited in: 02-interpreters

  • [Ers58] Andrei P. Ershov. On Programming of Arithmetic Operations. Communications of the ACM 1(8), pp. 3–6, 1958. doi:10.1145/368892.368907
    Why and when: Origin of the numbering that gives the minimum number of registers (or stack slots) to evaluate an expression tree — Theorem 0.2.14.
    Cited in: 02-interpreters

  • [ES70] Jay Earley and Howard Sturgis. A Formalism for Translator Interactions. Communications of the ACM 13(10), pp. 607–617, 1970. doi:10.1145/355598.362740
    Why and when: Extends T-diagrams to interpreters and machines and gives composition rules — the algebra of Theorems 0.5.4 and 0.5.5.
    Cited in: 05-bootstrapping-and-trust

  • [FL87] Marc Feeley and Guy Lapalme. Using Closures for Code Generation. Computer Languages 12(1), pp. 47–66, 1987. doi:10.1016/0096-0551(87)90012-9
    Why and when: The origin of closure compilation: each syntax-tree node is turned once into a host-language closure that calls its children's closures, removing the per-node dispatch of a tree walker (Lesson 0.2 §6). Read §2–3 after Lesson 0.2's tree-walker section.
    Cited in: 02-interpreters

  • [Fut71] Yoshihiko Futamura. Partial Evaluation of Computation Process — An Approach to a Compiler-Compiler. Systems, Computers, Controls 2(5), pp. 45–50; reprinted in Higher-Order and Symbolic Computation 12(4), 1999, pp. 381–391, 1971. doi:10.1023/A:1010095604496
    Why and when: Core reading. The three projections relating interpreters, compilers and compiler generators through a partial evaluator (Theorem 0.3.15). The DOI is the 1999 reprint; read it after Lesson 0.3 §4.
    Cited in: overview, 03-aot-jit-transpilers

  • [Gal+09] Andreas Gal, Brendan Eich, Mike Shaver, David Anderson, David Mandelin, Mohammad R. Haghighat, Blake Kaplan, Graydon Hoare, Boris Zbarsky, Jason Orendorff, Jesse Ruderman, Edwin W. Smith, Rick Reitmaier, Michael Bebenita, Mason Chang, and Michael Franz. Trace-based Just-in-Time Type Specialization for Dynamic Languages. PLDI 2009, pp. 465–478, 2009. doi:10.1145/1542476.1542528
    Why and when: TraceMonkey: recording type-specialized traces with guards and trace trees in a JavaScript engine — Algorithm 0.3.5 in production.
    Cited in: 03-aot-jit-transpilers

  • [HCU92] Urs Hölzle, Craig Chambers, and David Ungar. Debugging Optimized Code with Dynamic Deoptimization. PLDI 1992, pp. 32–43, 1992. doi:10.1145/143095.143114
    Why and when: Introduces deoptimization (reconstructing unoptimized frames from optimized ones) in Self — the mechanism behind Definition 0.3.6 and Theorem 0.3.13.
    Cited in: 03-aot-jit-transpilers

  • [HPHF14] Matthew A. Hammer, Khoo Yit Phang, Michael Hicks, and Jeffrey S. Foster. Adapton: Composable, Demand-Driven Incremental Computation. PLDI 2014, 2014. doi:10.1145/2594291.2594324
    Why and when: Demand-driven incremental computation with dependency graphs, the research lineage of salsa and rustc's query system (Lesson 0.1, query-based compilation).
    Cited in: 01-pipeline-shapes

  • [HU94] Urs Hölzle and David Ungar. Optimizing Dynamically-Dispatched Calls with Run-Time Type Feedback. PLDI 1994, 1994. doi:10.1145/178243.178478
    Why and when: Type feedback from the running program drives recompilation: the adaptive-optimization core of tiered compilation (Lesson 0.3).
    Cited in: 03-aot-jit-transpilers

  • [IdFC05] Roberto Ierusalimschy, Luiz Henrique de Figueiredo, and Waldemar Celes. The Implementation of Lua 5.0. Journal of Universal Computer Science 11(7), pp. 1159–1176, 2005. pdf
    Why and when: Why Lua moved from a stack VM to a register VM, with the instruction formats; the most readable production register-VM design (Lesson 0.2). Read §7 on the virtual machine.
    Cited in: 02-interpreters

  • [KMRS88] Anna R. Karlin, Mark S. Manasse, Larry Rudolph, and Daniel D. Sleator. Competitive Snoopy Caching. Algorithmica 3, pp. 79–119, 1988. doi:10.1007/BF01762111
    Why and when: The competitive analysis behind the rent-or-buy (ski-rental) argument; Theorem 0.3.14 applies it to "interpret or compile". Read the rent-or-buy part only.
    Cited in: 03-aot-jit-transpilers

  • [KWM+08] Thomas Kotzmann, Christian Wimmer, Hanspeter Mössenböck, Thomas Rodriguez, Kenneth Russell, and David Cox. Design of the Java HotSpot Client Compiler for Java 6. ACM Transactions on Architecture and Code Optimization 5(1), Article 7, 2008. doi:10.1145/1369396.1370017
    Why and when: HotSpot's C1, the fast lower tier (levels 1–3 in -XX:+PrintCompilation), with its SSA-based HIR and linear-scan register allocation.
    Cited in: 03-aot-jit-transpilers

  • [LA04] Chris Lattner and Vikram Adve. LLVM: A Compilation Framework for Lifelong Program Analysis & Transformation. CGO 2004, pp. 75–86, 2004. doi:10.1109/CGO.2004.1281665
    Why and when: Core reading. The original LLVM paper: a typed SSA IR shared across compile, link and run time. Read §2–3 after Lesson 0.4 to see which design goals of 2004 survive in LLVM 23.
    Cited in: 04-retargeting

  • [LAB+21] Chris Lattner, Mehdi Amini, Uday Bondhugula, Albert Cohen, Andy Davis, Jacques Pienaar, River Riddle, Tatiana Shpeisman, Nicolas Vasilache, and Oleksandr Zinenko. MLIR: Scaling Compiler Infrastructure for Domain Specific Computation. CGO 2021, pp. 2–14, 2021. doi:10.1109/CGO51591.2021.9370308
    Why and when: Core reading. The MLIR paper: dialects, operations with regions, progressive lowering and the rationale for a multi-level IR. The origin of the multi-level IR shape in Lesson 0.1.
    Cited in: 01-pipeline-shapes

  • [McC60] John McCarthy. Recursive Functions of Symbolic Expressions and Their Computation by Machine, Part I. Communications of the ACM 3(4), pp. 184–195, 1960. doi:10.1145/367177.367199
    Why and when: LISP and its eval: the first tree-walking interpreter, defined as a program over the program's own syntax. Origin of the tree-walking technique in Lesson 0.2.
    Cited in: 02-interpreters

  • [MMP18] Andrey Mokhov, Neil Mitchell, and Simon Peyton Jones. Build Systems à la Carte. Proc. ACM Program. Lang. 2 (ICFP), Article 79, 2018. doi:10.1145/3236774
    Why and when: A precise framework for build systems (and, by extension, query-based compilers): verifying traces and early cutoff are exactly Algorithm 0.1.12 and Theorem 0.1.19. Read §3–4 after the lesson's query section.
    Cited in: 01-pipeline-shapes

  • [PH90] Karl Pettis and Robert C. Hansen. Profile Guided Code Positioning. PLDI 1990, 1990. doi:10.1145/93542.93550
    Why and when: Profile-driven layout of functions and basic blocks: the classic profile-guided AOT optimization (Algorithm 0.3.2).
    Cited in: 03-aot-jit-transpilers

  • [PVC01] Michael Paleczny, Christopher Vick, and Cliff Click. The Java HotSpot Server Compiler. USENIX Java Virtual Machine Research and Technology Symposium (JVM '01), 2001. pdf
    Why and when: HotSpot's C2: an optimizing JIT with speculation and uncommon traps — the top tier in the HotSpot box of Lesson 0.3.
    Cited in: 03-aot-jit-transpilers

  • [RSS15] Erven Rohou, Bharath Narasimha Swamy, and André Seznec. Branch Prediction and the Performance of Interpreters — Don't Trust Folklore. CGO 2015, pp. 103–114, 2015. doi:10.1109/CGO.2015.7054191
    Why and when: Shows that modern predictors (ITTAGE-like, with global history) predict switch dispatch far better than the last-target model of Definition 0.2.9 assumes; a necessary correction to 2000s folklore. Read after measuring the lab's switch vs threaded VMs.
    Cited in: 02-interpreters

  • [SGBE05] Yunhe Shi, David Gregg, Andrew Beatty, and M. Anton Ertl. Virtual Machine Showdown: Stack Versus Registers. VEE 2005, pp. 153–163, 2005. doi:10.1145/1064979.1065001 · pdf
    Why and when: Translates JVM stack code to register code; the abstract reports that more than 47% of executed VM instructions are eliminated, code grows by roughly 25%, and execution time with switch dispatch on a Pentium 4 drops by 32.3% — the empirical basis of Lesson 0.2's register-VM comparison (the 2008 TACO journal version reports 46% and 26%). Read the abstract and the dynamic instruction counts in the evaluation.
    Cited in: overview, 02-interpreters

  • [SU70] Ravi Sethi and Jeffrey D. Ullman. The Generation of Optimal Code for Arithmetic Expressions. Journal of the ACM 17(4), pp. 715–728, 1970. doi:10.1145/321607.321620
    Why and when: Proves optimality of Ershov-style evaluation order for register machines; the rigorous background of Theorem 0.2.14 (and of instruction selection in Ch 21).
    Cited in: 02-interpreters

  • [SWT+58] J. Strong, J. Wegstein, A. Tritter, J. Olsztyn, O. Mock, and T. Steel. The Problem of Programming Communication with Changing Machines: A Proposed Solution. Communications of the ACM 1(8), pp. 12–18, 1958. doi:10.1145/368892.368915
    Why and when: The UNCOL proposal: one universal intermediate language turns m × n translators into m + n (Theorem 0.4.2). A historical read for Lesson 0.4 §1.
    Cited in: 04-retargeting

  • [Tho84] Ken Thompson. Reflections on Trusting Trust. Communications of the ACM 27(8), pp. 761–763, 1984. doi:10.1145/358198.358210
    Why and when: Core reading. The Turing Award lecture describing a compiler binary that backdoors login and reinserts itself when compiling the compiler (Definition 0.5.9). Three pages; read it in full.
    Cited in: overview, 05-bootstrapping-and-trust

  • [VPF+24] Alexa VanHattum, Monica Pardeshi, Chris Fallin, Adrian Sampson, and Fraser Brown. Lightweight, Modular Verification for WebAssembly-to-Native Instruction Selection. ASPLOS 2024, 2024. doi:10.1145/3617232.3624862
    Why and when: Verifies Cranelift's ISLE lowering rules with an SMT solver; the full version of Theorem 0.4.10's argument that rule-based lowering is correct when each rule is. Read after Lesson 0.4's Cranelift section.
    Cited in: 04-retargeting

  • [Whe05] David A. Wheeler. Countering Trusting Trust through Diverse Double-Compiling. Annual Computer Security Applications Conference (ACSAC 2005), 2005. doi:10.1109/CSAC.2005.17
    Why and when: Introduces diverse double-compiling and demonstrates it on the Tiny C Compiler — the experiment the Lesson 0.5 box reproduces with gcc and clang.
    Cited in: 05-bootstrapping-and-trust

  • [Wir71] Niklaus Wirth. The Design of a PASCAL Compiler. Software: Practice and Experience 1(4), pp. 309–333, 1971. doi:10.1002/spe.4380010403
    Why and when: Wirth's one-pass Pascal compiler and the language-design choices (declare before use) that make one-pass compilation possible — Definition 0.1.9. Read after Lesson 0.1 §2.
    Cited in: 01-pipeline-shapes

  • [WWH+17] Thomas Würthinger, Christian Wimmer, Christian Humer, Andreas Wöß, Lukas Stadler, Chris Seaton, Gilles Duboscq, Doug Simon, and Matthias Grimmer. Practical Partial Evaluation for High-Performance Dynamic Language Runtimes. PLDI 2017, 2017. doi:10.1145/3062341.3062381
    Why and when: Truffle/Graal: self-specializing AST interpreters turned into machine code by partial evaluation at run time, with deoptimization. The first Futamura projection in production (Lessons 0.2 §6 and 0.3).
    Cited in: 02-interpreters, 03-aot-jit-transpilers

  • [Zak11] Alon Zakai. Emscripten: An LLVM-to-JavaScript Compiler. SPLASH '11 Companion (OOPSLA), pp. 301–312, 2011. doi:10.1145/2048147.2048224
    Why and when: Transpiling LLVM IR to JavaScript, including the relooper that rebuilds structured control flow from a CFG — a transpiler whose target lacks goto (Lesson 0.3).
    Cited in: 03-aot-jit-transpilers

Textbooks and monographs

  • [Dragon2] Alfred V. Aho, Monica S. Lam, Ravi Sethi, and Jeffrey D. Ullman. Compilers: Principles, Techniques, and Tools, 2nd ed.. Addison-Wesley, 2006. Read: §1.1–1.2 (language processors, the structure of a compiler), §2.8 (intermediate code generation), §8.10 (optimal code generation for expressions: Ershov numbers).
    Why and when: Core reading. The classical picture of a compiler as a sequence of phases, which Lesson 0.1 starts from; §2.8 introduces stack and three-address code (Lesson 0.2) and §8.10 proves the Ershov-number optimality result behind Theorem 0.2.14. Read §1.2 before Lesson 0.1.
    Cited in: overview, 02-interpreters

  • [EaC3] Keith D. Cooper and Linda Torczon. Engineering a Compiler, 3rd ed.. Morgan Kaufmann, 2022. Read: Ch. 1 (Overview of compilation).
    Why and when: A gentler overview of front end, optimizer and back end than the Dragon book, with the engineering trade-offs this chapter's comparison tables are about; read Ch. 1 alongside the chapter README before the lessons.
    Cited in: overview

  • [JGS93] Neil D. Jones, Carsten K. Gomard, and Peter Sestoft. Partial Evaluation and Automatic Program Generation. Prentice Hall, 1993. Read: Ch. 1 (Introduction: specializers, interpreters, compilers and the Futamura projections), Ch. 4 (Partial evaluation for a flow chart language). link
    Why and when: The standard text on partial evaluation, freely available. Ch. 1 states the Futamura projections exactly as Theorem 0.3.15 and discusses when the compiled programs are actually fast; Ch. 4 builds a self-applicable specializer. Read after Lesson 0.3 §4.
    Cited in: 03-aot-jit-transpilers

  • [Lat11] Chris Lattner. LLVM, in: The Architecture of Open Source Applications, Vol. 1 (A. Brown, G. Wilson, eds.). aosabook.org (lulu.com), 2011. Read: §11.1 (classical compiler design), §11.3 (LLVM IR), §11.4 (LLVM's implementation of the three-phase design). link
    Why and when: Lattner's own account of the three-phase design and of why LLVM IR is the only interface between front end, optimizer and back end. The best short companion to Lesson 0.4; read it right after the lesson's §1.
    Cited in: overview, 04-retargeting

  • [Lev00] John R. Levine. Linkers and Loaders. Morgan Kaufmann, 2000. Read: Ch. 3 (Object files), Ch. 5 (Symbol management), Ch. 7 (Relocation), Ch. 10 (Dynamic linking and loading).
    Why and when: The standard book on the tools of Lesson 0.6. Ch. 5 and 7 are the long versions of Algorithm 0.6.6 (resolution, relocation); Ch. 10 explains PLT/GOT and lazy binding.
    Cited in: 06-toolchain

  • [Nys21] Robert Nystrom. Crafting Interpreters. Genever Benning, 2021. Read: Part II, ch. 4–13 (a tree-walk interpreter, jlox); Part III, ch. 14–30 (a bytecode virtual machine, clox). link
    Why and when: A book-length build of exactly the two interpreter designs of Lesson 0.2 (tree walker, then stack VM with a switch loop), freely readable online. Useful while doing lab milestones L2–L4.
    Cited in: 02-interpreters

Surveys and tutorials

  • [Ayc03] John Aycock. A Brief History of Just-In-Time. ACM Computing Surveys 35(2), pp. 97–113, 2003. doi:10.1145/857076.857077
    Why and when: Survey of JIT compilation from the 1960s to Java: method JITs, mixed-mode execution, and the terminology of Lesson 0.3. Read before the lesson for the history.
    Cited in: 03-aot-jit-transpilers

Theses and technical reports

  • [Whe09] David A. Wheeler. Fully Countering Trusting Trust through Diverse Double-Compiling. PhD dissertation, George Mason University, 2009. link
    Why and when: Core reading. The formal proof of DDC (Theorem 0.5.16) with all assumptions made explicit, and demonstrations on four compilers (a small C compiler, a small Lisp compiler, a trojaned Lisp compiler, GCC). Also arXiv:1004.5534. Read the introduction and the proof's assumptions after Lesson 0.5.
    Cited in: 05-bootstrapping-and-trust

Source code (pinned versions)

  • [CL-Src] Cranelift's x86-64 lowering rules in ISLE — cranelift/codegen/src/isa/x64/lower.isle in bytecodealliance/wasmtime at v37.0.2. Symbols: iadd_base_case_32_or_64_lea.
    Why and when: Real ISLE rules with priorities (Algorithm 0.4.7); cranelift/codegen/src/egraph.rs (EgraphPass) is the e-graph mid-end.
    Cited in: 04-retargeting

  • [CLANG-Driver] The clang driver — builds the phase graph (-ccc-print-phases) and the jobs (-###) — clang/lib/Driver/Driver.cpp in llvm/llvm-project at llvmorg-23.1.2. Symbols: Driver::BuildActions, Driver::BuildJobs, Driver::PrintActions.
    Why and when: Read BuildActions after Lesson 0.1's first box to see how a command line becomes preprocess/compile/backend/assemble/link actions; clang/lib/Driver/ToolChains/Gnu.cpp adds the crt files of Lesson 0.6.
    Cited in: 01-pipeline-shapes, 05-bootstrapping-and-trust, 06-toolchain

  • [CLANG-Lex] Clang's numeric-literal parser (lexer side) — clang/lib/Lex/LiteralSupport.cpp in llvm/llvm-project at llvmorg-23.1.2. Symbols: NumericLiteralParser::NumericLiteralParser.
    Why and when: Reports invalid suffix 'x' on integer constant — a lexical error in the phases drill; clang/lib/Lex/PPDirectives.cpp implements #include/#define/#if (Lesson 0.6).
    Cited in: 06-toolchain

  • [CPY-Ceval] CPython's bytecode interpreter loop — Python/ceval.c in python/cpython at v3.11.15. Symbols: _PyEval_EvalFrameDefault, USE_COMPUTED_GOTOS.
    Why and when: A production stack VM with token-threaded dispatch through Python/opcode_targets.h when the compiler supports computed goto (Lesson 0.2).
    Cited in: 02-interpreters

  • [CPY-Spec] CPython 3.11's specializing adaptive interpreter (PEP 659) — Python/specialize.c in python/cpython at v3.11.15. Symbols: _PyCode_Quicken, _Py_Specialize_BinaryOp.
    Why and when: Quickening and specialization (BINARY_OP_ADD_INT) seen in Lesson 0.2's dis box.
    Cited in: 02-interpreters

  • [GCC-i386md] GCC's x86 machine description (RTL instruction patterns) — gcc/config/i386/i386.md in gcc-mirror/gcc at releases/gcc-15.1.0. Symbols: *add<mode>_1<nf_name>.
    Why and when: The target description GCC's back end is generated from; the add patterns carry the flags clobber visible in Lesson 0.4's RTL dump (which was produced by gcc 13.3, where the pattern is named *add<mode>_1; gcc 15 adds the APX <nf_name> suffix).
    Cited in: 04-retargeting

  • [GCC-Passes] GCC's pass list (GIMPLE, IPA and RTL passes in order) — gcc/passes.def in gcc-mirror/gcc at releases/gcc-15.1.0. Symbols: pass_build_ssa_passes, pass_expand.
    Why and when: GCC's pipeline as data: every NEXT_PASS line is one pass; pass_expand is the GIMPLE-to-RTL boundary of Lesson 0.4. gcc/passes.cc runs the list.
    Cited in: 01-pipeline-shapes, 03-aot-jit-transpilers, 04-retargeting

  • [GO-Dist] Go's bootstrap driver (toolchain1, toolchain2, toolchain3) — src/cmd/dist/build.go in golang/go at go1.24.7. Symbols: cmdbootstrap.
    Why and when: The multi-stage self-hosting build of Lesson 0.5 as real code; src/cmd/dist/buildtool.go sets minBootstrap.
    Cited in: 05-bootstrapping-and-trust

  • [HS-Interp] HotSpot's template interpreter — src/hotspot/share/interpreter/templateInterpreter.cpp in openjdk/jdk at jdk-21+35.
    Why and when: The JVM's stack-bytecode interpreter, whose handlers are machine-code templates generated at VM startup (Lesson 0.2 §7).
    Cited in: 02-interpreters

  • [HS-Tiered] HotSpot's tiered compilation policy (tier transitions 0–4, OSR) — src/hotspot/share/compiler/compilationPolicy.cpp in openjdk/jdk at jdk-21+35. Symbols: CompilationPolicy::event, CompilationPolicy::call_event, CompilationPolicy::common.
    Why and when: Where HotSpot decides to move a method between the interpreter, C1 levels and C2 — the policy behind the -XX:+PrintCompilation levels in Lesson 0.3.
    Cited in: 03-aot-jit-transpilers

  • [LLD-ELF] lld's symbol resolution (strong, weak, lazy archive members, shared symbols) — lld/ELF/Symbols.cpp in llvm/llvm-project at llvmorg-23.1.2. Symbols: Symbol::resolve, elf::reportDuplicate.
    Why and when: The production version of Algorithm 0.6.6's Add; continue with lld/ELF/InputSection.cpp (InputSectionBase::getRelocTargetVA: S + A − P) and lld/ELF/Arch/X86_64.cpp (X86_64::relocate).
    Cited in: 06-toolchain

  • [LLD-GC] lld's --gc-sections, a mark-sweep collector over input sections — lld/ELF/MarkLive.cpp in llvm/llvm-project at llvmorg-23.1.2. Symbols: elf::markLive.
    Why and when: The file comment explains section garbage collection in a dozen lines: sections reachable from GC roots (the entry symbol, exported symbols) through relocations are kept, the rest dropped. Identical code folding is lld/ELF/ICF.cpp (elf::doIcf). Read after Lesson 0.6 §6.
    Cited in: 06-toolchain

  • [LLVM-CodeGen] The machine-code pipeline shared by LLVM back ends — llvm/lib/CodeGen/TargetPassConfig.cpp in llvm/llvm-project at llvmorg-23.1.2. Symbols: TargetPassConfig::addMachinePasses.
    Why and when: The back-end phase of the three-phase design: instruction selection, register allocation, scheduling and emission, customized per target (Chapters 21–23).
    Cited in: 04-retargeting

  • [LLVM-Interp] LLVM's IR interpreter (lli -force-interpreter) — llvm/lib/ExecutionEngine/Interpreter/Execution.cpp in llvm/llvm-project at llvmorg-23.1.2. Symbols: Interpreter::run, Interpreter::visitBinaryOperator.
    Why and when: A walker over LLVM IR with one visit method per instruction kind — a real tree/graph-walking interpreter, 110× slower than LLVM's JIT in Lesson 0.2's box.
    Cited in: 02-interpreters

  • [LLVM-JIT] ORC's LLJIT — the method JIT the lab uses — llvm/lib/ExecutionEngine/Orc/LLJIT.cpp in llvm/llvm-project at llvmorg-23.1.2. Symbols: LLJITBuilderState::prepareForConstruction, LLJIT::addIRModule, LLJIT::lookupLinkerMangled.
    Why and when: What happens between addIRModule and a callable function pointer in the lab's JIT engine; prepareForConstruction picks defaults for the host.
    Cited in: 03-aot-jit-transpilers

  • [LLVM-MC] LLVM's integrated assembler — ELF object writer — llvm/lib/MC/ELFObjectWriter.cpp in llvm/llvm-project at llvmorg-23.1.2. Symbols: ELFWriter, ELFWriter::computeSymbolTable, ELFWriter::writeSectionHeaders.
    Why and when: Where sections, symbols and relocations of Lesson 0.6's main.o are written; llvm/lib/MC/MCParser/AsmParser.cpp parses assembly text and inline asm.
    Cited in: 06-toolchain

  • [LLVM-Pipelines] The -O0/-O1/-O2/-O3 pass pipelines of LLVM's new pass manager — llvm/lib/Passes/PassBuilderPipelines.cpp in llvm/llvm-project at llvmorg-23.1.2. Symbols: PassBuilder::buildPerModuleDefaultPipeline, PassBuilder::buildO0DefaultPipeline.
    Why and when: The composition of passes that default<O2> means (Lesson 0.1's box prints it); come back in Ch 12 and Ch 24 when you build Pebble's own pipeline.
    Cited in: 01-pipeline-shapes, 03-aot-jit-transpilers

  • [LLVM-TargetRegistry] The registry of back ends that makes LLVM retargetable — llvm/include/llvm/MC/TargetRegistry.h in llvm/llvm-project at llvmorg-23.1.2. Symbols: TargetRegistry::lookupTarget, Target::createTargetMachine.
    Why and when: Algorithm 0.4.4's registry: each back end registers itself; a triple selects it. Read with llvm/tools/llc/llc.cpp.
    Cited in: 04-retargeting

  • [LUA-Src] Lua 5.4's register-based VM — lvm.c in lua/lua at v5.4.6. Symbols: luaV_execute, vmdispatch.
    Why and when: luaV_execute is a switch-dispatched register VM (lopcodes.h documents the iABC format); ljumptab.h provides a computed-goto variant.
    Cited in: 02-interpreters

  • [LUAJIT-Src] LuaJIT's trace recorder entry points and side exits — src/lj_trace.c in LuaJIT/LuaJIT at v2.1. Symbols: lj_trace_hot, lj_trace_ins, lj_trace_exit.
    Why and when: Hot-loop detection, recording and side exits (Algorithm 0.3.5); src/lj_record.c (lj_record_ins) records each bytecode.
    Cited in: 03-aot-jit-transpilers

  • [MLIR-DialectConversion] MLIR's dialect-conversion driver — mlir/lib/Transforms/Utils/DialectConversion.cpp in llvm/llvm-project at llvmorg-23.1.2. Symbols: mlir::applyPartialConversion, OperationLegalizer.
    Why and when: The implementation of Algorithm 0.1.14; mlir/lib/Conversion/SCFToControlFlow/SCFToControlFlow.cpp (ForLowering) is the pattern that lowered scf.for in Lesson 0.1's box.
    Cited in: 01-pipeline-shapes

  • [PYPY-Src] PyPy's meta-interpreter (the meta-tracing JIT) — rpython/jit/metainterp/pyjitpl.py in pypy/pypy at release-pypy3.10-v7.3.17. Symbols: MetaInterp.
    Why and when: Traces the RPython interpreter; rpython/rlib/jit.py defines JitDriver and jit_merge_point, the hints that mark the user program's loops (Lesson 0.3).
    Cited in: 03-aot-jit-transpilers

  • [RUSTC-DepGraph] rustc's dependency graph and red-green marking — compiler/rustc_query_system/src/dep_graph/graph.rs in rust-lang/rust at 1.94.1. Symbols: DepGraph, DepGraph::try_mark_green.
    Why and when: try_mark_green is Algorithm 0.1.12's TryMarkGreen; read it after the query worked example in Lesson 0.1.
    Cited in: 01-pipeline-shapes

  • [TCC-Src] TCC's parser-and-code-generator in one — tccgen.c in TinyCC/tinycc at release_0_9_27. Symbols: gen_op, vpushi, block.
    Why and when: The single pass of Lesson 0.1: block parses a statement and emits its code, gen_op emits an operator as soon as its operands are known.
    Cited in: 01-pipeline-shapes

  • [V8-Src] V8's tiering decisions (Ignition → Sparkplug → Maglev → TurboFan) — src/execution/tiering-manager.cc in v8/v8 at 12.4.254.21. Symbols: TieringManager::MaybeOptimizeFrame.
    Why and when: Where V8 decides a function is "hot and stable"; src/baseline/baseline-compiler.cc (Sparkplug), src/maglev/maglev-compiler.cc and src/compiler/pipeline.cc (TurboFan) are the tiers.
    Cited in: 03-aot-jit-transpilers

Official documentation and specifications

  • [C-Std] ISO/IEC 9899 (C) working draft N3220. C2y working draft, 2024. link
    Why and when: §6.10 specifies preprocessing directives and macro replacement, including rescanning and the rule that a macro is not replaced inside its own expansion (the hide sets of Lesson 0.6).
    Cited in: 06-toolchain

  • [CL-Docs] Cranelift IR reference (cranelift/docs/ir.md) and project documentation. Wasmtime v37.0.2. link
    Why and when: CLIF: SSA with block parameters, types, instructions. Read it with the CLIF box of Lesson 0.4 open.
    Cited in: 04-retargeting

  • [CLANG-Modules] Standard C++ Modules (Clang documentation). LLVM 23.1.2. link
    Why and when: How clang compiles C++20 module interfaces once into BMI files that importers load instead of re-preprocessing headers — the "modules" variant of Lesson 0.6 §6. Read the introduction and the "Background and terminology" section.
    Cited in: 06-toolchain

  • [CLANG-PCH] Precompiled Header and Modules Internals (Clang documentation). LLVM 23.1.2. link
    Why and when: What a precompiled header stores (the serialized AST after a common include prefix) and how clang loads it lazily; the end-user view is the "Precompiled Headers" section of clang/docs/UsersManual.md. Read after Lesson 0.6 §6.
    Cited in: 06-toolchain

  • [Dre11] How To Write Shared Libraries (Ulrich Drepper). 4.1.2, 2011. link
    Why and when: The glibc maintainer's explanation of dynamic linking: symbol lookup order, relocations processed at startup, PLT/GOT, lazy binding and its cost. The source for Algorithm 0.6.8.
    Cited in: 06-toolchain

  • [ELF] Tool Interface Standard (TIS) Executable and Linking Format (ELF) Specification, Version 1.2. link
    Why and when: The ELF container: sections, symbol tables, relocations, program headers. Keep open while reading the llvm-readelf output in Lesson 0.6.
    Cited in: 06-toolchain

  • [GCC-Install] Installing GCC: Building (gcc/doc/install.texi) — 3-stage bootstrap and stage comparison. gcc-15.1.0. link
    Why and when: States the default native build: a 3-stage bootstrap followed by "a comparison test of the stage2 and stage3 compilers" — Algorithm 0.5.8 in GCC's words.
    Cited in: 05-bootstrapping-and-trust

  • [GCC-Int] GNU Compiler Collection (GCC) Internals — Passes and Files of the Compiler; GENERIC; GIMPLE; RTL. GCC 15. link
    Why and when: GCC's three IRs and its pass structure, from the maintainers. Read the "Passes and Files of the Compiler" chapter after Lesson 0.4's GCC section. The manual is generated from gcc/doc/passes.texi, generic.texi, gimple.texi and rtl.texi in the GCC source tree (e.g. the gcc-mirror/gcc tag releases/gcc-15.1.0), a mirror if the web page is unreachable.
    Cited in: 04-retargeting

  • [LLVM-AdvBuilds] Advanced Build Configurations: Bootstrap Builds. LLVM 23.1.2. link
    Why and when: CLANG_ENABLE_BOOTSTRAP, ninja stage2, and the stage3 build that "should be bit-for-bit identical" to stage2 (Theorem 0.5.14). Read when you build LLVM yourself.
    Cited in: 05-bootstrapping-and-trust

  • [LLVM-NPM] Using the New Pass Manager. LLVM 23.1.2. link
    Why and when: How LLVM's pass manager caches analyses and how passes report PreservedAnalyses — Algorithm 0.1.8 in LLVM's vocabulary; essential again in Ch 12.
    Cited in: 01-pipeline-shapes

  • [LLVM-ORC] ORC Design and Implementation (ORCv2). LLVM 23.1.2. link
    Why and when: The JIT APIs the lab uses: LLJIT, ThreadSafeModule, symbol lookup, lazy compilation. Read before milestone L5 of the lab.
    Cited in: 03-aot-jit-transpilers

  • [MLIR-Conv] MLIR Dialect Conversion. LLVM 23.1.2. link
    Why and when: Conversion targets, legality and rewrite patterns — the documentation of Algorithm 0.1.14.

  • [ROSLYN-Overview] .NET Compiler Platform (Roslyn) Overview — immutable syntax trees and compilations. dotnet/roslyn commit e0882714833e0943485161b1a01d52e835baf3eb. link
    Why and when: Roslyn's compiler-as-API design: immutable, snapshot-based syntax trees and compilations that share unchanged parts between versions (red-green trees; Lesson 0.1 §6).
    Cited in: 01-pipeline-shapes

  • [RUSTC-Query] rustc dev guide: Queries — demand-driven compilation; Incremental compilation (red-green algorithm). link
    Why and when: How rustc is organized as queries and how the red-green algorithm decides what to reuse (Algorithm 0.1.12). Read both pages after Lesson 0.1's query section.
    Cited in: 01-pipeline-shapes

  • [Salsa] Salsa — a generic framework for on-demand, incrementalized computation. link
    Why and when: The query engine of rust-analyzer: tracked functions, revisions, durability, early cutoff. The overview chapter of its book is a compact second explanation of Algorithm 0.1.12.
    Cited in: 01-pipeline-shapes

  • [SWIFT-Req] Swift Request-Evaluator (docs/RequestEvaluator.md). swift-6.1-RELEASE. link
    Why and when: Swift's incremental adoption of demand-driven requests with cycle detection inside a traditional compiler (Lesson 0.1 §6); implemented by class Evaluator in include/swift/AST/Evaluator.h.
    Cited in: 01-pipeline-shapes

  • [SysV-ABI] System V Application Binary Interface, AMD64 Architecture Processor Supplement. link
    Why and when: Defines the x86-64 relocation types, the calling convention and the PLT/GOT layout used in Lesson 0.6 §3. The "Relocation Types" table of the Object Files chapter (x86-64-ABI/object-files.tex) gives R_X86_64_PC32 = S + A − P and R_X86_64_PLT32 = L + A − P, where L is the PLT entry — S itself when the symbol is defined in the output.
    Cited in: 06-toolchain

  • [TCC-Doc] Tiny C Compiler Reference Documentation (tcc-doc.texi), "Parser" section. tcc 0.9.27. link
    Why and when: "The parser is hardcoded … It does only one pass, except" for two cases — the primary source for TCC's single-pass design (Lesson 0.1).
    Cited in: 01-pipeline-shapes

  • [TS-Handbook] TypeScript Handbook and compiler options (target, downlevel emit). link
    Why and when: What --target does: which constructs are erased and which are downleveled — the transpiler box of Lesson 0.3.
    Cited in: 03-aot-jit-transpilers

Blog posts and articles

  • [V8-Ignition] Ross McIlroy. Firing up the Ignition interpreter. 2016. link
    Why and when: Why V8 introduced a register-based bytecode interpreter with an accumulator, and how its handlers are generated — the design behind Lesson 0.2's --print-bytecode box.
    Cited in: 02-interpreters

  • [V8-Maglev] Toon Verwaest, Leszek Swirski, Victor Gomes, Olivier Flückiger, Darius Mercadier, and Camillo Bruni. Maglev — V8's Fastest Optimizing JIT. 2023. link
    Why and when: The mid tier introduced in Chrome M117, between Sparkplug and TurboFan; explains why a fast optimizing tier pays off (Lesson 0.3 §6).
    Cited in: 03-aot-jit-transpilers

  • [V8-Sparkplug] Leszek Swirski. Sparkplug — a non-optimizing JavaScript compiler. 2021. link
    Why and when: V8's baseline tier: compiles Ignition bytecode to machine code in one linear pass, reusing the interpreter's frame layout. Read with Lesson 0.3's tiering box.
    Cited in: 01-pipeline-shapes, 03-aot-jit-transpilers