Lesson 10.8 — Running IR and consuming LLVM: interpreter, MCJIT, ORC, CMake¶
Techniques: the IR interpreter (
ExecutionEngineInterpreter,lli -force-interpreter), MCJIT (the legacy JIT), ORCLLJIT(layers, JITDylibs, symbol resolution,ThreadSafeModule), ORCLLLazyJIT(per-function lazy compilation), consuming LLVM from a build (llvm-config,LLVMConfig.cmake, dylib vs components, RTTI/EH flags) · Pebble uses:LLJITin Lab 10.1's tests and Ch 0's lab; the CMake scheme of this repository (cmake/PebbleLLVM.cmake) · Lab: Lab 10.3 (interpreter vs LLJIT) · Prerequisites: Lessons 10.1, 10.7; Ch 0 (interpreters vs JITs) · Time: 3 hours
You have built @square_plus in memory with IRBuilder. How do you run it? You can walk the IR with an interpreter, compile the whole module to machine code in memory and call it, or compile each function only when it is first called. And how does your program find LLVM's headers and libraries, and with which compiler flags? This lesson compares the four execution engines LLVM 23 offers, with measurements, and makes the build-integration rules precise.
1. Problem and motivation¶
Running freshly built IR is needed by test harnesses (Lab 10.1 calls your functions hundreds of times), REPLs and scripting languages, database query compilers, and GPU/ML runtimes. The engines trade startup latency against peak speed against API complexity, the same trade-off as Ch 0's interpreter vs JIT discussion, now with LLVM's concrete APIs [LLVM-ORC, BAJIT].
Interpreter¶
ExecutionEngine's Interpreter executes IR directly: a big-step loop over instructions with a value map (GenericValue) per stack frame. It has no code generation, so it starts instantly, runs slowly, supports few intrinsics and calls external functions only through a small table or libffi [LLVM-Interp].
MCJIT¶
MCJIT (2013) compiles a whole module to an object file in memory with the regular code generator, links it with RuntimeDyld, and hands out function addresses. It is the older JIT API (EngineBuilder(…).setEngineKind(EngineKind::JIT)). It is still in the tree but superseded by ORC: new features (concurrency, lazy compilation, JITLink) go to ORC [LLVM-MCJIT, LLVM-ORC].
ORC LLJIT¶
ORC (On Request Compilation) is a set of composable layers (IR transform → IR compile → object linking) on top of an ExecutionSession that owns JITDylibs, symbol tables that behave like dynamic libraries. LLJIT assembles a default stack. You add ThreadSafeModules (a module bundled with its context and a lock) and lookup symbols, and a lookup compiles the module that defines the symbol (the whole module) [LLVM-ORC, Ham16, Ham18].
ORC LLLazyJIT¶
LLLazyJIT adds two layers on top of LLJIT's: an IRPartitionLayer, which splits a module into per-function partitions, and a CompileOnDemandLayer, which replaces each function by a stub. A function's partition is compiled the first time its stub is called. Functions that never run are never compiled [LLVM-ORC, BAJIT].
Consuming LLVM from CMake¶
A client must compile against LLVM's headers with compatible flags and link the right libraries. LLVM exports this through llvm-config (a command-line oracle) and LLVMConfig.cmake (find_package(LLVM CONFIG)), which provides include directories, definitions, LLVM_ENABLE_RTTI and LLVM_ENABLE_EH, and either the single shared libLLVM or per-component static libraries [LLVM-CMake].
2. Definitions and algorithms¶
Definition 10.8.1 (Execution engine, materialization)
An execution engine maps a module \(M\) and a symbol name \(s\) to a callable entity with \(s\)'s semantics. Materializing a definition means producing its executable form: interpreter state (nothing to do), or machine code in executable memory. An engine is eager if adding \(M\) (or the first lookup in \(M\)) materializes every definition in \(M\), and lazy if it materializes each function only when the function is first executed.
Interpreter¶
Algorithm 10.8.2 (Interpreter::run, outline)
- Input: a function \(F\) and argument values \(\vec{a}\) (
GenericValues). - Output: the return value.
- Precondition: \(F\) has a body; every called intrinsic can be lowered by
IntrinsicLowering; external calls are to known library functions. - Postcondition: the result equals the IR semantics of \(F(\vec{a})\) (Ch 9).
- Invariant: the top stack frame holds (current block, next instruction, map from SSA values to
GenericValues), and every value used by the next instruction is in the map (SSA dominance).
function Run(F, a⃗):
push frame(F, entry block, {args ↦ a⃗})
while the stack is non-empty:
fr ← top; I ← fr.next; fr.next ← next(I)
visit(I) # InstVisitor<Interpreter> (Lesson 10.4)
# binary op: fr.vals[I] ← op(fr.vals[op0], fr.vals[op1])
# br: SwitchToNewBasicBlock(dest) — evaluates the dest's phis with the
# predecessor's values *simultaneously*, then continues at dest
# call: push a new frame (or call the external-function table)
# ret: pop, write the value into the caller's frame
# intrinsic: IntrinsicLowering, or report_fatal_error if unsupported
return the value of the final ret
lli: the interpreter against three JITs on the same IR
Reproduce (lli 23.1.2, Linux x86-64; bash's time):
cd chapters/10-llvm-cpp-api/examples # fib.ll: recursive fib(30); main returns fib(30) mod 256
bash -c 'TIMEFORMAT=" %R s wall"; for k in -force-interpreter -jit-kind=mcjit -jit-kind=orc -jit-kind=orc-lazy; do echo "lli $k"; time lli $k fib.ll; echo " exit code $?"; done'
Output (complete; times vary by machine and run):
lli -force-interpreter
1.521 s wall
exit code 40
lli -jit-kind=mcjit
0.051 s wall
exit code 40
lli -jit-kind=orc
0.038 s wall
exit code 40
lli -jit-kind=orc-lazy
0.025 s wall
exit code 40
What to notice: all four agree (\(\mathrm{fib}(30) = 832040\), and \(832040 \bmod 256 = 40\)). The interpreter executes \(2F(31) - 1 \approx 2.7\) million calls of @fib, each walked instruction by instruction, and is ~30–60× slower than the JITs, whose times are mostly process start-up and compilation. The orders of magnitude, not the milliseconds, are the point.
What the interpreter cannot run
Reproduce (lli 23.1.2):
cd chapters/10-llvm-cpp-api/examples
lli -force-interpreter abs.ll 2>&1 | head -1; lli abs.ll; echo "exit $?"
Output (complete):
What to notice: the interpreter lowers intrinsics through IntrinsicLowering::LowerIntrinsicCall, which knows only a few of them, and llvm.abs is not among them. The failure is report_fatal_error, which aborts the process instead of returning an Error. The JIT uses the real code generator and returns \(\max(|{-7}|, 3) = 7\). That is why Lab 10.3's runInterpreter must refuse modules declaring intrinsics before running them (R5).
MCJIT¶
Algorithm 10.8.3 (MCJIT lookup and finalization, outline)
- Input: one or more modules added to an MCJIT
ExecutionEngine, and a name \(s\). - Output: the address of \(s\) (0 if no added module defines it).
- Precondition: the native target is initialized; no module is modified after it is added.
- Postcondition: the module defining \(s\) is compiled and linked, and so is every added module that defines a symbol it (transitively) references;
getFunctionAddress(s)returns executable code.finalizeObject()compiles every added module. - Invariant: a module is either not yet compiled, or compiled whole and loaded into RuntimeDyld.
function GetFunctionAddress(EE, s): # MCJIT::getFunctionAddress
a ← FindSymbol(EE, s)
if a ≠ 0: FinalizeLoaded(EE)
return a
function FindSymbol(EE, s): # MCJIT::findSymbol
if s is already loaded: return its address
M ← the added module that defines s # findModuleForSymbol
if M = none: return 0
obj ← codegen(M) # generateCodeForModule: the whole module
RuntimeDyld.load(obj)
return address of s
function FinalizeLoaded(EE): # finalizeLoadedModules
for each unresolved external symbol u of the loaded objects:
FindSymbol(EE, u) # via LinkingSymbolResolver: may compile another module
apply relocations; set memory permissions (RW → RX)
MCJIT compiles whole modules, and the modules they reference
Reproduce (clang 23.1.2, LLVM 23.1.2, Linux x86-64):
cd chapters/10-llvm-cpp-api/examples
clang++-23 $(llvm-config --cxxflags) -std=c++23 mcjit.cpp $(llvm-config --ldflags --libs) \
-Wl,-rpath,$(llvm-config --libdir) -o mcjit
for a in 5 500; do echo "== ./mcjit m1.ll m2.ll m3.ll $a"; ./mcjit m1.ll m2.ll m3.ll $a; done
m1.ll and m2.ll are the modules of the LLLazyJIT box below, and m3.ll defines other, which nothing refers to. An ObjectCache whose notifyObjectCompiled prints each module MCJIT compiles is attached with setObjectCache.
Output (complete):
== ./mcjit m1.ll m2.ll m3.ll 5
getFunctionAddress(entry)
[compile] m1.ll: entry used never_called
[compile] m2.ll: helper unrelated
call entry(5)
result 15
== ./mcjit m1.ll m2.ll m3.ll 500
getFunctionAddress(entry)
[compile] m1.ll: entry used never_called
[compile] m2.ll: helper unrelated
call entry(500)
result 501
What to notice: everything happens inside getFunctionAddress, before the first call (Algorithm 10.8.3). m1 is compiled whole, never_called included, and m2 follows because finalization must resolve m1's reference to helper, even in the run where helper is never called. m3 is never compiled. Module granularity and reference-driven compilation are the same as LLJIT's (next box), and this is the eager behavior the jit-compile-set drill asks about.
ORC LLJIT¶
Definition 10.8.4 (JITDylib, link order, resolution)
A JITDylib \(D\) is a symbol table from names to (address or materializer, flags). Its link order is a sequence \(\langle D, D_1, \dots, D_r \rangle\) of JITDylibs (by default \(D\) first, then the process's symbols for LLJIT's main dylib, via a generator). Resolving a name \(s\) referenced from code in \(D\) finds the first \(D_i\) in the link order that defines \(s\). If that definition is not yet materialized, its materialization unit (for IR: the whole ThreadSafeModule that defines it) is run first.
Algorithm 10.8.5 (LLJIT lookup)
- Input: an
LLJIT\(J\) and a name \(s\). - Output:
Expected<ExecutorAddr>. - Precondition: every added module has been verified; names are mangled for the target (
lookupmangles). - Postcondition: \(s\)'s module is compiled and linked, and so is every module that it (transitively) references; the result is \(s\)'s address. A missing symbol returns an
Error. - Invariant: each materialization unit runs at most once (the session tracks symbol states).
function Lookup(J, s):
(D_i, def) ← first definition of s in link order of Main
if def is unmaterialized:
MU ← materialization unit of def # the whole module
TSM ← MU.module
TSM ← IRTransformLayer.transform(TSM) # optional hook (the box uses it to log)
obj ← IRCompileLayer.compile(TSM) # codegen, holding TSM's context lock
ObjectLinkingLayer.link(obj):
for each undefined symbol u in obj: Lookup(J, u) # may materialize other modules
apply relocations; set memory permissions
mark all of MU's symbols materialized
return address(s)
LLJIT: build IR, add it, look it up, call it
Reproduce (clang 23.1.2, LLVM 23.1.2, Linux x86-64):
cd chapters/10-llvm-cpp-api/examples
clang++-23 $(llvm-config --cxxflags) -std=c++23 lljit.cpp $(llvm-config --ldflags --libs) \
-Wl,-rpath,$(llvm-config --libdir) -o lljit && ./lljit
Output (complete):
triple x86_64-conda-linux-gnu
square_plus(3, 1) = 10
square_plus(-4, 1) = 17
square_plus(1000000, 1) = 1000000000001
What to notice: the program never touches a file. It moves the module and its context into a ThreadSafeModule (ownership passes to the JIT), lookup returns an ExecutorAddr, and toPtr<int64_t (*)(int64_t, int64_t)>() turns it into a callable C++ function pointer. Every step returns Expected/Error, handled with ExitOnError (Lesson 10.7). The triple is this conda-built LLVM's host triple.
ORC LLLazyJIT¶
Algorithm 10.8.6 (Per-function lazy compilation)
- Input: a module \(M\) added with
addLazyIRModule. - Output: callable stubs for \(M\)'s functions.
- Precondition: as for Algorithm 10.8.5.
- Postcondition: Theorem 10.8.10.
- Invariant: each function's stub points either to a compile callback (not yet compiled) or to the compiled body.
function AddLazy(J, M):
for each function f defined in M:
partition P_f ← a module containing f's body (+ declarations it needs)
define symbol f in Main as a reexport of a stub → compile-callback(P_f)
on first call through f's stub: # compile-callback(P_f)
compile P_f through the IR layers (Algorithm 10.8.5's layers)
update f's stub to jump to the compiled body
continue into the body
Eager vs lazy: what gets compiled, and when
Reproduce (clang 23.1.2, LLVM 23.1.2, Linux x86-64):
cd chapters/10-llvm-cpp-api/examples
clang++-23 $(llvm-config --cxxflags) -std=c++23 eager.cpp $(llvm-config --ldflags --libs) \
-Wl,-rpath,$(llvm-config --libdir) -o eager
for m in lljit lazy; do for a in 5 500; do echo "== ./eager $m m1.ll m2.ll $a"; ./eager $m m1.ll m2.ll $a; done; done
Module 1 defines entry (which calls used if \(x \le 100\), else helper), used and never_called. Module 2 defines helper and unrelated. An IRTransformLayer hook prints the functions of each module that reaches the compiler.
Output (complete):
== ./eager lljit m1.ll m2.ll 5
lookup(entry)
[compile] entry used never_called
[compile] helper unrelated
call entry(5)
result 15
== ./eager lljit m1.ll m2.ll 500
lookup(entry)
[compile] entry used never_called
[compile] helper unrelated
call entry(500)
result 501
== ./eager lazy m1.ll m2.ll 5
lookup(entry)
call entry(5)
[compile] entry
[compile] used
result 15
== ./eager lazy m1.ll m2.ll 500
lookup(entry)
call entry(500)
[compile] entry
[compile] helper
result 501
What to notice: LLJIT compiled all of module 1 at the lookup, never_called included, and then all of module 2, because module 1 references helper (Algorithm 10.8.5's recursive lookup). LLLazyJIT compiled nothing at lookup and then exactly the functions that ran, in call order (Theorem 10.8.10). The drill jit-compile-set asks for these two answers on random programs.
Consuming LLVM from CMake¶
Definition 10.8.7 (Component closure)
LLVM's static libraries form a DAG under "requires". A set \(C\) of components (core, support, orcjit, native, …; llvm-config --components lists 226 here) denotes libraries; its closure \(\overline{C}\) adds every library reachable in the DAG. A link line is complete if it contains \(\overline{C}\) in a topological order (users before dependencies). With a shared build (LLVM_LINK_LLVM_DYLIB), all components are in one library, libLLVM, and the closure is trivial.
Algorithm 10.8.8 (llvm_map_components_to_libnames + expansion, outline)
- Input: component names.
- Output: an ordered list of library targets.
- Precondition:
find_package(LLVM CONFIG)succeeded (so the exported targets and their dependencies are known). - Postcondition: the list is complete for the requested components (Definition 10.8.7).
- Invariant: each library appears once.
function MapComponents(C):
expand keywords: "native" → the host target's CodeGen/AsmParser/Desc/Info, "all" → everything, …
L ← [ "LLVM" + capitalize(c) for c in C ]
return ExpandTopologically(L) # llvm_expand_dependencies in LLVM-Config.cmake
function ExpandTopologically(L):
visit each l in L depth-first along its INTERFACE_LINK_LIBRARIES, appending in post-order
return the reverse of that order
What llvm-config and LLVMConfig.cmake say about this LLVM
Reproduce (llvm-config 23.1.2 and CMake 3.28.3; CMakeLists.txt and lljit.cpp are in examples/):
for f in --version --shared-mode --has-rtti --assertion-mode --build-mode; do printf "%-17s %s\n" "$f" "$(llvm-config $f)"; done
echo "--cxxflags $(llvm-config --cxxflags)"
echo "--libs core $(llvm-config --libs core)"
echo "--link-static --libs core: $(llvm-config --link-static --libs core 2>&1 | cut -c1-80)"
cd chapters/10-llvm-cpp-api/examples && cmake -S . -B build -DLLVM_DIR=$(llvm-config --cmakedir) | grep -E "LLVM|linking"
Output (complete for llvm-config; the CMake run abridged to the three message lines):
--version 23.1.2
--shared-mode shared
--has-rtti YES
--assertion-mode OFF
--build-mode Release
--cxxflags -I/opt/llvm-23/include -std=c++17 -D_GNU_SOURCE -D_GLIBCXX_USE_CXX11_ABI=1 -D__STDC_CONSTANT_MACROS -D__STDC_FORMAT_MACROS -D__STDC_LIMIT_MACROS -fno-exceptions
--libs core -lLLVM-23
--link-static --libs core: -lLLVMCore -lLLVMRemarks -lLLVMBitstreamReader -lLLVMBinaryFormat -lLLVMTargetPa
-- LLVM 23.1.2 from /opt/llvm-23/lib/cmake/llvm
-- LLVM_ENABLE_RTTI=ON LLVM_ENABLE_EH=OFF LLVM_LINK_LLVM_DYLIB=yes
-- linking: LLVM
What to notice:
- Shared build: this LLVM is a shared-library build, so
--libs coreis the single-lLLVM-23and CMake links theLLVMtarget. The static closure ofcorealone is already a long list (Definition 10.8.7). - Flags: it has RTTI on (conda's choice; Homebrew and apt builds differ, so read the flag rather than assuming), exceptions off, and assertions off (so no must-check aborts, Lesson 10.7).
- The
-std=c++17trap:--cxxflagscontains-std=c++17, which is why this chapter's compile commands put-std=c++23after it. Your flag must come last to win.
3. Worked example¶
The instance is exec.ll's iterative @fib called with \(n = 10^6\) (Lab 10.3), and the same module with @fibrec called with 24.
Interpreter¶
Algorithm 10.8.2 on fib(3). Frame values after each block transition:
| step | block entered | phi evaluation (simultaneous) | map after |
|---|---|---|---|
| 1 | entry |
— | n=3, small=false |
| 2 | loop from entry |
i=1, a=0, b=1 |
s=1, i.next=2, more=(1<3)=true |
| 3 | loop from loop |
i=2, a=1, b=1 (old b, old s) |
s=2, i.next=3, more=true |
| 4 | loop from loop |
i=3, a=1, b=2 |
s=3, i.next=4, more=false |
| 5 | exit from loop |
r=b=2 |
return 2 |
The simultaneous phi evaluation matters at step 3: a ← b must read the old b. Lab 10.3 measured the interpreter at 412.8 ms for \(n = 10^6\): about 7 instructions per iteration, so roughly 60 ns per interpreted instruction.
MCJIT¶
lli -jit-kind=mcjit fib.ll: Algorithm 10.8.3 compiles the whole module (@fib, @main) at the first getFunctionAddress("main"), links it with RuntimeDyld and runs it. That is 0.051 s in the box, most of it process start-up and code generation.
ORC LLJIT¶
runLLJIT(exec.ll, "fib", {1000000}): lookup triggers Algorithm 10.8.5 for the module. The IR compile layer compiles all six functions of exec.ll, the object layer links them, and the call runs natively. Lab 10.3 measured 13.3 ms in total, dominated by compilation (the loop itself takes about 1 ms).
ORC LLLazyJIT¶
The same call under LLLazyJIT: lookup returns fib's stub. The first call compiles only the fib partition, and the other five functions are never compiled. With fibrec(24), the first call compiles fibrec, and the recursive calls go through the already-updated stub.
Consuming LLVM from CMake¶
For the jitdemo target with components core orcjit native: on this shared build, Algorithm 10.8.8 is skipped (LLVM_LINK_LLVM_DYLIB is true) and the link line is -lLLVM. On a static build, orcjit alone expands to LLVMOrcJIT, LLVMJITLink, LLVMOrcTargetProcess, LLVMOrcShared, LLVMExecutionEngine, LLVMRuntimeDyld, LLVMPasses, … down to LLVMSupport and LLVMDemangle, in topological order.
Try it
./course drill jit-compile-set --seed 2 --difficulty medium: two modules, an execution trace; give LLJIT's compile set and LLLazyJIT's compile order.
4. Invariants and correctness¶
Interpreter¶
The interpreter is correct exactly when every visit* method implements its instruction's semantics and SwitchToNewBasicBlock evaluates all phis of the destination with the values of the predecessor before assigning any of them. The invariant of Algorithm 10.8.2 (all operands of the next instruction are in the frame's map) holds by the SSA dominance property: a definition dominating a use was executed before it on every path (Ch 9). The precondition on intrinsics is not checked. Violating it aborts the process (the box above).
MCJIT¶
MCJIT's correctness reduces to the code generator's (the same one llc uses) and RuntimeDyld's relocation processing. The precondition "no module is modified after it is added" matters: MCJIT compiles lazily at the first address request, so a later change to the module would be compiled in, or not, depending on timing.
ORC LLJIT¶
Proposition 10.8.9 (Resolution follows the link order)
If a name \(s\) is defined in several JITDylibs, a reference to \(s\) from code in \(D\) resolves to the definition in the first JITDylib of \(D\)'s link order that defines \(s\), and that definition's materialization unit runs at most once.
Proof sketch (full description: [LLVM-ORC] §'Design Overview', JITDylib::setLinkOrder in Core.h)
The linker's lookup for an undefined symbol queries the JITDylibs of the link order in sequence and stops at the first one with a definition (the search is ordered). The session records a per-symbol state (NeverSearched → Materializing → Resolved → Emitted → Ready). The first query that finds an unmaterialized symbol moves it to Materializing and runs the unit. Concurrent and later queries attach to the pending state instead of starting another materialization, so each unit runs once.
ORC LLLazyJIT¶
Theorem 10.8.10 (Lazy compilation compiles exactly the executed functions)
Under LLLazyJIT with per-function partitions, each function \(f\) of an added module is compiled at most once, and it is compiled if and only if some call to \(f\) (through its stub, from the host or from compiled code) is executed. The functions are compiled in the order of their first calls.
Proof
At most once: compiling \(f\) is the materialization of \(f\)'s partition, which runs at most once (Proposition 10.8.9), after which \(f\)'s stub points to the body. If: the first executed call to \(f\) goes through the stub, which, until \(f\) is compiled, points to the compile callback (the invariant of Algorithm 10.8.6), and the callback compiles \(f\) before continuing. Only if: a partition is materialized only when its symbol is looked up. addLazyIRModule defines \(f\) in the main JITDylib as a reexport of the stub, so lookup(f) and the linking of other partitions resolve to the stub and do not materialize \(f\)'s body. Only the callback does that, and the callback runs only when a call through the stub executes. Order: compilation happens at the moment of each function's first call, so the compile order is the order of first calls. The box's lazy runs show all three properties.
Consuming LLVM from CMake¶
Proposition 10.8.11 (Flag compatibility requirements)
A client translation unit that includes LLVM headers and links LLVM must (a) use the same RTTI setting as LLVM if it uses LLVM classes with virtual functions whose type information LLVM emits, (b) compile with the same LLVM_ENABLE_ABI_BREAKING_CHECKS value (enforced at link time), and (c) not throw exceptions across LLVM frames when LLVM is built with -fno-exceptions.
Proof sketch (full rules: [LLVM-CMake], [LLVM-CS] 'Do not use RTTI or Exceptions', llvm/include/llvm/Config/abi-breaking.h)
(a) A client built with RTTI that derives from an LLVM class built without it references the base's typeinfo symbol, which a -fno-rtti LLVM never emitted, so the link fails. Conversely, -fno-rtti code cannot dynamic_cast its own classes derived from LLVM's (Lesson 10.4). Both are avoided by matching LLVM_ENABLE_RTTI. (b) abi-breaking.h defines a weak pointer variable that refers to EnableABIBreakingChecks or DisableABIBreakingChecks, and the library defines only the one matching its own setting, so a mismatched client fails to link. The shared library here exports DisableABIBreakingChecks. (c) Frames of code compiled with -fno-exceptions have no cleanup landing pads (Lesson 10.7, Definition 10.7.4), so unwinding through them skips their destructors.
5. Complexity¶
Variables: \(I\) = instructions executed, \(S\) = static size of the compiled code, \(F\) = functions in the module, \(F_{\mathrm{run}}\) = functions executed.
| Technique | Startup | Per executed instruction | Memory | Notes |
|---|---|---|---|---|
| Interpreter | \(O(1)\) | \(\Theta(1)\) but ~60 ns (measured) | frames + a value map per frame | no codegen; few intrinsics |
| MCJIT | \(\Theta(S)\) codegen of the module defining the symbol and of the modules it references | native | object code + RuntimeDyld | eager, whole module |
| ORC LLJIT | \(\Theta(S)\) of the looked-up modules and their references | native | object code; JITLink | eager per module; concurrent compilation possible |
| ORC LLLazyJIT | \(O(F)\) stubs; codegen of \(F_{\mathrm{run}}\) functions only | native (+ one stub jump per call) | stubs + compiled partitions | best when \(F_{\mathrm{run}} \ll F\) |
| CMake integration | configure once | — | shared: one library; static: only the closure | link time: shared \(\ll\) static |
Proposition 10.8.12 (Interpreter/JIT crossover)
If the interpreter takes \(t_i\) per executed instruction, compiled code \(t_c \ll t_i\), and JIT compilation costs \(K\) up front, then the JIT is faster for a run of \(I\) instructions iff \(I > K / (t_i - t_c)\).
Proof
The interpreter's total is \(I t_i\) and the JIT's is \(K + I t_c\). Then \(K + I t_c < I t_i \iff I (t_i - t_c) > K\).
Measured in Lab 10.3 (ch10-apibench): \(K \approx 12\) ms and \(t_i \approx 60\) ns, so the crossover is at \(I \approx 2 \times 10^5\) instructions. collatz(837799) (524 iterations, a few thousand instructions) took 0.7 ms interpreted and 12.0 ms with LLJIT. fib(10^6) (about \(7 \times 10^6\) instructions) took 412.8 ms interpreted and 13.3 ms with LLJIT. Pathological input for lazy JITs: a program that calls each of \(10^5\) small functions once pays one compile callback, one partition extraction and one codegen invocation per function, all fixed costs, and can be slower than compiling the module eagerly in one batch.
6. Variants and refinements¶
Interpreter¶
lli -force-interpreteris the command-line front of the same engine. The interpreter uses libffi, when LLVM is built with it, to call arbitrary external functions.- Other IR interpreters:
llubi(the UB-aware interpreter used in Ch 9), and Alive2's semantics-based evaluator.
MCJIT¶
EngineBuilderoptions: optimization level, target options, custom memory managers (SectionMemoryManager).- Migration: ORC's documentation lists the MCJIT features and their ORC equivalents [LLVM-ORC].
ORC LLJIT¶
- Custom layer stacks (the BuildingAJIT tutorial chapters 1–3 build one by hand) [BAJIT].
LLJITBuilderoptions: number of compile threads (setNumCompileThreads), object linking layer (JITLink by default on most platforms), process-symbol generators, platform support for static initializers.- Remote execution (
ExecutorProcessControl): compile in one process, run in another.
ORC LLLazyJIT¶
- Partition functions (
LLLazyJIT::setPartitionFunction, forwarded to theIRPartitionLayer):IRPartitionLayer::compileRequested(the default: only the requested function) vscompileWholeModule, or a custom partitioner that groups call-graph neighbours. - Re-optimization and speculation (
Speculation.h): tiered compilation on top of lazy stubs, the path toward a real tiered JIT (Ch 0).
Consuming LLVM from CMake¶
llvm-configfor Makefile or shell builds,LLVMConfig.cmakefor CMake, and in-tree builds (add_llvm_executable,LLVM_LINK_COMPONENTS).- Pass plugins link nothing and take LLVM's symbols from the host tool (
-undefined dynamic_lookupon macOS). This repository'spebble_add_pass_plugindoes exactly that (FOUNDATION §2).
7. In real compilers¶
Interpreter¶
LLVM
llvm/lib/ExecutionEngine/Interpreter/Execution.cpp (Interpreter::run, SwitchToNewBasicBlock, visitIntrinsicInst → IntrinsicLowering), llvm/lib/ExecutionEngine/Interpreter/Interpreter.h (class Interpreter : public ExecutionEngine, public InstVisitor<Interpreter>), llvm/tools/lli/lli.cpp (-force-interpreter) [LLVM-Interp].
- CPython, Lua and V8's Ignition are bytecode interpreters in the same role (fast startup) for their languages (Ch 0).
MCJIT¶
LLVM
llvm/lib/ExecutionEngine/MCJIT/MCJIT.cpp (MCJIT::getFunctionAddress, generateCodeForModule, finalizeObject), llvm/lib/ExecutionEngine/RuntimeDyld/; design notes in llvm/docs/MCJITDesignAndImplementation.rst [LLVM-MCJIT].
ORC LLJIT¶
LLVM
llvm/include/llvm/ExecutionEngine/Orc/LLJIT.h (LLJIT, getMainJITDylib, getIRTransformLayer, lookup), llvm/lib/ExecutionEngine/Orc/LLJIT.cpp (LLJITBuilderState::prepareForConstruction, LLJIT::addIRModule), llvm/include/llvm/ExecutionEngine/Orc/Core.h (ExecutionSession, JITDylib::setLinkOrder), llvm/include/llvm/ExecutionEngine/Orc/ThreadSafeModule.h (withModuleDo) [LLVM-LLJIT, LLVM-ORC].
- Users: clang-repl and Cling (C++ REPLs), Julia (ORC-based JIT), PostgreSQL's
llvmjit(expression compilation), and the Ch 0 lab's JIT engine.
Find where LLVM does it. In LLJIT.h, find class LLLazyJIT. Question: which method adds a module lazily, and which JITDylib does its one-argument overload use? (Quiz llvm-where-addlazy.)
ORC LLLazyJIT¶
LLVM
llvm/include/llvm/ExecutionEngine/Orc/CompileOnDemandLayer.h (CompileOnDemandLayer: the stubs), llvm/include/llvm/ExecutionEngine/Orc/IRPartitionLayer.h (IRPartitionLayer, the partition functions compileRequested (the default) and compileWholeModule), llvm/lib/ExecutionEngine/Orc/LLJIT.cpp (LLLazyJIT construction: IPLayer, then CODLayer with a LazyCallThroughManager) [LLVM-LLJIT].
lli -jit-kind=orc-lazyuses it. The BuildingAJIT tutorial chapter 3 builds the same thing by hand [BAJIT].
Consuming LLVM from CMake¶
LLVM
llvm/cmake/modules/LLVMConfig.cmake.in (the installed LLVMConfig.cmake: LLVM_ENABLE_RTTI, LLVM_ENABLE_EH, LLVM_LINK_LLVM_DYLIB, LLVM_DEFINITIONS), llvm/cmake/modules/LLVM-Config.cmake (llvm_map_components_to_libnames, llvm_expand_dependencies), llvm/docs/CMake.md "Embedding LLVM in your project" [LLVM-CMake].
- This course:
cmake/PebbleLLVM.cmake(the major-version check, and dylib vs components inpebble_link_llvm), and thelinux/macospresets.
8. Comparison¶
| Technique | Power / precision | Speed | Output / error quality | Implementation effort | Typical use |
|---|---|---|---|---|---|
| Interpreter | IR semantics; few intrinsics, limited external calls | instant start; ~30–60× slower (measured) | unsupported features abort (report_fatal_error) |
EngineBuilder + runFunction |
quick checks of tiny functions, short runs below the crossover |
| MCJIT | full codegen; whole modules | codegen of whole modules at the first lookup | ExecutionEngine string errors |
EngineBuilder |
legacy code; superseded |
| ORC LLJIT | full codegen; JITDylibs, link order, concurrency, remote execution | per-module codegen at first lookup | Error/Expected everywhere |
LLJITBuilder, ThreadSafeModule, lookup |
test harnesses (Lab 10.1), REPLs, embedding |
| ORC LLLazyJIT | as LLJIT, compiling only executed functions (Theorem 10.8.10) | stubs up front; codegen on first call | as LLJIT | LLLazyJITBuilder, addLazyIRModule |
large modules where little code runs |
| CMake integration | shared or static linking with the right flags | shared: fast links | configure-time errors are clear; flag mismatches are link errors | find_package + four lines |
every out-of-tree LLVM client |
Choose LLJIT by default for running IR in-process. Choose LLLazyJIT when modules are large and only part of them runs, such as a REPL or a whole-program JIT. Choose the interpreter only for tiny, intrinsic-free code or to debug semantics without the code generator. Do not start new code on MCJIT. For builds, use find_package(LLVM 23.1 CONFIG) and read LLVM_ENABLE_RTTI/LLVM_ENABLE_EH instead of assuming them.
9. Assessment¶
| Technique | Quiz ids (solutions/quizzes/ch10.yaml) |
Drill | Flashcard tag | Exercises |
|---|---|---|---|---|
| Interpreter | interp-crossover, interp-intrinsics |
— (see note) | interpreter |
Lab 10.3 R5 |
| MCJIT | mcjit-status, mcjit-compile-set |
./course drill jit-compile-set (eager behavior) |
mcjit |
— |
| ORC LLJIT | lljit-compile-set, llvm-where-addlazy |
./course drill jit-compile-set |
lljit |
Lab 10.1 tests, Lab 10.3 R6 |
| ORC LLLazyJIT | lazy-order, llvm-where-addlazy |
./course drill jit-compile-set |
lazyjit |
Lab 10.3 ★ |
| CMake integration | cmake-rtti, cxxflags-std, cmake-closure |
— (see note) | cmake-llvm |
this repository's build |
The interpreter's behavior is a crossover formula (Proposition 10.8.12, quiz interp-crossover), and build integration is a set of flags, so neither has a randomizable trace worth a drill. Lab 10.3's measurement and the quiz cover both.
Pitfall
A ThreadSafeModule owns its LLVMContext: once you hand a module to ORC, do not touch the module or its context from your own code without withModuleDo, and never pass a module whose context you also use elsewhere. Conversely, the interpreter (EngineBuilder) takes the Module but not the context, so the context must outlive the engine. Lab 10.3 R8 tests exactly this ordering.
References¶
See the chapter references.