Skip to content

Lesson 10.8 — Running IR and consuming LLVM: interpreter, MCJIT, ORC, CMake

Techniques: the IR interpreter (ExecutionEngine Interpreter, lli -force-interpreter), MCJIT (the legacy JIT), ORC LLJIT (layers, JITDylibs, symbol resolution, ThreadSafeModule), ORC LLLazyJIT (per-function lazy compilation), consuming LLVM from a build (llvm-config, LLVMConfig.cmake, dylib vs components, RTTI/EH flags) · Pebble uses: LLJIT in Lab 10.1's tests and Ch 0's lab; the CMake scheme of this repository (cmake/PebbleLLVM.cmake) · Lab: Lab 10.3 (interpreter vs LLJIT) · Prerequisites: Lessons 10.1, 10.7; Ch 0 (interpreters vs JITs) · Time: 3 hours

You have built @square_plus in memory with IRBuilder. How do you run it? You can walk the IR with an interpreter, compile the whole module to machine code in memory and call it, or compile each function only when it is first called. And how does your program find LLVM's headers and libraries, and with which compiler flags? This lesson compares the four execution engines LLVM 23 offers, with measurements, and makes the build-integration rules precise.

1. Problem and motivation

Running freshly built IR is needed by test harnesses (Lab 10.1 calls your functions hundreds of times), REPLs and scripting languages, database query compilers, and GPU/ML runtimes. The engines trade startup latency against peak speed against API complexity, the same trade-off as Ch 0's interpreter vs JIT discussion, now with LLVM's concrete APIs [LLVM-ORC, BAJIT].

Interpreter

ExecutionEngine's Interpreter executes IR directly: a big-step loop over instructions with a value map (GenericValue) per stack frame. It has no code generation, so it starts instantly, runs slowly, supports few intrinsics and calls external functions only through a small table or libffi [LLVM-Interp].

MCJIT

MCJIT (2013) compiles a whole module to an object file in memory with the regular code generator, links it with RuntimeDyld, and hands out function addresses. It is the older JIT API (EngineBuilder(…).setEngineKind(EngineKind::JIT)). It is still in the tree but superseded by ORC: new features (concurrency, lazy compilation, JITLink) go to ORC [LLVM-MCJIT, LLVM-ORC].

ORC LLJIT

ORC (On Request Compilation) is a set of composable layers (IR transform → IR compile → object linking) on top of an ExecutionSession that owns JITDylibs, symbol tables that behave like dynamic libraries. LLJIT assembles a default stack. You add ThreadSafeModules (a module bundled with its context and a lock) and lookup symbols, and a lookup compiles the module that defines the symbol (the whole module) [LLVM-ORC, Ham16, Ham18].

ORC LLLazyJIT

LLLazyJIT adds two layers on top of LLJIT's: an IRPartitionLayer, which splits a module into per-function partitions, and a CompileOnDemandLayer, which replaces each function by a stub. A function's partition is compiled the first time its stub is called. Functions that never run are never compiled [LLVM-ORC, BAJIT].

Consuming LLVM from CMake

A client must compile against LLVM's headers with compatible flags and link the right libraries. LLVM exports this through llvm-config (a command-line oracle) and LLVMConfig.cmake (find_package(LLVM CONFIG)), which provides include directories, definitions, LLVM_ENABLE_RTTI and LLVM_ENABLE_EH, and either the single shared libLLVM or per-component static libraries [LLVM-CMake].

2. Definitions and algorithms

Definition 10.8.1 (Execution engine, materialization)

An execution engine maps a module \(M\) and a symbol name \(s\) to a callable entity with \(s\)'s semantics. Materializing a definition means producing its executable form: interpreter state (nothing to do), or machine code in executable memory. An engine is eager if adding \(M\) (or the first lookup in \(M\)) materializes every definition in \(M\), and lazy if it materializes each function only when the function is first executed.

Interpreter

Algorithm 10.8.2 (Interpreter::run, outline)

  • Input: a function \(F\) and argument values \(\vec{a}\) (GenericValues).
  • Output: the return value.
  • Precondition: \(F\) has a body; every called intrinsic can be lowered by IntrinsicLowering; external calls are to known library functions.
  • Postcondition: the result equals the IR semantics of \(F(\vec{a})\) (Ch 9).
  • Invariant: the top stack frame holds (current block, next instruction, map from SSA values to GenericValues), and every value used by the next instruction is in the map (SSA dominance).
function Run(F, a⃗):
    push frame(F, entry block, {args ↦ a⃗})
    while the stack is non-empty:
        fr ← top; I ← fr.next; fr.next ← next(I)
        visit(I)                                    # InstVisitor<Interpreter> (Lesson 10.4)
            # binary op: fr.vals[I] ← op(fr.vals[op0], fr.vals[op1])
            # br: SwitchToNewBasicBlock(dest) — evaluates the dest's phis with the
            #     predecessor's values *simultaneously*, then continues at dest
            # call: push a new frame (or call the external-function table)
            # ret: pop, write the value into the caller's frame
            # intrinsic: IntrinsicLowering, or report_fatal_error if unsupported
    return the value of the final ret

lli: the interpreter against three JITs on the same IR

Reproduce (lli 23.1.2, Linux x86-64; bash's time):

cd chapters/10-llvm-cpp-api/examples      # fib.ll: recursive fib(30); main returns fib(30) mod 256
bash -c 'TIMEFORMAT="  %R s wall"; for k in -force-interpreter -jit-kind=mcjit -jit-kind=orc -jit-kind=orc-lazy; do echo "lli $k"; time lli $k fib.ll; echo "  exit code $?"; done'

Output (complete; times vary by machine and run):

lli -force-interpreter
  1.521 s wall
  exit code 40
lli -jit-kind=mcjit
  0.051 s wall
  exit code 40
lli -jit-kind=orc
  0.038 s wall
  exit code 40
lli -jit-kind=orc-lazy
  0.025 s wall
  exit code 40

What to notice: all four agree (\(\mathrm{fib}(30) = 832040\), and \(832040 \bmod 256 = 40\)). The interpreter executes \(2F(31) - 1 \approx 2.7\) million calls of @fib, each walked instruction by instruction, and is ~30–60× slower than the JITs, whose times are mostly process start-up and compilation. The orders of magnitude, not the milliseconds, are the point.

What the interpreter cannot run

Reproduce (lli 23.1.2):

cd chapters/10-llvm-cpp-api/examples
lli -force-interpreter abs.ll 2>&1 | head -1; lli abs.ll; echo "exit $?"

Output (complete):

LLVM ERROR: Code generator does not support intrinsic function 'llvm.abs.i32'!
exit 7

What to notice: the interpreter lowers intrinsics through IntrinsicLowering::LowerIntrinsicCall, which knows only a few of them, and llvm.abs is not among them. The failure is report_fatal_error, which aborts the process instead of returning an Error. The JIT uses the real code generator and returns \(\max(|{-7}|, 3) = 7\). That is why Lab 10.3's runInterpreter must refuse modules declaring intrinsics before running them (R5).

MCJIT

Algorithm 10.8.3 (MCJIT lookup and finalization, outline)

  • Input: one or more modules added to an MCJIT ExecutionEngine, and a name \(s\).
  • Output: the address of \(s\) (0 if no added module defines it).
  • Precondition: the native target is initialized; no module is modified after it is added.
  • Postcondition: the module defining \(s\) is compiled and linked, and so is every added module that defines a symbol it (transitively) references; getFunctionAddress(s) returns executable code. finalizeObject() compiles every added module.
  • Invariant: a module is either not yet compiled, or compiled whole and loaded into RuntimeDyld.
function GetFunctionAddress(EE, s):              # MCJIT::getFunctionAddress
    a ← FindSymbol(EE, s)
    if a ≠ 0: FinalizeLoaded(EE)
    return a
function FindSymbol(EE, s):                      # MCJIT::findSymbol
    if s is already loaded: return its address
    M ← the added module that defines s          # findModuleForSymbol
    if M = none: return 0
    obj ← codegen(M)                             # generateCodeForModule: the whole module
    RuntimeDyld.load(obj)
    return address of s
function FinalizeLoaded(EE):                     # finalizeLoadedModules
    for each unresolved external symbol u of the loaded objects:
        FindSymbol(EE, u)                        # via LinkingSymbolResolver: may compile another module
    apply relocations; set memory permissions (RW → RX)

MCJIT compiles whole modules, and the modules they reference

Reproduce (clang 23.1.2, LLVM 23.1.2, Linux x86-64):

cd chapters/10-llvm-cpp-api/examples
clang++-23 $(llvm-config --cxxflags) -std=c++23 mcjit.cpp $(llvm-config --ldflags --libs) \
  -Wl,-rpath,$(llvm-config --libdir) -o mcjit
for a in 5 500; do echo "== ./mcjit m1.ll m2.ll m3.ll $a"; ./mcjit m1.ll m2.ll m3.ll $a; done

m1.ll and m2.ll are the modules of the LLLazyJIT box below, and m3.ll defines other, which nothing refers to. An ObjectCache whose notifyObjectCompiled prints each module MCJIT compiles is attached with setObjectCache.

Output (complete):

== ./mcjit m1.ll m2.ll m3.ll 5
getFunctionAddress(entry)
  [compile] m1.ll: entry used never_called
  [compile] m2.ll: helper unrelated
call entry(5)
  result 15
== ./mcjit m1.ll m2.ll m3.ll 500
getFunctionAddress(entry)
  [compile] m1.ll: entry used never_called
  [compile] m2.ll: helper unrelated
call entry(500)
  result 501

What to notice: everything happens inside getFunctionAddress, before the first call (Algorithm 10.8.3). m1 is compiled whole, never_called included, and m2 follows because finalization must resolve m1's reference to helper, even in the run where helper is never called. m3 is never compiled. Module granularity and reference-driven compilation are the same as LLJIT's (next box), and this is the eager behavior the jit-compile-set drill asks about.

ORC LLJIT

Definition 10.8.4 (JITDylib, link order, resolution)

A JITDylib \(D\) is a symbol table from names to (address or materializer, flags). Its link order is a sequence \(\langle D, D_1, \dots, D_r \rangle\) of JITDylibs (by default \(D\) first, then the process's symbols for LLJIT's main dylib, via a generator). Resolving a name \(s\) referenced from code in \(D\) finds the first \(D_i\) in the link order that defines \(s\). If that definition is not yet materialized, its materialization unit (for IR: the whole ThreadSafeModule that defines it) is run first.

Algorithm 10.8.5 (LLJIT lookup)

  • Input: an LLJIT \(J\) and a name \(s\).
  • Output: Expected<ExecutorAddr>.
  • Precondition: every added module has been verified; names are mangled for the target (lookup mangles).
  • Postcondition: \(s\)'s module is compiled and linked, and so is every module that it (transitively) references; the result is \(s\)'s address. A missing symbol returns an Error.
  • Invariant: each materialization unit runs at most once (the session tracks symbol states).
function Lookup(J, s):
    (D_i, def) ← first definition of s in link order of Main
    if def is unmaterialized:
        MU ← materialization unit of def             # the whole module
        TSM ← MU.module
        TSM ← IRTransformLayer.transform(TSM)         # optional hook (the box uses it to log)
        obj ← IRCompileLayer.compile(TSM)             # codegen, holding TSM's context lock
        ObjectLinkingLayer.link(obj):
            for each undefined symbol u in obj: Lookup(J, u)   # may materialize other modules
            apply relocations; set memory permissions
        mark all of MU's symbols materialized
    return address(s)

LLJIT: build IR, add it, look it up, call it

Reproduce (clang 23.1.2, LLVM 23.1.2, Linux x86-64):

cd chapters/10-llvm-cpp-api/examples
clang++-23 $(llvm-config --cxxflags) -std=c++23 lljit.cpp $(llvm-config --ldflags --libs) \
  -Wl,-rpath,$(llvm-config --libdir) -o lljit && ./lljit

Output (complete):

triple x86_64-conda-linux-gnu
square_plus(3, 1) = 10
square_plus(-4, 1) = 17
square_plus(1000000, 1) = 1000000000001

What to notice: the program never touches a file. It moves the module and its context into a ThreadSafeModule (ownership passes to the JIT), lookup returns an ExecutorAddr, and toPtr<int64_t (*)(int64_t, int64_t)>() turns it into a callable C++ function pointer. Every step returns Expected/Error, handled with ExitOnError (Lesson 10.7). The triple is this conda-built LLVM's host triple.

ORC LLLazyJIT

Algorithm 10.8.6 (Per-function lazy compilation)

  • Input: a module \(M\) added with addLazyIRModule.
  • Output: callable stubs for \(M\)'s functions.
  • Precondition: as for Algorithm 10.8.5.
  • Postcondition: Theorem 10.8.10.
  • Invariant: each function's stub points either to a compile callback (not yet compiled) or to the compiled body.
function AddLazy(J, M):
    for each function f defined in M:
        partition P_f ← a module containing f's body (+ declarations it needs)
        define symbol f in Main as a reexport of a stub → compile-callback(P_f)
on first call through f's stub:                  # compile-callback(P_f)
    compile P_f through the IR layers (Algorithm 10.8.5's layers)
    update f's stub to jump to the compiled body
    continue into the body

Eager vs lazy: what gets compiled, and when

Reproduce (clang 23.1.2, LLVM 23.1.2, Linux x86-64):

cd chapters/10-llvm-cpp-api/examples
clang++-23 $(llvm-config --cxxflags) -std=c++23 eager.cpp $(llvm-config --ldflags --libs) \
  -Wl,-rpath,$(llvm-config --libdir) -o eager
for m in lljit lazy; do for a in 5 500; do echo "== ./eager $m m1.ll m2.ll $a"; ./eager $m m1.ll m2.ll $a; done; done

Module 1 defines entry (which calls used if \(x \le 100\), else helper), used and never_called. Module 2 defines helper and unrelated. An IRTransformLayer hook prints the functions of each module that reaches the compiler.

Output (complete):

== ./eager lljit m1.ll m2.ll 5
lookup(entry)
  [compile] entry used never_called
  [compile] helper unrelated
call entry(5)
  result 15
== ./eager lljit m1.ll m2.ll 500
lookup(entry)
  [compile] entry used never_called
  [compile] helper unrelated
call entry(500)
  result 501
== ./eager lazy m1.ll m2.ll 5
lookup(entry)
call entry(5)
  [compile] entry
  [compile] used
  result 15
== ./eager lazy m1.ll m2.ll 500
lookup(entry)
call entry(500)
  [compile] entry
  [compile] helper
  result 501

What to notice: LLJIT compiled all of module 1 at the lookup, never_called included, and then all of module 2, because module 1 references helper (Algorithm 10.8.5's recursive lookup). LLLazyJIT compiled nothing at lookup and then exactly the functions that ran, in call order (Theorem 10.8.10). The drill jit-compile-set asks for these two answers on random programs.

Consuming LLVM from CMake

Definition 10.8.7 (Component closure)

LLVM's static libraries form a DAG under "requires". A set \(C\) of components (core, support, orcjit, native, …; llvm-config --components lists 226 here) denotes libraries; its closure \(\overline{C}\) adds every library reachable in the DAG. A link line is complete if it contains \(\overline{C}\) in a topological order (users before dependencies). With a shared build (LLVM_LINK_LLVM_DYLIB), all components are in one library, libLLVM, and the closure is trivial.

Algorithm 10.8.8 (llvm_map_components_to_libnames + expansion, outline)

  • Input: component names.
  • Output: an ordered list of library targets.
  • Precondition: find_package(LLVM CONFIG) succeeded (so the exported targets and their dependencies are known).
  • Postcondition: the list is complete for the requested components (Definition 10.8.7).
  • Invariant: each library appears once.
function MapComponents(C):
    expand keywords: "native" → the host target's CodeGen/AsmParser/Desc/Info, "all" → everything, …
    L ← [ "LLVM" + capitalize(c) for c in C ]
    return ExpandTopologically(L)       # llvm_expand_dependencies in LLVM-Config.cmake
function ExpandTopologically(L):
    visit each l in L depth-first along its INTERFACE_LINK_LIBRARIES, appending in post-order
    return the reverse of that order

What llvm-config and LLVMConfig.cmake say about this LLVM

Reproduce (llvm-config 23.1.2 and CMake 3.28.3; CMakeLists.txt and lljit.cpp are in examples/):

for f in --version --shared-mode --has-rtti --assertion-mode --build-mode; do printf "%-17s %s\n" "$f" "$(llvm-config $f)"; done
echo "--cxxflags        $(llvm-config --cxxflags)"
echo "--libs core       $(llvm-config --libs core)"
echo "--link-static --libs core: $(llvm-config --link-static --libs core 2>&1 | cut -c1-80)"
cd chapters/10-llvm-cpp-api/examples && cmake -S . -B build -DLLVM_DIR=$(llvm-config --cmakedir) | grep -E "LLVM|linking"

Output (complete for llvm-config; the CMake run abridged to the three message lines):

--version         23.1.2
--shared-mode     shared
--has-rtti        YES
--assertion-mode  OFF
--build-mode      Release
--cxxflags        -I/opt/llvm-23/include -std=c++17  -D_GNU_SOURCE -D_GLIBCXX_USE_CXX11_ABI=1 -D__STDC_CONSTANT_MACROS -D__STDC_FORMAT_MACROS -D__STDC_LIMIT_MACROS -fno-exceptions
--libs core       -lLLVM-23
--link-static --libs core: -lLLVMCore -lLLVMRemarks -lLLVMBitstreamReader -lLLVMBinaryFormat -lLLVMTargetPa
-- LLVM 23.1.2 from /opt/llvm-23/lib/cmake/llvm
-- LLVM_ENABLE_RTTI=ON LLVM_ENABLE_EH=OFF LLVM_LINK_LLVM_DYLIB=yes
-- linking: LLVM

What to notice:

  • Shared build: this LLVM is a shared-library build, so --libs core is the single -lLLVM-23 and CMake links the LLVM target. The static closure of core alone is already a long list (Definition 10.8.7).
  • Flags: it has RTTI on (conda's choice; Homebrew and apt builds differ, so read the flag rather than assuming), exceptions off, and assertions off (so no must-check aborts, Lesson 10.7).
  • The -std=c++17 trap: --cxxflags contains -std=c++17, which is why this chapter's compile commands put -std=c++23 after it. Your flag must come last to win.

3. Worked example

The instance is exec.ll's iterative @fib called with \(n = 10^6\) (Lab 10.3), and the same module with @fibrec called with 24.

Interpreter

Algorithm 10.8.2 on fib(3). Frame values after each block transition:

step block entered phi evaluation (simultaneous) map after
1 entry — n=3, small=false
2 loop from entry i=1, a=0, b=1 s=1, i.next=2, more=(1<3)=true
3 loop from loop i=2, a=1, b=1 (old b, old s) s=2, i.next=3, more=true
4 loop from loop i=3, a=1, b=2 s=3, i.next=4, more=false
5 exit from loop r=b=2 return 2

The simultaneous phi evaluation matters at step 3: a ← b must read the old b. Lab 10.3 measured the interpreter at 412.8 ms for \(n = 10^6\): about 7 instructions per iteration, so roughly 60 ns per interpreted instruction.

MCJIT

lli -jit-kind=mcjit fib.ll: Algorithm 10.8.3 compiles the whole module (@fib, @main) at the first getFunctionAddress("main"), links it with RuntimeDyld and runs it. That is 0.051 s in the box, most of it process start-up and code generation.

ORC LLJIT

runLLJIT(exec.ll, "fib", {1000000}): lookup triggers Algorithm 10.8.5 for the module. The IR compile layer compiles all six functions of exec.ll, the object layer links them, and the call runs natively. Lab 10.3 measured 13.3 ms in total, dominated by compilation (the loop itself takes about 1 ms).

ORC LLLazyJIT

The same call under LLLazyJIT: lookup returns fib's stub. The first call compiles only the fib partition, and the other five functions are never compiled. With fibrec(24), the first call compiles fibrec, and the recursive calls go through the already-updated stub.

Consuming LLVM from CMake

For the jitdemo target with components core orcjit native: on this shared build, Algorithm 10.8.8 is skipped (LLVM_LINK_LLVM_DYLIB is true) and the link line is -lLLVM. On a static build, orcjit alone expands to LLVMOrcJIT, LLVMJITLink, LLVMOrcTargetProcess, LLVMOrcShared, LLVMExecutionEngine, LLVMRuntimeDyld, LLVMPasses, … down to LLVMSupport and LLVMDemangle, in topological order.

Try it

./course drill jit-compile-set --seed 2 --difficulty medium: two modules, an execution trace; give LLJIT's compile set and LLLazyJIT's compile order.

4. Invariants and correctness

Interpreter

The interpreter is correct exactly when every visit* method implements its instruction's semantics and SwitchToNewBasicBlock evaluates all phis of the destination with the values of the predecessor before assigning any of them. The invariant of Algorithm 10.8.2 (all operands of the next instruction are in the frame's map) holds by the SSA dominance property: a definition dominating a use was executed before it on every path (Ch 9). The precondition on intrinsics is not checked. Violating it aborts the process (the box above).

MCJIT

MCJIT's correctness reduces to the code generator's (the same one llc uses) and RuntimeDyld's relocation processing. The precondition "no module is modified after it is added" matters: MCJIT compiles lazily at the first address request, so a later change to the module would be compiled in, or not, depending on timing.

ORC LLJIT

Proposition 10.8.9 (Resolution follows the link order)

If a name \(s\) is defined in several JITDylibs, a reference to \(s\) from code in \(D\) resolves to the definition in the first JITDylib of \(D\)'s link order that defines \(s\), and that definition's materialization unit runs at most once.

Proof sketch (full description: [LLVM-ORC] §'Design Overview', JITDylib::setLinkOrder in Core.h)

The linker's lookup for an undefined symbol queries the JITDylibs of the link order in sequence and stops at the first one with a definition (the search is ordered). The session records a per-symbol state (NeverSearched → Materializing → Resolved → Emitted → Ready). The first query that finds an unmaterialized symbol moves it to Materializing and runs the unit. Concurrent and later queries attach to the pending state instead of starting another materialization, so each unit runs once.

ORC LLLazyJIT

Theorem 10.8.10 (Lazy compilation compiles exactly the executed functions)

Under LLLazyJIT with per-function partitions, each function \(f\) of an added module is compiled at most once, and it is compiled if and only if some call to \(f\) (through its stub, from the host or from compiled code) is executed. The functions are compiled in the order of their first calls.

Proof

At most once: compiling \(f\) is the materialization of \(f\)'s partition, which runs at most once (Proposition 10.8.9), after which \(f\)'s stub points to the body. If: the first executed call to \(f\) goes through the stub, which, until \(f\) is compiled, points to the compile callback (the invariant of Algorithm 10.8.6), and the callback compiles \(f\) before continuing. Only if: a partition is materialized only when its symbol is looked up. addLazyIRModule defines \(f\) in the main JITDylib as a reexport of the stub, so lookup(f) and the linking of other partitions resolve to the stub and do not materialize \(f\)'s body. Only the callback does that, and the callback runs only when a call through the stub executes. Order: compilation happens at the moment of each function's first call, so the compile order is the order of first calls. The box's lazy runs show all three properties.

Consuming LLVM from CMake

Proposition 10.8.11 (Flag compatibility requirements)

A client translation unit that includes LLVM headers and links LLVM must (a) use the same RTTI setting as LLVM if it uses LLVM classes with virtual functions whose type information LLVM emits, (b) compile with the same LLVM_ENABLE_ABI_BREAKING_CHECKS value (enforced at link time), and (c) not throw exceptions across LLVM frames when LLVM is built with -fno-exceptions.

Proof sketch (full rules: [LLVM-CMake], [LLVM-CS] 'Do not use RTTI or Exceptions', llvm/include/llvm/Config/abi-breaking.h)

(a) A client built with RTTI that derives from an LLVM class built without it references the base's typeinfo symbol, which a -fno-rtti LLVM never emitted, so the link fails. Conversely, -fno-rtti code cannot dynamic_cast its own classes derived from LLVM's (Lesson 10.4). Both are avoided by matching LLVM_ENABLE_RTTI. (b) abi-breaking.h defines a weak pointer variable that refers to EnableABIBreakingChecks or DisableABIBreakingChecks, and the library defines only the one matching its own setting, so a mismatched client fails to link. The shared library here exports DisableABIBreakingChecks. (c) Frames of code compiled with -fno-exceptions have no cleanup landing pads (Lesson 10.7, Definition 10.7.4), so unwinding through them skips their destructors.

5. Complexity

Variables: \(I\) = instructions executed, \(S\) = static size of the compiled code, \(F\) = functions in the module, \(F_{\mathrm{run}}\) = functions executed.

Technique Startup Per executed instruction Memory Notes
Interpreter \(O(1)\) \(\Theta(1)\) but ~60 ns (measured) frames + a value map per frame no codegen; few intrinsics
MCJIT \(\Theta(S)\) codegen of the module defining the symbol and of the modules it references native object code + RuntimeDyld eager, whole module
ORC LLJIT \(\Theta(S)\) of the looked-up modules and their references native object code; JITLink eager per module; concurrent compilation possible
ORC LLLazyJIT \(O(F)\) stubs; codegen of \(F_{\mathrm{run}}\) functions only native (+ one stub jump per call) stubs + compiled partitions best when \(F_{\mathrm{run}} \ll F\)
CMake integration configure once — shared: one library; static: only the closure link time: shared \(\ll\) static

Proposition 10.8.12 (Interpreter/JIT crossover)

If the interpreter takes \(t_i\) per executed instruction, compiled code \(t_c \ll t_i\), and JIT compilation costs \(K\) up front, then the JIT is faster for a run of \(I\) instructions iff \(I > K / (t_i - t_c)\).

Proof

The interpreter's total is \(I t_i\) and the JIT's is \(K + I t_c\). Then \(K + I t_c < I t_i \iff I (t_i - t_c) > K\).

Measured in Lab 10.3 (ch10-apibench): \(K \approx 12\) ms and \(t_i \approx 60\) ns, so the crossover is at \(I \approx 2 \times 10^5\) instructions. collatz(837799) (524 iterations, a few thousand instructions) took 0.7 ms interpreted and 12.0 ms with LLJIT. fib(10^6) (about \(7 \times 10^6\) instructions) took 412.8 ms interpreted and 13.3 ms with LLJIT. Pathological input for lazy JITs: a program that calls each of \(10^5\) small functions once pays one compile callback, one partition extraction and one codegen invocation per function, all fixed costs, and can be slower than compiling the module eagerly in one batch.

6. Variants and refinements

Interpreter

  • lli -force-interpreter is the command-line front of the same engine. The interpreter uses libffi, when LLVM is built with it, to call arbitrary external functions.
  • Other IR interpreters: llubi (the UB-aware interpreter used in Ch 9), and Alive2's semantics-based evaluator.

MCJIT

  • EngineBuilder options: optimization level, target options, custom memory managers (SectionMemoryManager).
  • Migration: ORC's documentation lists the MCJIT features and their ORC equivalents [LLVM-ORC].

ORC LLJIT

  • Custom layer stacks (the BuildingAJIT tutorial chapters 1–3 build one by hand) [BAJIT].
  • LLJITBuilder options: number of compile threads (setNumCompileThreads), object linking layer (JITLink by default on most platforms), process-symbol generators, platform support for static initializers.
  • Remote execution (ExecutorProcessControl): compile in one process, run in another.

ORC LLLazyJIT

  • Partition functions (LLLazyJIT::setPartitionFunction, forwarded to the IRPartitionLayer): IRPartitionLayer::compileRequested (the default: only the requested function) vs compileWholeModule, or a custom partitioner that groups call-graph neighbours.
  • Re-optimization and speculation (Speculation.h): tiered compilation on top of lazy stubs, the path toward a real tiered JIT (Ch 0).

Consuming LLVM from CMake

  • llvm-config for Makefile or shell builds, LLVMConfig.cmake for CMake, and in-tree builds (add_llvm_executable, LLVM_LINK_COMPONENTS).
  • Pass plugins link nothing and take LLVM's symbols from the host tool (-undefined dynamic_lookup on macOS). This repository's pebble_add_pass_plugin does exactly that (FOUNDATION §2).

7. In real compilers

Interpreter

LLVM

llvm/lib/ExecutionEngine/Interpreter/Execution.cpp (Interpreter::run, SwitchToNewBasicBlock, visitIntrinsicInst → IntrinsicLowering), llvm/lib/ExecutionEngine/Interpreter/Interpreter.h (class Interpreter : public ExecutionEngine, public InstVisitor<Interpreter>), llvm/tools/lli/lli.cpp (-force-interpreter) [LLVM-Interp].

  • CPython, Lua and V8's Ignition are bytecode interpreters in the same role (fast startup) for their languages (Ch 0).

MCJIT

LLVM

llvm/lib/ExecutionEngine/MCJIT/MCJIT.cpp (MCJIT::getFunctionAddress, generateCodeForModule, finalizeObject), llvm/lib/ExecutionEngine/RuntimeDyld/; design notes in llvm/docs/MCJITDesignAndImplementation.rst [LLVM-MCJIT].

ORC LLJIT

LLVM

llvm/include/llvm/ExecutionEngine/Orc/LLJIT.h (LLJIT, getMainJITDylib, getIRTransformLayer, lookup), llvm/lib/ExecutionEngine/Orc/LLJIT.cpp (LLJITBuilderState::prepareForConstruction, LLJIT::addIRModule), llvm/include/llvm/ExecutionEngine/Orc/Core.h (ExecutionSession, JITDylib::setLinkOrder), llvm/include/llvm/ExecutionEngine/Orc/ThreadSafeModule.h (withModuleDo) [LLVM-LLJIT, LLVM-ORC].

  • Users: clang-repl and Cling (C++ REPLs), Julia (ORC-based JIT), PostgreSQL's llvmjit (expression compilation), and the Ch 0 lab's JIT engine.

Find where LLVM does it. In LLJIT.h, find class LLLazyJIT. Question: which method adds a module lazily, and which JITDylib does its one-argument overload use? (Quiz llvm-where-addlazy.)

ORC LLLazyJIT

LLVM

llvm/include/llvm/ExecutionEngine/Orc/CompileOnDemandLayer.h (CompileOnDemandLayer: the stubs), llvm/include/llvm/ExecutionEngine/Orc/IRPartitionLayer.h (IRPartitionLayer, the partition functions compileRequested (the default) and compileWholeModule), llvm/lib/ExecutionEngine/Orc/LLJIT.cpp (LLLazyJIT construction: IPLayer, then CODLayer with a LazyCallThroughManager) [LLVM-LLJIT].

  • lli -jit-kind=orc-lazy uses it. The BuildingAJIT tutorial chapter 3 builds the same thing by hand [BAJIT].

Consuming LLVM from CMake

LLVM

llvm/cmake/modules/LLVMConfig.cmake.in (the installed LLVMConfig.cmake: LLVM_ENABLE_RTTI, LLVM_ENABLE_EH, LLVM_LINK_LLVM_DYLIB, LLVM_DEFINITIONS), llvm/cmake/modules/LLVM-Config.cmake (llvm_map_components_to_libnames, llvm_expand_dependencies), llvm/docs/CMake.md "Embedding LLVM in your project" [LLVM-CMake].

  • This course: cmake/PebbleLLVM.cmake (the major-version check, and dylib vs components in pebble_link_llvm), and the linux/macos presets.

8. Comparison

Technique Power / precision Speed Output / error quality Implementation effort Typical use
Interpreter IR semantics; few intrinsics, limited external calls instant start; ~30–60× slower (measured) unsupported features abort (report_fatal_error) EngineBuilder + runFunction quick checks of tiny functions, short runs below the crossover
MCJIT full codegen; whole modules codegen of whole modules at the first lookup ExecutionEngine string errors EngineBuilder legacy code; superseded
ORC LLJIT full codegen; JITDylibs, link order, concurrency, remote execution per-module codegen at first lookup Error/Expected everywhere LLJITBuilder, ThreadSafeModule, lookup test harnesses (Lab 10.1), REPLs, embedding
ORC LLLazyJIT as LLJIT, compiling only executed functions (Theorem 10.8.10) stubs up front; codegen on first call as LLJIT LLLazyJITBuilder, addLazyIRModule large modules where little code runs
CMake integration shared or static linking with the right flags shared: fast links configure-time errors are clear; flag mismatches are link errors find_package + four lines every out-of-tree LLVM client

Choose LLJIT by default for running IR in-process. Choose LLLazyJIT when modules are large and only part of them runs, such as a REPL or a whole-program JIT. Choose the interpreter only for tiny, intrinsic-free code or to debug semantics without the code generator. Do not start new code on MCJIT. For builds, use find_package(LLVM 23.1 CONFIG) and read LLVM_ENABLE_RTTI/LLVM_ENABLE_EH instead of assuming them.

9. Assessment

Technique Quiz ids (solutions/quizzes/ch10.yaml) Drill Flashcard tag Exercises
Interpreter interp-crossover, interp-intrinsics — (see note) interpreter Lab 10.3 R5
MCJIT mcjit-status, mcjit-compile-set ./course drill jit-compile-set (eager behavior) mcjit —
ORC LLJIT lljit-compile-set, llvm-where-addlazy ./course drill jit-compile-set lljit Lab 10.1 tests, Lab 10.3 R6
ORC LLLazyJIT lazy-order, llvm-where-addlazy ./course drill jit-compile-set lazyjit Lab 10.3 ★
CMake integration cmake-rtti, cxxflags-std, cmake-closure — (see note) cmake-llvm this repository's build

The interpreter's behavior is a crossover formula (Proposition 10.8.12, quiz interp-crossover), and build integration is a set of flags, so neither has a randomizable trace worth a drill. Lab 10.3's measurement and the quiz cover both.

Pitfall

A ThreadSafeModule owns its LLVMContext: once you hand a module to ORC, do not touch the module or its context from your own code without withModuleDo, and never pass a module whose context you also use elsewhere. Conversely, the interpreter (EngineBuilder) takes the Module but not the context, so the context must outlive the engine. Lab 10.3 R8 tests exactly this ordering.

References

See the chapter references.