Lesson 24.3 — JIT designs on ORC: eager LLJIT, lazy stubs, re-optimization¶
Techniques: ORC
LLJIT, eager — modules are added to a JITDylib and compiled, whole, when one of their symbols is first looked up; lazy compilation by reexport stubs — a module is split into one materialization unit per function, and a call to a function not yet compiled lands in a stub that compiles it on first use (LLLazyJIT,LazyReexports); tiered re-optimization — functions start at a cheap tier, a counter finds the hot ones, and a re-optimize layer recompiles them at a higher tier and redirects their symbol · Pebble implements: ★pebble-jit(lab J1: eager and lazy sessions, a REPL); tiering is studied on ORC'sReOptimizeLayerand measured withpebble-jit -O0/-O2· Drill: none (see §9) · Prerequisites: Lesson 0.3 (method JITs, tiers, OSR, deoptimization: the theory this lesson instantiates), Lesson 10.8 (a firstLLJIT), Lesson 11.9 (the runtime ABI) · Time: 5 hours
pebblec writes an object file and calls the linker. A JIT does the linker's job in memory, at run time, for the same LLVM module, and gets to choose when: everything at load, each function at its first call, or hot functions twice. ORC (On-Request Compilation), LLVM's JIT infrastructure since LLVM 7, expresses the three choices with one vocabulary: a JITDylib is a symbol table, a materialization unit is a promise to define some symbols when asked, and layers turn IR into objects, objects into linked memory, and (lazily) symbols into stubs that trigger the promise. This lesson defines that vocabulary, gives the lookup and lazy-call algorithms with their invariants, and runs pebble-jit in both modes on the same program, so that you can watch 4 functions compile eagerly and 2 lazily. Lesson 0.3 explained why a runtime would tier; this lesson shows what the tiering machinery is made of in ORC.
1. Problem and motivation¶
ORC LLJIT (eager)¶
MCJIT, LLVM's first production JIT, compiled a module when it was added and could not add a second module that referenced the first without finalizing everything. ORC (Hames, 2014–; the "ORCv2" design of LLVM 7 [LLVM-ORC]) replaced it with a symbol-table-driven design: adding a module defines its symbols in a JITDylib without compiling anything, and a lookup triggers compilation of exactly the materialization units that define the requested symbols and their dependencies. LLJIT is the pre-assembled stack for the common case: an IRCompileLayer over an ObjectLinkingLayer (JITLink) or RTDyldObjectLinkingLayer, a main JITDylib, and a JITTargetMachineBuilder for the host. pebble-jit is 150 lines on top of it: the runtime's symbols as absolute definitions, a JITDylib per program, a lookup of pebble_main.
Lazy compilation by reexport stubs¶
Deutsch and Schiffman's Smalltalk-80 implementation [DS84] compiled a method the first time it was called and cached the result: the origin of every "compile on demand" system. In ORC the mechanism is a lazy reexport: the JITDylib defines f as a stub whose body jumps through a pointer initialized to a trampoline; the trampoline's handler looks up the real f (materializing it now), patches the pointer, and resumes the call. LLLazyJIT adds a CompileOnDemandLayer that splits each module into one unit per function and defines every function as a lazy reexport, so startup compiles nothing and each function costs its compilation only when it runs.
Tiered re-optimization¶
Compiling at -O2 costs more than at -O0 (the box in §2: 3–5× for a small module) and pays only for hot code. Hölzle and Ungar's adaptive optimization for Self [HU94] introduced the pattern every modern VM uses: fast compilation first, counters, recompilation of hot methods with the profile in hand (Lesson 0.3, Algorithm 0.3.7). ORC's ReOptimizeLayer [LLVM-ReOpt] is the LLVM building block: it wraps the compile layer, gives every materialization unit a re-optimization ID, instruments functions with a call counter, and, when a threshold trips, recompiles the unit with a user-supplied transformation and redirects the symbol through a RedirectableSymbolManager, so that existing callers pick up the new code at their next call.
2. Definitions and algorithms¶
Definition 24.3.1 (JITDylib, materialization unit, lookup)
An execution session owns a set of JITDylibs \(\mathrm{JD}\), each a partial map from symbol names to states: never materialized, materializing, or ready with an address. A materialization unit \(\mathrm{MU}\) is a set of symbol names \(\mathrm{provides}(\mathrm{MU})\) together with a procedure that, when run, defines each of them (compiles, links, or assigns an address). Adding \(\mathrm{MU}\) to \(\mathrm{JD}\) registers its names as never materialized. A lookup of a set of names \(S\) in a search order \(\langle \mathrm{JD}_1, \dots, \mathrm{JD}_m \rangle\) finds, for each name, the first dylib defining it, runs the units of the names that are not ready, and returns addresses once all are ready.
Definition 24.3.2 (Layer)
A layer is a function from one representation of a unit to the next that materializes on demand:
IRTransformLayer (IR to IR), IRCompileLayer (IR to an object file), ObjectLinkingLayer (object to linked
memory with relocations applied). LLJIT chains them; addIRModule(JD, M) wraps \(M\) in one unit whose
procedure runs the chain.
ORC LLJIT (eager)¶
Algorithm 24.3.3 (Eager materialization by lookup: what pebble-jit does per program)
- Input: a verified LLVM module \(M\) following the runtime ABI; the runtime's symbol addresses \(R\).
- Output: the result of
pebble_main. - Precondition: every symbol \(M\) references is defined by \(M\) or by \(R\).
- Postcondition: every defined function of \(M\) has been compiled exactly once (Proposition 24.3.7);
pebble_mainran against the process's runtime. - Invariant: a symbol is compiled at most once per session, and no symbol is compiled before some lookup needs it.
function RunEager(M, R):
JD_rt ← main JITDylib; JD_rt.define(absoluteSymbols(R)) # pebble_print_int ↦ &pebble_print_int, ...
JD ← createJITDylib("pebble.N"); JD.linkOrder ← [JD, JD_rt]
JD.add(IRMaterializationUnit(M)) # provides every global of M; compiles nothing yet
addr ← lookup([JD], {pebble_main}) # runs M's unit: compile M, link, resolve R
return call addr as i64 ()
lookup(order, S):
for each name in S: find the first JD in order defining it; if its unit is not ready, queue the unit
run the queued units (each may look up further names, recursively: the runtime symbols here)
wait until every name in S is ready; return their addresses
Four functions compiled eagerly, two lazily
Reproduce (pebble-jit from a -DPEBBLE_USE_SOLUTION=all build, LLVM 23.1.2; SOL as in Lesson 24.1):
cat > lazy.pbl <<'EOF'
fn used(x: int) -> int { return x * 2; }
fn unused(x: int) -> int { return x * 3; }
fn also_unused() -> int { return unused(1); }
fn main() -> int {
print(used(21));
return 0;
}
EOF
$SOL/bin/pebble-jit --jit-stats lazy.pbl
$SOL/bin/pebble-jit --lazy --jit-trace --jit-stats lazy.pbl
Output (complete; stderr lines may interleave differently with stdout):
pebble-jit: compiled 4 function(s)
42
pebble-jit: compiled pebble_main
pebble-jit: compiled __orc_lcl.used.0
pebble-jit: compiled 2 function(s)
42
What to notice: eagerly, the one lookup of pebble_main materializes the whole module: 4 bodies. Lazily,
pebble_main is compiled by the lookup, and used only when pebble_main calls its stub; unused and
also_unused are never compiled. The name __orc_lcl.used.0 is what CompileOnDemandLayer calls the
implementation of used after extracting it into its own unit: the JITDylib's used is the stub, and the
module's internal linkage function got a unique local name.
Lazy compilation by reexport stubs¶
Definition 24.3.4 (Lazy reexport)
A lazy reexport of symbol \(f\) (implemented by \(f'\) in some JITDylib) is a definition of \(f\) as a
stub: a small code sequence that jumps through a pointer \(p_f\). Initially \(p_f\) points at a
trampoline owned by a LazyCallThroughManager; the trampoline saves the registers, calls the
manager with its own address, and the manager looks up \(f'\), stores its address in \(p_f\), restores the
registers and jumps to \(f'\). After the first call \(p_f\) is \(f'\) and the stub costs one indirect jump.
Algorithm 24.3.5 (Compile on demand: addLazyIRModule and the first call)
- Input: a module \(M\) with functions \(f_1, \dots, f_k\).
- Output: definitions of \(f_1, \dots, f_k\) in \(\mathrm{JD}\), each compiled at its first call.
- Precondition: the target supports stubs and trampolines (
OrcABISupport: x86-64, AArch64, ...). - Postcondition: every function that was called is compiled exactly once; functions never called are never compiled (Proposition 24.3.8).
- Invariant: for every \(f_i\), either \(p_{f_i}\) points at the trampoline and \(f_i'\) is not materialized, or \(p_{f_i}\) is the address of the materialized \(f_i'\).
function AddLazy(JD, M):
for each function f in M:
JD.define(LazyReexport(f ↦ impl(f))) # a stub per function, p_f ← trampoline
JD_impl.add(PartitioningMU(M)) # materializes one function's body per request
on a call reaching trampoline t (from stub of f): # LazyCallThroughManager::resolveTrampolineLandingAddress
f' ← findReexport(t)
addr ← lookup([JD_impl], {impl(f)}) # extracts f's body into its own module, compiles it
p_f ← addr # notifyResolved: patch the stub's pointer
jump addr # the original call proceeds
The stub's cost: what a lazy call looks like after resolution
Reproduce (lli 23.1.2; the option names are those of LLVM 23's lli):
Output (complete):
--compile-threads=<uint> - Choose the number of compile threads (jit-kind=orc-lazy only)
--disable-lazy-compilation - Disable JIT lazy compilation
--jit-kind=<value> - Choose underlying JIT kind.
=orc-lazy - Orc-based lazy JIT.
--per-module-lazy - Performs lazy compilation on whole module boundaries rather than individual functions
--thread-entry=<string> - calls the given entry-point on a new thread (jit-kind=orc-lazy only)
What to notice: lli -jit-kind=orc-lazy is LLLazyJIT; --per-module-lazy chooses the partition of
Algorithm 24.3.5 (one unit per module instead of per function), and --compile-threads shows why the
session's state machine has a materializing state: with several threads, two calls can race to the same
trampoline, and the manager must serialize the first materialization. pebble-jit --lazy is the per-function
partition with one thread.
Tiered re-optimization¶
Algorithm 24.3.6 (Re-optimization with symbol redirection, after ReOptimizeLayer)
- Input: a module \(M\); a tier-0 pipeline; a tier-1 transformation \(T\) (a higher
-Opipeline, possibly using profile facts); a threshold \(\theta\). - Output: every function of \(M\) runs at tier 0 until it has been called \(\theta\) times, then at tier 1.
- Precondition: symbols are redirectable: every call to \(f\) goes through an indirection the layer
controls (a
RedirectableSymbolManager), so that a new body can replace the old one for future calls. - Postcondition: the behavior is that of \(M\) (Proposition 24.3.9); each function is compiled at most twice.
- Invariant: at any moment, a function's redirect target is a body compiled from \(M\) (tier 0 or tier 1); calls in progress finish in the body they entered.
function AddTiered(JD, M, T, θ):
id ← fresh materialization-unit id
for each function f in M: insert at f's entry: counter[id, f] ← counter[id, f] + 1;
if counter[id, f] = θ: call __orc_rt_reoptimize(id, f)
JD.add(RedirectableMU(compile at tier 0)) # f's symbol = a redirectable stub → tier-0 body
__orc_rt_reoptimize(id, f): # ReOptimizeLayer::reoptimizeIfCallFrequent
M' ← the unit's IR (kept by the layer), transformed by T (e.g. -O2 with counters removed)
body' ← compile M'; redirect(f, body') # RedirectableSymbolManager::redirect
# in-flight calls continue in the tier-0 body; the next call enters body'
The price the tiers trade: compile time against run time in the JIT
Reproduce (pebble-jit from a -DPEBBLE_USE_SOLUTION=all build; bash's time; wall clock on a shared machine, so read the ratios):
for L in -O0 -O2; do TIMEFORMAT="pebble-jit $L bench-nbody: %R s"; time ($SOL/bin/pebble-jit $L labs/ch24-capstone/inputs/bench/bench-nbody.pbl > /dev/null); done
for L in -O0 -O2; do TIMEFORMAT="pebble-jit $L fn-unit (tiny): %R s"; time ($SOL/bin/pebble-jit $L tests/ch11/e2e/fn-unit.pbl > /dev/null); done
Output (one run):
pebble-jit -O0 bench-nbody: 0.100 s
pebble-jit -O2 bench-nbody: 0.062 s
pebble-jit -O0 fn-unit (tiny): 0.020 s
pebble-jit -O2 fn-unit (tiny): 0.019 s
What to notice: for the tiny program the two levels cost the same: process startup and LLVM's target
initialization dominate, and there is nothing to optimize. For bench-nbody the -O2 compile is slower but
the run is much faster, and the whole process finishes 40 % sooner. A tiered JIT wants both columns: the
tiny program's startup (compile everything at -O0) and the hot loop's speed (recompile step and accel
at -O2 once the counter trips). Without tiers a JIT must pick one level for all code.
3. Worked example¶
Running example (used for every technique in this lesson): the module of lazy.pbl (the box in §2), with call graph pebble_main → used, also_unused → unused, and the runtime symbols pebble_print_int, pebble_print_newline.
flowchart LR
M[pebble_main] --> U[used]
AU[also_unused] --> UU[unused]
M --> P[pebble_print_int]
M --> N[pebble_print_newline]
subgraph runtime JITDylib
P
N
end
ORC LLJIT (eager)¶
Algorithm 24.3.3 step by step:
| step | action | JITDylib state | compiled |
|---|---|---|---|
| 1 | define the runtime's 7 symbols as absolute | main: pebble_print_int … ready |
– |
| 2 | create pebble.1, link order [pebble.1, main] |
pebble.1: empty |
– |
| 3 | addIRModule(pebble.1, M) |
pebble.1: pebble_main, used, unused, also_unused never materialized |
– |
| 4 | lookup({pebble_main}) |
pebble_main materializing: the IR unit runs |
the whole module: 4 bodies |
| 5 | the object's relocations need pebble_print_int, pebble_print_newline |
looked up in main: ready (absolute) |
– |
| 6 | all four symbols ready; return the address | – | – |
| 7 | call pebble_main() |
– | prints 42 |
Lazy compilation by reexport stubs¶
Algorithm 24.3.5 on the same module:
| step | action | stub pointers | compiled |
|---|---|---|---|
| 1 | addLazyIRModule: four stubs |
all four \(p_f\) → trampoline | – |
| 2 | lookup({pebble_main}) |
pebble_main's body extracted and compiled; \(p_{\mathrm{pebble\_main}}\) patched |
pebble_main |
| 3 | call pebble_main(): call used enters used's stub, then the trampoline |
manager: findReexport → used |
– |
| 4 | lookup({__orc_lcl.used.0}) |
used's body compiled; \(p_{\mathrm{used}}\) patched |
used |
| 5 | the trampoline returns into used; the print calls are absolute symbols, no stub |
– | prints 42 |
| 6 | exit | unused, also_unused still → trampoline |
never |
Tiered re-optimization¶
Algorithm 24.3.6 with \(\theta = 3\) on a hypothetical driver that calls pebble_main five times (the REPL evaluates many modules; imagine one of them with a hot used):
call of used |
counter | tier | action |
|---|---|---|---|
| 1 | 1 | 0 | runs the -O0 body |
| 2 | 2 | 0 | runs the -O0 body |
| 3 | 3 = \(\theta\) | 0 → 1 | __orc_rt_reoptimize(id, used): recompile at -O2, redirect(used, body'); this call finishes in the old body |
| 4 | – | 1 | enters body' through the redirect |
| 5 | – | 1 | body' |
Try it
No drill: see §9. Instead, run pebble-jit --lazy --jit-trace on tests/ch11/e2e/fn-mutual.pbl and predict, before looking, the order in which its four functions compile (hint: is_even(10) is called before ackermann, and mutual recursion resolves each stub once).
4. Invariants and correctness¶
ORC LLJIT (eager)¶
Proposition 24.3.7 (Lookup materializes each unit at most once, and everything a symbol needs)
In Algorithm 24.3.3, every materialization unit runs at most once per session, and when lookup returns
an address for \(s\), every symbol that \(s\)'s code references is ready.
Proof
At most once: a unit's names move from never materialized to materializing when the unit is queued, and
a lookup only queues units whose names are never materialized; the state change happens under the session
lock before the unit runs, so a second lookup of the same name finds it materializing and waits instead of
queueing (this is the MaterializationResponsibility of ORC: exactly one owner per symbol). Completeness:
the object linking layer resolves relocations by looking up every undefined symbol of the object, recursively,
and a symbol is marked ready only after its unit reports that all of its dependencies are ready
(notifyEmitted with the dependency set). Hence the address returned points at code whose references are
all resolved. With one thread this is a plain recursion; with several it is why the session keeps
dependency sets per symbol rather than a global "done" flag.
Lazy compilation by reexport stubs¶
Proposition 24.3.8 (Laziness: called iff compiled)
Under Algorithm 24.3.5, after any execution: a function's body has been materialized if and only if the function was called at least once or looked up explicitly, and it was materialized exactly once.
Proof
If: a call enters the stub; if \(p_f\) is the trampoline, the manager's handler looks up \(\mathrm{impl}(f)\),
which materializes it (Proposition 24.3.7 gives exactly once) and patches \(p_f\). Only if: the only paths that
look up \(\mathrm{impl}(f)\) are the trampoline handler (after a call) and an explicit lookup; adding the module
defines stubs and queues nothing. Exactly once: after the patch, \(p_f\) is the body's address, and the
invariant of Algorithm 24.3.5 says the trampoline is never re-installed. Where it breaks: an object that
takes the address of unused (a function pointer) forces the stub's address to be defined, which is fine
(the stub exists), but a lookup that requests the implementation symbol directly compiles it without a call;
__orc_lcl names are private for this reason.
Tiered re-optimization¶
Proposition 24.3.9 (Redirection preserves behavior)
If the tier-0 and tier-1 bodies of \(f\) are both correct compilations of \(f\) (each refines \(M\)'s \(f\)) and calls in progress complete in the body they entered, then the program's behavior under Algorithm 24.3.6 is a behavior of \(M\).
Proof
Every execution of \(f\) runs one of the two bodies from entry to exit (the redirect changes only where the
next call lands). Each body refines \(f\), so each execution of \(f\) is an execution of \(M\)'s \(f\); the counter
and the reoptimize call have no observable effect (they touch only the layer's state). By induction over the
calls, the whole execution is a behavior of \(M\). Where it breaks: on-stack replacement (Lesson 0.3), which
transfers a running activation between bodies, needs the frame-mapping argument of Theorem 24.2.10;
ReOptimizeLayer does not do OSR, which keeps its proof this short and its hot loops unoptimized until the
next call.
5. Complexity¶
\(k\) = functions in the module, \(c_j(f)\) = the cost of compiling \(f\) at tier \(j\), \(r\) = the number of calls at run time, \(u\) = the number of distinct functions called, \(\theta\) = the tiering threshold.
| Technique | Startup | Per call | Total compile | Space | Justification |
|---|---|---|---|---|---|
| Eager LLJIT | \(\sum_{f} c_0(f)\) (the whole module at the first lookup) | a direct call | \(\sum_f c_0(f)\) | one body per function | Algorithm 24.3.3 materializes the unit as a whole |
| Lazy stubs | \(O(k)\) stubs, no compilation | first call: \(c_0(f)\) plus the trampoline round trip; later: one indirect jump | \(\sum_{f \text{ called}} c_0(f)\) | a stub and a pointer per function, plus bodies of called functions | Proposition 24.3.8 |
| Tiered | as tier 0 | a counter increment; at the \(\theta\)-th call, \(c_1(f)\) | \(\le \sum_f c_0(f) + \sum_{f \text{ hot}} c_1(f)\) | two bodies per hot function until the old one is released | Algorithm 24.3.6 compiles each function at most twice |
Pathological family (laziness). A program that calls each of its \(k\) functions exactly once pays \(k\) trampoline round trips (each: save registers, look up, extract the function into its own module, compile, patch) for no reuse; per-function partitioning costs \(\Theta(k)\) extractions where the eager JIT compiles once. lli --per-module-lazy exists for this case. Pathological family (tiering). A function called exactly \(\theta\) times pays \(c_1(f)\) and never runs the result; with \(c_1 \approx 5 c_0\) (the box: 3–5× for small modules) the threshold must exceed the number of calls that amortize it, which is why HotSpot's thresholds are in the thousands (Lesson 0.3).
At scale: on bench-nbody the JIT's whole process is 100 ms at -O0 and 62 ms at -O2 (the §2 box); startup and target initialization are 19–20 ms of both (the tiny program), so the loop's run time dominates and -O2 wins; on the tiny program neither level matters. LLVM's own lli -jit-kind=orc-lazy and Julia's JIT (which compiles each method at its first call with a per-signature specialization) are the production instances of Algorithm 24.3.5.
6. Variants and refinements¶
ORC LLJIT (eager)¶
RTDyldObjectLinkingLayervsObjectLinkingLayer(JITLink) [LLVM-ORC]: the old runtime dynamic loader vs JITLink's link-graph model with platform support (ELFNixPlatform,MachOPlatform: static initializers, TLS, EH frames); trade-off: JITLink is the default on macOS/arm64 and needed for lazy stubs there.- Out-of-process execution (
ExecutorProcessControl,llvm-jitlink -oop-executor): the compiled code runs in another process; trade-off: every lookup and memory write crosses a channel. - Symbols from the host process (
DynamicLibrarySearchGenerator::GetForCurrentProcess): resolve undefined symbols throughdlsym; trade-off: needs the host built with exported symbols (-rdynamic), which is whypebble-jitdefines the runtime as absolute symbols instead: portable to macOS without linker flags. - Resource trackers (
ResourceTracker,removeResources): free a module's memory when a REPL discards it; trade-off: symbols in use elsewhere must not be removed (the session checks).
Lazy compilation by reexport stubs¶
- Per-module laziness (
--per-module-lazy): one unit per module instead of per function; trade-off: fewer extractions, coarser laziness. - Speculative compilation (
Speculator,IRSpeculationLayerinllvm/lib/ExecutionEngine/Orc/Speculation.cpp): when \(f\) is materialized, start compiling the functions it is likely to call on another thread; trade-off: compile threads and a likely-callee analysis. - Concurrent materialization (
setNumCompileThreads): trampolines resolved on any thread; trade-off: the session's dependency tracking replaces a simple recursion (Proposition 24.3.7). - Lazy object files (
LazyObjectLinkingLayer): laziness at the object rather than the IR level, for precompiled code; trade-off: no per-function partitioning below the object's granularity.
Tiered re-optimization¶
- Profile-directed tier 1 [HU94, LLVM-ReOpt]: the counters of tier 0 can be extended to record call targets and branch frequencies for
T(ReOptimizeLayer::ReOptimizeFuncreceives the module and can consult a profile); trade-off: instrumentation cost in tier 0. - Deoptimization from tier 1 (Lesson 24.2, Theorem 24.2.10): needed once tier 1 speculates;
ReOptimizeLayerwithout it can only apply sound optimizations at tier 1. - More tiers (HotSpot's five levels, V8's three, Lesson 0.3): more thresholds, smoother cost curves; trade-off: several code bodies per method and OSR between more pairs.
- Background compilation: tier 1 on a compile thread while tier 0 keeps running (V8's concurrent TurboFan); trade-off: the redirect must be atomic (
JITLinkRedirectableSymbolManagerwrites a pointer, which is).
7. In real compilers¶
ORC LLJIT (eager)¶
LLVM
llvm/lib/ExecutionEngine/Orc/LLJIT.cpp — LLJIT::addIRModule, LLJIT::lookup (via ExecutionSession::lookup),
the builder's prepareForConstruction; llvm/lib/ExecutionEngine/Orc/Core.cpp — JITDylib::define,
ExecutionSession::lookup, JITDylib::addToLinkOrder (LLVM 23.1.2) [LLVM-LLJIT]. The textbook JIT compiles a
function and returns a pointer; ORC compiles symbols and returns addresses, which is why pebble-jit
never mentions llvm::Function after addIRModule.
- Pebble
solutions/labs/ch24-jit/src/Session.cpp—OrcSession::run,publishRuntime(absolute symbols). - Julia
src/jitlayers.cpp(JuliaOJIT, Julia 1.11): an LLJIT-style stack with a custom compile layer per method instance. - Numba, Zig, Cling (ROOT): LLJIT users; Cling's incremental C++ interpreter is the REPL pattern of
pebble-jit --replat scale.
Find where LLVM does it. In llvm/lib/ExecutionEngine/Orc/LLJIT.cpp, find LLJIT::addIRModule(JITDylib &, ThreadSafeModule). Question: which layer's add does it call, and what happens to the module's symbol names before they are defined (which function mangles them)?
Lazy compilation by reexport stubs¶
LLVM
llvm/lib/ExecutionEngine/Orc/LazyReexports.cpp — LazyCallThroughManager::resolveTrampolineLandingAddress,
findReexport, notifyResolved (LLVM 23.1.2) [LLVM-LazyReexports]; llvm/lib/ExecutionEngine/Orc/LLJIT.cpp —
LLLazyJIT::addLazyIRModule; llvm/lib/ExecutionEngine/Orc/CompileOnDemandLayer.cpp — CompileOnDemandLayer::emit
(the partition into per-function units). The stubs and trampolines themselves are target code in
llvm/lib/ExecutionEngine/Orc/OrcABISupport.cpp (OrcX86_64_Base::writeTrampolines, writeIndirectStubsBlock; OrcAArch64 likewise).
- Smalltalk-80 [DS84]: the origin, with a code cache instead of stubs.
- HotSpot
src/hotspot/share/interpreter/interpreterRuntime.cpp—InterpreterRuntime::resolve_invoke: call sites start unresolved and are patched at first execution; the interpreter plays the trampoline's role. - Wasmtime: lazy function-table initialization rather than lazy compilation (Cranelift compiles everything ahead), a design choice discussed in Lesson 0.3.
Find where LLVM does it. In llvm/lib/ExecutionEngine/Orc/LazyReexports.cpp, find LazyCallThroughManager::resolveTrampolineLandingAddress. Question: what does it do when the lookup of the implementation symbol fails (which address does the trampoline jump to)? (Quiz find-lazy-callthrough.)
Tiered re-optimization¶
LLVM
llvm/include/llvm/ExecutionEngine/Orc/ReOptimizeLayer.h — ReOptimizeLayer, reoptimizeIfCallFrequent,
setReoptimizeFunc; llvm/lib/ExecutionEngine/Orc/ReOptimizeLayer.cpp;
llvm/lib/ExecutionEngine/Orc/JITLinkRedirectableSymbolManager.cpp — JITLinkRedirectableSymbolManager::redirect
(LLVM 23.1.2) [LLVM-ReOpt]. Algorithm 24.3.6 is the layer's design; the default profiler (reoptimizeIfCallFrequent)
counts calls with a threshold and the transformation is whatever the user installs.
- HotSpot
src/hotspot/share/compiler/compilationPolicy.cpp—CompilationPolicy::event, the tier thresholdsTier3InvocationThreshold,Tier4InvocationThreshold(OpenJDK 21): the five-tier policy of Lesson 0.3. - V8
src/execution/tiering-manager.cc—TieringManager::OnInterruptTick(V8 13.6): Ignition → Sparkplug → Maglev → TurboFan decisions from feedback-vector budgets. - .NET
src/coreclr/vm/tieredcompilation.cpp(tiered compilation with a call-count threshold of 30 and a backedge-based OSR).
Find where LLVM does it. In llvm/include/llvm/ExecutionEngine/Orc/ReOptimizeLayer.h, find reoptimizeIfCallFrequent. Question: what is the default call-count threshold (CallCountThreshold) at which the function is re-optimized? (Quiz find-reoptimize-threshold.)
8. Comparison¶
| Technique | Power / precision | Speed | Output / error quality | Implementation effort | Typical use |
|---|---|---|---|---|---|
| ORC LLJIT (eager) | whole module compiled on first lookup; symbols resolved by JITDylib search order | compile everything up front (4 of 4 functions on the running example) | simplest; startup pays for unused code | low (LLJITBuilder) |
lli, pebble-jit, Julia's base JIT |
| Lazy compilation (stubs) | one materialization unit per function; a function compiles at its first call (2 of 4) | startup proportional to what runs | one indirection per not-yet-compiled call, then patched | low with LLLazyJIT; moderate by hand |
lli -jit-kind=orc-lazy, Kaleidoscope |
| Tiered re-optimization | recompile hot functions at higher -O with profile facts (Algorithm 24.3.6) |
pays compile time only where it matters (nbody: 100 ms at -O0 vs 62 ms at -O2 in the JIT) | best steady-state; needs counters and redirection | high (ReOptimizeLayer, RedirectableSymbolManager) |
HotSpot, V8, .NET, ORC's ReOptimizeLayer |
Choose eager when the module is small or all of it runs (a compiled program run once, tests); choose lazy when startup latency matters and most functions are cold (a REPL, a large library loaded for one entry point); choose tiers when the program runs long enough for hot code to appear and you can afford two bodies and a counter per function.
Measured: the compiled-function counts (4 vs 2) and the timings of the §2 boxes; ch24.JIT.* asserts the counts.
9. Assessment¶
- Quiz:
lookup-materializes(set),lazy-compiled-set(set),stub-pointer-states(sequence),find-lazy-callthrough(text),reopt-tier-table(mapping),find-reoptimize-threshold(number). Tagslljit,lazy-compilation,tiering. - Drill: none. The lazy and tiered traces of §3 are deterministic functions of the call graph and the call order, which the
gc-rootsdrill's program shape does not express; the quiz'slazy-compiled-setandstub-pointer-statesask exactly those traces on fresh instances, andpebble-jit --jit-traceis an oracle you can run on any program. - Flashcards: tags
lljit,lazy-compilation,tiering. - Exercises: ★ J1 (
pebblejit::Session, eager and lazy; the REPL is provided on top of it).
Pitfall
"Lazy compilation is free until the first call." The stub is not free: every call to a function that was
resolved still goes through an indirect jump (jmp *p_f), and the first call pays a full round trip through
the manager, including extracting the function into its own module and running the whole compile pipeline on
it. For a program that calls every function once, lazy is strictly slower than eager (the pathological
family of §5).
References¶
See the chapter references.