Lesson 21.9 — Calling conventions, frame lowering and prologue/epilogue insertion¶
Techniques: argument assignment (calling conventions); frame layout and frame lowering; prologue/epilogue insertion and shrink-wrapping · Lab:
labs/ch21-mir(task 7) · Prerequisites: Lesson 21.5 (MIR), Ch 15 (dominators and post-dominators), Ch 11 (the source-level ABI) · Time: 4–5 hours
Instruction selection turns operations into instructions, but a function also has an interface and a frame. Its arguments arrive in registers and stack slots fixed by the platform's calling convention, the callee-saved registers it uses must be restored before it returns, its locals and spills need stack slots at known offsets, and the stack pointer must stay aligned at every call. The back end handles this in three places. Call and argument lowering happens during selection: TableGen-generated CCAssignFns decide where each argument goes. Frame lowering lays out the frame after register allocation, when the spill slots are known. Prologue/epilogue insertion (PEI) writes the code that builds and tears down the frame, and shrink-wrapping moves that code off the paths that do not need it. The running example is f(x, y) = g(x) + g(y) + x, whose prologue on x86-64 is three pushes and nothing else.
1. Problem and motivation¶
The inputs are a function's signature, the target's calling convention and ABI rules, and, after register allocation, the set of callee-saved registers the function clobbers and the sizes and alignments of its stack objects. The outputs are the location of every argument and return value (register, register pair, or stack offset), a frame layout (an offset for every stack object), and prologue and epilogue code. Two independently compiled functions must agree on all of this, so the rules are fixed by platform documents: the System V AMD64 psABI [SysV-ABI] and Arm's AAPCS64 [AAPCS64] for this chapter's two targets.
Argument assignment (calling conventions)¶
A calling convention maps a signature to locations. Both conventions here classify each argument (integer, floating point, aggregate by size) and hand out registers of the matching class in order, spilling to the stack when they run out. They differ in the details: SysV treats a 128-bit integer as two eightbytes of class INTEGER that must both fit in registers, while AAPCS64 rounds the next register number up to an even one first and, once an argument spills, sends all later integer arguments to the stack. LLVM encodes conventions in TableGen (CallingConv<[CCIfType<…, CCAssignToReg<…>>, …]>) and generates the assignment functions [LLVM-CC].
Frame layout and frame lowering¶
The frame holds the return address (x86) or the saved link register (AArch64), callee-saved registers, locals, spill slots and the outgoing argument area. The target's TargetFrameLowering decides whether a frame pointer is needed, how to align the stack, and whether a leaf function may use the red zone: 128 bytes below %rsp that the SysV ABI guarantees signal handlers will not clobber [SysV-ABI, §3.2.2], so a leaf function can use it without adjusting %rsp.
Prologue/epilogue insertion and shrink-wrapping¶
PEI runs after register allocation. It saves the callee-saved registers the function uses, assigns final offsets to all frame objects, emits prologue and epilogue code, and replaces abstract frame indices by sp/fp-relative addresses. Chow observed that saving registers at entry is wasteful when only some paths use them, and proposed shrink-wrapping: place the saves and restores around the region that needs them [Cho88]. LLVM's ShrinkWrap pass computes such points for the whole prologue and epilogue.
2. Definitions and algorithms¶
Definition 21.9.1 (Calling convention as location assignment)
A location is a register, a pair of registers lo:hi, or a stack offset in the
outgoing argument area (offset 0 is the lowest address, the first stack argument; at the
callee's entry, SysV puts it at 8(%rsp) above the return address and AAPCS64 at [sp]). A
calling convention is a function \(\mathit{cc}(\tau_1, \dots, \tau_n) = (\ell_1, \dots,
\ell_n)\) from the argument types to locations such that (i) no register is assigned twice,
(ii) stack locations do not overlap and respect their alignment, and (iii) the function is
computed left to right with finite state (the counters below). Caller and callee apply the
same function, which is why separately compiled code interoperates.
Argument assignment (calling conventions)¶
Algorithm 21.9.2 (SysV x86-64 assignment for scalar arguments)
- Input: argument types from {
i32,i64,ptr,float,double,i128} (the scalar subset of [SysV-ABI, §3.2.3]). - Output: a location per argument.
- Precondition: a non-variadic function.
- Postcondition: properties (i)–(iii) of Definition 21.9.1 (Theorem 21.9.4).
- Invariant:
gGPRs ofrdi, rsi, rdx, rcx, r8, r9andfXMMs ofxmm0–xmm7are used, all by earlier arguments, andspis the next free stack offset.
function SysV(types):
g ← 0; f ← 0; sp ← 0
for each type τ, left to right:
if τ ∈ {i32, i64, ptr}: # class INTEGER, one eightbyte
if g < 6: loc ← GPR[g]; g ← g + 1
else: sp ← Align(sp, 8); loc ← stack+sp; sp ← sp + 8
elif τ ∈ {float, double}: # class SSE
if f < 8: loc ← XMM[f]; f ← f + 1
else: sp ← Align(sp, 8); loc ← stack+sp; sp ← sp + 8
elif τ = i128: # two INTEGER eightbytes
if g ≤ 4: loc ← GPR[g]:GPR[g+1]; g ← g + 2
else: sp ← Align(sp, 16); loc ← stack+sp; sp ← sp + 16
# g unchanged: a later INTEGER argument may still use GPR[g]
output loc
Algorithm 21.9.3 (AAPCS64 assignment for scalar arguments, stage C)
- Input: the same types (C.1, C.9–C.11, C.13–C.17 of [AAPCS64, §6.8.2]).
- Output: a location per argument (
xregisters for integers,vregisters for floating point). - Precondition: a non-variadic function; standard AAPCS64 (Linux), not Apple's variant, which packs stack arguments by their natural size.
- Postcondition: properties (i)–(iii) of Definition 21.9.1.
- Invariant: NGRN, NSRN and NSAA are the next general register, SIMD/FP register and stack address, as in the standard.
function AAPCS64(types):
NGRN ← 0; NSRN ← 0; NSAA ← 0 # stage A
for each type τ, left to right:
if τ ∈ {float, double}: # C.1, C.5, C.6
if NSRN < 8: loc ← v[NSRN]; NSRN ← NSRN + 1
else: NSAA ← Align(NSAA, 8); loc ← stack+NSAA; NSAA ← NSAA + 8
elif τ ∈ {i32, i64, ptr}: # C.9, then C.13, C.14, C.16, C.17
if NGRN < 8: loc ← x[NGRN]; NGRN ← NGRN + 1
else: NSAA ← Align(NSAA, 8); loc ← stack+NSAA; NSAA ← NSAA + 8
elif τ = i128: # C.10, C.11, else C.13, C.14, C.17
NGRN ← RoundUpToEven(NGRN)
if NGRN < 7: loc ← x[NGRN]:x[NGRN+1]; NGRN ← NGRN + 2
else: NGRN ← 8 # C.13: no more GPR arguments
NSAA ← Align(NSAA, 16); loc ← stack+NSAA; NSAA ← NSAA + 16
output loc
Theorem 21.9.4 (Both assignments are calling conventions)
Algorithms 21.9.2 and 21.9.3 satisfy properties (i)–(iii) of Definition 21.9.1: registers are
never assigned twice, stack slots are disjoint and aligned (8 bytes, or 16 for i128), and
the result depends only on the list of types.
Proof
(i) Every register assignment uses the current counter value(s) and then increases the counter
past them (by 1, or by 2 for a pair). AAPCS64's RoundUpToEven and NGRN ← 8 only increase
NGRN. The counters never decrease, so no register index is handed out twice. SysV's
"g unchanged" branch assigns no register at all. (ii) Each stack assignment first aligns the
offset to the slot's alignment (8 or 16), assigns the interval \([\mathit{sp}, \mathit{sp} +
\mathit{size})\), and then advances past it. The offset never decreases, so the intervals are
disjoint and in increasing order. (iii) The algorithms read only the types and their own
counters. They are deterministic single passes, so caller and callee, running them on the
same signature, compute the same locations. This is what makes the separately compiled code
agree.
The two conventions disagree on an i128
For (i64 × 5, i128, i64): SysV gives rdi, rsi, rdx, rcx, r8, stack+0, r9. The i128
needs two GPRs, only r9 is left, so it goes to the stack and r9 stays free for the last
i64. AAPCS64 gives x0–x4, x6:x7, stack+0. NGRN = 5 is rounded up to 6, the pair takes
x6:x7, and the last i64 finds NGRN = 8 and goes to the stack. Both results were checked
with llc 23.1.2 (tools/course/tests/test_ch21.py, CallingConventions).
Frame layout and frame lowering¶
Definition 21.9.5 (Frame, frame objects, frame index)
A frame object is a stack slot with a size and an alignment: a local (alloca), a spill
slot, a callee-saved register slot, or a fixed object at a known offset (an incoming stack
argument). Before layout, instructions refer to frame objects by frame index (%stack.0,
%fixed-stack.1 in MIR). The frame is the region between the incoming stack pointer and
the stack pointer after the prologue. Its size stackSize is chosen so that every object has
an offset aligned to its alignment and the stack pointer satisfies the ABI's alignment (16
bytes) at every call.
Algorithm 21.9.6 (Frame object layout, after calculateFrameObjectOffsets)
- Input: the frame objects; callee-saved slots; the stack alignment \(A\) (16); whether the function calls anything; red-zone eligibility.
- Output: an offset per object (negative, relative to the incoming stack pointer) and
stackSize. - Precondition: fixed objects have their ABI offsets already.
- Postcondition: objects do not overlap, each offset is a multiple of the object's alignment, and, if the function makes calls, the stack pointer is \(A\)-aligned at each call (Proposition 21.9.7).
- Invariant:
offis the lowest address allocated so far (it grows downwards).
function LayoutFrame(objects, A):
off ← the size of the fixed area (return address, pushed callee-saved registers)
for each callee-saved slot, then each local and spill slot (in an order that groups
objects of equal alignment to reduce padding):
off ← AlignUp(off + size(obj), align(obj)) # growing downwards
offset(obj) ← −off
stackSize ← off − fixed area
if the function has calls:
round stackSize up so that (incoming misalignment + fixed area + stackSize) ≡ 0 mod A
if the function is a leaf, needs no realignment, and stackSize ≤ 128 and the red zone
is allowed: use the red zone (do not move the stack pointer)
return offsets, stackSize
Proposition 21.9.7 (Layout invariants)
Algorithm 21.9.6 gives pairwise disjoint objects at aligned offsets. If the function makes a call, the stack pointer is \(A\)-aligned at the call instruction.
Proof
Disjointness and alignment. Each object gets the interval \([-\mathit{off}, -\mathit{off} +
\mathrm{size})\) after off has been increased past the previous objects and rounded up to a
multiple of its alignment. So the intervals are disjoint (decreasing addresses) and each start
is aligned. Call alignment. At entry the ABI guarantees a known misalignment (SysV: %rsp
≡ 8 mod 16, because the call pushed the return address; AAPCS64: sp ≡ 0 mod 16). The
prologue moves the stack pointer by the fixed area plus stackSize, and the rounding step
makes the sum of the misalignment and that movement a multiple of \(A\). Outgoing argument areas
are reserved inside stackSize (a reserved call frame), so the stack pointer does not move
between the prologue and the calls.
Prologue/epilogue insertion and shrink-wrapping¶
Algorithm 21.9.8 (Prologue/epilogue insertion)
- Input: a register-allocated function with frame indices and call-frame pseudos
(
ADJCALLSTACKDOWN/UP); the set \(\mathit{CSR}\) of callee-saved registers the function writes; save point \(S\) and restore point \(R\) (the entry and the exits unless shrink-wrapped). - Output: the function with prologue at \(S\), epilogue(s) at \(R\), and concrete addresses.
- Precondition: \(S\) dominates and \(R\) post-dominates every instruction that uses a register in \(\mathit{CSR}\) or a frame index (Theorem 21.9.11).
- Postcondition: every path from entry to exit through \(S\) saves each register of
\(\mathit{CSR}\) before using it and restores it before returning. Every frame index is
replaced by
sp- orfp-relative addressing. - Invariant: the offsets from Algorithm 21.9.6 are final before any frame index is rewritten.
function PEI(F):
CSR ← the callee-saved registers written in F
assign a spill slot to each r ∈ CSR (or use push/pop, x86)
LayoutFrame(frame objects of F, A) # Algorithm 21.9.6
at S: emit the prologue: save CSR (push or stp), set up fp if needed,
allocate the frame (sub sp, stackSize), emit CFI directives for unwinding
at each block in R: emit the epilogue: free the frame, restore CSR, (return)
replace each ADJCALLSTACKDOWN/UP by a stack adjustment, or delete it when the call frame
is reserved in the fixed frame
for each instruction with a frame index fi: rewrite fi as base + offset(fi), with
base = sp or fp, as the target's eliminateFrameIndex chooses
Definition 21.9.9 (Save and restore points)
Let \(U\) be the set of blocks that use a callee-saved register the function writes, or that touch the frame. A save point is a block \(S\) that dominates every block of \(U\). A restore point is a block \(R\) that post-dominates every block of \(U\). Additionally \(S\) dominates \(R\), \(R\) post-dominates \(S\), and neither lies inside a loop that does not contain all of \(U\), so the prologue and epilogue run once per execution of the region.
Algorithm 21.9.10 (Shrink-wrapping, after LLVM's ShrinkWrap)
- Input: a function after register allocation; its dominator and post-dominator trees and loop information.
- Output: save point \(S\) and restore point \(R\), or the entry and the exits (no change).
- Precondition: the target supports shrink-wrapping for this function (no stack realignment, no exception-handling funclets, and so on).
- Postcondition: \(S\) and \(R\) satisfy Definition 21.9.9 (Theorem 21.9.11).
- Invariant: after each update, \(S\) dominates and \(R\) post-dominates every block of \(U\) seen so far.
function ShrinkWrap(F):
S ← none; R ← none
for each block B that uses a CSR or the frame:
S ← (S = none) ? B : NearestCommonDominator(S, B)
R ← (R = none) ? B : NearestCommonPostDominator(R, B)
if S = none: return (no frame needed at all)
repeat until stable: # repair until A, B and C hold:
if not (S dominates R): S ← NearestCommonDominator(S, R)
if not (R post-dominates S): R ← NearestCommonPostDominator(R, S)
if S is in a loop L that does not contain R (or vice versa): move S to the
dominator of L's header outside L, R to the post-dominating exit of L
return (S, R)
Theorem 21.9.11 (Shrink-wrapped prologues and epilogues are correct)
If \(S\) and \(R\) satisfy Definition 21.9.9 and the prologue is placed at the start of \(S\) and the epilogue at the end of \(R\), then on every execution: (a) every use of a callee-saved register or the frame happens after the prologue and before the epilogue; (b) the prologue executes iff the epilogue executes, exactly once each; hence (c) every callee-saved register has its entry value when the function returns.
Proof
(a) A use in block \(B \in U\) is reached only by paths through \(S\), because \(S\) dominates \(B\), and every path from \(B\) to the exit passes through \(R\), because \(R\) post-dominates \(B\). So each use is between an execution of \(S\) and an execution of \(R\). (b) Every path through \(S\) continues to the exit through \(R\), since \(R\) post-dominates \(S\). Every path through \(R\) came through \(S\), since \(S\) dominates \(R\). Neither is inside a loop that leaves the region, so between entry and exit \(S\) and \(R\) each execute at most once, and they execute on the same paths. (c) The prologue saves the entry value of every callee-saved register the function writes, before any write (by (a), all writes are uses in \(U\)). The epilogue restores it after the last use. Paths that skip both never write those registers, so their values are unchanged.
3. Worked example¶
Argument assignment (calling conventions)¶
The signature void f(long a0, __int128 a1, double a2, long a3, double a4, long a5, long a6, char *a7, __int128 a8, int a9) (the cc3.ll box in §7). Its trace comes from the drill oracle (./course drill calling-convention uses the same code), and it was checked against llc:
| arg | type | SysV location | \(g, f, sp\) after | AAPCS64 location | NGRN, NSRN, NSAA after |
|---|---|---|---|---|---|
| a0 | i64 | rdi | 1, 0, 0 | x0 | 1, 0, 0 |
| a1 | i128 | rsi:rdx | 3, 0, 0 | x2:x3 (NGRN 1 → 2) | 4, 0, 0 |
| a2 | double | xmm0 | 3, 1, 0 | v0 | 4, 1, 0 |
| a3 | i64 | rcx | 4, 1, 0 | x4 | 5, 1, 0 |
| a4 | double | xmm1 | 4, 2, 0 | v1 | 5, 2, 0 |
| a5 | i64 | r8 | 5, 2, 0 | x5 | 6, 2, 0 |
| a6 | i64 | r9 | 6, 2, 0 | x6 | 7, 2, 0 |
| a7 | ptr | stack+0 | 6, 2, 8 | x7 | 8, 2, 0 |
| a8 | i128 | stack+16 (aligned 16) | 6, 2, 32 | stack+0 (NGRN 8, C.13) | 8, 2, 16 |
| a9 | i32 | stack+32 | 6, 2, 40 | stack+16 | 8, 2, 24 |
On SysV the pointer a7 already overflows to the stack, and a8 skips 8 bytes of padding to reach a 16-byte-aligned slot. On AAPCS64 a7 still gets x7, and the pair a1 skips x1 to start at an even register.
Try it
./course drill calling-convention --seed 5 --difficulty hard --solution generates signatures
like this one and grades both conventions.
Frame layout and frame lowering¶
leaf(i), which stores to and reads from a local long a[4] (the leaf.ll box):
| variant | frame decision | code |
|---|---|---|
| x86-64, default | leaf, 32-byte local ≤ 128: red zone; offset −40 keeps the array 16-aligned (entry %rsp ≡ 8 mod 16) |
movq $1, -40(%rsp): no %rsp adjustment |
x86-64, noredzone |
allocate: stackSize = 40 (32 + 8 to re-align) |
subq $40, %rsp … addq $40, %rsp |
x86-64, -frame-pointer=all |
push %rbp, %rbp = %rsp; the array at -32(%rbp) (entry misalignment 8 + push 8 = aligned) |
pushq %rbp; movq %rsp, %rbp … popq %rbp |
| AArch64 | no red zone in AAPCS64 on Linux: allocate 32 (a multiple of 16) | str x8, [sp, #-32]! … add sp, sp, #32 |
Prologue/epilogue insertion and shrink-wrapping¶
PEI on f(x, y) = g(x) + g(y) + x (x86-64). The register allocator used %rbx, %r14 and %r15 to keep y, x and g(x) alive across the calls. All three are callee-saved in SysV:
| step | what PEI does | result in the MIR (the call.ll box) |
|---|---|---|
| 1 | CSR = {rbx, r14, r15}; x86 saves them with push |
frame-setup PUSH64r killed $r15, $r14, $rbx |
| 2 | layout: 3 × 8 = 24 bytes of CSR, no locals; entry misalignment 8 + 24 = 32 ≡ 0 mod 16, so no sub is needed |
stackSize: 24, no SUB64ri |
| 3 | CFI for unwinding | CFI_INSTRUCTION def_cfa_offset 16/24/32, offset $rbx, -32, … |
| 4 | the call-frame pseudos around each call are deleted (reserved call frame of size 0) | no ADJCALLSTACKDOWN64/UP64 left |
| 5 | epilogue before RET: pops in reverse order |
$rbx = frame-destroy POP64r, $r14, $r15 |
Shrink-wrapping (the shrink.ll box): f(x) = x == 0 ? 0 : g(x) + x. Only the slow block uses a callee-saved register (%rbx keeps x across the call). So \(U = \{\mathit{slow}\}\) and Algorithm 21.9.10 gives \(S = R = \mathit{slow}\): slow dominates and post-dominates itself, and it is in no loop. The push/pop pair moves into slow, and the fast path returns without touching the stack. With -enable-shrink-wrap=false, the push is in the entry block and both return paths need a pop.
4. Invariants and correctness¶
Argument assignment (calling conventions)¶
Theorem 21.9.4 gives the structural properties. The semantic invariant is that both sides agree: the caller's call lowering and the callee's argument lowering must use the same function and the same interpretation of each location. The classic failure is a type that one side passes differently, for example i128 before LLVM's CCIfConsecutiveRegs fix, or _BitInt in older compilers. That produces a silent ABI mismatch that shows up only when the two sides are compiled by different compilers.
Frame layout and frame lowering¶
Proposition 21.9.7. Stack alignment is the invariant most often broken in practice: an unaligned stack at a call is harmless until the callee uses an aligned SSE store (movaps) to its frame, and then it crashes. The red zone is sound only in leaf functions (a call would push a return address into it) and only when signal handlers respect it. Kernels, which take interrupts on the current stack, compile with -mno-red-zone (the noredzone attribute).
Prologue/epilogue insertion and shrink-wrapping¶
Theorem 21.9.11 is the correctness condition. LLVM's ShrinkWrapImpl asserts exactly its hypotheses (MDT->dominates(Save, Restore) && MPDT->dominates(Restore, Save)) and repairs the points until they hold. The loop condition matters: a save point inside a loop would push the registers on every iteration, and the stack would grow without bound.
5. Complexity¶
Variables: \(n\) arguments, \(b\) blocks, \(k\) frame objects, \(c\) callee-saved registers.
| Technique | Time | Space | Justification |
|---|---|---|---|
| Argument assignment | \(\Theta(n)\) | \(O(1)\) counters | one pass (Theorem 21.9.4) |
| Frame layout and frame lowering | \(O(k \log k)\) with sorting by alignment, \(O(k)\) otherwise | \(O(k)\) | one pass over objects (Algorithm 21.9.6) |
| Prologue/epilogue insertion and shrink-wrapping | PEI \(O(\text{instructions})\); shrink-wrap \(O(b \log b)\) plus the dominator trees (Ch 15) | \(O(b)\) | nearest-common-dominator queries are \(O(\log b)\) with depth-first numbering, or amortized near constant |
A pathological family for shrink-wrapping. A function where a callee-saved register is used in the entry block and in a block deep inside a loop nest gets \(S\) = the entry: nothing is gained. Worse, the shrink-wrap benefit is zero whenever \(U\) includes a block that dominates everything. The pass then costs its dominator computations for nothing. That is why it bails out early when a CSR is used in the entry block.
Real-world scale. Chow's original paper motivated shrink-wrapping with the observation that many functions have a cheap early-exit path that never needs the callee-saved registers [Cho88]. The shrink.ll box is exactly such a function: the fast path goes from 4 instructions (push, test, branch, pop, plus the return) to 3.
6. Variants and refinements¶
Argument assignment (calling conventions)¶
- Aggregates (SysV classification of structs into eightbytes, AAPCS64 HFAs): more complex rules, usually handled in the front end (Clang's
ABIInfo) before LLVM sees the call. This splits the ABI between front end and back end, a well-known source of bugs (Ch 11). - Custom conventions (
fastcc,tailcc, GHC's convention,preserve_most): chosen per function by the compiler when both sides are known. More registers or tail calls, at the price of interoperability. - Apple arm64: stack arguments packed to their natural size, not 8 bytes. The same algorithm with a different C.16.
Frame layout and frame lowering¶
- Frame pointer or not (
-fomit-frame-pointervs-frame-pointer=all): one more free register, against simpler profiling and debugging. Distributions have been moving back to keeping frame pointers. - Stack realignment (over-aligned locals): the prologue aligns
spwithand, which forces a frame pointer (or base pointer) to reach incoming arguments. - Stack coloring and slot sharing (
StackColoring,StackSlotColoring): locals and spills with disjoint lifetimes share slots, giving smaller frames.
Prologue/epilogue insertion and shrink-wrapping¶
- Per-register shrink-wrapping (Chow's original [Cho88]) vs LLVM's whole-prologue shrink-wrapping: finer placement, more complex CFI and unwinding.
- Push/pop vs stores (x86
push, AArch64stppairs): smaller code vs better scheduling. Targets choose per register class. - Split restore points / multiple epilogues: LLVM can split a restore point to get tighter placement (
checkIfRestoreSplittable).
7. In real compilers¶
Argument assignment (calling conventions)¶
CC_X86_64_C in llvm/lib/Target/X86/X86CallingConv.td and CC_AArch64_AAPCS in llvm/lib/Target/AArch64/AArch64CallingConvention.td are compiled by TableGen's -gen-callingconv into CCAssignFns. CCState::AnalyzeFormalArguments and AnalyzeCallOperands in llvm/lib/CodeGen/CallingConvLower.cpp drive them, and the SysV i128 rule is the custom function CC_X86_64_I128 in X86CallingConv.cpp (LLVM 23.1.2) [LLVM-CC].
The SysV rules in TableGen, and where llc put the arguments
Reproduce (llc 23.1.2; the first command needs network access to GitHub):
curl -sL https://raw.githubusercontent.com/llvm/llvm-project/llvmorg-23.1.2/llvm/lib/Target/X86/X86CallingConv.td \
| sed -n '/^def CC_X86_64_C : CallingConv/,/^]>;/p' \
| grep -E 'CCIfType<\[i32\], CCAssignToReg|i128 can|not split|CCIfConsecutiveRegs|CCIfType<\[i64\], CCAssignToReg<\[RDI|CCAssignToReg<\[XMM0|CCIfType<\[i32, i64, f16, f32, f64\], CCAssignToStack'
cat > cc3.ll <<'EOF'
declare void @use(i64)
define void @f(i64 %a0, i128 %a1, double %a2, i64 %a3, double %a4, i64 %a5, i64 %a6, ptr %a7, i128 %a8, i32 %a9) {
%t = trunc i128 %a1 to i64
call void @use(i64 %t)
%u = lshr i128 %a1, 64
%u2 = trunc i128 %u to i64
call void @use(i64 %u2)
%p = ptrtoint ptr %a7 to i64
call void @use(i64 %p)
%t8 = trunc i128 %a8 to i64
call void @use(i64 %t8)
%z = zext i32 %a9 to i64
call void @use(i64 %z)
ret void
}
EOF
for t in x86_64-linux-gnu aarch64-linux-gnu; do
echo "== $t"
llc -O2 -mtriple=$t -stop-after=finalize-isel cc3.ll -o - \
| grep -E 'liveins: \$|offset: [0-9]+, size' | sed -E 's/, alignment.*//'
done
Output (complete):
CCIfType<[i32], CCAssignToReg<[EDI, ESI, EDX, ECX, R8D, R9D]>>,
// i128 can be either passed in two i64 registers, or on the stack, but
// not split across register and stack. Handle this with a custom function.
CCIfConsecutiveRegs<CCCustom<"CC_X86_64_I128">>>,
CCIfType<[i64], CCAssignToReg<[RDI, RSI, RDX, RCX, R8 , R9 ]>>,
CCAssignToReg<[XMM0, XMM1, XMM2, XMM3, XMM4, XMM5, XMM6, XMM7]>>>,
CCIfType<[i32, i64, f16, f32, f64], CCAssignToStack<8, 8>>,
== x86_64-linux-gnu
- { id: 0, type: default, offset: 32, size: 4
- { id: 1, type: default, offset: 24, size: 8
- { id: 2, type: default, offset: 16, size: 8
- { id: 3, type: default, offset: 0, size: 8
liveins: $rsi, $rdx
== aarch64-linux-gnu
- { id: 0, type: default, offset: 16, size: 4
- { id: 1, type: default, offset: 8, size: 8
- { id: 2, type: default, offset: 0, size: 8
liveins: $x2, $x3, $x7
What to notice: the TableGen rules are Algorithm 21.9.2's branches in order: i32/i64
to the six GPRs, FP to eight XMMs, the rest to 8-byte stack slots, and i128 through the
custom "not split across register and stack" rule. The fixedStack entries are the §3
table: on x86-64, a7 at 0, a8 at 16 and 24 (its two halves), a9 at 32. On AArch64, a8
at 0 and 8 and a9 at 16. The liveins list only the registers of arguments the function
reads: a1 in rsi:rdx versus x2:x3 (the even-register rule), and a7 in x7 on AArch64.
Frame layout and frame lowering¶
X86FrameLowering (llvm/lib/Target/X86/X86FrameLowering.cpp: hasFPImpl, has128ByteRedZone, emitPrologue, emitEpilogue) and AArch64FrameLowering implement TargetFrameLowering. The object layout is PEIImpl::calculateFrameObjectOffsets in llvm/lib/CodeGen/PrologEpilogInserter.cpp (LLVM 23.1.2) [LLVM-PEI].
Red zone, no red zone, frame pointer, and AArch64
Reproduce (llc 23.1.2):
cat > leaf.ll <<'EOF'
define i64 @leaf(i64 %i) {
%a = alloca [4 x i64], align 16
store volatile i64 1, ptr %a, align 16
%p = getelementptr inbounds i64, ptr %a, i64 %i
%v = load volatile i64, ptr %p, align 8
ret i64 %v
}
EOF
sed 's/define i64 @leaf(i64 %i) {/define i64 @leaf(i64 %i) noredzone {/' leaf.ll > leaf-nrz.ll
F='^\s*\.|^$|# -- |^#|// -- |^//'
llc -O2 -mtriple=x86_64-linux-gnu leaf.ll -o - | grep -vE "$F"
llc -O2 -mtriple=x86_64-linux-gnu leaf-nrz.ll -o - | grep -vE "$F"
llc -O2 -mtriple=x86_64-linux-gnu -frame-pointer=all leaf.ll -o - | grep -vE "$F"
llc -O2 -mtriple=aarch64-linux-gnu leaf.ll -o - | grep -vE "$F"
Output (complete):
leaf: # @leaf
movq $1, -40(%rsp)
movq -40(%rsp,%rdi,8), %rax
retq
leaf: # @leaf
subq $40, %rsp
movq $1, (%rsp)
movq (%rsp,%rdi,8), %rax
addq $40, %rsp
retq
leaf: # @leaf
pushq %rbp
movq %rsp, %rbp
movq $1, -32(%rbp)
movq -32(%rbp,%rdi,8), %rax
popq %rbp
retq
leaf: // @leaf
mov w8, #1 // =0x1
str x8, [sp, #-32]!
mov x8, sp
ldr x0, [x8, x0, lsl #3]
add sp, sp, #32
ret
What to notice: the four rows of the §3 table. In the red zone the array lives below
%rsp at −40, not −32, because %rsp is 8 mod 16 at entry and the array needs 16-byte
alignment (Proposition 21.9.7). Without the red zone, stackSize = 40 makes the array
16-aligned at (%rsp). With a frame pointer, the push %rbp already re-aligns, so −32 from
%rbp is aligned. AArch64 allocates 32 bytes with a pre-indexed store: its entry sp is
16-aligned and AAPCS64 on Linux has no red zone.
Prologue/epilogue insertion and shrink-wrapping¶
PEIImpl::run in llvm/lib/CodeGen/PrologEpilogInserter.cpp (spillCalleeSavedRegs, calculateFrameObjectOffsets, insertPrologEpilogCode, replaceFrameIndices), and ShrinkWrapImpl in llvm/lib/CodeGen/ShrinkWrap.cpp. The latter uses findNearestCommonDominator on the dominator and post-dominator trees and asserts MDT->dominates(Save, Restore) && MPDT->dominates(Restore, Save) (LLVM 23.1.2) [LLVM-PEI].
PEI on f(x, y), and shrink-wrapping on an early exit
Reproduce (llc 23.1.2; call.ll is in labs/ch21-mir/inputs/):
llc -O2 -mtriple=x86_64-linux-gnu -stop-after=prolog-epilog labs/ch21-mir/inputs/call.ll -o - \
| grep -E 'stackSize|PUSH64r|POP64r'
cat > shrink.ll <<'EOF'
declare i64 @g(i64)
define i64 @f(i64 %x) {
entry:
%z = icmp eq i64 %x, 0
br i1 %z, label %fast, label %slow
fast:
ret i64 0
slow:
%r = call i64 @g(i64 %x)
%s = add i64 %r, %x
ret i64 %s
}
EOF
F='^\s*\.(file|text|globl|type|size|section|p2align|ident|att_syntax|prefalign)|^\s*\.cfi|^$|# -- |^#'
llc -O2 -mtriple=x86_64-linux-gnu shrink.ll -o - | grep -vE "$F"
llc -O2 -mtriple=x86_64-linux-gnu -enable-shrink-wrap=false shrink.ll -o - | grep -vE "$F"
Output (complete):
stackSize: 24
frame-setup PUSH64r killed $r15, implicit-def $rsp, implicit $rsp
frame-setup PUSH64r killed $r14, implicit-def $rsp, implicit $rsp
frame-setup PUSH64r killed $rbx, implicit-def $rsp, implicit $rsp
$rbx = frame-destroy POP64r implicit-def $rsp, implicit $rsp
$r14 = frame-destroy POP64r implicit-def $rsp, implicit $rsp
$r15 = frame-destroy POP64r implicit-def $rsp, implicit $rsp
f: # @f
testq %rdi, %rdi
je .LBB0_1
pushq %rbx
movq %rdi, %rbx
callq g@PLT
addq %rbx, %rax
popq %rbx
retq
.LBB0_1: # %fast
xorl %eax, %eax
retq
.Lfunc_end0:
f: # @f
pushq %rbx
testq %rdi, %rdi
je .LBB0_1
movq %rdi, %rbx
callq g@PLT
addq %rbx, %rax
popq %rbx
retq
.LBB0_1: # %fast
xorl %eax, %eax
popq %rbx
retq
.Lfunc_end0:
What to notice: the PEI table of §3: stackSize: 24, the three pushes marked
frame-setup and the pops marked frame-destroy, in reverse order. In the shrink-wrapped f,
the pushq %rbx/popq %rbx pair sits inside the slow path (save point = restore point =
slow, Theorem 21.9.11), and the fast path is test, je, xor, ret. Without
shrink-wrapping the entry block saves %rbx and the fast path must restore it too.
8. Comparison¶
| Technique | Power / precision | Speed (asymptotic · practical) | Output / error quality | Implementation effort | Typical use |
|---|---|---|---|---|---|
| Argument assignment (calling conventions) | exact by construction (Theorem 21.9.4); aggregates need front-end help | \(\Theta(n)\) · negligible | ABI mismatches are silent until run time | declarative in TableGen; custom C++ for odd cases (i128) |
every call and function entry; drill calling-convention |
| Frame layout and frame lowering | aligned, disjoint slots (Proposition 21.9.7); red zone for leaves | \(O(k \log k)\) · negligible | smaller frames with slot coloring; alignment bugs crash late | moderate per target (TargetFrameLowering) |
every function; -frame-pointer, noredzone |
| Prologue/epilogue insertion and shrink-wrapping | correct under dominance conditions (Theorem 21.9.11) | linear plus dominator trees | faster early-exit paths (4 → 3 instructions in the box) | high: CFI, unwinding, funclets, stack probing | PEI in every function; shrink-wrapping at -O1 and above on most targets |
- Choose custom conventions (
fastcc) when both caller and callee are internal: more registers, tail calls. Keep the platform convention at every external boundary. - Keep frame pointers when profiling and debugging in production matter more than one register.
- Rely on shrink-wrapping when functions have cheap early exits. It costs nothing when it does not apply.
9. Assessment¶
| Technique | Quiz ids (solutions/quizzes/ch21.yaml) |
Drill | Flashcard tag | Exercises |
|---|---|---|---|---|
| Argument assignment (calling conventions) | cc-sysv-i128, cc-aapcs-even |
./course drill calling-convention |
calling-convention |
— |
| Frame layout and frame lowering | frame-red-zone, frame-align |
none: layouts depend on target details; the quiz uses the §3 table | frame-lowering |
E5 task 7 |
| Prologue/epilogue insertion and shrink-wrapping | pei-csr, shrinkwrap-points, llvm-where-shrinkwrap |
none: save/restore points are a dominator computation already drilled in Ch 15 (./course drill dominators, post-dominance) |
pei-shrinkwrap |
E5 task 7 |
Counting registers per argument instead of per class
SysV and AAPCS64 keep separate counters for integer and floating-point registers. In
f(double, long) the long goes to rdi/x0, not the second register, because the
double used an XMM/V register. The other classic slip is forgetting that an AAPCS64 i128
that does not fit sets NGRN to 8, so no later integer argument can use a register. SysV
leaves the remaining register available.
References¶
See the chapter references.