Skip to content

Lesson 21.9 — Calling conventions, frame lowering and prologue/epilogue insertion

Techniques: argument assignment (calling conventions); frame layout and frame lowering; prologue/epilogue insertion and shrink-wrapping · Lab: labs/ch21-mir (task 7) · Prerequisites: Lesson 21.5 (MIR), Ch 15 (dominators and post-dominators), Ch 11 (the source-level ABI) · Time: 4–5 hours

Instruction selection turns operations into instructions, but a function also has an interface and a frame. Its arguments arrive in registers and stack slots fixed by the platform's calling convention, the callee-saved registers it uses must be restored before it returns, its locals and spills need stack slots at known offsets, and the stack pointer must stay aligned at every call. The back end handles this in three places. Call and argument lowering happens during selection: TableGen-generated CCAssignFns decide where each argument goes. Frame lowering lays out the frame after register allocation, when the spill slots are known. Prologue/epilogue insertion (PEI) writes the code that builds and tears down the frame, and shrink-wrapping moves that code off the paths that do not need it. The running example is f(x, y) = g(x) + g(y) + x, whose prologue on x86-64 is three pushes and nothing else.

1. Problem and motivation

The inputs are a function's signature, the target's calling convention and ABI rules, and, after register allocation, the set of callee-saved registers the function clobbers and the sizes and alignments of its stack objects. The outputs are the location of every argument and return value (register, register pair, or stack offset), a frame layout (an offset for every stack object), and prologue and epilogue code. Two independently compiled functions must agree on all of this, so the rules are fixed by platform documents: the System V AMD64 psABI [SysV-ABI] and Arm's AAPCS64 [AAPCS64] for this chapter's two targets.

Argument assignment (calling conventions)

A calling convention maps a signature to locations. Both conventions here classify each argument (integer, floating point, aggregate by size) and hand out registers of the matching class in order, spilling to the stack when they run out. They differ in the details: SysV treats a 128-bit integer as two eightbytes of class INTEGER that must both fit in registers, while AAPCS64 rounds the next register number up to an even one first and, once an argument spills, sends all later integer arguments to the stack. LLVM encodes conventions in TableGen (CallingConv<[CCIfType<…, CCAssignToReg<…>>, …]>) and generates the assignment functions [LLVM-CC].

Frame layout and frame lowering

The frame holds the return address (x86) or the saved link register (AArch64), callee-saved registers, locals, spill slots and the outgoing argument area. The target's TargetFrameLowering decides whether a frame pointer is needed, how to align the stack, and whether a leaf function may use the red zone: 128 bytes below %rsp that the SysV ABI guarantees signal handlers will not clobber [SysV-ABI, §3.2.2], so a leaf function can use it without adjusting %rsp.

Prologue/epilogue insertion and shrink-wrapping

PEI runs after register allocation. It saves the callee-saved registers the function uses, assigns final offsets to all frame objects, emits prologue and epilogue code, and replaces abstract frame indices by sp/fp-relative addresses. Chow observed that saving registers at entry is wasteful when only some paths use them, and proposed shrink-wrapping: place the saves and restores around the region that needs them [Cho88]. LLVM's ShrinkWrap pass computes such points for the whole prologue and epilogue.

2. Definitions and algorithms

Definition 21.9.1 (Calling convention as location assignment)

A location is a register, a pair of registers lo:hi, or a stack offset in the outgoing argument area (offset 0 is the lowest address, the first stack argument; at the callee's entry, SysV puts it at 8(%rsp) above the return address and AAPCS64 at [sp]). A calling convention is a function \(\mathit{cc}(\tau_1, \dots, \tau_n) = (\ell_1, \dots, \ell_n)\) from the argument types to locations such that (i) no register is assigned twice, (ii) stack locations do not overlap and respect their alignment, and (iii) the function is computed left to right with finite state (the counters below). Caller and callee apply the same function, which is why separately compiled code interoperates.

Argument assignment (calling conventions)

Algorithm 21.9.2 (SysV x86-64 assignment for scalar arguments)

  • Input: argument types from {i32, i64, ptr, float, double, i128} (the scalar subset of [SysV-ABI, §3.2.3]).
  • Output: a location per argument.
  • Precondition: a non-variadic function.
  • Postcondition: properties (i)–(iii) of Definition 21.9.1 (Theorem 21.9.4).
  • Invariant: g GPRs of rdi, rsi, rdx, rcx, r8, r9 and f XMMs of xmm0–xmm7 are used, all by earlier arguments, and sp is the next free stack offset.
function SysV(types):
    g ← 0; f ← 0; sp ← 0
    for each type τ, left to right:
        if τ ∈ {i32, i64, ptr}:                      # class INTEGER, one eightbyte
            if g < 6: loc ← GPR[g]; g ← g + 1
            else:     sp ← Align(sp, 8); loc ← stack+sp; sp ← sp + 8
        elif τ ∈ {float, double}:                     # class SSE
            if f < 8: loc ← XMM[f]; f ← f + 1
            else:     sp ← Align(sp, 8); loc ← stack+sp; sp ← sp + 8
        elif τ = i128:                                # two INTEGER eightbytes
            if g ≤ 4: loc ← GPR[g]:GPR[g+1]; g ← g + 2
            else:     sp ← Align(sp, 16); loc ← stack+sp; sp ← sp + 16
                      # g unchanged: a later INTEGER argument may still use GPR[g]
        output loc

Algorithm 21.9.3 (AAPCS64 assignment for scalar arguments, stage C)

  • Input: the same types (C.1, C.9–C.11, C.13–C.17 of [AAPCS64, §6.8.2]).
  • Output: a location per argument (x registers for integers, v registers for floating point).
  • Precondition: a non-variadic function; standard AAPCS64 (Linux), not Apple's variant, which packs stack arguments by their natural size.
  • Postcondition: properties (i)–(iii) of Definition 21.9.1.
  • Invariant: NGRN, NSRN and NSAA are the next general register, SIMD/FP register and stack address, as in the standard.
function AAPCS64(types):
    NGRN ← 0; NSRN ← 0; NSAA ← 0                        # stage A
    for each type τ, left to right:
        if τ ∈ {float, double}:                          # C.1, C.5, C.6
            if NSRN < 8: loc ← v[NSRN]; NSRN ← NSRN + 1
            else:        NSAA ← Align(NSAA, 8); loc ← stack+NSAA; NSAA ← NSAA + 8
        elif τ ∈ {i32, i64, ptr}:                        # C.9, then C.13, C.14, C.16, C.17
            if NGRN < 8: loc ← x[NGRN]; NGRN ← NGRN + 1
            else:        NSAA ← Align(NSAA, 8); loc ← stack+NSAA; NSAA ← NSAA + 8
        elif τ = i128:                                   # C.10, C.11, else C.13, C.14, C.17
            NGRN ← RoundUpToEven(NGRN)
            if NGRN < 7: loc ← x[NGRN]:x[NGRN+1]; NGRN ← NGRN + 2
            else:        NGRN ← 8                        # C.13: no more GPR arguments
                         NSAA ← Align(NSAA, 16); loc ← stack+NSAA; NSAA ← NSAA + 16
        output loc

Theorem 21.9.4 (Both assignments are calling conventions)

Algorithms 21.9.2 and 21.9.3 satisfy properties (i)–(iii) of Definition 21.9.1: registers are never assigned twice, stack slots are disjoint and aligned (8 bytes, or 16 for i128), and the result depends only on the list of types.

Proof

(i) Every register assignment uses the current counter value(s) and then increases the counter past them (by 1, or by 2 for a pair). AAPCS64's RoundUpToEven and NGRN ← 8 only increase NGRN. The counters never decrease, so no register index is handed out twice. SysV's "g unchanged" branch assigns no register at all. (ii) Each stack assignment first aligns the offset to the slot's alignment (8 or 16), assigns the interval \([\mathit{sp}, \mathit{sp} + \mathit{size})\), and then advances past it. The offset never decreases, so the intervals are disjoint and in increasing order. (iii) The algorithms read only the types and their own counters. They are deterministic single passes, so caller and callee, running them on the same signature, compute the same locations. This is what makes the separately compiled code agree.

The two conventions disagree on an i128

For (i64 × 5, i128, i64): SysV gives rdi, rsi, rdx, rcx, r8, stack+0, r9. The i128 needs two GPRs, only r9 is left, so it goes to the stack and r9 stays free for the last i64. AAPCS64 gives x0–x4, x6:x7, stack+0. NGRN = 5 is rounded up to 6, the pair takes x6:x7, and the last i64 finds NGRN = 8 and goes to the stack. Both results were checked with llc 23.1.2 (tools/course/tests/test_ch21.py, CallingConventions).

Frame layout and frame lowering

Definition 21.9.5 (Frame, frame objects, frame index)

A frame object is a stack slot with a size and an alignment: a local (alloca), a spill slot, a callee-saved register slot, or a fixed object at a known offset (an incoming stack argument). Before layout, instructions refer to frame objects by frame index (%stack.0, %fixed-stack.1 in MIR). The frame is the region between the incoming stack pointer and the stack pointer after the prologue. Its size stackSize is chosen so that every object has an offset aligned to its alignment and the stack pointer satisfies the ABI's alignment (16 bytes) at every call.

Algorithm 21.9.6 (Frame object layout, after calculateFrameObjectOffsets)

  • Input: the frame objects; callee-saved slots; the stack alignment \(A\) (16); whether the function calls anything; red-zone eligibility.
  • Output: an offset per object (negative, relative to the incoming stack pointer) and stackSize.
  • Precondition: fixed objects have their ABI offsets already.
  • Postcondition: objects do not overlap, each offset is a multiple of the object's alignment, and, if the function makes calls, the stack pointer is \(A\)-aligned at each call (Proposition 21.9.7).
  • Invariant: off is the lowest address allocated so far (it grows downwards).
function LayoutFrame(objects, A):
    off ← the size of the fixed area (return address, pushed callee-saved registers)
    for each callee-saved slot, then each local and spill slot (in an order that groups
    objects of equal alignment to reduce padding):
        off ← AlignUp(off + size(obj), align(obj))       # growing downwards
        offset(obj) ← −off
    stackSize ← off − fixed area
    if the function has calls:
        round stackSize up so that (incoming misalignment + fixed area + stackSize) ≡ 0 mod A
    if the function is a leaf, needs no realignment, and stackSize ≤ 128 and the red zone
       is allowed: use the red zone (do not move the stack pointer)
    return offsets, stackSize

Proposition 21.9.7 (Layout invariants)

Algorithm 21.9.6 gives pairwise disjoint objects at aligned offsets. If the function makes a call, the stack pointer is \(A\)-aligned at the call instruction.

Proof

Disjointness and alignment. Each object gets the interval \([-\mathit{off}, -\mathit{off} + \mathrm{size})\) after off has been increased past the previous objects and rounded up to a multiple of its alignment. So the intervals are disjoint (decreasing addresses) and each start is aligned. Call alignment. At entry the ABI guarantees a known misalignment (SysV: %rsp ≡ 8 mod 16, because the call pushed the return address; AAPCS64: sp ≡ 0 mod 16). The prologue moves the stack pointer by the fixed area plus stackSize, and the rounding step makes the sum of the misalignment and that movement a multiple of \(A\). Outgoing argument areas are reserved inside stackSize (a reserved call frame), so the stack pointer does not move between the prologue and the calls.

Prologue/epilogue insertion and shrink-wrapping

Algorithm 21.9.8 (Prologue/epilogue insertion)

  • Input: a register-allocated function with frame indices and call-frame pseudos (ADJCALLSTACKDOWN/UP); the set \(\mathit{CSR}\) of callee-saved registers the function writes; save point \(S\) and restore point \(R\) (the entry and the exits unless shrink-wrapped).
  • Output: the function with prologue at \(S\), epilogue(s) at \(R\), and concrete addresses.
  • Precondition: \(S\) dominates and \(R\) post-dominates every instruction that uses a register in \(\mathit{CSR}\) or a frame index (Theorem 21.9.11).
  • Postcondition: every path from entry to exit through \(S\) saves each register of \(\mathit{CSR}\) before using it and restores it before returning. Every frame index is replaced by sp- or fp-relative addressing.
  • Invariant: the offsets from Algorithm 21.9.6 are final before any frame index is rewritten.
function PEI(F):
    CSR ← the callee-saved registers written in F
    assign a spill slot to each r ∈ CSR (or use push/pop, x86)
    LayoutFrame(frame objects of F, A)                        # Algorithm 21.9.6
    at S: emit the prologue: save CSR (push or stp), set up fp if needed,
          allocate the frame (sub sp, stackSize), emit CFI directives for unwinding
    at each block in R: emit the epilogue: free the frame, restore CSR, (return)
    replace each ADJCALLSTACKDOWN/UP by a stack adjustment, or delete it when the call frame
    is reserved in the fixed frame
    for each instruction with a frame index fi: rewrite fi as base + offset(fi), with
    base = sp or fp, as the target's eliminateFrameIndex chooses

Definition 21.9.9 (Save and restore points)

Let \(U\) be the set of blocks that use a callee-saved register the function writes, or that touch the frame. A save point is a block \(S\) that dominates every block of \(U\). A restore point is a block \(R\) that post-dominates every block of \(U\). Additionally \(S\) dominates \(R\), \(R\) post-dominates \(S\), and neither lies inside a loop that does not contain all of \(U\), so the prologue and epilogue run once per execution of the region.

Algorithm 21.9.10 (Shrink-wrapping, after LLVM's ShrinkWrap)

  • Input: a function after register allocation; its dominator and post-dominator trees and loop information.
  • Output: save point \(S\) and restore point \(R\), or the entry and the exits (no change).
  • Precondition: the target supports shrink-wrapping for this function (no stack realignment, no exception-handling funclets, and so on).
  • Postcondition: \(S\) and \(R\) satisfy Definition 21.9.9 (Theorem 21.9.11).
  • Invariant: after each update, \(S\) dominates and \(R\) post-dominates every block of \(U\) seen so far.
function ShrinkWrap(F):
    S ← none; R ← none
    for each block B that uses a CSR or the frame:
        S ← (S = none) ? B : NearestCommonDominator(S, B)
        R ← (R = none) ? B : NearestCommonPostDominator(R, B)
    if S = none: return (no frame needed at all)
    repeat until stable:                              # repair until A, B and C hold:
        if not (S dominates R): S ← NearestCommonDominator(S, R)
        if not (R post-dominates S): R ← NearestCommonPostDominator(R, S)
        if S is in a loop L that does not contain R (or vice versa): move S to the
           dominator of L's header outside L, R to the post-dominating exit of L
    return (S, R)

Theorem 21.9.11 (Shrink-wrapped prologues and epilogues are correct)

If \(S\) and \(R\) satisfy Definition 21.9.9 and the prologue is placed at the start of \(S\) and the epilogue at the end of \(R\), then on every execution: (a) every use of a callee-saved register or the frame happens after the prologue and before the epilogue; (b) the prologue executes iff the epilogue executes, exactly once each; hence (c) every callee-saved register has its entry value when the function returns.

Proof

(a) A use in block \(B \in U\) is reached only by paths through \(S\), because \(S\) dominates \(B\), and every path from \(B\) to the exit passes through \(R\), because \(R\) post-dominates \(B\). So each use is between an execution of \(S\) and an execution of \(R\). (b) Every path through \(S\) continues to the exit through \(R\), since \(R\) post-dominates \(S\). Every path through \(R\) came through \(S\), since \(S\) dominates \(R\). Neither is inside a loop that leaves the region, so between entry and exit \(S\) and \(R\) each execute at most once, and they execute on the same paths. (c) The prologue saves the entry value of every callee-saved register the function writes, before any write (by (a), all writes are uses in \(U\)). The epilogue restores it after the last use. Paths that skip both never write those registers, so their values are unchanged.

3. Worked example

Argument assignment (calling conventions)

The signature void f(long a0, __int128 a1, double a2, long a3, double a4, long a5, long a6, char *a7, __int128 a8, int a9) (the cc3.ll box in §7). Its trace comes from the drill oracle (./course drill calling-convention uses the same code), and it was checked against llc:

arg type SysV location \(g, f, sp\) after AAPCS64 location NGRN, NSRN, NSAA after
a0 i64 rdi 1, 0, 0 x0 1, 0, 0
a1 i128 rsi:rdx 3, 0, 0 x2:x3 (NGRN 1 → 2) 4, 0, 0
a2 double xmm0 3, 1, 0 v0 4, 1, 0
a3 i64 rcx 4, 1, 0 x4 5, 1, 0
a4 double xmm1 4, 2, 0 v1 5, 2, 0
a5 i64 r8 5, 2, 0 x5 6, 2, 0
a6 i64 r9 6, 2, 0 x6 7, 2, 0
a7 ptr stack+0 6, 2, 8 x7 8, 2, 0
a8 i128 stack+16 (aligned 16) 6, 2, 32 stack+0 (NGRN 8, C.13) 8, 2, 16
a9 i32 stack+32 6, 2, 40 stack+16 8, 2, 24

On SysV the pointer a7 already overflows to the stack, and a8 skips 8 bytes of padding to reach a 16-byte-aligned slot. On AAPCS64 a7 still gets x7, and the pair a1 skips x1 to start at an even register.

Try it

./course drill calling-convention --seed 5 --difficulty hard --solution generates signatures like this one and grades both conventions.

Frame layout and frame lowering

leaf(i), which stores to and reads from a local long a[4] (the leaf.ll box):

variant frame decision code
x86-64, default leaf, 32-byte local ≤ 128: red zone; offset −40 keeps the array 16-aligned (entry %rsp ≡ 8 mod 16) movq $1, -40(%rsp): no %rsp adjustment
x86-64, noredzone allocate: stackSize = 40 (32 + 8 to re-align) subq $40, %rsp … addq $40, %rsp
x86-64, -frame-pointer=all push %rbp, %rbp = %rsp; the array at -32(%rbp) (entry misalignment 8 + push 8 = aligned) pushq %rbp; movq %rsp, %rbp … popq %rbp
AArch64 no red zone in AAPCS64 on Linux: allocate 32 (a multiple of 16) str x8, [sp, #-32]! … add sp, sp, #32

Prologue/epilogue insertion and shrink-wrapping

PEI on f(x, y) = g(x) + g(y) + x (x86-64). The register allocator used %rbx, %r14 and %r15 to keep y, x and g(x) alive across the calls. All three are callee-saved in SysV:

step what PEI does result in the MIR (the call.ll box)
1 CSR = {rbx, r14, r15}; x86 saves them with push frame-setup PUSH64r killed $r15, $r14, $rbx
2 layout: 3 × 8 = 24 bytes of CSR, no locals; entry misalignment 8 + 24 = 32 ≡ 0 mod 16, so no sub is needed stackSize: 24, no SUB64ri
3 CFI for unwinding CFI_INSTRUCTION def_cfa_offset 16/24/32, offset $rbx, -32, …
4 the call-frame pseudos around each call are deleted (reserved call frame of size 0) no ADJCALLSTACKDOWN64/UP64 left
5 epilogue before RET: pops in reverse order $rbx = frame-destroy POP64r, $r14, $r15

Shrink-wrapping (the shrink.ll box): f(x) = x == 0 ? 0 : g(x) + x. Only the slow block uses a callee-saved register (%rbx keeps x across the call). So \(U = \{\mathit{slow}\}\) and Algorithm 21.9.10 gives \(S = R = \mathit{slow}\): slow dominates and post-dominates itself, and it is in no loop. The push/pop pair moves into slow, and the fast path returns without touching the stack. With -enable-shrink-wrap=false, the push is in the entry block and both return paths need a pop.

4. Invariants and correctness

Argument assignment (calling conventions)

Theorem 21.9.4 gives the structural properties. The semantic invariant is that both sides agree: the caller's call lowering and the callee's argument lowering must use the same function and the same interpretation of each location. The classic failure is a type that one side passes differently, for example i128 before LLVM's CCIfConsecutiveRegs fix, or _BitInt in older compilers. That produces a silent ABI mismatch that shows up only when the two sides are compiled by different compilers.

Frame layout and frame lowering

Proposition 21.9.7. Stack alignment is the invariant most often broken in practice: an unaligned stack at a call is harmless until the callee uses an aligned SSE store (movaps) to its frame, and then it crashes. The red zone is sound only in leaf functions (a call would push a return address into it) and only when signal handlers respect it. Kernels, which take interrupts on the current stack, compile with -mno-red-zone (the noredzone attribute).

Prologue/epilogue insertion and shrink-wrapping

Theorem 21.9.11 is the correctness condition. LLVM's ShrinkWrapImpl asserts exactly its hypotheses (MDT->dominates(Save, Restore) && MPDT->dominates(Restore, Save)) and repairs the points until they hold. The loop condition matters: a save point inside a loop would push the registers on every iteration, and the stack would grow without bound.

5. Complexity

Variables: \(n\) arguments, \(b\) blocks, \(k\) frame objects, \(c\) callee-saved registers.

Technique Time Space Justification
Argument assignment \(\Theta(n)\) \(O(1)\) counters one pass (Theorem 21.9.4)
Frame layout and frame lowering \(O(k \log k)\) with sorting by alignment, \(O(k)\) otherwise \(O(k)\) one pass over objects (Algorithm 21.9.6)
Prologue/epilogue insertion and shrink-wrapping PEI \(O(\text{instructions})\); shrink-wrap \(O(b \log b)\) plus the dominator trees (Ch 15) \(O(b)\) nearest-common-dominator queries are \(O(\log b)\) with depth-first numbering, or amortized near constant

A pathological family for shrink-wrapping. A function where a callee-saved register is used in the entry block and in a block deep inside a loop nest gets \(S\) = the entry: nothing is gained. Worse, the shrink-wrap benefit is zero whenever \(U\) includes a block that dominates everything. The pass then costs its dominator computations for nothing. That is why it bails out early when a CSR is used in the entry block.

Real-world scale. Chow's original paper motivated shrink-wrapping with the observation that many functions have a cheap early-exit path that never needs the callee-saved registers [Cho88]. The shrink.ll box is exactly such a function: the fast path goes from 4 instructions (push, test, branch, pop, plus the return) to 3.

6. Variants and refinements

Argument assignment (calling conventions)

  • Aggregates (SysV classification of structs into eightbytes, AAPCS64 HFAs): more complex rules, usually handled in the front end (Clang's ABIInfo) before LLVM sees the call. This splits the ABI between front end and back end, a well-known source of bugs (Ch 11).
  • Custom conventions (fastcc, tailcc, GHC's convention, preserve_most): chosen per function by the compiler when both sides are known. More registers or tail calls, at the price of interoperability.
  • Apple arm64: stack arguments packed to their natural size, not 8 bytes. The same algorithm with a different C.16.

Frame layout and frame lowering

  • Frame pointer or not (-fomit-frame-pointer vs -frame-pointer=all): one more free register, against simpler profiling and debugging. Distributions have been moving back to keeping frame pointers.
  • Stack realignment (over-aligned locals): the prologue aligns sp with and, which forces a frame pointer (or base pointer) to reach incoming arguments.
  • Stack coloring and slot sharing (StackColoring, StackSlotColoring): locals and spills with disjoint lifetimes share slots, giving smaller frames.

Prologue/epilogue insertion and shrink-wrapping

  • Per-register shrink-wrapping (Chow's original [Cho88]) vs LLVM's whole-prologue shrink-wrapping: finer placement, more complex CFI and unwinding.
  • Push/pop vs stores (x86 push, AArch64 stp pairs): smaller code vs better scheduling. Targets choose per register class.
  • Split restore points / multiple epilogues: LLVM can split a restore point to get tighter placement (checkIfRestoreSplittable).

7. In real compilers

Argument assignment (calling conventions)

CC_X86_64_C in llvm/lib/Target/X86/X86CallingConv.td and CC_AArch64_AAPCS in llvm/lib/Target/AArch64/AArch64CallingConvention.td are compiled by TableGen's -gen-callingconv into CCAssignFns. CCState::AnalyzeFormalArguments and AnalyzeCallOperands in llvm/lib/CodeGen/CallingConvLower.cpp drive them, and the SysV i128 rule is the custom function CC_X86_64_I128 in X86CallingConv.cpp (LLVM 23.1.2) [LLVM-CC].

The SysV rules in TableGen, and where llc put the arguments

Reproduce (llc 23.1.2; the first command needs network access to GitHub):

curl -sL https://raw.githubusercontent.com/llvm/llvm-project/llvmorg-23.1.2/llvm/lib/Target/X86/X86CallingConv.td \
  | sed -n '/^def CC_X86_64_C : CallingConv/,/^]>;/p' \
  | grep -E 'CCIfType<\[i32\], CCAssignToReg|i128 can|not split|CCIfConsecutiveRegs|CCIfType<\[i64\], CCAssignToReg<\[RDI|CCAssignToReg<\[XMM0|CCIfType<\[i32, i64, f16, f32, f64\], CCAssignToStack'
cat > cc3.ll <<'EOF'
declare void @use(i64)
define void @f(i64 %a0, i128 %a1, double %a2, i64 %a3, double %a4, i64 %a5, i64 %a6, ptr %a7, i128 %a8, i32 %a9) {
  %t = trunc i128 %a1 to i64
  call void @use(i64 %t)
  %u = lshr i128 %a1, 64
  %u2 = trunc i128 %u to i64
  call void @use(i64 %u2)
  %p = ptrtoint ptr %a7 to i64
  call void @use(i64 %p)
  %t8 = trunc i128 %a8 to i64
  call void @use(i64 %t8)
  %z = zext i32 %a9 to i64
  call void @use(i64 %z)
  ret void
}
EOF
for t in x86_64-linux-gnu aarch64-linux-gnu; do
  echo "== $t"
  llc -O2 -mtriple=$t -stop-after=finalize-isel cc3.ll -o - \
    | grep -E 'liveins: \$|offset: [0-9]+, size' | sed -E 's/, alignment.*//'
done

Output (complete):

  CCIfType<[i32], CCAssignToReg<[EDI, ESI, EDX, ECX, R8D, R9D]>>,
  // i128 can be either passed in two i64 registers, or on the stack, but
  // not split across register and stack. Handle this with a custom function.
           CCIfConsecutiveRegs<CCCustom<"CC_X86_64_I128">>>,
  CCIfType<[i64], CCAssignToReg<[RDI, RSI, RDX, RCX, R8 , R9 ]>>,
            CCAssignToReg<[XMM0, XMM1, XMM2, XMM3, XMM4, XMM5, XMM6, XMM7]>>>,
  CCIfType<[i32, i64, f16, f32, f64], CCAssignToStack<8, 8>>,
== x86_64-linux-gnu
  - { id: 0, type: default, offset: 32, size: 4
  - { id: 1, type: default, offset: 24, size: 8
  - { id: 2, type: default, offset: 16, size: 8
  - { id: 3, type: default, offset: 0, size: 8
    liveins: $rsi, $rdx
== aarch64-linux-gnu
  - { id: 0, type: default, offset: 16, size: 4
  - { id: 1, type: default, offset: 8, size: 8
  - { id: 2, type: default, offset: 0, size: 8
    liveins: $x2, $x3, $x7

What to notice: the TableGen rules are Algorithm 21.9.2's branches in order: i32/i64 to the six GPRs, FP to eight XMMs, the rest to 8-byte stack slots, and i128 through the custom "not split across register and stack" rule. The fixedStack entries are the §3 table: on x86-64, a7 at 0, a8 at 16 and 24 (its two halves), a9 at 32. On AArch64, a8 at 0 and 8 and a9 at 16. The liveins list only the registers of arguments the function reads: a1 in rsi:rdx versus x2:x3 (the even-register rule), and a7 in x7 on AArch64.

Frame layout and frame lowering

X86FrameLowering (llvm/lib/Target/X86/X86FrameLowering.cpp: hasFPImpl, has128ByteRedZone, emitPrologue, emitEpilogue) and AArch64FrameLowering implement TargetFrameLowering. The object layout is PEIImpl::calculateFrameObjectOffsets in llvm/lib/CodeGen/PrologEpilogInserter.cpp (LLVM 23.1.2) [LLVM-PEI].

Red zone, no red zone, frame pointer, and AArch64

Reproduce (llc 23.1.2):

cat > leaf.ll <<'EOF'
define i64 @leaf(i64 %i) {
  %a = alloca [4 x i64], align 16
  store volatile i64 1, ptr %a, align 16
  %p = getelementptr inbounds i64, ptr %a, i64 %i
  %v = load volatile i64, ptr %p, align 8
  ret i64 %v
}
EOF
sed 's/define i64 @leaf(i64 %i) {/define i64 @leaf(i64 %i) noredzone {/' leaf.ll > leaf-nrz.ll
F='^\s*\.|^$|# -- |^#|// -- |^//'
llc -O2 -mtriple=x86_64-linux-gnu leaf.ll -o - | grep -vE "$F"
llc -O2 -mtriple=x86_64-linux-gnu leaf-nrz.ll -o - | grep -vE "$F"
llc -O2 -mtriple=x86_64-linux-gnu -frame-pointer=all leaf.ll -o - | grep -vE "$F"
llc -O2 -mtriple=aarch64-linux-gnu leaf.ll -o - | grep -vE "$F"

Output (complete):

leaf:                                   # @leaf
    movq    $1, -40(%rsp)
    movq    -40(%rsp,%rdi,8), %rax
    retq
leaf:                                   # @leaf
    subq    $40, %rsp
    movq    $1, (%rsp)
    movq    (%rsp,%rdi,8), %rax
    addq    $40, %rsp
    retq
leaf:                                   # @leaf
    pushq   %rbp
    movq    %rsp, %rbp
    movq    $1, -32(%rbp)
    movq    -32(%rbp,%rdi,8), %rax
    popq    %rbp
    retq
leaf:                                   // @leaf
    mov w8, #1                          // =0x1
    str x8, [sp, #-32]!
    mov x8, sp
    ldr x0, [x8, x0, lsl #3]
    add sp, sp, #32
    ret

What to notice: the four rows of the §3 table. In the red zone the array lives below %rsp at −40, not −32, because %rsp is 8 mod 16 at entry and the array needs 16-byte alignment (Proposition 21.9.7). Without the red zone, stackSize = 40 makes the array 16-aligned at (%rsp). With a frame pointer, the push %rbp already re-aligns, so −32 from %rbp is aligned. AArch64 allocates 32 bytes with a pre-indexed store: its entry sp is 16-aligned and AAPCS64 on Linux has no red zone.

Prologue/epilogue insertion and shrink-wrapping

PEIImpl::run in llvm/lib/CodeGen/PrologEpilogInserter.cpp (spillCalleeSavedRegs, calculateFrameObjectOffsets, insertPrologEpilogCode, replaceFrameIndices), and ShrinkWrapImpl in llvm/lib/CodeGen/ShrinkWrap.cpp. The latter uses findNearestCommonDominator on the dominator and post-dominator trees and asserts MDT->dominates(Save, Restore) && MPDT->dominates(Restore, Save) (LLVM 23.1.2) [LLVM-PEI].

PEI on f(x, y), and shrink-wrapping on an early exit

Reproduce (llc 23.1.2; call.ll is in labs/ch21-mir/inputs/):

llc -O2 -mtriple=x86_64-linux-gnu -stop-after=prolog-epilog labs/ch21-mir/inputs/call.ll -o - \
  | grep -E 'stackSize|PUSH64r|POP64r'
cat > shrink.ll <<'EOF'
declare i64 @g(i64)
define i64 @f(i64 %x) {
entry:
  %z = icmp eq i64 %x, 0
  br i1 %z, label %fast, label %slow
fast:
  ret i64 0
slow:
  %r = call i64 @g(i64 %x)
  %s = add i64 %r, %x
  ret i64 %s
}
EOF
F='^\s*\.(file|text|globl|type|size|section|p2align|ident|att_syntax|prefalign)|^\s*\.cfi|^$|# -- |^#'
llc -O2 -mtriple=x86_64-linux-gnu shrink.ll -o - | grep -vE "$F"
llc -O2 -mtriple=x86_64-linux-gnu -enable-shrink-wrap=false shrink.ll -o - | grep -vE "$F"

Output (complete):

  stackSize:       24
    frame-setup PUSH64r killed $r15, implicit-def $rsp, implicit $rsp
    frame-setup PUSH64r killed $r14, implicit-def $rsp, implicit $rsp
    frame-setup PUSH64r killed $rbx, implicit-def $rsp, implicit $rsp
    $rbx = frame-destroy POP64r implicit-def $rsp, implicit $rsp
    $r14 = frame-destroy POP64r implicit-def $rsp, implicit $rsp
    $r15 = frame-destroy POP64r implicit-def $rsp, implicit $rsp
f:                                      # @f
    testq   %rdi, %rdi
    je  .LBB0_1
    pushq   %rbx
    movq    %rdi, %rbx
    callq   g@PLT
    addq    %rbx, %rax
    popq    %rbx
    retq
.LBB0_1:                                # %fast
    xorl    %eax, %eax
    retq
.Lfunc_end0:
f:                                      # @f
    pushq   %rbx
    testq   %rdi, %rdi
    je  .LBB0_1
    movq    %rdi, %rbx
    callq   g@PLT
    addq    %rbx, %rax
    popq    %rbx
    retq
.LBB0_1:                                # %fast
    xorl    %eax, %eax
    popq    %rbx
    retq
.Lfunc_end0:

What to notice: the PEI table of §3: stackSize: 24, the three pushes marked frame-setup and the pops marked frame-destroy, in reverse order. In the shrink-wrapped f, the pushq %rbx/popq %rbx pair sits inside the slow path (save point = restore point = slow, Theorem 21.9.11), and the fast path is test, je, xor, ret. Without shrink-wrapping the entry block saves %rbx and the fast path must restore it too.

8. Comparison

Technique Power / precision Speed (asymptotic · practical) Output / error quality Implementation effort Typical use
Argument assignment (calling conventions) exact by construction (Theorem 21.9.4); aggregates need front-end help \(\Theta(n)\) · negligible ABI mismatches are silent until run time declarative in TableGen; custom C++ for odd cases (i128) every call and function entry; drill calling-convention
Frame layout and frame lowering aligned, disjoint slots (Proposition 21.9.7); red zone for leaves \(O(k \log k)\) · negligible smaller frames with slot coloring; alignment bugs crash late moderate per target (TargetFrameLowering) every function; -frame-pointer, noredzone
Prologue/epilogue insertion and shrink-wrapping correct under dominance conditions (Theorem 21.9.11) linear plus dominator trees faster early-exit paths (4 → 3 instructions in the box) high: CFI, unwinding, funclets, stack probing PEI in every function; shrink-wrapping at -O1 and above on most targets
  • Choose custom conventions (fastcc) when both caller and callee are internal: more registers, tail calls. Keep the platform convention at every external boundary.
  • Keep frame pointers when profiling and debugging in production matter more than one register.
  • Rely on shrink-wrapping when functions have cheap early exits. It costs nothing when it does not apply.

9. Assessment

Technique Quiz ids (solutions/quizzes/ch21.yaml) Drill Flashcard tag Exercises
Argument assignment (calling conventions) cc-sysv-i128, cc-aapcs-even ./course drill calling-convention calling-convention —
Frame layout and frame lowering frame-red-zone, frame-align none: layouts depend on target details; the quiz uses the §3 table frame-lowering E5 task 7
Prologue/epilogue insertion and shrink-wrapping pei-csr, shrinkwrap-points, llvm-where-shrinkwrap none: save/restore points are a dominator computation already drilled in Ch 15 (./course drill dominators, post-dominance) pei-shrinkwrap E5 task 7

Counting registers per argument instead of per class

SysV and AAPCS64 keep separate counters for integer and floating-point registers. In f(double, long) the long goes to rdi/x0, not the second register, because the double used an XMM/V register. The other classic slip is forgetting that an AAPCS64 i128 that does not fit sets NGRN to 8, so no later integer argument can use a register. SysV leaves the remaining register available.

References

See the chapter references.