Lesson 9.5 — Calls, calling conventions, attributes, memory effects and intrinsics¶
Techniques: calls and calling conventions (including varargs and tail calls); parameter, return and function attributes, with memory effects; intrinsics · Pebble uses:
callto runtime functions,noundef/nonnullfrom the lowering, overflow intrinsics for checked arithmetic · Lab: E5, E6, E8, F6 · Prerequisites: Lesson 9.3 (invoke), Lesson 9.4 · Time: 3–4 hours
clang compiles long first(const long *p) { return p[0]; } at -O2 to
define dso_local i64 @first(ptr nofree noundef readonly captures(none) %0) local_unnamed_addr #1
attributes #1 = { mustprogress nofree norecurse nosync nounwind willreturn memory(argmem: read) … }
Most of that line is not the function but facts about it: the pointer is never freed, never written, never captured; the function only reads memory reachable from its argument, always returns, never throws. None of it changes what @first computes; all of it changes what a caller may do around a call to it (keep a value in a register across the call, delete the call if its result is unused, move it out of a loop). This lesson covers the call instruction, the conventions that make caller and callee agree, the attribute system, and intrinsics: calls to functions whose meaning LLVM knows.
1. Problem and motivation¶
A call is where an optimizer loses sight of the program. Without information about the callee, a call may read or write any memory, loop forever, throw, or return anything, so the optimizer must assume the worst on both sides of it. LLVM's answer has three parts. The calling convention fixes the machine-level contract. Attributes record facts about parameters, return values and whole functions, either from the front end (the ABI's signext, C's restrict → noalias) or inferred by passes (function-attrs). Intrinsics give names and exact semantics to operations the core instruction set lacks (overflow arithmetic, memcpy, population count), without adding instructions [LLVM-LangRef, §Intrinsic Functions].
Calls and calling conventions¶
In LLVM IR the callee of a call is just a pointer value, and since opaque pointers the call itself spells the function type (call i32 (ptr, ...) @printf(…)). How arguments reach registers and stack is the calling convention's job, chosen per function and per call: ccc (the platform C ABI, the default), fastcc (anything faster, usable when all callers are known), coldcc, tailcc/swifttailcc (guaranteed tail calls), ghccc and more [LLVM-LangRef, §Calling Conventions]. The C ABI itself is only partly expressed in IR: the front end lowers struct arguments and varargs according to the target's ABI document, and attributes such as zeroext, byval, sret tell the back end the rest.
Parameter and function attributes¶
Attributes were in LLVM from the start (nounwind, readonly); the modern design adds precise, composable facts: memory(argmem: read) replaces the old readonly/argmemonly flags, captures(…) replaced nocapture in LLVM 21, and value facts such as range(i32 0, 10), nofpclass(nan), noundef let callers and callees exchange what they know. Violating an attribute is undefined behavior or makes the value poison (Lesson 9.7); that is what makes them usable for optimization.
Intrinsics¶
Intrinsics are functions named llvm.* whose semantics the LangRef defines, which the optimizer understands and the back end lowers, often to one instruction. They extend the IR without touching every pass and the bitcode format [LLVM-LangRef, §Intrinsic Functions]: llvm.smul.with.overflow.i64 instead of a new instruction, llvm.memcpy instead of a loop, llvm.x86.* for target-specific operations.
2. Definitions and algorithms¶
Calls and calling conventions¶
Definition 9.5.1 (Call)
%r = [tail | musttail | notail] call [fast-math flags] [cc] [ret attrs] τ | τ(τ1, …[, ...]) %f(args) [fn attrs] [operand bundles]
calls the function at pointer \(f\) with the listed arguments, each \(\tau_i\ [\text{param attrs}]\ v_i\). The spelled function type must match the callee's definition in the number and types of fixed parameters; for a variadic type τ (τ1, …, τk, ...) the call may pass extra arguments, whose types it chooses. The calling convention \(cc\) of the call and of the callee must be equal: otherwise the behavior is undefined. tail marks a call that does not access the caller's stack objects (a tail-call candidate); musttail requires a tail call and constrains the signature so one is always possible.
Definition 9.5.2 (Variadic calls)
A call to a variadic function passes each extra argument as a first-class value of the type the caller writes. C's default argument promotions are therefore applied by the front end: float becomes double (fpext), integers narrower than int become i32 (sext or zext by the C type's signedness). Inside a variadic function, llvm.va_start, va_arg and llvm.va_end walk the arguments through a target-specific va_list object; on targets where va_arg cannot be lowered generically (x86-64 System V), the front end expands it into explicit loads from the va_list structure.
Parameter and function attributes¶
Definition 9.5.3 (Attributes and their kinds of guarantee)
An attribute is attached to a parameter, a return value, or a function, at a definition, a declaration or a call site (call-site attributes add to the callee's). Each attribute states a property \(P\) and what happens if \(P\) fails:
| kind | examples | if violated |
|---|---|---|
| ABI | zeroext, signext, inreg, byval(τ), sret(τ), swiftself |
wrong code (the attribute defines the ABI) |
| value facts | noundef, nonnull, align N, dereferenceable(N), range(τ lo, hi), nofpclass(…) |
noundef, dereferenceable: UB; the others: the value is poison |
| pointer use | noalias, captures(…), readonly, writeonly, nofree, returned, initializes(…), dead_on_unwind, writable |
UB |
| function behavior | memory(…), nounwind, willreturn, mustprogress, nosync, norecurse, speculatable, noreturn, convergent |
UB |
| code generation | noinline, alwaysinline, optnone, cold, minsize, uwtable, string attributes "target-cpu"=… |
none (hints and settings) |
Definition 9.5.4 (Memory effects)
Let \(\mathit{Loc} = \{\texttt{argmem}, \texttt{inaccessiblemem}, \texttt{errnomem}, \texttt{target\_mem0}, \texttt{target\_mem1}, \texttt{other}\}\) and \(\mathit{Acc} = \{\texttt{none} \sqsubset \texttt{read}, \texttt{write} \sqsubset \texttt{readwrite}\}\) (the lattice \(\mathcal{P}(\{R, W\})\) under \(\subseteq\)). A memory effect is a function \(e : \mathit{Loc} \to \mathit{Acc}\), ordered pointwise (\(e \sqsubseteq e'\) iff \(e(l) \sqsubseteq e'(l)\) for all \(l\)), with join \(\sqcup\) pointwise. The attribute memory(…) on \(f\) states that every execution of a call to \(f\) makes only accesses allowed by \(e_f\), where argmem means "based on a pointer argument", other everything not otherwise listed; memory(none) is \(\bot\), the absent attribute is \(\top\) = memory(readwrite). A read location that is written is immediate UB; a write location read before being written yields poison.
Algorithm 9.5.5 (Inferring memory effects for a call-graph SCC)
- Input: a strongly connected component \(C\) of the call graph, whose callees outside \(C\) already have memory effects (or \(\top\) if unknown).
- Output: \(e_f\) for every \(f \in C\).
- Precondition: every \(f \in C\) has an exact definition (not interposable, Lesson 9.1), so its body is the one that runs.
- Postcondition: \(e_f\) over-approximates the accesses of every execution of \(f\) (Theorem 9.5.8), and all members of \(C\) get the same effect.
- Invariant: after visiting instructions \(I\), \(E = \bigsqcup_{i \in I} \mathrm{eff}(i)\).
function InferSCC(C):
E ← ⊥ # memory(none)
for f in C:
for i in instructions(f):
E ← E ⊔ eff(i)
for f in C: e_f ← E
return e
function eff(i):
case i of
load/store/atomic through pointer q:
acc ← read (load) | write (store) | readwrite (atomicrmw, cmpxchg)
if q is based only on a local alloca or a constant global: return ⊥ # not observable
if q is based only on pointer arguments: return [argmem ↦ acc]
return [other ↦ acc, argmem ↦ acc] # unknown pointer: may alias anything
call g with arguments a1..ak:
if g ∈ C: return ⊥ # accounted for by g's own instructions
e ← e_g (⊤ if unknown)
replace e's argmem part by the effect on the locations of a1..ak (as for a pointer q above)
e(argmem) ← e(argmem) ⊔ e(other) # g's "other" memory may be what our arguments point to
return e
anything else: return ⊥
Intrinsics¶
Definition 9.5.6 (Intrinsic, overloading and mangling)
An intrinsic is a function whose name starts with llvm. and is listed in the LangRef (and llvm/include/llvm/IR/Intrinsics*.td), with a fixed signature scheme and attributes. It may only be called, never defined or address-taken. An overloaded intrinsic has type parameters; its name is the base name followed by . + the mangled name of each overloaded type: i64, f32, p0 for ptr, v4i32 for <4 x i32>, nxv4i32 for <vscale x 4 x i32>. For example llvm.smul.with.overflow.i64 : (i64, i64) → { i64, i1 } and llvm.memcpy.p0.p0.i64. A declaration whose type does not fit the scheme is rejected by the verifier (lab F6).
3. Worked example¶
A variadic call, by hand (lab E5). printf("#%d %c %s: %.2f\n", id, grade, name, score) with int id, char grade, const char *name, float score:
| C argument | C type | default promotion | IR argument |
|---|---|---|---|
| format | const char * |
none | ptr @fmt |
id |
int |
none | i32 %id |
grade |
char (signed on x86-64) |
to int: sign-extend |
i32 (sext i8 %grade) |
name |
const char * |
none | ptr %name |
score |
float |
to double |
double (fpext float %score) |
The call is call i32 (ptr, ...) @printf(ptr @fmt, i32 %id, i32 %g, ptr %name, double %s). Forgetting the fpext passes a float in a register printf reads as a double: the lab's mutation test prints garbage for exactly this.
Inferring memory effects (Algorithm 9.5.5) on four functions (the real-world box in "Parameter and function attributes" runs opt -passes=function-attrs on them). The call graph has no cycles, so each function is its own SCC, processed callees first:
| SCC | instruction | \(\mathrm{eff}(i)\) | \(E\) after | result |
|---|---|---|---|---|
@reads_arg |
load i32, ptr %p |
argmem ↦ read | argmem: read | memory(argmem: read) |
@writes_global |
store i32 %v, ptr @g |
other ↦ write (a global that is not constant) | other: write | memory(write, argmem: none, …) |
@pure |
add |
⊥ | ⊥ | memory(none) |
@calls_both |
call @reads_arg(ptr %p) |
argmem ↦ read (callee's argmem mapped to %p, an argument) |
argmem: read | |
call @writes_global(i32 %v) |
other ↦ write; and since %p might point to @g, argmem ↦ write |
argmem: readwrite, other: write | memory(write, argmem: readwrite, …) |
LLVM prints the default access first (write for all locations not listed), then the exceptions. inaccessiblemem, errnomem and the target locations stay none because no instruction touches them.
Remangling an intrinsic. A declaration i64 @llvm.ctpop.i32(i64) has the right shape (ctpop is overloaded on one integer type) but the wrong suffix; LLVM 23's reader renames it to @llvm.ctpop.i64 instead of rejecting it (box below). A declaration whose shape is wrong, such as i64 @llvm.sadd.with.overflow.i64(i64, i64) (the result must be { i64, i1 }), is rejected: that is lab task F6.
Try it
./course drill flags --difficulty medium asks which flags an intrinsic-free rewrite may keep; the overflow intrinsics of lab E6 are the alternative when you need the overflow bit rather than a poison value.
4. Invariants and correctness¶
Calls and calling conventions¶
Proposition 9.5.7 (A call with a mismatched convention may be replaced by unreachable)
If a call site's calling convention differs from its callee's, replacing the call and everything after it in the block by unreachable (or any code) is a refinement.
Proof
By the LangRef, "the calling convention of any pair of dynamic caller/callee must match, or the behavior of the program is undefined" [LLVM-LangRef, §Calling Conventions]. When the callee is known statically, every execution that reaches the call has undefined behavior from that point on, and any behavior refines UB (Lesson 9.7). InstCombine does this, marking the spot with store i1 true, ptr poison (its canonical "this is unreachable" store) and returning poison (box below).
Parameter and function attributes¶
Theorem 9.5.8 (Soundness of Algorithm 9.5.5)
Assume every call to a function outside \(C\) respects that function's memory effect. Then for every \(f \in C\), every execution of a call to \(f\) only makes accesses allowed by the computed \(e_f\).
Proof
By induction on the length of the execution of the call (number of executed instructions, including those in nested calls to members of \(C\)). Every access is made either (i) directly by an instruction \(i\) of some \(g \in C\) on the call stack, or (ii) inside a call from some \(g \in C\) to a function outside \(C\). In case (i), \(\mathrm{eff}(i) \sqsubseteq E\) by the invariant, and \(\mathrm{eff}(i)\) covers the access: a pointer based on an argument of \(g\) is argument memory for \(g\); if \(g \ne f\), \(g\) was called (transitively) from \(f\) and that argument is itself based on one of \(f\)'s arguments, a local of some frame (unobservable), or an arbitrary pointer, which other/argmem covers in \(E\) as computed for the call instruction in the caller. In case (ii), the callee respects its effect by assumption, and \(\mathrm{eff}\) of that call instruction maps it into \(E\). So every access is allowed by \(E = e_f\).
Proposition 9.5.9 (SCC members share one effect)
The system \(e_f = \mathrm{local}(f) \sqcup \bigsqcup_{g \in \mathrm{callees}(f)} e_g\) restricted to one SCC \(C\) (with outside callees fixed) has the least solution \(e_f = \bigsqcup_{h \in C} \mathrm{local}(h)\) for all \(f \in C\), which is what Algorithm 9.5.5 computes.
Proof
Let \(U = \bigsqcup_{h \in C} \mathrm{local}(h)\). \(U\) is a solution: for \(f \in C\), \(\mathrm{local}(f) \sqsubseteq U\) and every \(e_g\) with \(g \in C\) equals \(U\), while callees outside \(C\) contribute terms already included in the \(\mathrm{local}\) parts (Algorithm 9.5.5 maps outside calls into \(\mathrm{eff}\) of the call instruction); so the right-hand side is \(U\). \(U\) is least: in any solution, \(e_f \sqsupseteq e_g\) whenever \(g\) is reachable from \(f\) (by induction on the path length, using \(e_f \sqsupseteq e_{g'}\) for each direct callee \(g'\)); in an SCC every \(h\) is reachable from every \(f\), so \(e_f \sqsupseteq \mathrm{local}(h)\) for all \(h \in C\), i.e. \(e_f \sqsupseteq U\).
When it breaks. The precondition "exact definition" is essential: for a linkonce_odr or weak function, the body in this module may not be the one that runs (Lesson 9.1), and a less optimized copy elsewhere may access more memory. FunctionAttrs therefore infers attributes for such functions only from call sites or not at all (F->hasExactDefinition()).
Intrinsics¶
Proposition 9.5.10 (Mangling is injective on signatures)
For a fixed base name, two declarations of an overloaded intrinsic with different overloaded types have different mangled names.
Proof
The suffix is the concatenation of . + \(\mathrm{mangle}(\tau_i)\) over the overloaded types in a fixed order, and \(\mathrm{mangle}\) is injective on types: integers i\(N\), pointers p\(a\), fixed vectors v\(n\,\mathrm{mangle}(e)\), scalable vectors nxv\(n\,\mathrm{mangle}(e)\), floats by name, literal structs by sl_ + the mangled fields + s, and each of these forms starts with a distinct prefix and determines its components, so a mangled string parses back uniquely. (Unnamed identified structs get an extra numeric suffix to keep the names distinct.) Different overloaded types therefore give different suffixes. Conversely, LLVM can recompute the correct name from a declaration's type, which is how the reader repairs a wrong suffix (Section 3).
5. Complexity¶
| Technique | Time (worst) | Time (typical) | Space | Variables |
|---|---|---|---|---|
| Call-site / callee attribute lookup | \(O(\log a)\) | \(O(1)\) (attribute lists are uniqued) | \(O(a)\) | \(a\) = attributes |
| Memory-effect inference (Algorithm 9.5.5) | \(O(i)\) per SCC, \(O(I)\) for the module | linear | \(O(1)\) per function | \(i\) = instructions of the SCC, \(I\) = of the module |
| Intrinsic lookup by name | \(O(\lvert\text{name}\rvert \log n)\) | cached IntrinsicID per function |
— | \(n\) = intrinsics |
Proposition 9.5.11 (Cost of attribute inference over the call graph)
Running Algorithm 9.5.5 on every SCC in bottom-up order costs \(O(I + E_{cg})\) plus the cost of the "based on" queries, where \(E_{cg}\) is the number of call edges.
Proof
Tarjan's algorithm yields the SCCs in reverse topological order in \(O(F + E_{cg})\) (\(F\) = functions). Each instruction is visited once, in its function's SCC, and does \(O(1)\) lattice work: a join of two maps over the 6-element set \(\mathit{Loc}\). Proposition 9.5.9 removes any need to iterate inside an SCC. The "based on" test (getUnderlyingObject) is bounded by a fixed search depth in LLVM, so it is \(O(1)\) per access.
Pathological input. A module in which every function may call every other (a single SCC of size \(F\), e.g. an interpreter whose opcodes call a common dispatch) gets one effect for all functions: one function that writes a global makes all of them memory(write), and precision collapses. Splitting such SCCs needs context-sensitive analysis (Ch 20).
6. Variants and refinements¶
Calls and calling conventions¶
- Guaranteed tail calls (
musttail,tailcc,swifttailcc): required by languages with proper tail calls (Scheme, Swift async); trade-off: the signature constraints ofmusttail. - Operand bundles (
[ "deopt"(…) ],[ "funclet"(token %p) ],[ "convergencectrl"(…) ]): extra operands with their own semantics, for deoptimization state in JITs and for EH funclets. - Calling-convention rewriting: GlobalOpt turns internal functions whose address is not taken to
fastcc(box below); trade-off: none for internal functions, impossible for exported ones.
Parameter and function attributes¶
- Attributor (
-passes=attributor): a fixpoint framework that infers many attributes together (and with more precision thanfunction-attrs), at higher compile time. captures(…)with components (address,provenance,read_provenance,ret: …) replaced the booleannocapturein LLVM 21:captures(ret: address, provenance)in the boxes says "the pointer escapes only through the return value".- Value-range and FP-class attributes (
range,nofpclass) let callers use callee knowledge without inlining (Ch 20).
Intrinsics¶
- Constrained FP intrinsics (
llvm.experimental.constrained.fadd): FP operations that respect rounding modes and exceptions, for#pragma STDC FENV_ACCESS. - Vector-predicated intrinsics (
llvm.vp.*): vector operations with a mask and an explicit vector length, for RISC-V V and SVE. - Target intrinsics (
llvm.x86.*,llvm.aarch64.*): a direct line to machine instructions; trade-off: opaque to most target-independent passes.
7. In real compilers¶
Calls and calling conventions¶
LLVM
llvm/include/llvm/IR/InstrTypes.h — CallBase (common base of CallInst, InvokeInst, CallBrInst: callee, arguments, attributes, bundles); llvm/include/llvm/IR/CallingConv.h — the numbered conventions (C = 0, Fast = 8, Cold = 9, Tail = 18, SwiftTail = 20, …); llvm/lib/Transforms/IPO/GlobalOpt.cpp — ChangeCalleesToFastCall and hasChangeableCC (LLVM 23.1.2) [LLVM-CallingConv].
- clang lowers the C ABI in
clang/lib/CodeGen/Targets/X86.cpp(X86_64ABIInfo::classifyArgumentType, andEmitVAArg, which expandsva_arginto loads from%struct.__va_list_taginstead of emitting the IR instruction). - rustc chooses the Rust ABI in
compiler/rustc_target/src/callconv/mod.rsand marksextern "C"functions with LLVM'sccc.
Find where LLVM does it. In llvm/lib/Transforms/IPO/GlobalOpt.cpp, find hasChangeableCCImpl. Question: name one condition that prevents a function's convention from being changed to fastcc.
Varargs: promotions at the call, two ways to walk the arguments
Reproduce (clang 23.1.2):
cat > va.c <<'EOF'
int printf(const char *, ...);
void report(int id, float score, char grade) { printf("#%d %c %.2f\n", id, grade, score); }
long sum_ints(int n, ...) {
__builtin_va_list ap;
__builtin_va_start(ap, n);
long s = 0;
for (int i = 0; i < n; i++) s += __builtin_va_arg(ap, int);
__builtin_va_end(ap);
return s;
}
EOF
clang-23 --target=x86_64-unknown-linux-gnu -O1 -S -emit-llvm va.c -o - \
| grep -E "call i32 \(ptr, \.\.\.\)|fpext|sext i8|va_list_tag = |llvm.va_start|llvm.va_end| va_arg |^define"
echo ---
clang-23 --target=aarch64-apple-macosx14 -O1 -S -emit-llvm va.c -o - | grep -E "va_arg|llvm.va_start|^define"
Output (complete):
%struct.__va_list_tag = type { i32, i32, ptr, ptr }
define dso_local void @report(i32 noundef %0, float noundef %1, i8 noundef signext %2) local_unnamed_addr #0 {
%4 = sext i8 %2 to i32
%5 = fpext float %1 to double
%6 = tail call i32 (ptr, ...) @printf(ptr noundef nonnull dereferenceable(1) @.str, i32 noundef %0, i32 noundef %4, double noundef %5)
define dso_local i64 @sum_ints(i32 noundef %0, ...) local_unnamed_addr #2 {
call void @llvm.va_start.p0(ptr nonnull %2)
call void @llvm.va_end.p0(ptr %2)
declare void @llvm.va_start.p0(ptr) #4
declare void @llvm.va_end.p0(ptr) #4
---
define void @report(i32 noundef %0, float noundef %1, i8 noundef signext %2) local_unnamed_addr #0 {
define i64 @sum_ints(i32 noundef %0, ...) local_unnamed_addr #2 {
call void @llvm.va_start.p0(ptr nonnull %2)
%9 = va_arg ptr %2, i32
declare void @llvm.va_start.p0(ptr) #4
What to notice: the call spells the callee's type (ptr, ...) and passes the promoted i32 and double (Definition 9.5.2, the Section 3 table). On x86-64 clang expands va_arg itself using the ABI's __va_list_tag (register save area + overflow area); on AArch64 macOS, where va_list is a plain pointer, it emits the IR va_arg instruction and leaves the lowering to LLVM.
Conventions: fastcc for internal functions, UB for mismatches
Reproduce (opt 23.1.2):
cat > cc.ll <<'EOF'
define internal i32 @helper(i32 %x) noinline {
%y = mul i32 %x, 3
ret i32 %y
}
define i32 @api(i32 %x) {
%r = call i32 @helper(i32 %x)
ret i32 %r
}
define fastcc i32 @callee_fast(i32 %x) local_unnamed_addr {
ret i32 %x
}
define i32 @mismatch(i32 %x) {
%r = call i32 @callee_fast(i32 %x)
ret i32 %r
}
EOF
opt -mtriple=x86_64-unknown-linux-gnu -passes='globalopt,instcombine' -S cc.ll | grep -v "^;\|^$\|source_filename\|^target"
Output (complete):
define internal fastcc i32 @helper(i32 %x) unnamed_addr #0 {
%y = mul i32 %x, 3
ret i32 %y
}
define i32 @api(i32 %x) local_unnamed_addr {
%r = call fastcc i32 @helper(i32 %x)
ret i32 %r
}
define fastcc i32 @callee_fast(i32 %x) local_unnamed_addr {
ret i32 %x
}
define i32 @mismatch(i32 %x) local_unnamed_addr {
store i1 true, ptr poison, align 1
ret i32 poison
}
attributes #0 = { noinline }
What to notice: every caller of the internal @helper is visible, so GlobalOpt switches both the definition and the call to fastcc. The ccc call to a fastcc function is undefined behavior (Proposition 9.5.7); InstCombine replaces it by its canonical unreachable marker, a store to a poison pointer.
Parameter and function attributes¶
LLVM
llvm/include/llvm/IR/Attributes.td — every attribute (def NoUndef, def Memory, def Captures, def Range, def ZExt, …) with where it may appear; llvm/include/llvm/Support/ModRef.h — MemoryEffects and IRMemLocation (Definition 9.5.4); llvm/lib/Transforms/IPO/FunctionAttrs.cpp — addMemoryAttrs and checkFunctionMemoryAccess implement Algorithm 9.5.5 over each SCC (LLVM 23.1.2) [LLVM-Attributes, LLVM-FunctionAttrs].
- GCC 15: the same facts as function attributes and flags (
pure,const,ECF_NOTHROW,ECF_LEAF) computed byipa-pure-const.ccand, for parameters, by IPA mod/ref (ipa-modref.cc), which records per-argument access summaries much likeargmem. - rustc emits
noaliasfor&mut Tand&Twithout interior mutability,noundef,nonnull,alignanddereferenceablefor references (compiler/rustc_codegen_llvm/src/abi.rs), which is where many of Rust's aliasing guarantees reach LLVM.
Find where LLVM does it. In llvm/include/llvm/Support/ModRef.h, find enum class IRMemLocation. Question: which locations does it list (ignore the First/Last helpers), and which one stands for everything the others do not cover?
Attributes clang infers, and opt -passes=function-attrs by itself
Reproduce (clang 23.1.2, opt 23.1.2):
cat > attrs.c <<'EOF'
int square(int x) { return x * x; }
long first(const long *p) { return p[0]; }
void zero(long *p, long n) { for (long i = 0; i < n; i++) p[i] = 0; }
int counter;
void tick(void) { counter++; }
char *id(char *p) { return p; }
EOF
clang-23 --target=x86_64-unknown-linux-gnu -O2 -S -emit-llvm attrs.c -o - | grep -E "^define|^attributes #[0-3]" \
| sed 's/ uwtable.*}/ … }/'
cat > fa.ll <<'EOF'
@g = global i32 0
define i32 @reads_arg(ptr %p) {
%v = load i32, ptr %p
ret i32 %v
}
define void @writes_global(i32 %v) {
store i32 %v, ptr @g
ret void
}
define i32 @calls_both(ptr %p) {
%v = call i32 @reads_arg(ptr %p)
call void @writes_global(i32 %v)
ret i32 %v
}
define i32 @pure(i32 %x) {
%y = add i32 %x, 1
ret i32 %y
}
EOF
opt -passes=function-attrs -S fa.ll | grep -E "^define|^attributes"
Output (complete):
define dso_local i32 @square(i32 noundef %0) local_unnamed_addr #0 {
define dso_local i64 @first(ptr nofree noundef readonly captures(none) %0) local_unnamed_addr #1 {
define dso_local void @zero(ptr nofree noundef writeonly captures(none) %0, i64 noundef %1) local_unnamed_addr #2 {
define dso_local void @tick() local_unnamed_addr #3 {
define dso_local noundef ptr @id(ptr nofree noundef readnone returned captures(ret: address, provenance) %0) local_unnamed_addr #0 {
attributes #0 = { mustprogress nofree norecurse nosync nounwind willreturn memory(none) … }
attributes #1 = { mustprogress nofree norecurse nosync nounwind willreturn memory(argmem: read) … }
attributes #2 = { mustprogress nofree norecurse nosync nounwind willreturn memory(argmem: write) … }
attributes #3 = { mustprogress nofree norecurse nosync nounwind willreturn memory(readwrite, argmem: none, inaccessiblemem: none, target_mem: none) … }
define i32 @reads_arg(ptr nofree readonly captures(none) %p) #0 {
define void @writes_global(i32 %v) #1 {
define i32 @calls_both(ptr nofree readonly captures(none) %p) #2 {
define i32 @pure(i32 %x) #3 {
attributes #0 = { mustprogress nofree norecurse nosync nounwind willreturn memory(argmem: read) }
attributes #1 = { mustprogress nofree norecurse nosync nounwind willreturn memory(write, argmem: none, inaccessiblemem: none, target_mem: none) }
attributes #2 = { mustprogress nofree norecurse nosync nounwind willreturn memory(write, argmem: readwrite, inaccessiblemem: none, target_mem: none) }
attributes #3 = { mustprogress nofree norecurse nosync nounwind willreturn memory(none) }
What to notice: the second half is the Section 3 trace: each function's memory(…) is the join of its instructions' effects, and @calls_both inherits @reads_arg's argmem: read mapped to its own argument plus @writes_global's write of other memory (which may be what %p points to, hence argmem: readwrite). The printer groups the two target locations as target_mem. @id returns its argument: returned and captures(ret: address, provenance) say so, which lets callers replace the result by the argument.
Intrinsics¶
LLVM
llvm/include/llvm/IR/Intrinsics.td — def int_sadd_with_overflow, def int_smul_with_overflow, … (signatures and attributes); llvm/lib/IR/AutoUpgrade.cpp — UpgradeIntrinsicFunction rewrites outdated intrinsic declarations and calls (Lesson 9.6); the verifier's signature check is Verifier::visitIntrinsicCall (LLVM 23.1.2) [LLVM-Intrinsics, LLVM-AutoUpgrade].
- GCC 15: built-in functions (
BUILT_IN_POPCOUNT,BUILT_IN_MEMCPY,gcc/builtins.def) and internal functions (IFN_ADD_OVERFLOW,gcc/internal-fn.def) play the role of intrinsics in GIMPLE. - rustc 1.94:
core::intrinsicsfunctions map to LLVM intrinsics incompiler/rustc_codegen_llvm/src/intrinsic.rs, and checked arithmetic (checked_add, debug-mode overflow checks) goes throughchecked_binopincompiler/rustc_codegen_llvm/src/builder.rs, which builds the namellvm.{s,u}{add,sub,mul}.with.overflow(and for unsigned subtraction deliberately emitssub+icmp, LLVM's canonical form).
Find where LLVM does it. In llvm/include/llvm/IR/Intrinsics.td, find int_smul_with_overflow. Question: what is its return type, written in the TableGen notation?
C builtins become intrinsics; the reader repairs a wrong name and declares missing ones
Reproduce (clang 23.1.2, opt 23.1.2):
cat > intr.c <<'EOF'
int checked(long a, long b, long *out) { return __builtin_mul_overflow(a, b, out); }
int bits(unsigned x) { return __builtin_popcount(x) + __builtin_clz(x | 1); }
unsigned swap(unsigned x) { return __builtin_bswap32(x); }
void copy(char *d, const char *s, unsigned long n) { __builtin_memcpy(d, s, n); }
EOF
clang-23 --target=x86_64-unknown-linux-gnu -O1 -S -emit-llvm intr.c -o - | grep -oE "@llvm\.[a-z0-9_.]+" | sort -u
printf 'declare i64 @llvm.ctpop.i32(i64)\ndefine i64 @f(i64 %%x) {\n %%c = call i64 @llvm.ctpop.i32(i64 %%x)\n ret i64 %%c\n}\n' > mangle.ll
opt -S mangle.ll | grep -E "^declare| %c = call"
printf 'define i32 @g(i32 %%x) {\n %%c = call i32 @llvm.ctpop.i32(i32 %%x)\n ret i32 %%c\n}\n' > nodecl.ll
opt -S nodecl.ll | grep -E "^declare"
Output (complete):
@llvm.bswap.i32
@llvm.ctlz.i32
@llvm.ctpop.i32
@llvm.memcpy.p0.p0.i64
@llvm.smul.with.overflow.i64
%c = call i64 @llvm.ctpop.i64(i64 %x)
declare i64 @llvm.ctpop.i64(i64) #0
declare i32 @llvm.ctpop.i32(i32) #0
What to notice: each builtin is an overloaded intrinsic with its mangled suffix (Definition 9.5.6): three integer suffixes and, for memcpy, two pointer address spaces and the length type. The misnamed @llvm.ctpop.i32(i64) was renamed to the mangling of its actual type (Proposition 9.5.10 run backwards), and a call to an undeclared intrinsic got a declaration with the intrinsic's attributes from Intrinsics.td.
8. Comparison¶
| Technique | Power / precision | Speed (asymptotic · practical) | Output / error quality | Implementation effort | Typical use |
|---|---|---|---|---|---|
| Calls and calling conventions | any callee pointer; per-call convention; varargs with caller-chosen types; guaranteed tail calls | \(O(1)\) per call in the IR; ABI lowering in front and back ends | mismatches are UB, not errors | high (ABI lowering is the front end's job) | every call; fastcc for internal functions |
| Parameter and function attributes | precise facts per parameter, return and function; memory effects per location kind | uniqued lists, \(O(1)\); inference linear per SCC | wrong attributes are UB or poison, silently | moderate (inference), low (use) | ABI (zeroext), alias info (noalias), IPO facts |
| Intrinsics | new operations without new instructions; exact semantics; overloading by type | lowered to 1–few machine instructions or libcalls | verifier checks declarations; reader remangles names | low per intrinsic (TableGen) | overflow, bit manipulation, memcpy, SIMD, target ops |
Choose an intrinsic over a hand-written pattern when one exists (llvm.umul.with.overflow over a division check): passes understand it and back ends select the best instruction. Choose attributes you can justify: noundef, nonnull and range on values you have checked let callers drop checks, but a wrong one is a miscompile. Choose fastcc only for functions all of whose callers you control; the optimizer does it for you for internal functions.
9. Assessment¶
| Technique | Quiz ids (solutions/quizzes/ch09.yaml) |
Drill | Flashcard tag | Exercises |
|---|---|---|---|---|
| Calls and calling conventions | vararg-promotion, cc-mismatch |
— (see below) | calls |
E5, E8 |
| Parameter and function attributes | memory-effects-join, find-memloc |
— (see below) | attributes |
E6 (zeroext) |
| Intrinsics | intrinsic-mangling, find-smul-overflow |
./course drill ir-validity (types of calls) |
intrinsics |
E6, E7, F6 |
Calls and attributes are exercised by the lab (E5, E6, E8) and by computational quiz questions (memory-effects-join is a small run of Algorithm 9.5.5); a random-generation drill would mostly test ABI trivia, so the chapter uses its drill budget on GEP, poison and validity.
An attribute is a promise, and the optimizer believes it
Writing nonnull or noundef on a parameter that can be null or poison is not a "hint that might be wrong"; it makes the call undefined behavior, and the optimizer deletes checks on that basis. When you are not sure, leave the attribute out: its absence is always correct.
References¶
See the chapter references.