
Rileva i caricamenti di memoria inventati dal compilatore che trasformano C sicuro in vulnerabilità TOCTOU. Include audit automatici del codice sorgente, analisi binaria basata su Unicorn e sweep di compilatori/architetture/flag su oltre 100 progetti.
Minimal proofs-of-concept for compiler-invented loads — situations where a single C-level read of memory compiles to two or more architectural reads of the same address.
We call these cat-states – if the pattern is used across a check/use boundary, the compiler may emit a time-of-check time-of-use vulnerability (TOCTOU). Similar to the Schrödinger's cat thought experiment, the C code could either be vulnerable, or not vulnerable – the only way to decide which is to "measure" it with compilation.
Whether that second load actually appears is a decision of the compiler's
backend cost model, not of the C language — it turns on the exact compiler,
version, target architecture, and optimization flags, none of which the source
reveals. The same file reads clean under one toolchain and double-reads under
another (Clang loads once where GCC loads twice; -flto revives a double-read a
module boundary had suppressed). There is no source-level signal to audit for —
the only way to know whether a given build is vulnerable is to compile it and
look at the loads it emits.
Each .c file is one minimal cat-state, annotated with the exact assembly it
emits and the compiler / arch / flags that trigger it. The files are driven by
alpha-lab/matrix_runner.py through the // @key: value front-matter described
in SPEC.md.
m0_signext_reload.c is the example that opens the
top-level README. The C source reads *p exactly once:
unsigned int g(unsigned short *p)
{
short t = *p; /* copy p into a local for safekeeping */
return (unsigned short)t - t;
}
The C code appears to have no TOCTOU: t is a local variable and cannot be
modified by an attacker. The developer's mental model is that short t = *p
takes a single, stable snapshot of the memory, and that every later use of
t — both the zero-extended (unsigned short)t and the sign-extended t —
reads back that one captured value.
But the compiler quietly breaks that model:
GCC targeting the 32-bit ARM family (and MIPS) at -O1/-O2/-O3 instead emits
two loads of the same address:
g:
ldrh r2, [r0] ; load *p, once — zero-extended view
ldrsh r0, [r0] ; load *p, twice — sign-extended view (INVENTED)
subs r0, r2, r0
bx lr
*p is read twice — once by the ldrh, once by the standalone ldrsh — with a
double-read (divergence) window between them. The two views of t are no longer
guaranteed to agree: if the memory at *p changes between the two loads (another
thread, DMA, attacker-controlled MMIO), the zero-extended half holds the old
value and the sign-extended half the low halfword of the new one.
This minimal example only subtracts them, so the inconsistency is harmless — it has no check/use boundary to straddle. But drop the same reload into a validate-then-use shape — snapshot a value, check a field of the snapshot, then act on the snapshot — and the two loads land on opposite sides of the check, introducing a TOCTOU.
Invented-loads are observed in all major C compilers, with some compilers more likely to emit the vulnerable pattern than others:
| # | Mechanism | Compiler | Targets |
|---|---|---|---|
| 1 | Rematerialization | GCC, Clang, ICX, ICC, MSVC | x86-64, i386, m68k, VAX, MSP430 |
| 2 | Width-mismatch reload | GCC | ARM, MIPS, RISC-V (RV64), MIPS64, s390x |
| 3 | Bulk-vs-scalar overlap | GCC (broad); Clang / ICX (SIMD, SLP-coalesce, memcmp-seam, AltiVec-realign & sub-word-atomic forms); MSVC (benign tearing) | x86-64, i386, ARM, AArch64, AVR, Xtensa, SPARC, PPC64, s390x, MIPS64, RV64, RV32, LoongArch64, m68k, MSP430, VAX, HPPA |
| 4 | Cross-class reload | GCC | x86-64, s390x |
| 7 | CISC mem-op fold | GCC; Clang on 6502 | m68k, MSP430, s390x, VAX, 6502 |
| 8 | Byte-order reload | GCC | s390x |
Invented-loads impact most CPU architectures:
| Architecture | Invented-load classes | ISA property that invites them |
|---|---|---|
| x86-64 | 1, 3, 4 | register-rich, cheap RIP-relative global reload, split FP/GPR files |
| i386 | 1, 3 | single-instruction absolute global reload + register-poor 8-GPR file (1); bulk copy overlaps a scalar field load (3) |
| s390x | 2, 3, 4, 7, 8 | signed+unsigned 32→64 widening loads lgf/algf (2); an FP/GPR split, memory-operand ALU, and a byte-reversed load |
| ARM (32-bit family) | 2, 3 | rich narrow-load variants + pipelined loads (2); ldm bulk copy overlaps a scalar ldr (3) |
| AArch64 | 3 | wide ldp/ldr q bulk copy and NEON ld2 de-interleave overlap a scalar field load (3); adrp+ldr addressing and free extend operand-modifiers suppress 1 and 2 |
| MIPS | 2, 3 | rich narrow loads, like ARM (2); word-granular atomics word-load a neighbor on MIPS64 (3) |
| RISC-V (RV64) | 2, 3 | lw+ld bitfield reload (2); wide ld bulk copy vs scalar lw (3) |
| LoongArch64 | 3 | wide ldptr.d bulk copy vs scalar ldptr.w field load |
| PPC64 (pre-VSX AltiVec) | 3 | no unaligned vector load — a realigned wide load reloads each straddling aligned block (clang; GCC rotates) |
| SPARC | 3 | word-granular atomics only — a sub-word _Atomic RMW word-loads a neighbor via an ld+cas loop |
| m68k / MSP430 / VAX | 1, 3, 7 | single-instruction global addressing (1); CISC memory-operand ALU add.l x,%d0 (7) |
| 6502 | 7 | memory-operand ALU; the one non-GCC instance |
The classes group by the ISA property that makes the second load look cheap: