
A very very very very very very very long interrupt
Exploiting System Management Mode with a very very very very very very very long interrupt.
It turns out that you can break SMM — the secure, ultra privileged execution environment running invisibly in the background of every x86 CPU — with nothing more than an obscenely long-running machine instruction.
SMM requires that all cores are either in SMM or out of SMM at the same time. Its security model doesn't work without this - when one thread enters SMM, it makes all the others enter too.
To break this, all we need is someone too busy to notice they're supposed to join SMM.
It works something like this:
core 0 - start a long instruction
|
|
|
core 1 - invite core 0 to smm
|
|
|
core 1 - enter smm
|
|
|
core 1 - wait for core 0
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
core 1 - wait for core 0
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
core 1 - wait for core 0
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
core 1 - wait for core 0
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
core 1 - wait for core 0
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
core 1 - wait for core 0
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
core 1 - wait for core 0
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
core 1 - wait for core 0
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
core 1 - give up
core 1 - do secret smm stuff
core 1 - finish smm
|
|
|
core 0 - join smm
At this point, core 1 is out of SMM while core 0 is in, letting core 1 attack
core 0. Here's the catch: for this to work we need a very, very, verrrrry long
instruction — longer than any instruction was ever supposed to take. Most
machine instructions on a modern CPU are fast: add takes 1 cycle. To get core
1 to give up waiting on core 0, we need an instruction on core 0 that takes
around 4,000,000,000 cycles — over 1 second of wall-clock time.
x86 firmware runs the following code when a CPU core enters SMM:
for (Timer = StartSyncTimer ();
!IsSyncTimerTimeout (Timer, mTimeoutTicker) && SyncNeeded;
)
{
mSmmMpSyncData->AllApArrivedWithException = AllCpusInSmmExceptBlockedDisabled ();
if (mSmmMpSyncData->AllApArrivedWithException) {
break;
}
CpuPause ();
}
The code waits for all cores to enter SMM, or for up to 1 second, whichever occurs first. To get a core to execute SMM code while another core stays executing outside SMM, we need that outside core to stay uninterruptible for the entire second — an SMI is taken at an instruction boundary, so any gap between two instructions lets the pending SMI pull the core into SMM. The delay therefore has to be a single instruction: one uninterruptible op that outlasts the one-second rendezvous.
There are many ways to reach the forbidden 1-second instruction, and the exact approach will vary platform-to-platform. But, roughly: find a high-latency MMIO address, and then convince the CPU to read from it as slowly as possible — abuse an undocumented region that answers reads at a crawl, use the widest load the ISA will give you to move as many bytes as possible across it in a single instruction, and let the other cores contend for the same bus to slow it further. One read, one instruction, and the CPU is stuck holding it for the better part of a second.
The provided proof-of-concept
is tuned for a Zen 3 Ryzen 7
5800H, where a wide xmm load from slow MMIO at
0xfcc68860 stalls long enough to break the all-cores rendezvous:
mov $0xfcc68860, %rsi ; the target MMIO address
vmovdqu (%rsi), %xmm0 ; the very, very long load
The PoC exploits this by pitting two cores against each other. One core is held outside SMM by the long instruction — a tight loop on the very slow load:
/* the victim core: spin on the ~1-second load, too busy to answer the SMI */
for (;;)
asm volatile ("vmovdqu (%0), %%xmm0" :: "r"(mmio) : "xmm0");
Meanwhile another core arms the per-core SMI counters:
#define MSR_PERF_CTL0 0xc0010200 /* AMD core perf event-select MSR */
#define MSR_PERF_CTR0 0xc0010201 /* the paired 48-bit counter */
for (int cpu = 0; cpu < ACTIVE_CPUS; cpu++) {
msr_write(cpu, MSR_PERF_CTL0, 0x43002b); /* EN | OS | USR | event 0x2b */
msr_write(cpu, MSR_PERF_CTR0, 0); /* zero the count */
}
Then fires a storm of SMIs:
asm volatile ("outb %%al, $0xb2" :: "a"(0)); /* kick port 0xb2 -> #SMI */
And reads every core's tally back:
/* ...fire the storm, then read every core's tally back... */
uint64_t delta = smi_max - smi_min;
if (delta)
puts("!!! a core ran outside SMM");
If the counts diverge, it means a core kept running outside SMM while the others were pulled in — it missed the SMIs the rest of them serviced.
To better illustrate this, we can run the proof-of-concept behind a needlessly flashy and entirely pointless GUI, tracking SMI counters on each core, to watch them execute in perfect lockstep until a wildly delayed core breaks their required synchronization:
