
Proof-of-concept exploit that bypasses ASLR on Intel CPUs by abusing branch target buffer and speculative execution to leak randomized addresses via side channel.
Address Space Layout Randomization is a mitigation used to make harder to exploit memory corruption attacks. In a scenario of a buffer overflow vulnerability, for example, an attacker that tries to make a Return Oriented Programing exploit needs to know the addresses of the gadgets in the chain. If the code segment of the exploited binary is randomized, then it's much harder for an attacker to pick the correct address for the exploit, making the exploitation unfeasible.
The following example shows how an address is randomized:
#include <stdio.h>
void DoNothing();
void (*codePtr)() = DoNothing;
void DoNothing(){}
int main(int argc,char **argv){
printf("Destination %p\n",codePtr);
DoNothing();
}
Each execution the value is randomized:
Destination 0x563714256149
Destination 0x556d8e2f1149
Destination 0x5618c8bdd149
Destination 0x55ee623b0149
The last 12 bits 149 are always the same, but the location of the function can be roughly anywhere between 0x550000000000 and 0x570000000000, meaning that 29 bits are randomized, occupying a possible address space of 0x200 0000 0000 or 2.2 TB of size.
The processing of each instruction is a hard task. Some stages of the processing of a single instruction are:
In order to increase the throughput of instructions in the CPU, each task of the instruction is performed by a specific unity of the processor. With all the unities working in parallel, it allows the CPU to execute at much higher clock speeds, that's the idea of a pipeline.
| Operation \ clock cycle | 1 | 2 | 3 | 4 | 5 |
|---|---|---|---|---|---|
| Fetch | A | B | C | ||
| Decoding | A | B | C | ||
| Execution | A | B | C |
Execution of instructions A B and C over the cycles 1-5. In cycle 3, for example, the read, decoding and execution unities are simultaneously active
However, instructions are not completely independent from each other. For example, the following sequence:
A. add ax,[bx]
B. jz $+1
C. mov dl,[rsi]
D. nop
In this case, the A instruction in the best case will only finish on the cycle 3 in the execution. However, the fetch unity needs to decide what's the next instruction to be fetched from memory, whether instruction C (mov dl,[rsi]) should be skipped.
In this scenario, the CPU has the option to wait for the instruction A to finish, that will only happen in the third clock cycle to then fetch the correct instruction on the memory, if the add operation returns 0, for example:
| Operation \ clock cycle | 1 | 2 | 3 | 4 | 5 | 6 |
|---|---|---|---|---|---|---|
| Fetch | A | B | D | |||
| Decoding | A | B | D | |||
| Execution | A | B | D |
That implies in a delay in the pipeline because the CPU must wait for the instruction to get executed. In this example the delay is a single clock cycle, but the instruction add ax,[bx] requires a memory operation, which, as seen before, can take up to hundreds of cycles to complete, thus leveraging a significant performance cost on the processor.
A faster option would be trying to "guess" the correct execution path. The CPU can speculate if the branch is taken or not. After that point, the execution continues from the speculated path and the values are only committed if the path is proven correct after the finish of the A instruction. If the path is proven wrong, the results are discarded and the state is reverted to before the speculated point.
| Operation \ clock cycle | 1 | 2 | 3 | 4 | 5 | 6 |
|---|---|---|---|---|---|---|
| Fetch | A | B | (S) C | |||
| Decoding | A | B | (S) C | |||
| Execution | A | B | (S) C |
The only problem in reverting the path taken is that the microarchitectural state of the CPU cannot be reverted. So if the CPU speculates to execute the instruction C (mov dl,[rsi]) the data pointed by rsi will be moved to the cache. This effect can be measured later using a side channel attack.

2 Bit conditional predictor. https://en.wikipedia.org/wiki/Branch_predictor
Not only conditional instructions must be predicted, but also indirect branches. The CPU must have a mechanism for guessing the destinations of an instruction as call [rdi].
The spectre v2 vulnerability shows that it's possible to exploit the indirect predictor to achieve transient execution in other processes:
Extracted from https://spectreattack.com/spectre.pdf
When placing a call instruction on context A in the same virtual address of another call in context B, the attacker can train the CPU to execute code on a position chose by the attacker on context B, in a code reuse attack, similar as Return Oriented Programing (ROP).
The targeted victim must have a piece of code known as "spectre gadget" that is able to leak a secret using a side channel attack. For a successful spectre attack, the attacker must also know the location of the spectre gadget. Therefore, in user-user attacks, protecting the victim with ASLR used to be a mitigation for this kind of attack. However, there are also techniques for extracting ASLR using microarchitectural attacks such as Jump Over ASLR. However, this technique has some limitations about the amount of bits leaked, since it relies collision on the direct predictor to bypass ASLR.
The internals mechanisms of this predictor are shown below:

Extracted from https://spectreattack.com/spectre.pdf
Some of these components are:
The classical spectre v2 attack layout looks like this:
