
CVE-2026-31413: BPF verifier soundness bug - container escape
I found a soundness bug in the Linux BPF verifier - a + 1 in a push_stack()
call that causes the verifier to skip an ALU instruction on a forked path. For
BPF_OR, this means the verifier tracks dst = 0 while the CPU computes
0 | K = K. I wrote a full container escape: OOB read/write from a BPF map,
vtable hijack, modprobe_path overwrite, root on the host. Then I authored a
two-patch series - a one-character verifier fix and 90 lines of selftests - and
got it merged into mainline.
| CVE | CVE-2026-31413 |
| Bug class | Verifier soundness - register value divergence |
| Root cause | push_stack(env, env->insn_idx + 1, ...) skips ALU insn on forked path |
| Introduced | bffacdb80b93 - Linux 7.0-rc1 (Jan 14, 2026) |
| Fixed | c845894ebd6f - Linux 7.0-rc5 (Mar 22, 2026) |
| Affected | 6.12.75+ (stable backport dea9989a3f) through 7.0-rc4 |
| Impact | Arbitrary kernel R/W → container escape → host root |
| Required | CAP_BPF + CAP_PERFMON + CAP_NET_ADMIN |
| Fix | One character: insn_idx + 1 → insn_idx |
maybe_fork_scalars() forks verifier state when it sees ARSH + AND/OR with a
constant. The pushed path gets dst = 0 and skips the ALU instruction. For AND
that's fine: 0 & K = 0. For OR it's wrong: 0 | K = K, not 0.
The verifier thinks the register is zero. The CPU has K. I used that to build
arbitrary OOB read/write from a BPF map value, leaked the map's kernel address,
built a fake bpf_map_ops vtable, redirected map_push_elem through
array_map_get_next_key for arbitrary write, and overwrote modprobe_path.
Trigger an unknown binary format, kernel runs my script as root. In a container,
full host escape.
Two-patch series: a one-character verifier fix plus 90 lines of BPF selftests covering OR vs AND forking. Merged by Alexei Starovoitov on March 22. CVE-2026-31413 assigned by Greg Kroah-Hartman on April 12.
eBPF lets you load small programs into the kernel - packet filters, tracing hooks, security policies - without compiling a kernel module. The catch is that you're injecting code into ring 0. If that code has a bug, it's a kernel bug.
So before any BPF program runs, the kernel's verifier simulates every possible execution path. It tracks what each register holds (a pointer? a scalar? what range?), checks every memory access against map bounds, and rejects anything that could read or write out of bounds. If the verifier says a program is safe, the JIT compiles it to native machine code and runs it at full kernel privilege. There are no runtime bounds checks after that point. The verifier is the security boundary.
This is why verifier soundness bugs are different from normal memory corruption. With a heap overflow or UAF, you get one corruption primitive and have to work from there - spray the heap, groom objects, race a window. With a verifier bug, you get the kernel to believe a lie about a register's value. Every bounds check that depends on that register passes. The kernel approved your OOB access. It runs it without question. If you can line the register state up correctly, you get a clean and reliable primitive out of it.
I was auditing maybe_fork_scalars() - new code, added January 2026 in
bffacdb80b93. State forking is always interesting because it's where the
verifier splits into parallel exploration paths, and if any path tracks an
incorrect value, everything downstream of that path is unsound.
The function forks when it sees ARSH + AND/OR with a constant source. Pushed
path gets dst = 0, skips the ALU instruction. I was reading the
push_stack(env, env->insn_idx + 1, ...) line and it clicked immediately - the
+ 1 means the pushed path never executes the ALU op. For AND, 0 & K = 0, so
skipping is fine. For OR, 0 | K = K. The pushed path thinks the result is 0
when it's actually K.
I wrote a BPF program that evening. ARSH 63 to get {0, -1}, OR with a
constant, conditional branch to separate the verifier paths, then add the
"zero" register to a map pointer. The verifier approved map_value + 0. The
CPU accessed map_value + K. KASAN confirmed the out-of-bounds access in
testing.
OOB read/write by the next morning. Container escape by the next night. I used
Claude (Opus 4.5) throughout - for working through the verifier's state
forking logic, brainstorming exploitation primitives, and turning the OOB
into a full escape chain. The vtable hijack approach came out of a back-and-forth
where Claude walked through the bpf_map_ops function pointers looking for
callable gadgets.
Commit bffacdb80b93 ("bpf: Recognize special arithmetic shift in the
verifier") landed January 14, 2026 in 7.0-rc1. Alexei Starovoitov, co-developed
by Puranjay Mohan. It added maybe_fork_scalars() to handle an LLVM
DAGCombiner pattern:
w2 s>>= 31 // arithmetic shift right: w2 becomes 0 or -1
w2 &= -134 // AND with constant K
LLVM lowers select_cc setlt X, 0, A, 0 to sra + and. After the arithmetic
right shift, the register is either 0 (non-negative input) or -1 (all ones).
AND with a constant gives 0 or K.
The verifier can't track {0, K} in a single bpf_reg_state - its signed range
[0, K] over-approximates, and that was causing it to reject valid Cilium
programs. The fix: fork the verifier state. One path explores dst = 0, the
other dst = -1, each tracking the precise value.
The implementation:
static int maybe_fork_scalars(struct bpf_verifier_env *env,
struct bpf_insn *insn,
struct bpf_reg_state *dst_reg)
{
// ... condition check: dst range is [-1, 0], src is constant ...
branch = push_stack(env, env->insn_idx + 1, env->insn_idx, false);
// ^^^^^^^^^^^^
// pushed path resumes AFTER the ALU insn
if (IS_ERR(branch))
return PTR_ERR(branch);
regs = branch->frame[branch->curframe]->regs;
__mark_reg_known(®s[insn->dst_reg], 0); // pushed: dst = 0
__mark_reg_known(dst_reg, -1ull); // current: dst = -1
return 0;
}
Two things happen on the pushed path:
0insn_idx + 1 - the instruction after the ALU opFor BPF_AND: dst = 0, skip the AND. Runtime: 0 & K = 0. Match. Sound.
For BPF_OR: dst = 0, skip the OR. Runtime: 0 | K = K. Mismatch.
The verifier sees 0. The CPU has K. Unsound.
The function doesn't check the opcode. It was written for AND - where skipping
the instruction is the same as executing it with dst = 0 - and got applied to
OR too. For OR, that equivalence doesn't hold.
The trigger pattern is five instructions:
r6 = *(u64*)(map_value + 0) // load a positive value (guaranteed by map init)
r6 s>>= 63 // arithmetic shift: r6 = 0 (positive input)
r6 |= K // BUG: verifier forks, pushed path gets r6=0
if r6 s< 0 goto exit // steers verifier paths
r9 += r6 // verifier: r9 += 0 (in-bounds)
// runtime: r9 += K (OOB)
The verifier explores two paths:
Current path (dst = -1): The OR executes, -1 | K is still -1. The
branch r6 s< 0 is taken. The verifier follows the exit. This path is safe and
the verifier confirms it.
Pushed path (dst = 0, skipped OR): r6 = 0. The branch r6 s< 0 is not
taken. The verifier falls through to r9 += r6, sees r9 += 0, and approves the
subsequent memory access as in-bounds.