
Proof-of-concept local privilege escalation exploit for a Linux qdisc rate-table race condition, using BPF heap grooming and a pipe leak to gain root.

This repository contains a proof-of-concept local privilege-escalation exploit for CVE-2026-68138, a race in the Linux traffic-control rate-table code. In the tested QEMU environment, the PoC escalates from an ordinary process with outer UID 1000 to a shell with UID 0 in the initial user namespace.
Warning
This code intentionally corrupts kernel heap state. Use it only in an isolated, disposable VM that you own. A missed race or premature cleanup can panic the guest. Do not run it on a workstation, server, or third-party system.
qdisc_get_rtab() and qdisc_put_rtab() manage a process-global singly linked
list, qdisc_rtab_list, and a plain non-atomic int refcnt. Historically,
callers held the RTNL mutex, which serialized access to the list and reference
count.
The flower classifier sets TCF_PROTO_OPS_DOIT_UNLOCKED. An
RTM_NEWTFILTER request for a flower rule can consequently reach a police
action and call the qdisc rate-table helpers without RTNL:
tc_new_tfilter()
-> fl_change()
-> tcf_exts_validate_ex()
-> tcf_action_init()
-> tcf_police_init()
-> qdisc_get_rtab()/qdisc_put_rtab()
Concurrent requests using the same rate table can race the global list and its
reference count. The result is a use-after-free or double-free of
struct qdisc_rate_table, a 1056-byte object allocated from kmalloc-2k on
the tested x86-64 kernel.
Because the list is global rather than per-network-namespace, requests from separate network namespaces still race the same object.
The Linux CNA record identifies introduction commit
470502de5bdb,
released in Linux 5.1.
| Line | Status | Commit/version |
|---|---|---|
| Linux before 5.1 | Not affected | The introducing change is absent |
| Linux 5.1 through 7.1.5 | Affected unless a vendor backport is present | 470502de5bdb through the commit before the stable fix |
| Linux 7.1.y | Fixed | 7.1.6, fb29e1b41052 |
| Linux 7.2 development series | Affected before rc5 | rc1 through rc4 |
| Mainline | Fixed | 7.2-rc5, f43ee0c0730d |
Distribution kernels frequently backport fixes without changing to the upstream version shown above. Check whether either fixing commit—or the equivalent qdisc rate-table spinlock change—is present in the exact kernel source used by the system.
The exploit was developed and validated against vulnerable commit
92d3817649df2b0b6a008a686c8275c88d7ef594, the direct parent of the mainline
fix. The fixed-kernel control used
f43ee0c0730d6191629b5ee1ceae27b1ebfdc047.
As of 2026-08-12, the
Ubuntu CVE tracker search
did not return an entry for this CVE. The table below is therefore a direct
source inspection, not a Canonical security-status determination. Each linked
Ubuntu tag still has the unlocked qdisc_rtab_list and lacks the fixing
qdisc_rtab_lock.
| Ubuntu line | Inspected package/tag | Source result |
|---|---|---|
| Ubuntu 22.04 GA | 5.15.0-187.197 | Vulnerable code present; full exploit reproduced in QEMU |
| Ubuntu 22.04 HWE | 6.8.0-136.136~22.04.1 | Vulnerable code present; exploit chain not tested |
| Ubuntu 24.04 HWE | 7.0.0-28.28~24.04.1 | Vulnerable code present; this exploit chain is not compatible with its allocator hardening |
| Ubuntu 26.04 | 7.0.0-28.28 | Vulnerable code present; this exploit chain is not compatible with its allocator hardening |
No corrected Ubuntu package was identified at that inspection date. Future Ubuntu packages should be checked for a backport equivalent to the linked upstream fixes rather than judged only by their version number.
The build-specific Ubuntu 22.04 exploit, QEMU lab, exact image checksum, and
conditions are documented in ubuntu/README.md. It was
validated against the official 5.15.0-187-generic #197-Ubuntu kernel with
memory-cgroup accounting enabled.
The underlying bug and this particular exploit chain have different requirements. The PoC was validated with:
CONFIG_USER_NS=y and CONFIG_NET_NS=y;CONFIG_NET_CLS=y, CONFIG_NET_CLS_FLOWER=y,
CONFIG_NET_CLS_ACT=y, and CONFIG_NET_ACT_POLICE=y;CONFIG_TMPFS_XATTR=y for the simple_xattr heap spray;CONFIG_MODULES=y and a usable /sbin/modprobe for the final root helper;CONFIG_MEMCG=n, so the qdisc/BPF/pipe/xattr allocations used by this chain
share the expected kmalloc-2k cache;CONFIG_SLAB_BUCKETS=y, freelist randomization, and freelist hardening were
enabled in the successful test kernel. KASLR is not inherently bypassed with a
hard-coded address: the PoC obtains the required page and operations pointers
from the pipe leak. The supplied lab used nokaslr to simplify debugging.
Kernels with memory-cgroup accounting enabled or different allocator/cache layouts require a different reclaim strategy. The PoC deliberately refuses to claim portability across arbitrary distribution configurations.
The separate Ubuntu variant implements that different reclaim strategy; its
requirements are intentionally narrower and are listed in
ubuntu/README.md.
Four worker threads enter distinct network namespaces and create flower filters with police actions. Three workers take the successful action path; one supplies a deliberately invalid estimator after acquiring both rate-table references, forcing the cleanup path. This combination makes concurrent reference-count and list manipulation reproducible without sharing flower classifier state between workers.
qdisc_rate_table with classic BPFAfter every netlink request, the same CPU immediately attaches a 133-instruction
classic-BPF filter. Its 1064-byte instruction array is allocated from
kmalloc-2k and is shaped so the bytes overlapping qdisc_rate_table.next and
qdisc_rate_table.refcnt initially remain valid.
The final instruction contains a per-socket marker. SO_GET_FILTER lets the
PoC detect when one socket's orig_prog->filter pointer has been redirected to
another live BPF allocation. This yields two socket objects referring to the
same instruction buffer.
Closing one owner defers the actual free through sk_filter_release_rcu().
After the grace period, the PoC allocates 32-slot pipe rings:
32 * sizeof(struct pipe_buffer) = 32 * 40 = 1280 bytes -> kmalloc-2k