
PoC CVE-2017-5123 - LPE - Bypassing SMEP/SMAP. No KASLR
PoC CVE-2017-5123 - LPE - Bypassing SMEP/SMAP. No KASLR
In this little writeup, I will analyze a kernel vulnerability that allow us to obtain root privilege.
This file is divided into four parts:
I want to point out that there are a lot of better ways to exploit this CVE (indeed, this is just a PoC for learning the kernel, it can't be used in the wild) but I think that this methodology can be useful as an introduction to kernel exploitation.
This vulnerability was introduced in 4c48abe91be0 so we need to build that version of the kernel.
This can be a little tricky because this is an old version and the code should be patched.
I made a repository with an already patched kernel code and a .config file so you can clone and build.
git clone https://github.com/c3r34lk1ll3r/kernel_mirror.git
cd kernel_mirror
git checkout origin/modified_v4.14
wget https://gist.githubusercontent.com/c3r34lk1ll3r/c9c34ae86140cc7a24d0d90141686ee8/raw/52431b577a71e3fe8f89d6ce355ce9c1c54c53b6/.config
make -j 8 --output-sync=recurse
Note that this kernel will be built with virtio drivers so you can use virtio disk for sharing file from/to VM.
Now, we will create the initial rootfs:
qemu-img create -f raw hda.raw 10G
# Format the disk to ext4
mkfs.ext4 ./hda.raw
# Make a mountpoint for the image
mkdir /tmp/mount1
# Mount the disk
sudo mount -o loop ./hda.raw /tmp/mount1
Then, we should install a basic Linux distribution, for example using pacstrap or debootstrap.
sudo pacstrap /tmp/mount1 base base-devel vim
Finally, we can modify the system:
# Add a 'test' user
echo 'test:x:1000:1000::/home/test:/bin/bash' | sudo tee -a /tmp/mount1/etc/passwd
# without password
echo 'test::14871::::::' | sudo tee -a /tmp/mount1/etc/shadow
# we can mount a virtio disk in order to share files between host and guest
echo '/transient /home/test/shared 9p trans=virtio,version=9p2000.L,rw,user,exec 0 0' | sudo tee -a /tmp/mount1/etc/fstab
sudo mkdir -p /tmp/mount1/home/test/shared
# It is usefull to have sudo permission
echo '%wheel ALL=(ALL) NOPASSWD: ALL' | sudo tee -a /tmp/mount1/etc/sudoers
echo 'wheel:x:998:test' | sudo tee -a /tmp/mount1/etc/group
sudo chown -R 1000:1000 /tmp/mount1/home/test
sudo umount /tmp/mount1
If everything is in order, we can now try our testing system with qemu:
qemu-system-x86_64 \
-kernel ./kernel_mirror/arch/x86_64/boot/bzImage \
-hda ./hda.raw \
-m 4G \
-cpu "Skylake-Client-IBRS,ss=on,vmx=on,hypervisor=on,tsc-adjust=on,clflushopt=on,umip=on,md-clear=on,stibp=on,arch-capabilities=on,ssbd=on,xsaves=on,pdpe1gb=on,ibpb=on,amd-ssbd=on,skip-l1dfl-vmentry=on,hle=off,rtm=off" \
-smp 4 \
-vga virtio \
-enable-kvm \
-nographic \
-machine type=q35,accel=kvm \
-virtfs "fsdriver=local,id=fs.1,path=./trans_fs,security_model=mapped,writeout=immediate,mount_tag=/transient" \
-append "root=/dev/sda rw noquiet nokaslr console=ttyS0 loglevel=5" \
-chardev "vc,id=vc.0,cols=1920,rows=1080" \
-net "user,hostfwd=tcp::10022-:22" \
-net "nic" \
-s
The description of the CVE says that there is an unrestricted write operation during the waitid system call.
Let's open kernel/exit.c and look the code:
SYSCALL_DEFINE5(waitid, int, which, pid_t, upid, struct siginfo __user *,
infop, int, options, struct rusage __user *, ru)
{
struct rusage r;
struct waitid_info info = {.status = 0};
long err = kernel_waitid(which, upid, &info, options, ru ? &r : NULL);
int signo = 0;
if (err > 0) {
signo = SIGCHLD;
err = 0;
if (ru && copy_to_user(ru, &r, sizeof(struct rusage)))
return -EFAULT;
}
if (!infop)
return err;
user_access_begin();
unsafe_put_user(signo, &infop->si_signo, Efault);
unsafe_put_user(0, &infop->si_errno, Efault);
unsafe_put_user(info.cause, &infop->si_code, Efault);
unsafe_put_user(info.pid, &infop->si_pid, Efault);
unsafe_put_user(info.uid, &infop->si_uid, Efault);
unsafe_put_user(info.status, &infop->si_status, Efault);
user_access_end();
return err;
Efault:
user_access_end();
return -EFAULT;
}
This function is pretty straightforward: after few checks, there are various call to unsafe_put_user(...) and the function returns.
The main part of this function is composed by unsafe_put_user(...) function so let's move there (arch/x86/include/asm/uaccess.h):
/*
* The "unsafe" user accesses aren't really "unsafe", but the naming
* is a big fat warning: you have to not only do the access_ok()
* checking before using them, but you have to surround them with the
* user_access_begin/end() pair.
*/
#define user_access_begin() __uaccess_begin()
#define user_access_end() __uaccess_end()
#define unsafe_put_user(x, ptr, err_label) \
do { \
int __pu_err; \
__typeof__(*(ptr)) __pu_val = (x); \
__put_user_size(__pu_val, (ptr), sizeof(*(ptr)), __pu_err, -EFAULT); \
if (unlikely(__pu_err)) goto err_label; \
} while (0)
#define unsafe_get_user(x, ptr, err_label) \
do { \
int __gu_err; \
__inttype(*(ptr)) __gu_val; \
__get_user_size(__gu_val, (ptr), sizeof(*(ptr)), __gu_err, -EFAULT); \
(x) = (__force __typeof__(*(ptr)))__gu_val; \
if (unlikely(__gu_err)) goto err_label; \
} while (0)
There is a big fat warning in the comment: if you want to use unsafe_put/get_user you should first call access_ok() and surround them with user_access_begin/end().
If we take a look at the previous code (waitid) we can see that access_ok() is never called so the system call violates this warning.
But what are those macros?
SMAP and SMEP are two security features introduced in the kernel in order to makes harder to write exploits. To be noted that those features are enforced by the CPU.
SMEP prevents to execute userspace code while the CPU is in supervisor mode; SMAP, instead, blocks read/write access to user memory.
The kernel needs to write/read data to/from user memory and this can be accomplished in two ways:
copy_from_user) that allows to copy the memory in kernel space;As we can see in the definition of unsafe_put_user, this function will only copy the value of x in memory pointed by ptr (and jump to err_label if there was an error). We have just said that the kernel can't access to userspace because SMAP and this is why those functions should be wrapped between user_access_begin/end().
#define __uaccess_begin() stac()
#define __uaccess_end() clac()
As we can see, user_access_begin/end simply are the ASM instruction stac and clac.