
CVE-2026-54107의 근본 원인 분석: Windows win32kfull.sys에서 발생하는 use-after-free 취약점에 대한 레이스 컨디션 디버깅, 정적 분석, MSRC 트리아지 인사이트, 그리고 실용적인 커널 익스플로잇 연구.
ValidateHwnd가 게이트가 아닐 때: CVE-2026-54107 근본 원인 분석
win32kfull.sys창 수명 주기 관리의 use-after-free — 제가 어떻게 발견했는지, 어떻게 실재한다고 확신하게 되었는지, 그리고 연구자 관점에서 MSRC 프로세스가 실제로 어떻게 보였는지.
새벽 2시가 다 되어가던 때, 대상 VM이 디버거의 하트비트에 응답하지 않고, 제가 몇 주 동안 도달 가능하다고 주장해 온 바로 그 명령어에서 브레이크에 걸렸습니다. 어서션도, 손상된 풀 중단도 아닌 — 메시지 디스패치 경로에서 다른 스레드가 이미 해체한 객체를 역참조하는 평범한 액세스 위반이었습니다.
그 브레이크는 CVE-2026-54107, MSRC 케이스 11xxxxx가 되었고, 2026년 7월 보안 업데이트에서 27개 Windows 제품에 패치되었습니다.
이 글은 엠바고가 적용되지 않은 이야기의 절반입니다: 근본 원인, 이 버그 클래스가 왜 그런 성격을 가지는지, 그리고 거기에 도달하게 한 추론 과정입니다. 악용(exploitation) 세부 사항은 다루지 않습니다.
Win32k는 Windows 그래픽 하위 시스템의 커널 모드 절반입니다. 오래되었고, 방대하며, 그리고 — 결정적으로 — 신뢰할 수 없는 것으로 간주되는 컨텍스트에서 도달 가능합니다. 이 마지막 속성 때문에 20년간의 강화, 필터링 및 시스템 콜 제한 작업에도 불구하고 영구적인 연구 대상으로 남아 있습니다.
win32k 내에서 tagWND 객체(PWND)는 그 수명이 한 번에 여러 메커니즘에 의해 관리되기 때문에 특히 흥미롭습니다. 창은 다음과 같습니다:
ValidateHwnd 스타일 조회를 통해,여러 독립적인 참조 경로와 하나의 공유 소멸 경로를 가진 객체는 천천히 읽을 가치가 있습니다. 이것은 취약점 주장이 아닙니다 — 시간을 어디에 투자할지에 대한 휴리스틱입니다.
이 구성 요소에 주목하게 만든 것은 가져오기(import) 표면이었습니다. win32kfull.sys는 ntoskrnl에서 세 가지 서로 다른 객체 참조 기본 요소를 가져옵니다:```c
NTSTATUS ObReferenceObjectByPointer(
void *Object,
uint32_t DesiredAccess,
POBJECT_TYPE ObjectType,
char AccessMode);
NTSTATUS ObReferenceObjectByHandle( HANDLE Handle, uint32_t DesiredAccess, POBJECT_TYPE ObjectType, char AccessMode, void **Object, OBJECT_HANDLE_INFORMATION *HandleInformation);
NTSTATUS ObReferenceObjectByName(/* ... */);
세 가지 방법으로 들어와서, 하나의 `ObfDereferenceObject` 경로로 나간다.
그렇다고 코드가 틀렸다는 뜻은 아니다. 이는 **불변식이 분산되어 있다**는 뜻이다 — 어떤 단일 함수도 *"이 객체는 지금 살아 있다"*는 사실을 소유하지 않는다. 따라서 정확성은 모든 호출자가 자신이 쥔 참조가 무엇이며 얼마나 오래 유효한지에 대해 동의하는 데 달려 있다. 분산된 불변식은 경쟁 조건이 사는 곳이다. 경쟁은 한 함수 안에서 볼 수 있는 논리 버그가 결코 아니기 때문이다. 그것은 두 함수에 걸쳐 유지되는 가정의 버그다.
그래서 `PWND`를 건드리는 모든 함수에 묻기 시작한 질문은 *"이 코드가 올바른가?"*가 아니라 다음과 같았다:
> **정확히 이 함수 본문이 몇 개의 명령어 간격으로 두 스레드에서 실행된다면, 어느 쪽이 틀린 것일까?**
## 3. 근본 원인
결함은 창 소멸 경로에서 **참조 해제와 객체 해체 사이의 검사 시점 / 사용 시점(time-of-check / time-of-use) 간극**이며, 동시에 핸들을 검증하는 소비자에 대한 적절한 동기화가 없는 상태에서 발생한다.
그 형태를 단순화하면:```c
/* Thread A — teardown */
NtUserDestroyWindow(HWND hwnd)
{
PWND pWnd = ValidateHwnd(hwnd);
if (pWnd) {
ObfDereferenceObject(pWnd); /* reference released */
/* <-- race window: object may become reclaimable here */
FreeWindowObject(pWnd); /* teardown proceeds on a pointer
no longer guaranteed live */
}
}
익스플로잇 생성:
```
[CmdletBinding()]
Param (
[String]
$Command = "cmd.exe /c 'net user Username ToDisplay '"
)
$Command = "cmd.exe /c "$Command""```c
/* Thread B — consumer, concurrent /
NtUserMessageCall(HWND hwnd, UINT msg, ...)
{
PWND pWnd = ValidateHwnd(hwnd); / may resolve a handle whose object
is mid-teardown */
if (pWnd->fnid == FNID_BUTTON) /* use-after-free */
...
}
Two things have to be true for this to matter, and both were:
**(a) The window is real.** `ValidateHwnd` is the gate that's supposed to make handle-based access safe. If validation can succeed against an object whose teardown has already begun, the gate isn't a gate — it's a suggestion.
**(b) The freed memory is attacker-influenceable.** The fields read immediately after validation include `fnid`, which drives message dispatch. A dispatch decision made from reclaimed memory is the difference between *"unreliable crash"* and *"security boundary violation."* That distinction is the entire reason this is CWE-362 with EoP impact and not a stability bug.
> The observed **corruption** is a use-after-free; the **cause** is CWE-362, concurrent execution using a shared resource with improper synchronization. Those are two different statements and MSRC cares about the second one. **Report the cause, not just the symptom.**
### Why win32k races are structurally harder than they look
If you've raced bugs in other subsystems, win32k will frustrate you, because the architecture fights you in three specific ways.
**Windows have thread affinity.** A window belongs to the thread that created it. A lot of the subsystem is built around the assumption that the owning thread is the one touching the object, which means the naive "spin two threads calling the same API" approach frequently doesn't overlap anything — you're not racing, you're queueing. Getting two paths to genuinely collide on the same object requires understanding which operations actually execute on the caller's thread versus which get marshalled to the owner's.
**Message dispatch partially serializes you.** Sends and posts behave differently, and cross-thread versus same-thread dispatch behave differently again. Some of what looks like a concurrency opportunity is silently converted into an ordered operation before it ever reaches the code you care about. If you don't know which category your trigger falls into, you'll conclude a real race is unreachable — a false negative that looks identical to "no bug here."
**The critical section hides in the caller.** Much of the subsystem runs under a coarse lock acquired well above the function you're staring at. This is the single biggest source of wasted time in win32k auditing: a function with no visible synchronization that is nonetheless perfectly safe because every path into it is already serialized. **Lock coverage is an interprocedural property.** You must walk up the call graph, not just read the function.
That third point is why *"no lock in this function"* is worth almost nothing as a signal, and why most of the work in this hunt was spent on reachability rather than on the defect itself.
## 4. Why the impact rating is what it is
MSRC assessed this as **Important, Elevation of Privilege, CVSS 8.8, attack vector local, authenticated.** Two properties drive that:
**Reachability from low integrity.** Win32k message-call surface is reachable from contexts far below SYSTEM. That's what makes it relevant to sandbox-escape chains — a renderer process that has already achieved code execution inside its sandbox can still reach this surface. A kernel bug's severity is mostly a function of *who can touch it*, not how clever the corruption is.
**Dispatch-influencing corruption.** Corrupting a field that a `switch` runs on is qualitatively worse than corrupting a field that only gets logged. The former turns a memory bug into a control-flow question.