Skip to content
KitploitKITPLOIT
ToolsExploitsBlog
Log in
Submit
ToolsExploitsBlog
Submit

Hacking, PenTest, and Cybersecurity Tools for Your Security Arsenal!

Kitploit is a directory of hacking, cybersecurity, and pentesting tools. Discover the latest project updates to find vulnerabilities, analyze systems, automate testing, and strengthen your security.

FeedsContactPrivacy© 2026 Kitploit

Tool Directory

Categories

View all categories
Loading categories
linux-kernel-codex-harness-v2 — Provenance-aware Linux kernel vulnerability research harness used in the investigation of CVE-2026-53075 | Kitploit
Tools/GitHubGitHub/foxirain/linux-kernel-codex-harness-v2
Static AnalysisVulnerability AnalysisThreat IntelligenceAI Security
GitHubfoxirain/linux-kernel-codex-harness-v2

linux-kernel-codex-harness-v2

Provenance-aware Linux kernel vulnerability research harness used in the investigation of CVE-2026-53075

View Repository

Most Popular

View all →

Discover the most used tools by our community.

Explore all tools

Browse our collection of tools

View all tools →
432 months agoNot yet reviewed
Share

Kernel Codex Harness v2

한국어 | English

CI

Research Tool · Original Import: 3 April 2026 · v2 Documentation Revision: 11 July 2026

External Signal: From Attention Allocation to Provenance-Aware Triage
Use reproducible observations to guide model attention, then use repository provenance to organize review queues—never to claim proof.

Project Lineage— Kernel Codex Harness v1 · Attention Allocation → Kernel Codex Harness v2 · Provenance-Aware Triage

Project status. This repository preserves an LLM-assisted research harness that evolved v1's attention-allocation workflow into provenance-aware triage for real Linux kernel vulnerability research. This version was used in the investigation that discovered the vulnerability published as CVE-2026-53075. It is not an automatic vulnerability detector, novelty classifier, exploit verifier, or kernel security assurance tool; final validation and reporting remain human responsibilities.

Abstract

Abstract— When an LLM is asked to explore a codebase as large as the Linux kernel without structure, context disperses and the presence of dangerous APIs is easily confused with actual exploitability. Kernel Codex Harness v2 defines this problem as two stages of External Signal processing. Before model invocation, path weights, lexical hits, and cached syzbot overlap rank candidate files and allocate attention. After the model responds, Git branch, HEAD, and dirty-state facts are combined with CVE, commit, and known markers extracted from the response to classify strong findings into provenance-aware review buckets. The harness was used in a real Linux kernel investigation that discovered a missing authorization check against the target network namespace in PPP, later published as CVE-2026-53075. Triage is a heuristic for organizing investigation queues; in particular, new_candidate means only that no known clue or provenance problem was detected, not that novelty has been proved. Every finding requires human revalidation of userspace reachability, the invariant break, and concrete impact.

Index Terms— Linux kernel, vulnerability research, external signal, provenance, heuristic triage, LLM orchestration, syzbot, Codex.

I. Introduction

Kernel security review contains two distinct kinds of uncertainty.

  1. Where should the review look first? The full source tree is too large for a single model context.
  2. How should a strong model finding be handled? Local modifications, existing fixes, known CVEs, or incomplete repository state can contaminate the conclusion.

The central problem in v1 was the first one: attention allocation. v2 preserves that principle and extends it to the second problem through provenance-aware triage. Both versions were used in real investigations: a v1-assisted investigation led to CVE-2026-31720, and a v2-assisted investigation led to CVE-2026-53075.

Narrow the investigation with observations outside the model, then attach verifiable repository provenance after the model responds. Signals at neither stage prove vulnerability or novelty.

II. External Signal and Design Principles

A. Stage 1 — Attention Allocation Before Inference

Pre-inference External Signal is not an LLM judgment, but an observation computed before model execution.

  • kernel paths and subsystem weights
  • lexical hits such as usercopy, allocator, refcount, size, and lock operations
  • file and subsystem overlap from stored syzbot JSON

Candidate ranks can be recomputed from the same source tree, profile, and cached syzbot JSON. The score is not a probability or exploitability measure, but a relative order for deciding where to look first.

B. Stage 2 — Provenance-Aware Triage After Inference

The post-inference stage combines a strong model verdict with the following information:

  • whether the target is a Git repository and whether status collection succeeded
  • branch and HEAD
  • dirty state of the repository and target file
  • CVEs, commit hashes, and known-issue markers extracted from the response
  • negation or unrelated-reference language in the response

In this document, post-inference External Signal refers only to provenance collected independently of the model, such as Git repository/status facts, branch, HEAD, dirty state, and local commit ancestry. CVE, commit, and known markers extracted from the model response are model-derived references, not External Signal or authoritative facts. Triage combines both kinds of input while recording their provenance separately.

C. Heuristic Buckets, Not Novelty Proof

A strong finding is operationally assigned to one of the following review buckets.

BucketMeaning
new_candidateProvenance was verified and no dirty or known blocking signal was found
known_issueA non-negated known reference was found, or a commit identified by the response as a fix/upstream relation is confirmed in the current HEAD
dirty_tree_suspectThe influence of a dirty repository or dirty target cannot be excluded
provenance_unknownThe Git repository, status, or HEAD could not be verified reliably

Every classification sets novelty_proven to false. new_candidate does not mean a new vulnerability; it means a queue in which a human should continue the novelty investigation first.

D. Reachability Before Bug Class

The audit first identifies boundaries that originate in userspace, such as a syscall, ioctl, netlink, procfs, filesystem, BPF, or driver hook. Only then does it evaluate bug classes such as UAF, OOB access, refcount errors, races, information leaks, or capability-check failures.

E. One Investigation Branch at a Time

An investigation unit is limited to one file and its nearby caller, teardown, and free paths. Model-proposed manual follow-ups are capped at two, preserving a short, verifiable path instead of broad exploration.

F. Evidence Over Confidence

A strong finding must explain at least the following:

  1. attacker-reachable entrypoint,
  2. an attacker-controlled field or lifetime transition,
  3. the object, length, or state invariant that breaks,
  4. a concrete impact such as corruption, leakage, or privilege escalation,
  5. why existing checks do not block the attack.

The parser normalizes the verdict and next target; it does not automatically prove the completeness of this evidence.

G. Design Lineage

The initial flow drew inspiration from the file-level analysis, bounded context expansion, and structured outputs used by Protect AI's vulnhuntr [1]. This project redesigned those ideas around userspace-reachable kernel surfaces, kernel object lifetimes, teardown paths, and syzbot overlap. v2's additional contribution is a finding-triage stage that uses repository provenance after attention allocation.

III. System Architecture

Two-stage External Signal architecture for Kernel Codex Harness v2

Download Tool