Skip to content
KitploitKITPLOIT
ToolsExploitsBlog
Log in
Submit
ToolsExploitsBlog
Submit

Hacking, PenTest, and Cybersecurity Tools for Your Security Arsenal!

Kitploit is a directory of hacking, cybersecurity, and pentesting tools. Discover the latest project updates to find vulnerabilities, analyze systems, automate testing, and strengthen your security.

FeedsContactPrivacy© 2026 Kitploit

Tool Directory

Categories

View all categories
Loading categories
agentic-workflow-injection — Reproducible vulnerable and fixed GitHub Actions fixtures for agentic workflow injection (CVE-2026-44246), with measured detector coverage and mitigation guidance. | Kitploit
Tools/GitHubGitHub/sushant-me/agentic-workflow-injection
Static AnalysisVulnerability ScannersVulnerability AnalysisDevSecOpsPapers & ResearchLearning & EducationAI Security
GitHubsushant-me/agentic-workflow-injection

Most Popular

View all →

Discover the most used tools by our community.

Explore all tools

Browse our collection of tools

View all tools →

agentic-workflow-injection

Reproducible vulnerable and fixed GitHub Actions fixtures for agentic workflow injection (CVE-2026-44246), with measured detector coverage and mitigation guidance.

View Repository
1921 days agoNot yet reviewed
Share

Agentic workflow injection: fixtures and a coverage comparison

Reproducible vulnerable/fixed fixtures for the class of bug published as CVE-2026-44246 (nnU-Net, agentic workflow injection, CVSS 7.2, CWE-1427), plus the measured output of three detectors against them.

It exists because there was nothing to test a detector against. Writing a rule for this class means writing a fixture by hand and hoping it is faithful; the published revisions were only available by knowing which commit to look at, which turned out to be the interesting part.

The class

A GitHub Actions workflow hands untrusted text to an AI agent that has the repository's credentials.

Four conditions, all necessary:

  1. An untrusted trigger. issues, issue_comment, pull_request_review — events whose text anyone with a GitHub account can author.
  2. A write scope on the job. issues: write, contents: write, and so on.
  3. An opt-out of the action's own actor check. claude-code-action and codex-action refuse a run actor without write permission unless an input (allowed_non_write_users, allow-users) explicitly opts them in. Without that input an untrusted author never reaches the agent, and the workflow is the safe configuration rather than this one.
  4. The text reaching the agent. Either interpolated into the workflow (${{ github.event.issue.body }} inside prompt:) or fetched at run time (gh issue view) with the token the job hands over.

It is not template injection. Nothing is evaluated as shell; the payload is prose, and the interpreter is the model. That is why the usual advice — quote your variables, do not eval — does not apply, and why the fix below is about what the agent can reach rather than about escaping input.

The part worth knowing: the fix commit is not the fix

nnU-Net hardened this workflow in two commits, and a detector that treats "before" and "after" as a binary gets the middle one wrong.

revisioncommitdateagent's --allowedToolsverdict
vulnerable94300b49e7162026-04-13gh issue comment, gh issue editreachable write
"the fix"4e4770b0b0e62026-04-24gh issue comment only; labelling moved to a wrapper scriptstill reachable write
later11bd8746fc062026-04-27neither; a later step posts from a file the agent writesnot reachable

The commit that everyone would call the fix — its message is "hardened issue and PR agents" — removed gh issue edit and routed labels through .github/scripts/safe-label.sh, but left Bash(gh issue comment:*) with the agent. The agent could still be steered into commenting, and the allowlist pattern gh issue comment:* is not scoped to the triggering issue, so the target was the model's to choose. Only the third commit closed it: the agent now writes /tmp/issue-comment.md and a later, non-agent step posts it with ISSUE: ${{ github.event.issue.number }} taken from the event.

Two things follow, and they are why this repository is not just two files:

"Fixed" is a claim about a revision, not a version. v2.4.1 is cited as the fixed release, but its .github/workflows/ contains only codespell.yml — the agent workflows are not in that tag at all. You cannot verify the fix from the tag; you have to pin the commit.

The permissions block never changes. issues: write is present and correct in all three revisions, including the last one, because a later step needs it. A detector that keys on permissions alone cannot separate revision 1 from revision 3. What separates them is what the agent may call.

Measured coverage

Three detectors, run over both fixtures and all three real revisions. Full raw output and tool versions in results.md.

revisionagentbound 0.1.3sisakulint v0.3.7zizmor 1.30.1
1-vulnerableHIGH write-scope, HIGH untrusted-contentai-action-prompt-injection—
2-fix-commitHIGH write-scope, CRITICAL author-association— (generic only)—
3-laterLOW write-scope, CRITICAL author-associationai-action-excessive-tools, ai-action-execution-order—

No detector is simply "wrong" here — they answer different questions:

  • zizmor is a general-purpose Actions auditor. It reports pinned-action and credential-persistence hygiene identically on all three revisions. Out of scope for this class by design, and included because "the well-known linter was clean on the vulnerable file" is easily mistaken for a clean bill of health.
  • sisakulint has purpose-built AI rules and is the only tool here that names the CVE's own mechanism: ai-action-prompt-injection fires on the vulnerable revision, where the issue body is interpolated into prompt:, and stops once the interpolation is gone. Correct. Its findings on the later revision are worth reading before acting on them — ai-action-excessive-tools flags Write, which this workflow uses on purpose to write the file that a later step posts; and ai-action-execution-order wants the agent last, which is exactly what this design avoids, since the privileged steps come after the model's turn ends.
  • agentbound is the only tool that separates revision 2 from revision 3, by reading the agent's tool allowlist rather than the job's permissions. Its ci-agent-missing-author-association finding is reported at critical on the later revision and its own README describes that as over-severe rather than false: auto-triage is meant to be reachable by anyone.

Every tool here has a finding it reports on a revision whose author had already reasoned about that exact condition. That is the normal state of a heuristic, and the reason a coverage table is more useful than a pass/fail column.

Using it

git clone https://github.com/sushant-me/agentic-workflow-injection
cd agentic-workflow-injection

bash scripts/fetch-revisions.sh   # the three real revisions, pinned by SHA
bash scripts/benchmark.sh         # runs every detector that is installed

benchmark.sh reports a detector that is missing as missing rather than skipping it silently, because a coverage table that quietly omits a tool proves nothing about it.

The fixtures are the reduced shape, written for this repository and MIT-licensed like the rest of it. They are mechanism-faithful rather than copies: fixtures/vulnerable.yml carries the four conditions and nothing else, and fixtures/fixed.yml is the same file with the six changes that matter, each annotated. If you are adding a rule, those two files are the smallest inputs it has to get right.

Two interface quirks, handled in the script but worth knowing:

  • sisakulint does not analyse a file outside a git repository. Handed one, it prints not found and exits 3. Unwatched, that looks identical to a clean scan.
  • agentbound's JSON path is a basename, so a directory scan over files with the same name is ambiguous.

Fixing it

Detection is the easy half. docs/mitigations.md covers the other half: how to bound an agent's authority, with the layers ranked by one question — does this control depend on the model choosing to comply?

That ranking is the whole point. An instruction in a prompt is input, not a guard clause; a tool pattern that names its target, or a token the agent never holds, is a bound. The guide works down from "the model cannot do the wrong thing" to "the model was asked not to", with a checklist and the nnU-Net evolution as the worked example.

Download Tool