
Fix-Like Artifacts With Embedded Defects
Fix Like Artifacts with Embedded Defects
A research harness for measuring how well AI agents patch vulnerabilities.
Quickstart · Datasets · Documentation · Security model · Contributing
FLAWED checks out an open-source project at a known-vulnerable commit, hands an AI agent a description of the bug, and asks it to write a patch. The agent never sees the real upstream fix.
Every patch is then validated, audited, and graded in isolated containers, so you can see not just whether the model fixed the bug, but whether it introduced new ones along the way.
flowchart LR
spec[bug spec] --> clone
clone["clone<br/>(open net)"] --> generate
generate["generate<br/>(offline)"] --> validate
validate["validate<br/>(offline)"] --> ast["ast<br/>(offline)"]
generate -.->|"patch.diff"| store[(Postgres)]
validate -.->|"verdict"| store
ast -.->|"summary"| store
store --> ui[web UI + notebook]
Each stage runs in its own container. Only clone has network access. Every
later stage is locked down to the LLM provider's API, so agents can't fetch
hints or the upstream fix mid-run.
| Feature | What it gives you |
|---|---|
| Repeated sampling | A run executes N iterations of the same input, so results are distributions, not anecdotes. |
| Campaigns | Sweep one bug across patcher variants and prompt styles, from a vague "fix this plz" to a full advisory, and compare outcomes on a live dashboard. |
| Outcome grading | Every patch lands in one of five scenarios, from S1 (clean fix) to S5 (didn't fix the bug and introduced a new vulnerability). |
| Cross-validation | Patches are re-judged by other models, and the headline numbers average the self- and cross-validation lenses so no single judge's bias dominates. |
| Cheat detection | An auditor flags iterations where the agent found the upstream fix instead of solving the bug itself. |
FLAWED drives the Claude, Codex, and Gemini CLIs head-to-head with the same inputs.
[!WARNING] FLAWED mounts the Docker socket (root-equivalent on the host) and executes untrusted third-party code inside its stage containers. Run it on a machine you trust to carry that workload. See
docs/security-model.md.
You need Docker, with the daemon socket accessible.
# 1. Configure. Writes .env for you (data dir + provider API keys)
./setup.sh
# 2. Bring up the stack (Postgres, API + worker, web UI, notebook)
docker compose up --build
# 3. Open http://127.0.0.1:8080
Then take your first run.
bugs/. Drag one
into the web UI's Bug Specs page, or use the CLI.
./scripts/import-all-bugs.sh
[!NOTE] The first run against a large upstream (e.g. Chromium) is slow. The validate stage clones the whole repo once to diff against the real upstream patch. The clone is cached and reused afterwards.
You need Docker, Node 20+, pnpm, uv, and
@devcontainers/cli (npm i -g @devcontainers/cli). Nix users can
nix-shell for everything except Docker.
make dev # uv sync + web deps
docker compose up -d postgres # FLAWED needs a Postgres to talk to
cp .env.example .env # points FLAWED_DB_URL at it
uv run flawed init # builds base images, creates the schema
uv run flawed serve # API + worker + webapp on port 8080
For the web dev loop, run make web-dev in a second terminal. It serves the
UI on port 5173 and proxies /api to flawed serve.
An isolated devcontainer for running AI coding agents against this repo safely
is documented in .devcontainer/README.md.
A bug spec is the unit of input. It carries a repo, a vulnerable commit, a bug description, an optional reproducer, and a set of prompt variants that model how the bug might realistically be reported (SAST finding, bug-bounty report, embargoed advisory, raw PoC, …). Specs are versioned and immutable, so a historical run's prompt and verdict never silently change.
The JSON contract lives in bugs.schema.json and is
documented in docs/bug-specs.md. All bundled specs
describe publicly disclosed, upstream-fixed vulnerabilities.
Postgres is the single source of truth. Run metadata and artifact bytes
(patches, verdicts, transcripts, logs) live in the database, so a deployment is
fully captured by its DB. The data/ directory is transient scratch that
stages mount at runtime.
scripts/export_dataset.py and
scripts/import_dataset.py (also available from the web UI).scripts/export_artifacts.py.uv run alembic upgrade head
runs automatically at startup).We publish pre-built datasets so you can load completed campaigns instead of
running everything yourself. Each dataset is a .tar.gz of one full snapshot,
hosted at https://flawed.s3.us-east-1.amazonaws.com.
All campaigns in a single bundle.
| Bundle | Archive |
|---|---|
| All campaigns | full.tar.gz |