AI-native code security auditor on AgentField that proves exploitability with verdicts, traces, and actionable evidence.
Output • Benchmark • How It Works • Comparison • Quick Start • API
Other tools flag patterns. SEC-AF proves exploitability: every finding ships with a verdict, a data flow trace, and evidence you can act on. Free, open source, one API call. A full audit with 30 verified findings costs about $1.40 in LLM calls.
Trigger it with the af CLI (requires af ≥ 0.1.87) — it streams live progress and prints the result:
af call sec-af.audit --in '{"repo_url": "https://github.com/dolevf/Damn-Vulnerable-GraphQL-Application"}'
Prefer raw HTTP? Hit the API directly with curl:
curl -X POST http://localhost:8080/api/v1/execute/async/sec-af.audit \
-H "Content-Type: application/json" \
-d '{"input": {"repo_url": "https://github.com/dolevf/Damn-Vulnerable-GraphQL-Application"}}'
This is a real finding from SEC-AF auditing DVGA (a deliberately vulnerable GraphQL app):
{
"title": "OS Command Injection in run_cmd Helper Function",
"severity": "critical",
"verdict": "confirmed", // not "maybe" — confirmed exploitable
"evidence_level": 5,
"cwe_id": "CWE-78",
"rationale": "Tracer confirms complete data flow from GraphQL parameters
(host, port, path, scheme, cmd, arg) to os.popen(cmd).read() sink.
Sanitization functions are bypassable in Easy mode...",
"proof": {
"verification_method": "composite_subagent_chain:sast",
"data_flow_trace": [
{ "description": "core/views.py:203 — GraphQL args defined (host, port, path, scheme)", "tainted": true },
{ "description": "core/views.py:210 — URL constructed from user input", "tainted": true },
{ "description": "core/views.py:211 — helpers.run_cmd(f'curl {url}') called", "tainted": true },
{ "description": "core/helpers.py:9 — os.popen(cmd).read() executes input", "tainted": true }
]
},
"location": {
"file_path": "core/helpers.py",
"start_line": 9,
"code_snippet": "def run_cmd(cmd):\n return os.popen(cmd).read()"
}
}
Every finding includes a verdict (confirmed / likely / inconclusive / not_exploitable), a proof object with the full taint trace, and the exact code location. Not "this might be a problem." SEC-AF traces data from source to sink and proves whether it's actually exploitable.
Full benchmark output (30 findings):
exampl/dvga-benchmark-result.json| Performance analysis:exampl/benchmark-analysis.json
We run SEC-AF against Damn Vulnerable GraphQL Application, a deliberately vulnerable app with 21 documented security scenarios.
| Metric | Value |
|---|---|
| Raw findings discovered | 106 |
| After AI deduplication | 61 |
| After adversarial verification | 28 confirmed |
| Inconclusive (needs manual review) | 1 |
| Not exploitable (correctly rejected) | 1 |
| Noise reduction | 94% |
| DAG edges (reasoner calls) | 82 |
| Agent calls | ~166–255 |
| Strategies run | 11 |
| Wall-clock time | ~78 min |
| Estimated cost (Kimi K2.5) | ~$0.18–$0.90 |
| Category | Count | Examples |
|---|---|---|
| Missing Authentication | 8 | ImportPaste, delete_all_pastes, system_debug, CreateUser, file upload |
| Command Injection | 4 | os.popen(cmd) via ImportPaste, system_debug, system_diagnostics |
| SQL Injection | 3 | Unsanitized filter in resolve_pastes, LIKE pattern injection, login |
| Authentication Bypass | 3 | JWT signature disabled, JWT authorization bypass, broken password auth |
| Plaintext Credentials | 3 | Cleartext password storage, plaintext comparison, password in diagnostics |
| SSRF | 2 | ImportPaste mutation follows user-supplied URLs server-side |
| Business Logic / URL Sanitization | 2 | Inadequate URL sanitization, unauthenticated mass deletion |
| DoS / Resource Exhaustion | 3 | Missing pagination on users/audits queries, uncontrolled simulate_load |
| Config / Secrets | 2 | Hardcoded JWT/Flask secrets, debug mode enabled in production |
SEC-AF applies several architectural patterns that are uniquely enabled by composing many focused AI agents instead of running one monolithic scan. These patterns address fundamental challenges in AI-driven security analysis.
1. Adversarial agent tension (HUNT vs. PROVE)
Most AI security tools ask a single model "is this vulnerable?" and accept the answer. SEC-AF structurally separates the finding agents from the disproving agents. Hunters are incentivized to find vulnerabilities; provers are incentivized to disprove them. Each finding passes through a 4-agent verification chain — a tracer reconstructs the data flow, a sanitization analyzer looks for mitigations the hunter may have missed, an exploit hypothesizer constructs a concrete attack scenario, and a verdict agent weighs all the conflicting evidence. This adversarial tension between agents is what drives the 94% noise reduction — the architecture itself encodes skepticism.
2. Signal cascade with progressive narrowing
Instead of dumping all findings on the user, the pipeline compresses signal at every stage: 106 raw findings → 61 after AI deduplication → 30 after adversarial verification. Each phase is a filter. This mirrors how human security teams triage — broad discovery first, then progressively stricter scrutiny. The key insight is that each filter is a different kind of AI reasoning: semantic similarity for dedup, taint analysis for verification, exploit construction for confirmation.
3. Information economy via context pruning
LLMs hallucinate more when given irrelevant context. SEC-AF routes only the information each agent needs: an injection hunter receives the recon context pruned to data flow maps and input entry points, while a crypto hunter receives dependency trees and key management patterns. Verifiers receive projected finding views with only the fields needed for their specific verification method. This per-strategy context pruning reduces both hallucination and cost — agents can't confuse themselves with information they never see.
4. Streaming phase overlap