Skip to content
KitploitKITPLOIT
ToolsExploitsBlog
Log in
Submit
ToolsExploitsBlog
Submit

Hacking, PenTest, and Cybersecurity Tools for Your Security Arsenal!

Kitploit is a directory of hacking, cybersecurity, and pentesting tools. Discover the latest project updates to find vulnerabilities, analyze systems, automate testing, and strengthen your security.

FeedsContactPrivacy© 2026 Kitploit

Tool Directory

Categories

View all categories
Loading categories
superpowers-evals — Behavioral eval lab (Quorum) for the superpowers project that drives real coding-agent CLIs (Claude, Codex, Gemini, Kimi, and more) through a QA agent and grades them on workflow compliance against scenario criteria and deterministic post-checks. | Kitploit
Tools/GitHubGitHub/prime-radiant-inc/superpowers-evals
Scripting & AutomationPenetration TestingUtilities & FrameworksLearning & EducationAI SecurityLabs & Practice
GitHubprime-radiant-inc/superpowers-evals

superpowers-evals

Behavioral eval lab (Quorum) for the superpowers project that drives real coding-agent CLIs (Claude, Codex, Gemini, Kimi, and more) through a QA agent and grades them on workflow compliance against scenario criteria and deterministic post-checks.

View Repository
97122019 days agoReviewed by Kitploit

Most Popular

View all →

Discover the most used tools by our community.

Explore all tools

Browse our collection of tools

View all tools →
Share

Superpowers Evals

Behavioral eval lab for superpowers. Quorum drives real coding-agent CLIs (Claude, Codex, Antigravity, Gemini, Kimi, OpenCode, Pi, and Copilot) through a Gauntlet QA agent and grades them against scenario acceptance criteria plus deterministic post-checks.

Code, CLI, paths, and inline prose all use lowercase quorum; the capitalized form Quorum appears in headings and the actor table.

This is not a generic benchmark suite. It is an eval lab for workflow compliance: skill triggering, worktree behavior, subagent coordination, verification reflexes, review quality, and cost-shaping patterns.

Safety Model

quorum has two very different execution modes:

  • Static/unit checks are safe for public CI. They run biome, tsc, and bun test. They do not call model APIs and do not launch agent CLIs.
  • Live evals are trusted-maintainer operations. They launch Claude Code, Codex CLI, Antigravity CLI, Gemini CLI, Kimi Code, OpenCode CLI, Pi CLI, or Copilot CLI in permissive modes and collect raw transcripts, tool calls, filesystem state, and session logs.

Public CI must stay on the static/unit side of that line. Never add API keys, live quorum run … invocations, or dangerous-mode agent launches to public CI.

Live Eval Risk

Live evals run the Coding-Agent under test with broad execution power:

  • Claude uses --dangerously-skip-permissions.
  • Codex uses --dangerously-bypass-approvals-and-sandbox.
  • Antigravity uses --dangerously-skip-permissions and relies on local browser/keyring auth for agy.
  • Gemini uses --skip-trust --approval-mode=yolo; API-key auth is default, with opt-in OAuth auth for trusted local runs.
  • Kimi uses --yolo.
  • OpenCode uses --dangerously-skip-permissions.
  • Pi uses explicit tool allowlists and API-key auth in a run-local config dir.
  • Copilot uses --allow-all.

quorum pins each Coding-Agent's HOME (plus the XDG base dirs and TMPDIR) to a throwaway per-run home at <run>/home — the launcher splices in the $QUORUM_HOME_ENV token built by src/agents/home-env.ts (xdgHomeEnv, the single source of truth). Each agent's config dir is collapsed under that home (Claude .claude, Codex .codex, Gemini ., OpenCode ., Antigravity ., Copilot .copilot, Kimi .kimi-code, Pi .pi/agent), so the Coding-Agent finds its config via its own $HOME default and never sees the host's real ~/.claude, ~/.codex, ~/.gemini, ~/.kimi-code, ~/.pi, ~/.copilot, ~/.config, or other home-relative state, installed plugins, or prior sessions. Provisioning seeds the config — and the host OAuth creds each agent needs — into that throwaway home before launch, so there is no run-time login. Copilot also stages the local Superpowers plugin under the isolated home, uses an allowlisted outer environment, and writes a secret-bearing chmod-0600 .copilot-env inside the run dir. That narrows the blast radius but is not a sandbox. OpenCode and Copilot launchers additionally use allowlisted environments, but live Coding-Agents still run with broad filesystem and command execution power.

Run live evals only from a trusted local environment:

  • Export only the API key needed for the selected Coding-Agent.
  • Avoid running with broad production or personal secrets in the environment.
  • Treat results/, raw session logs, session-state/tool-call artifacts, and Gauntlet-Agent inputs as sensitive.
  • Do not commit or paste raw run artifacts without checking them first.

Quick Start

Install and run the static gates:

bun install
bun run check
bun run quorum check

Run one local or break-glass scenario outside the container:

export SUPERPOWERS_ROOT=/path/to/superpowers
export ANTHROPIC_API_KEY=...
bun run quorum run scenarios/triggering-writing-plans --coding-agent claude
bun run quorum show <run-dir>

The Gauntlet-Agent (QA driver) authenticates to Anthropic with ANTHROPIC_API_KEY by default. To drive it from a logged-in Claude subscription instead, set CLAUDE_CODE_OAUTH_TOKEN (from claude setup-token) in the environment (e.g. .env); the harness passes it through and gauntlet prefers it over the API key. Note: a subscription has usage caps sized for interactive use — high-concurrency run-all batches can hit them, so the API key remains the better fit for heavy load.

Agent names are claude, codex, antigravity, gemini, kimi, opencode, pi, and copilot. Not every scenario is valid for every agent.

BREAKING (credential axis): claude-haiku and claude-sonnet are no longer separate agent names. To run the Claude harness against Sonnet or Haiku:

bun run quorum run scenarios/<name> --coding-agent claude --credential sonnet
bun run quorum run scenarios/<name> --coding-agent claude --credential haiku

The claude agent's default credential is opus.

Shared Eval Appliance

Shared remote live evals are designed to run from a trusted appliance host with one blessed credential bundle, exact repo/ref provenance, host locks, and recoverable job records. Agents should use the appliance helper once it exists on the configured host:

evals-appliance doctor --json
evals-appliance prepare --json --superpowers-ref <branch-tag-or-sha>
evals-appliance run-all --json --detach \
  --superpowers-ref <branch-tag-or-sha> \
  -- --tier sentinel \
     --coding-agents claude,codex,kimi \
     --jobs 4
evals-appliance status --json <job-id>
evals-appliance show --json <job-id>
evals-appliance costs --json <job-id>
evals-appliance cancel --json <job-id>

The target interface and operating rules are in docs/appliance-runbook.md, backed by docs/superpowers/specs/2026-06-18-shared-eval-appliance-design.md. doctor is read-only. prepare returns lock_busy rather than changing refs while a live job is active. Host access and provider-specific break-glass procedures are intentionally kept out of this public repo; use the private ops runbook for those details. Raw bun run quorum ... and scripts/evals-container exec quorum ... remain local or trusted break-glass workflows for shared live evals.

Container Runtime

The Docker runtime is the primary recipe for real suite runs. It keeps the evals checkout, the Superpowers checkout under test, credentials, auth sources, and all run artifacts on the host while quorum runs inside a rich Ubuntu workspace container.

Create .env.container or pass an explicit env file to up:

ANTHROPIC_API_KEY=...
OPENAI_API_KEY=...
OPENROUTER_API_KEY=...      # Pi default: OpenRouter GLM 5.2
GEMINI_API_KEY=...          # or GEMINI_AUTH_TYPE=oauth-personal
KIMI_MODEL_API_KEY=...      # unless using mounted Kimi OAuth
PI_PROVIDER=...             # only for raw/custom Pi env auth outside the default credential
PI_MODEL=...
PI_API_KEY=...
COPILOT_GITHUB_TOKEN=...

Then build, start, and validate the container:

scripts/evals-container build
scripts/evals-container down || true
scripts/evals-container --env-file .env.container up
scripts/evals-container exec evals-tool-versions
scripts/evals-container exec quorum check

The wrapper mounts this evals checkout at /workspace/evals, the parent Superpowers checkout at /workspace/superpowers, and host results/ at /workspace/evals/results. Override the Superpowers checkout with --superpowers-root <dir> when the default parent path is not the system under test.

The image build needs a local Gauntlet checkout. The wrapper discovers it from GAUNTLET_ROOT or a Bun global bun link install; use --gauntlet-root <dir> with build to choose explicitly.

Download Tool