
Behavioral eval lab (Quorum) for the superpowers project that drives real coding-agent CLIs (Claude, Codex, Gemini, Kimi, and more) through a QA agent and grades them on workflow compliance against scenario criteria and deterministic post-checks.
Behavioral eval lab for superpowers. Quorum drives real coding-agent CLIs (Claude, Codex, Antigravity, Gemini, Kimi, OpenCode, Pi, and Copilot) through a Gauntlet QA agent and grades them against scenario acceptance criteria plus deterministic post-checks.
Code, CLI, paths, and inline prose all use lowercase quorum; the capitalized
form Quorum appears in headings and the actor table.
This is not a generic benchmark suite. It is an eval lab for workflow compliance: skill triggering, worktree behavior, subagent coordination, verification reflexes, review quality, and cost-shaping patterns.
quorum has two very different execution modes:
biome, tsc, and
bun test. They do not call model APIs and do not launch agent CLIs.Public CI must stay on the static/unit side of that line. Never add API keys,
live quorum run … invocations, or dangerous-mode agent launches to public
CI.
Live evals run the Coding-Agent under test with broad execution power:
--dangerously-skip-permissions.--dangerously-bypass-approvals-and-sandbox.--dangerously-skip-permissions and relies on local
browser/keyring auth for agy.--skip-trust --approval-mode=yolo; API-key auth is default,
with opt-in OAuth auth for trusted local runs.--yolo.--dangerously-skip-permissions.--allow-all.quorum pins each Coding-Agent's HOME (plus the XDG base dirs and TMPDIR)
to a throwaway per-run home at <run>/home — the launcher splices in the
$QUORUM_HOME_ENV token built by src/agents/home-env.ts (xdgHomeEnv, the
single source of truth). Each agent's config dir is collapsed under that home
(Claude .claude, Codex .codex, Gemini ., OpenCode ., Antigravity .,
Copilot .copilot, Kimi .kimi-code, Pi .pi/agent), so the Coding-Agent
finds its config via its own $HOME default and never sees the host's real
~/.claude, ~/.codex, ~/.gemini, ~/.kimi-code, ~/.pi, ~/.copilot,
~/.config, or other home-relative state, installed plugins, or prior sessions.
Provisioning seeds the config — and the host OAuth creds each agent needs — into
that throwaway home before launch, so there is no run-time login. Copilot also
stages the local Superpowers plugin under the isolated home, uses an
allowlisted outer environment, and writes a secret-bearing chmod-0600
.copilot-env inside the run dir. That narrows the blast radius but is not a
sandbox. OpenCode and Copilot launchers additionally use allowlisted
environments, but live Coding-Agents still run with broad filesystem and
command execution power.
Run live evals only from a trusted local environment:
results/, raw session logs, session-state/tool-call artifacts, and
Gauntlet-Agent inputs as
sensitive.Install and run the static gates:
bun install
bun run check
bun run quorum check
Run one local or break-glass scenario outside the container:
export SUPERPOWERS_ROOT=/path/to/superpowers
export ANTHROPIC_API_KEY=...
bun run quorum run scenarios/triggering-writing-plans --coding-agent claude
bun run quorum show <run-dir>
The Gauntlet-Agent (QA driver) authenticates to Anthropic with ANTHROPIC_API_KEY
by default. To drive it from a logged-in Claude subscription instead, set
CLAUDE_CODE_OAUTH_TOKEN (from claude setup-token) in the environment (e.g.
.env); the harness passes it through and gauntlet prefers it over the API key.
Note: a subscription has usage caps sized for interactive use — high-concurrency
run-all batches can hit them, so the API key remains the better fit for heavy load.
Agent names are claude, codex, antigravity, gemini, kimi,
opencode, pi, and copilot. Not every scenario is valid for every agent.
BREAKING (credential axis): claude-haiku and claude-sonnet are no
longer separate agent names. To run the Claude harness against Sonnet or Haiku:
bun run quorum run scenarios/<name> --coding-agent claude --credential sonnet
bun run quorum run scenarios/<name> --coding-agent claude --credential haiku
The claude agent's default credential is opus.
Shared remote live evals are designed to run from a trusted appliance host with one blessed credential bundle, exact repo/ref provenance, host locks, and recoverable job records. Agents should use the appliance helper once it exists on the configured host:
evals-appliance doctor --json
evals-appliance prepare --json --superpowers-ref <branch-tag-or-sha>
evals-appliance run-all --json --detach \
--superpowers-ref <branch-tag-or-sha> \
-- --tier sentinel \
--coding-agents claude,codex,kimi \
--jobs 4
evals-appliance status --json <job-id>
evals-appliance show --json <job-id>
evals-appliance costs --json <job-id>
evals-appliance cancel --json <job-id>
The target interface and operating rules are in
docs/appliance-runbook.md, backed by
docs/superpowers/specs/2026-06-18-shared-eval-appliance-design.md.
doctor is read-only. prepare returns lock_busy rather than changing refs
while a live job is active.
Host access and provider-specific break-glass procedures are intentionally kept
out of this public repo; use the private ops runbook for those details.
Raw bun run quorum ... and scripts/evals-container exec quorum ... remain
local or trusted break-glass workflows for shared live evals.
The Docker runtime is the primary recipe for real suite runs. It keeps the evals checkout, the Superpowers checkout under test, credentials, auth sources, and all run artifacts on the host while quorum runs inside a rich Ubuntu workspace container.
Create .env.container or pass an explicit env file to up:
ANTHROPIC_API_KEY=...
OPENAI_API_KEY=...
OPENROUTER_API_KEY=... # Pi default: OpenRouter GLM 5.2
GEMINI_API_KEY=... # or GEMINI_AUTH_TYPE=oauth-personal
KIMI_MODEL_API_KEY=... # unless using mounted Kimi OAuth
PI_PROVIDER=... # only for raw/custom Pi env auth outside the default credential
PI_MODEL=...
PI_API_KEY=...
COPILOT_GITHUB_TOKEN=...
Then build, start, and validate the container:
scripts/evals-container build
scripts/evals-container down || true
scripts/evals-container --env-file .env.container up
scripts/evals-container exec evals-tool-versions
scripts/evals-container exec quorum check
The wrapper mounts this evals checkout at /workspace/evals, the parent
Superpowers checkout at /workspace/superpowers, and host results/ at
/workspace/evals/results. Override the Superpowers checkout with
--superpowers-root <dir> when the default parent path is not the system under
test.
The image build needs a local Gauntlet checkout. The wrapper discovers it
from GAUNTLET_ROOT or a Bun global bun link install; use
--gauntlet-root <dir> with build to choose explicitly.