
分散システムおよびステートフルシステム向けに、クレーム駆動のテストを設計・実行するAIコーディングエージェントのための2つのスキル。 この2つを組み合わせることで、構造化されたMarkdownテスト計画と、10状態の判定結果およびSUT/ハーネス/チェッカー/環境の明示的な責務分類を備えた調査レポートを生成します。レビュー担当者はこの2つの成果物を読んでリリース可否を判断します。その他の再実行は不要です。
Claude Code、Codex、Copilot CLI、Cursor、Gemini、またはMarkdownを読み取ってシェルを実行できるあらゆるエージェントで動作します。このスキルはプレーンなSKILL.mdファイルです。エージェントがこれを実行し、計画と調査レポートが出力となります。
一方のスキルが計画を設計し、もう一方がそれを実行します。計画は製品のクレームから始まり、そのクレームに紐づく仮説を生成し、それぞれが反証しようとするクレームにちなんで命名されたシナリオを記述します。一貫性が重要なシナリオでは、各シナリオは抽象モデル(register | queue | log | lock | lease | ledger | …)を操作履歴スキーマ、名前付きチェッカー、および観測可能な着地証拠を伴うネメシスに結び付けます。計画はカバレッジ妥当性の議論と保守的な信頼度ステートメントで締めくくられます。
分散システムおよびステートフルシステムのテストのデフォルト——統合テストをいくつか書いて完了とする——では、本番環境で実際にこれらのシステムを壊すバグのごく一部しか発見できません。部分的なネットワーク分断、非決定的な並行性、クラッシュリカバリ、アップグレード/ロールバック、リプレイ時の冪等性、タイミングに敏感な順序性などです。
これらのスキルは、この分野で苦労して得られた知見を活かした、信念に基づいたワークフローを強制します。
エンドツーエンドで、2つのスキルは以下を生成します:
docs/testing-plans/<slug>.md ← plan with §0–§9 (see below)
test-sessions/<slug>/<UTC>/
├── session-log.md ← timeline + toolbox + env probe
├── logs/ ← per-scenario stdout/stderr
├── metrics/ ← metric snapshots
├── artifacts/ ← ephemeral harnesses, dumps
└── findings/
├── <scenario>.md ← per-scenario verdict (written as run proceeds)
└── report.md ← summary + adequacy + confidence delta
計画の構造(レビュー担当者はこれを読めば、テストを再実行せずにリリース可否を判断できます):
0. Architectural summary — system as it actually exists
1. Scope
1b. Claims under test — the spine
1c. Missing claims discovered — docs ↔ code drift
2. SUT model
3. Existing test inventory — what's already covered
4. Failure-mode hypotheses — tied to claim IDs
5. Coverage matrix — claim × hypothesis
6. Technique selection — from the catalog
6b. Environment requirements
7. Scenarios — each named after the claim, with
Target test file + Skeleton
7.M Model / history / — mandatory when the scenario falsifies
checker discipline a claim in {safety, durability,
idempotency, isolation, ordering,
membership}: model under test,
operation-history schema, named
checker, nemesis + landing evidence,
ambiguous-outcome handling, reduction
plan (SUT/harness/checker/env blame)
7b. Coverage adequacy argument — why these tests are enough
7c. Residual uncertainty — what stays unverified, and why ok
7d. Confidence statement — the reviewer's verdict
8. What this plan does NOT cover
9. Open questions / followups
### Scenario S3: linearizable_append_under_partition
- Falsifies if it FAILs: C1 (every acknowledged append is durable
and linearisable), C5 (leader election completes within 5s)
- Workload: 8 clients, 70% append / 30% read, 5min, key-skew zipf
- Faults: asymmetric partition isolating current leader at T+60s
for 30s
- Oracle: linearizability via Porcupine over per-key histories
§7.M (model / history / checker discipline)
- Model under test: log
- Operation history: default 11-field schema (op id, process id,
invoke/complete ts, op type, key, input,
output, error, timeout marker, node seen,
fault epoch). Recorded in-process + server-
side audit.
- Checker: linearizability (Porcupine) per-key, then
no-lost-ack against final state
- Nemesis + landing: asymmetric-partition (iptables drop one
direction). Landing evidence = iptables drop
counter goes 0 → 14,712 over the 30s window
AND raft log emits "leader-lost; starting
election" within 2s of injection.
- Ambiguous outcomes: timeouts → timeout_marker=true, complete_ts
=null, treated as could-have-succeeded;
retries are separate ops sharing input
- Reduction plan: if FAIL, bisect fault window + fix seed, then
classify SUT / harness / checker / environment
per references/test-case-reduction.md
| ID | 判定 | ネメシスの着地証拠 | 削減クラス |
|---|---|---|---|
| S3 | PASS-hardening | iptables ctr 0→14,712; raft re-election at T+1.8s | n/a |
| S4 | FAIL-reproducible | partition landed; Elle: G2-item anomaly on key K17 | SUT |
| S7 | INCONCLUSIVE-fault-not-proven | iptables rule installed but counter stayed 0 — wrong chain | harness |
| S9 | PARTIAL-model | landing ok; checker covered per-key, not cross-key | n/a |
(完全な調査結果テンプレートには、Oracle、Oracle実行証拠、成果物リンク、計画との妥当性比較セクション、および信頼度デルタが含まれます — skills/executing-distributed-system-tests/assets/findings-report-template.md を参照してください。)
これを任意のAIコーディングエージェント(Claude Code、Codex、Copilot CLI、Cursor、Gemini、またはMarkdownを読み取ってシェルを実行するその他のツール)に貼り付けます:
Read https://raw.githubusercontent.com/shenli/distributed-system-testing/main/INSTALL.md
and follow the instructions to install and configure
distributed-testing-skills for this agent.
エージェントはINSTALL.mdを取得し、リポジトリを~/.local/share/distributed-testing-skills/にクローンして、スキルを組み込みます(Claude Codeでは~/.claude/skills/配下にシンボリックリンク、その他のエージェントでは~/AGENTS.md内のポインターブロック)。
その後、マシン上の任意のエージェントに「このシステムのテスト計画を設計して」または「Xで計画を実行して」と依頼すると、SKILL.mdのワークフローに従います。
同じ1行を再度貼り付けます。 INSTALL.mdは冪等です。インストールパスが存在すればgit pull --ff-onlyを実行し、存在しなければgit cloneを実行します。シンボリックリンクは常にクローンされたコンテンツを指すため、新しいバージョンを自動的に取り込みます。~/AGENTS.mdのポインターブロックはHTMLマーカーを使用しており、実行のたびに重複なくきれいに置き換えられます。
クローンしたスキルにローカルでの編集がある場合、git pull --ff-onlyは失敗します。エージェントは破棄する前に停止して確認を求めます。
git clone https://github.com/shenli/distributed-system-testing.git \
~/.local/share/distributed-testing-skills
# Claude Code: symlink under ~/.claude/skills/
mkdir -p ~/.claude/skills
ln -snf ~/.local/share/distributed-testing-skills/skills/designing-distributed-system-tests \
~/.claude/skills/designing-distributed-system-tests
ln -snf ~/.local/share/distributed-testing-skills/skills/executing-distributed-system-tests \
~/.claude/skills/executing-distributed-system-tests
# Codex / Copilot CLI / Cursor / Gemini / others: see INSTALL.md
リポジトリには.claude-plugin/配下にプラグインマニフェストとマーケットプレイスマニフェストが含まれているため、Claude Codeはシンボリックリンクの代わりにプラグインとしてインストールできます:
/plugin marketplace add shenli/distributed-system-testing
/plugin install distributed-testing-skills@distributed-testing-skills
両方のスキルはskills/から自動検出されます。上記の1行INSTALL.mdフローは、エージェント非依存の経路(Codex、Copilot CLI、Cursor、Gemini)として残ります。
スキルがインストールされると、2つの方法でこれらを駆動できます:
カジュアルな依頼(自動トリガー付きClaude Code):