
분산 시스템 테스트를 위한 AI 에이전트 스킬
분산 및 상태 저장 시스템을 대상으로 클레임 기반 테스트를 설계·실행하는 AI 코딩 에이전트용 두 가지 스킬. 두 스킬은 함께 10가지 상태 판정과 명시적인 SUT / 하니스 / 체커 / 환경 책임 분류가 포함된 구조화된 Markdown 테스트 계획과 발견 보고서를 생성한다. 리뷰어는 두 산출물을 읽고 출시 여부를 결정한다. 그 외에 다시 실행할 것은 없다.
Claude Code, Codex, Copilot CLI, Cursor, Gemini 또는 Markdown을 읽고 셸을 실행하는 모든 에이전트에서 동작한다. 스킬은 단순한 SKILL.md 파일이다. 에이전트가 이를 실행하며, 계획과 발견 보고서가 출력물이다.
한 스킬은 계획을 설계한다. 다른 스킬은 이를 실행한다. 계획은 제품의 클레임에서 시작하여 해당 클레임에 연결된 가설을 생성하고, 각각이 반증하려는 클레임의 이름을 딴 시나리오를 작성한다. 일관성 중요 시나리오의 경우 각 시나리오는 추상 모델(register | queue | log | lock | lease | ledger | …)을 작업 이력 스키마, 명명된 체커, 관찰 가능한 착지 증거가 있는 네미시스에 바인딩한다. 계획은 커버리지 적정성 논증과 보수적인 신뢰도 진술로 끝난다.
분산 및 상태 저장 시스템 테스트의 기본 방식 — 통합 테스트 몇 개를 작성하고 끝내는 것 — 은 실제 프로덕션에서 이 시스템을 망가뜨리는 버그의 극히 일부만 찾아낸다: 부분 네트워크 분할, 비결정적 동시성, 크래시 복구, 업그레이드/롤백, 리플레이 하에서의 멱등성, 타이밍에 민감한 순서.
이 스킬들은 현장의 값진 경험에서 얻은 지식을 활용하는 확고한 방식의 워크플로우를 강제한다:
엔드투엔드로 두 스킬은 다음을 생성한다:
docs/testing-plans/<slug>.md ← plan with §0–§9 (see below)
test-sessions/<slug>/<UTC>/
├── session-log.md ← timeline + toolbox + env probe
├── logs/ ← per-scenario stdout/stderr
├── metrics/ ← metric snapshots
├── artifacts/ ← ephemeral harnesses, dumps
└── findings/
├── <scenario>.md ← per-scenario verdict (written as run proceeds)
└── report.md ← summary + adequacy + confidence delta
계획 구조(리뷰어는 테스트를 다시 실행하지 않고 이 내용만으로 출시 여부를 결정할 수 있다):
0. Architectural summary — system as it actually exists
1. Scope
1b. Claims under test — the spine
1c. Missing claims discovered — docs ↔ code drift
2. SUT model
3. Existing test inventory — what's already covered
4. Failure-mode hypotheses — tied to claim IDs
5. Coverage matrix — claim × hypothesis
6. Technique selection — from the catalog
6b. Environment requirements
7. Scenarios — each named after the claim, with
Target test file + Skeleton
7.M Model / history / — mandatory when the scenario falsifies
checker discipline a claim in {safety, durability,
idempotency, isolation, ordering,
membership}: model under test,
operation-history schema, named
checker, nemesis + landing evidence,
ambiguous-outcome handling, reduction
plan (SUT/harness/checker/env blame)
7b. Coverage adequacy argument — why these tests are enough
7c. Residual uncertainty — what stays unverified, and why ok
7d. Confidence statement — the reviewer's verdict
8. What this plan does NOT cover
9. Open questions / followups
### Scenario S3: linearizable_append_under_partition
- Falsifies if it FAILs: C1 (every acknowledged append is durable
and linearisable), C5 (leader election completes within 5s)
- Workload: 8 clients, 70% append / 30% read, 5min, key-skew zipf
- Faults: asymmetric partition isolating current leader at T+60s
for 30s
- Oracle: linearizability via Porcupine over per-key histories
§7.M (model / history / checker discipline)
- Model under test: log
- Operation history: default 11-field schema (op id, process id,
invoke/complete ts, op type, key, input,
output, error, timeout marker, node seen,
fault epoch). Recorded in-process + server-
side audit.
- Checker: linearizability (Porcupine) per-key, then
no-lost-ack against final state
- Nemesis + landing: asymmetric-partition (iptables drop one
direction). Landing evidence = iptables drop
counter goes 0 → 14,712 over the 30s window
AND raft log emits "leader-lost; starting
election" within 2s of injection.
- Ambiguous outcomes: timeouts → timeout_marker=true, complete_ts
=null, treated as could-have-succeeded;
retries are separate ops sharing input
- Reduction plan: if FAIL, bisect fault window + fix seed, then
classify SUT / harness / checker / environment
per references/test-case-reduction.md
| ID | 판정 | 네미시스 착지 증거 | 축소 분류 |
|---|---|---|---|
| S3 | PASS-hardening | iptables ctr 0→14,712; T+1.8s에 raft 재선거 | n/a |
| S4 | FAIL-reproducible | 분할 착지; Elle: key K17에서 G2-item anomaly | SUT |
| S7 | INCONCLUSIVE-fault-not-proven | iptables 규칙은 설치됐지만 카운터가 0에 머물음 — 잘못된 체인 | harness |
| S9 | PARTIAL-model | 착지 정상; 체커는 key별로만 적용, key 간 미적용 | n/a |
(전체 발견 템플릿은 Oracle, Oracle 실행 증거, 아티팩트 링크, 계획 대비 적정성 섹션, 신뢰도 델타를 포함한다 — skills/executing-distributed-system-tests/assets/findings-report-template.md 참조.)
이 한 줄을 아무 AI 코딩 에이전트(Claude Code, Codex, Copilot CLI, Cursor, Gemini 또는 Markdown을 읽고 셸을 실행하는 모든 도구)에 붙여넣기:
Read https://raw.githubusercontent.com/shenli/distributed-system-testing/main/INSTALL.md
and follow the instructions to install and configure
distributed-testing-skills for this agent.
에이전트는 INSTALL.md를 가져와 저장소를 ~/.local/share/distributed-testing-skills/에 클론하고 스킬을 연결한다(Claude Code는 ~/.claude/skills/ 아래 심링크, 다른 에이전트는 ~/AGENTS.md의 포인터 블록).
그 후 머신의 어떤 에이전트에게든 "이 시스템을 위한 테스트 계획을 설계해" 또는 "X에 있는 계획을 실행해"라고 요청하면 SKILL.md 워크플로우를 따른다.
같은 한 줄 명령을 다시 붙여넣기하라. INSTALL.md는 멱등적이다: 설치 경로가 이미 있으면 git pull --ff-only를 실행하고, 없으면 git clone을 실행한다. 심링크는 항상 클론된 콘텐츠를 가리키므로 새 버전을 자동으로 반영한다. ~/AGENTS.md 포인터 블록은 HTML 마커를 사용하며 실행할 때마다 깔끔하게 교체된다 — 중복이 없다.
클론된 스킬에 로컬 수정 사항이 있으면 git pull --ff-only는 실패한다. 에이전트는 수정 사항을 폐기하기 전에 멈추고 확인을 요청한다.
git clone https://github.com/shenli/distributed-system-testing.git \
~/.local/share/distributed-testing-skills
# Claude Code: symlink under ~/.claude/skills/
mkdir -p ~/.claude/skills
ln -snf ~/.local/share/distributed-testing-skills/skills/designing-distributed-system-tests \
~/.claude/skills/designing-distributed-system-tests
ln -snf ~/.local/share/distributed-testing-skills/skills/executing-distributed-system-tests \
~/.claude/skills/executing-distributed-system-tests
# Codex / Copilot CLI / Cursor / Gemini / others: see INSTALL.md
저장소는 .claude-plugin/ 아래에 플러그인 매니페스트와 마켓플레이스 매니페스트를 포함하므로, Claude Code는 심링크 대신 플러그인으로 설치할 수 있다:
/plugin marketplace add shenli/distributed-system-testing
/plugin install distributed-testing-skills@distributed-testing-skills
두 스킬 모두 skills/에서 자동으로 발견된다. 위의 한 줄 INSTALL.md 흐름은 에이전트와 무관한 경로로 유지된다(Codex, Copilot CLI, Cursor, Gemini).
스킬이 설치되면 두 가지 방법으로 구동할 수 있다:
간편 요청(자동 트리거가 있는 Claude Code):