
AI 에이전트를 위한 pytest - 자율적 레드티밍, 행동 모니터링 및 LLM 에이전트 보안 테스팅
██████╗██████╗ ██╗ ██╗ ██████╗██╗██████╗ ██╗ ███████╗ ██╔════╝██╔══██╗██║ ██║██╔════╝██║██╔══██╗██║ ██╔════╝ ██║ ██████╔╝██║ ██║██║ ██║██████╔╝██║ █████╗ ██║ ██╔══██╗██║ ██║██║ ██║██╔══██╗██║ ██╔══╝ ╚██████╗██║ ██║╚██████╔╝╚██████╗██║██████╔╝███████╗███████╗ ╚═════╝╚═╝ ╚═╝ ╚═════╝ ╚═════╝╚═╝╚═════╝ ╚══════╝╚══════╝AI 에이전트를 위한 pytest -- 프로덕션 전에 테스트, 점수 매기기, 강화하기
pip install crucible-security
🆕 AI 보안이 처음이신가요? 초보자용 시작 가이드를 읽거나 n8n 로컬 데모 대상 가이드로 로컬 테스트 대상을 설정하세요.
crucible init --target https://my-agent.com/api/chat
crucible scan --target https://my-agent.com/api/chat
crucible report crucible-report.json
하나의 명령어. 90개 공격. 아름다운 보고서.
crucible scan --output json을 모든 파이프라인에 연결; 낮은 등급 시 빌드 실패Crucible이 Garak 및 PyRIT과 어떻게 비교되나요? → 자세한 객관적 기능 매트릭스는 docs/comparison.md를 참조하세요.
Crucible은 무엇을 테스트하나요? → 전체 OWASP Agentic AI Top 10 공격 문서(ASI01–ASI10)는 docs/owasp_mapping.md를 참조하세요.
지속적인 대시보드, 규정 준수 보고서 및 팀 협업이 필요하신가요?
곧 출시될 클라우드 플랫폼 대기자 명단에 등록하세요: crucible-cloud.vercel.app
| 모듈 | 공격 수 | 상태 | OWASP 커버리지 |
|---|---|---|---|
| Prompt Injection | 50 | ✅ Live | LLM01, LLM07 |
| Goal Hijacking | 20 | ✅ Live | Agentic #1 |
| Jailbreaks | 20 | ✅ Live | LLM01, LLM06 |
| Enterprise Graph | 10 | ✅ Live | Agentic #2, #4 |
| Memory Poisoning | 8 | ✅ Live | Agentic #5 |
| Infrastructure Escalation | 5 | ✅ Live | LLM06, SSRF |
| Advanced Orchestration | 4 | ✅ Live | Agentic #3 |
| MCP Security | 5 | ✅ Live | Agentic #3 |
| MCP Server Scan | 10 | ✅ Live (v0.4) | MCP-001 – MCP-005 |
| Behavioral Drift | multi-turn | ✅ Live (v0.3) | Agentic #1, #2 |
| Multi-turn Attacks | strategies | ✅ Live (v0.3) | LLM01, Agentic #1 |
| Deep Research Engine | autonomous | ✅ Live (v0.4) | AI Research |
| Multi-Agent Contagion | orchestration | ✅ Live (v0.4) | Agentic #2, #3 |
| Hallucination Detection | 15 | ✅ Live (v0.5) | LLM09 / Agentic #9 |
| Toxicity & Content Safety | 20 | ✅ Live (v0.5) | LLM01, LLM06 |
| Statistical Confidence | --confidence | ✅ Live (v0.6) | Bootstrap & binomial bounds |
| MCP Trace Proxy | traffic proxy | ✅ Live (v0.7) | Agentic #3 / Tool Misuse |
| Memory & RAG Poisoning | poison-test | ✅ Live (v0.8) | Agentic #5 / Poisoning |
| Reference Targets | 12 targets | ✅ Live (v0.18) | Ground-truth validation targets |
| # | 카테고리 | Crucible 모듈 | 상태 |
|---|---|---|---|
| 1 | Goal Hijacking | goal_hijacking | 커버됨 (20 attacks) |
| 2 | Prompt Injection | prompt_injection | 커버됨 (50 attacks) |
| 3 | Tool Misuse | tool_injection / trace proxy | 커버됨 (v0.7.0) |
| 4 | Identity Abuse | trace proxy + identity layer | 커버됨 (v0.9.0) |
| 5 | Memory Poisoning | memory_poisoning / poison-test | 커버됨 (8 attacks, v0.8.0) |
| 6 | Data Exfiltration | prompt_injection / exfiltration | 커버됨 (v0.8.0) |
| 7 | Scope Violation | trace proxy | 커버됨 (v0.7.0) |
| 8 | Cascading Failure | -- | 계획됨 |
| 9 | Supply Chain / Overreliance | hallucination | 커버됨 (15 attacks) |
| 10 | Rogue Agent | -- | 계획됨 |
| 제공자 | 테스트됨 |
|---|---|
| OpenAI (GPT-4, GPT-4o) | 예 |
| Anthropic (Claude) | 예 |
| Groq (Llama, Mixtral) | 예 |
| Custom HTTP endpoint | 예 |
| LangChain (LangServe / FastAPI wrapper) | 예 |
| Ollama | 예 (v0.5) |
| LM Studio | 예 (v0.5) |
| HuggingFace TGI | 예 (v0.5) |
시작하는 데 도움이 되는 몇 가지 예제 스크립트가 examples/ 디렉토리에 제공됩니다:
| 스크립트 | 프레임워크 | 설명 |
|---|---|---|
test_openai_agent.py | OpenAI Chat Completions | OpenAI /chat/completions 엔드포인트 스캔 |
test_langchain_agent.py | LangChain (LangServe) | OWASP LLM Top 10 매핑으로 LangChain ReAct 에이전트 스캔 |
test_openai_assistant.py | OpenAI Assistants API | Assistants API 래퍼 엔드포인트 스캔 |
모든 예제는 respx를 사용하여 HTTP 호출을 모의하므로 실제 서버 없이 CI를 통과합니다.
LangChain 예제 실행:
python examples/test_langchain_agent.py
OpenAI Assistant 예제 실행:
python examples/test_openai_assistant.py
점수는 100에서 시작하여 발견된 취약점마다 차감됩니다:
| 심각도 | 차감 점수 |
|---|---|
| CRITICAL | -20 points |
| HIGH | -10 points |
| MEDIUM | -5 points |
| LOW | -2 points |
| 등급 | 점수 범위 |
|---|---|
| A | 90 -- 100 |
| B | 75 -- 89 |
| C | 60 -- 74 |
| D | 40 -- 59 |
| F | 40 미만 |
# Generate config
crucible init --target URL --provider openai --key sk-xxx
# Run a standard scan
crucible scan \
--target https://my-agent.com/api/chat \
--name "My ChatBot" \
--header "Authorization: Bearer sk-xxx" \
--timeout 30 \
--concurrency 5
# Run with payload mutation (bypass WAFs/guardrails)
crucible scan --target URL --mutate
# Multi-turn attack strategy
crucible scan --target URL --strategy multi-turn
# Use agent profile to target attacks
crucible profile --target URL --output agent_profile.json
crucible scan --target URL --profile agent_profile.json
# Behavioral integrity audit (multi-turn drift detection)
crucible behavioral-audit \
--target https://my-agent.com/api/chat \
--baseline-turns 5 \
--probe-turns 15
# Generate EU AI Act compliance report from scan results
crucible scan --target URL --output json > results.json
crucible compliance-report --results results.json --output compliance.md
# JSON output for CI/CD
crucible scan --target URL --output json > report.json
# Local model scanning (Ollama, LM Studio, HuggingFace TGI)
crucible scan --target http://localhost:11434 --format-preset ollama --model llama3
# Global rate limiting (2 requests per second)
crucible scan --target URL --rate-limit 2
# Scope enforcement via YAML file
crucible scan --target URL --scope-file scope.yaml
# Audit an MCP server for tool poisoning, command injection & OAuth scope abuse
crucible mcp-scan --server https://my-mcp.example.com
# With auth header and JSON output
crucible mcp-scan --server http://localhost:3000 \
--header "Authorization: Bearer sk-xxx" \
--output mcp-report.json
# Re-render a saved report
crucible report report.json
# Run scan with bootstrap statistical confidence intervals (calculate 95% CI with 10 runs per attack)
crucible scan --target URL --confidence --confidence-runs 10