
giskard-oss giskard-checks/v1.0.4
π’ Open-Source Evaluation & Testing library for LLM Agents
Evals, Red Teaming and Test Generation for Agentic Systems
Modular, Lightweight, Dynamic and Async-first
Docs β’ Website β’ Community
[!IMPORTANT] Giskard v3 is a fresh rewrite designed for dynamic, multi-turn testing of AI agents. This release drops heavy dependencies for better efficiency while introducing a more powerful AI vulnerability scanner and enhanced RAG evaluation β both now shipping natively in
giskard-scan, with no dependency on v2. Only the legacy scan for tabular/ML models remains v2-only. Giskard v2 remains available but is no longer actively maintained. Follow progress β Read the v3 Announcement Β· Roadmap
Install
pip install giskard # checks (+ agents, llm, core)
pip install "giskard[scan]" # + vulnerability / quality scan
pip install "giskard[openai]" # provider SDK for LLM judges / generators
Requires Python 3.12+.
| Extra | Adds |
|---|---|
| (none) | giskard-checks and dependencies |
scan | giskard-scan |
openai / anthropic / β¦ | provider SDKs (see pyproject.toml optional deps) |
Telemetry: optional aggregated analytics via giskard-core. No prompts or outputs are sent.
Opt out before importing Giskard: export DO_NOT_TRACK=1 or export GISKARD_TELEMETRY_DISABLED=1.
Details: giskard-core README.
Giskard is an open-source Python library for testing and evaluating agentic systems. The v3 architecture is a modular set of focused packages β each carrying only the dependencies it needs β built from scratch to wrap anything: an LLM, a black-box agent, or a multi-step pipeline.
| Status | Package | Description |
|---|---|---|
| β Stable | giskard-checks | Testing & evaluation β scenario API, built-in checks, LLM-as-judge |
| β Stable | giskard-scan | Agent vulnerability scanner + RAG/quality evaluation β red teaming, prompt injection, jailbreaks & harmful content (vulnerability_scan, successor of v2 Scan), plus knowledge-base quality eval (quality_scan, successor of v2 RAGET) |
These build on three foundational libraries β giskard-core (shared utilities & telemetry), giskard-llm (provider-agnostic LLM routing), and giskard-agents (agent & workflow orchestration) β which are pulled in automatically and rarely used directly.
Giskard Checks β create and apply evals for testing agents
pip install giskard-checks
Giskard Checks is a lightweight library for creating evaluations (evals) that test LLM-based systems β from simple assertions to LLM-as-judge assessments. Unlike traditional unit tests, evals are designed for non-deterministic outputs where the same input can produce different valid responses.
Use Giskard Checks to:
- Catch regressions β verify your system still behaves correctly after changes
- Validate RAG quality β check if answers are grounded in retrieved context
- Enforce safety rules β ensure outputs conform to your content policies
- Evaluate multi-turn agents β test full conversations, not just single exchanges
Built-in evals include string matching, comparisons, regex, semantic similarity, and LLM-as-judge checks (Groundedness, Conformity, LLMJudge).
Concepts
- Target β your system under test: any sync/async callable
(inputs) -> outputs(optionally withtrace) - Scenario β one eval: interactions + checks
- Check β assertion or LLM judge over the trace
- Suite β many scenarios run together
giskard.agents.Generator is an LLM client for workflows/judges β not the same as
giskard.checks input generators (LLMGenerator) that synthesize user messages.
Quickstart
import asyncio
from giskard.checks import Scenario, Groundedness
def get_answer(inputs: str) -> str:
return "Paris" # replace with your model / agent
async def main() -> None:
scenario = (
Scenario("test_france_capital")
.interact(inputs="What is the capital of France?", outputs=get_answer)
.check(
Groundedness(
name="answer is grounded",
context="France is in Western Europe. Its capital is Paris.",
)
)
)
result = await scenario.run()
result.print_report()
asyncio.run(main())
Groundedness is an LLM judge β install a provider extra (e.g. pip install "giskard[openai]") and set the matching API key. Default model: openai/gpt-4o-mini.
See the full docs for Suites, LLMJudge, multi-turn scenarios, and more.
Giskard Scan β vulnerability scanner for AI agents
pip install "giskard[scan]" # or: pip install giskard-scan
Giskard Scan is the red-teaming and vulnerability scanning layer for agentic systems. It generates adversarial test suites automatically from a plain-language description of your agent, covering prompt injection, harmful content, stereotypes, misinformation, and more.
Use Giskard Scan to:
- Red-team your agent β automatically generate adversarial inputs across OWASP LLM Top-10 threat categories
- Run prompt-injection probes β built-in dataset of injection payloads ready to use
- Extend with custom generators β pass your own
ScenarioGeneratorinstances togenerate_suite, or register them onvulnerability_suite_generator_registry
Quickstart
import asyncio
from giskard.scan import vulnerability_scan