Back to updates
New releaseSep 15, 2026

giskard-oss giskard-checks/v1.0.4

🐒 Open-Source Evaluation & Testing library for LLM Agents

Share

giskardlogo giskardlogo

Evals, Red Teaming and Test Generation for Agentic Systems

Modular, Lightweight, Dynamic and Async-first

GitHub release License Downloads CI Giskard on Discord

Docs β€’ Website β€’ Community


[!IMPORTANT] Giskard v3 is a fresh rewrite designed for dynamic, multi-turn testing of AI agents. This release drops heavy dependencies for better efficiency while introducing a more powerful AI vulnerability scanner and enhanced RAG evaluation β€” both now shipping natively in giskard-scan, with no dependency on v2. Only the legacy scan for tabular/ML models remains v2-only. Giskard v2 remains available but is no longer actively maintained. Follow progress β†’ Read the v3 Announcement Β· Roadmap

Install

pip install giskard           # checks (+ agents, llm, core)
pip install "giskard[scan]"   # + vulnerability / quality scan
pip install "giskard[openai]" # provider SDK for LLM judges / generators

Requires Python 3.12+.

ExtraAdds
(none)giskard-checks and dependencies
scangiskard-scan
openai / anthropic / …provider SDKs (see pyproject.toml optional deps)

Telemetry: optional aggregated analytics via giskard-core. No prompts or outputs are sent. Opt out before importing Giskard: export DO_NOT_TRACK=1 or export GISKARD_TELEMETRY_DISABLED=1. Details: giskard-core README.


Giskard is an open-source Python library for testing and evaluating agentic systems. The v3 architecture is a modular set of focused packages β€” each carrying only the dependencies it needs β€” built from scratch to wrap anything: an LLM, a black-box agent, or a multi-step pipeline.

StatusPackageDescription
βœ… Stablegiskard-checksTesting & evaluation β€” scenario API, built-in checks, LLM-as-judge
βœ… Stablegiskard-scanAgent vulnerability scanner + RAG/quality evaluation β€” red teaming, prompt injection, jailbreaks & harmful content (vulnerability_scan, successor of v2 Scan), plus knowledge-base quality eval (quality_scan, successor of v2 RAGET)

These build on three foundational libraries β€” giskard-core (shared utilities & telemetry), giskard-llm (provider-agnostic LLM routing), and giskard-agents (agent & workflow orchestration) β€” which are pulled in automatically and rarely used directly.

Giskard Checks β€” create and apply evals for testing agents

pip install giskard-checks

Giskard Checks is a lightweight library for creating evaluations (evals) that test LLM-based systems β€” from simple assertions to LLM-as-judge assessments. Unlike traditional unit tests, evals are designed for non-deterministic outputs where the same input can produce different valid responses.

Use Giskard Checks to:

  • Catch regressions β€” verify your system still behaves correctly after changes
  • Validate RAG quality β€” check if answers are grounded in retrieved context
  • Enforce safety rules β€” ensure outputs conform to your content policies
  • Evaluate multi-turn agents β€” test full conversations, not just single exchanges

Built-in evals include string matching, comparisons, regex, semantic similarity, and LLM-as-judge checks (Groundedness, Conformity, LLMJudge).

Concepts

  • Target β€” your system under test: any sync/async callable (inputs) -> outputs (optionally with trace)
  • Scenario β€” one eval: interactions + checks
  • Check β€” assertion or LLM judge over the trace
  • Suite β€” many scenarios run together

giskard.agents.Generator is an LLM client for workflows/judges β€” not the same as giskard.checks input generators (LLMGenerator) that synthesize user messages.

Quickstart

import asyncio
from giskard.checks import Scenario, Groundedness


def get_answer(inputs: str) -> str:
    return "Paris"  # replace with your model / agent


async def main() -> None:
    scenario = (
        Scenario("test_france_capital")
        .interact(inputs="What is the capital of France?", outputs=get_answer)
        .check(
            Groundedness(
                name="answer is grounded",
                context="France is in Western Europe. Its capital is Paris.",
            )
        )
    )
    result = await scenario.run()
    result.print_report()


asyncio.run(main())

Groundedness is an LLM judge β€” install a provider extra (e.g. pip install "giskard[openai]") and set the matching API key. Default model: openai/gpt-4o-mini.

See the full docs for Suites, LLMJudge, multi-turn scenarios, and more.


Giskard Scan β€” vulnerability scanner for AI agents

pip install "giskard[scan]"   # or: pip install giskard-scan

Giskard Scan is the red-teaming and vulnerability scanning layer for agentic systems. It generates adversarial test suites automatically from a plain-language description of your agent, covering prompt injection, harmful content, stereotypes, misinformation, and more.

Use Giskard Scan to:

  • Red-team your agent β€” automatically generate adversarial inputs across OWASP LLM Top-10 threat categories
  • Run prompt-injection probes β€” built-in dataset of injection payloads ready to use
  • Extend with custom generators β€” pass your own ScenarioGenerator instances to generate_suite, or register them on vulnerability_suite_generator_registry

Quickstart

import asyncio
from giskard.scan import vulnerability_scan

Categories