
Research code for HARDE, an agent harness that probes and adaptively optimizes components for runtime risk detection and execution control across agent safety benchmarks.
Code release for HARDE: Optimizing Agent Harnesses for Runtime Risk Detection and Execution Control.
The repository contains two complementary entry-point families:
domains/: the Meta-Harness baseline and the two stages of HARDE for each benchmark.personalized_harness/: benchmark adapters, baseline harnesses, learned/personalized harnesses, and evaluation pipelines.HARDE/
├── domains/
│ ├── agentdyn/ # Meta-Harness baseline
│ ├── agentdyn_component_probe/ # HARDE Stage I
│ ├── agentdyn_adaptive_component_search/ # HARDE Stage II
│ ├── agentsafetybench/ # Meta-Harness baseline
│ ├── agentsafetybench_component_probe/ # HARDE Stage I
│ ├── agentsafetybench_adaptive_component_search/ # HARDE Stage II
│ ├── shade_arena/ # Meta-Harness baseline
│ ├── shade_arena_component_probe/ # HARDE Stage I
│ └── shade_arena_adaptive_component_search/ # HARDE Stage II
├── personalized_harness/
│ ├── AgentDyn/
│ ├── AgentSafetyBench/
│ └── SHADE_v2/
├── .env.example
├── .gitignore
└── requirements.txt
For a benchmark named benchmark, the directories follow this convention:
| Directory | Method stage | Main entry point |
|---|---|---|
domains/benchmark/ | Meta-Harness baseline | meta_harness.py |
domains/benchmark_component_probe/ | HARDE Stage I: component probing | component_probe.py |
domains/benchmark_adaptive_component_search/ | HARDE Stage II: adaptive component search | adaptive_component_search.py |
Stage I probes the harness components and produces a component-level rubric. Stage II consumes that rubric to select and optimize a component adaptively during search. In this repository, benchmark is one of agentdyn, agentsafetybench, or shade_arena.
Python 3.10 or newer is recommended.
conda create -n HARDE python=3.10
conda activate HARDE
pip install -r requirements.txt
cp .env.example .env
cp model_config.example.yaml model_config.yaml
# Edit .env, then export its values into the current shell.
set -a
source .env
set +a
Third-party benchmark code and datasets are not vendored. Clone the benchmark(s) you need into the repository root using the exact directory names below:
git clone https://github.com/leolee99/AgentDyn.git AgentDyn
git clone https://github.com/jkutaso/SHADE-Arena.git SHADE-Arena
git clone https://github.com/thu-coai/Agent-SafetyBench.git Agent-SafetyBench
Additional data sources:
The adapters resolve these repositories relative to the HARDE root. Expected layout:
HARDE/
├── AgentDyn/
├── SHADE-Arena/
├── Agent-SafetyBench/
├── domains/
└── personalized_harness/
Detailed benchmark-specific commands and options are documented in the README inside each corresponding directory under domains/. Fixed-harness evaluation examples are documented in personalized_harness/README.md.
The following SHADE-Arena examples show the complete execution flow for each method.
Meta-Harness directly optimizes the complete harness for three iterations and then evaluates the selected harness on the held-out test set:
python domains/shade_arena/meta_harness.py \
--run-name shade_meta_seed0 \
--fresh \
--iterations 3 \
--seed 0 \
--run-final-test
Results are written to domains/shade_arena/runs/shade_meta_seed0/.
This baseline evaluates the unmodified benchmark harness, proposes one personalized harness in a single LLM call, and evaluates that proposal on the same SHADE-Arena tasks. It does not perform iterative search:
python personalized_harness/shade_v2_pipeline.py \
--model_config model_config.yaml \
--tasks spam_filter_update \
--repeat 1 \
--output_harness personalized_harness/SHADE_v2/harnesses/one_shot_proposal.py \
--output_dir personalized_harness/SHADE_v2/output_one_shot
With no --skip_generate, --skip_base, or --skip_personalized flags, this command runs the full chain:
base-harness evaluation → one-shot harness proposal → proposed-harness evaluation → summary
HARDE first probes individual harness components in Stage I:
python domains/shade_arena_component_probe/component_probe.py \
--run-name shade_probe_seed0 \
--seed 0
Stage I writes the generated harness guide to:
domains/shade_arena_component_probe/runs/shade_probe_seed0/experience/module_rubric.md
Stage II consumes that guide, adaptively selects a component at each iteration, runs three search iterations, and evaluates the final selected harness:
python domains/shade_arena_adaptive_component_search/adaptive_component_search.py \
--run-name shade_adaptive_seed0 \
--seed 0 \
--iterations 3 \
--rubric domains/shade_arena_component_probe/runs/shade_probe_seed0/experience/module_rubric.md \
--initial-components initial_components/full_trajectory_monitor \
--run-final-test
Citation metadata will be added with the paper release.