
Fine-tuning framework that applies first-order optimal safety calibration and periodic recalibration to LLMs, preserving safety-compatible updates while improving downstream task utility.
First-Order Optimal Fine-Tuning with Recalibration for Safety–Utility Co-Enhancement
ASCENT derives a first-order optimal safety calibration update and the corresponding safety-related structure, optimizes downstream task updates to preserve safety-compatible components while suppressing safety-degrading ones, and periodically recalibrates this structure during fine-tuning to jointly improve safety and downstream utility.
Method · Quick start · Data · Configuration · Evaluation

config.toml.Requirements: Python 3.12 and local model checkpoints. Training uses two CUDA GPUs, one for the target model and one for Llama Guard; each model must fit on its GPU.
pip install -r requirements.txt
cp config.example.toml config.toml
Fill the empty values in config.toml, following its inline comments, then run:
python run.py \
--config config.toml \
--model-path /path/to/target-model \
--guard-model-path /path/to/Llama-Guard-3-8B \
--data-root data \
--output-dir /path/to/new-run \
--target-gpu 0 --guard-gpu 1
Add
--executeto train, generate responses, or save evaluation results.
Choose a new output directory outside the repository with space for checkpoints. Training prints the final merged-model path when it finishes.
Provide JSON arrays with the configured record counts:
<data-root>/calibration/prompts.json<data-root>/<task>/{train,test}.json, with disjoint train/test inputs.| Dataset | Required fields |
|---|---|
| SAMSum | dialogue, summary |
| AGNews | text, label_name: World, Sports, Business, or Sci/Tech |
| GSM8K | question, answer with a #### final answer |
| OpenBookQA | question_stem, choice_labels: ["A","B","C","D"], four choice_texts, answer_key: A–D |
| HarmBench | goal only; optional id, source; no stored responses |
Set your model, task, and hyperparameters in config.toml. Model loading
and matrix selection follow model.key.
Use evaluate.py for both modes; generation requires a fully merged local model.
Supply your own data and external safety scores. No datasets, online judge, or
API configuration are bundled.
python evaluate.py utility --task gsm8k --data /path/to/test.json \
--model-path /path/to/merged-model --output-dir /path/to/task-evaluation --execute
Tasks: samsum, agnews, gsm8k, openbookqa. Metrics are ROUGE-L for SAMSum
and accuracy/exact match for the others, reported as percentages. To score saved
responses, replace --model-path with --responses /path/to/responses.json.
Provide fixed prompts as records with id, goal, and optional prompt (defaults
to goal). Keep the original harmful goal separate from the attack prompt.
Generate responses:
python evaluate.py safety --data /path/to/prompts.json \
--model-path /path/to/merged-model --output-dir /path/to/safety-responses --execute
Aggregate external scores:
python evaluate.py safety \
--responses /path/to/safety-responses/responses.json --judgments /path/to/scores.json \
--output-dir /path/to/safety-metrics --execute
Scores are JSON records { "id": "...", "score": 1 }, using the response IDs and
scores 1–5 (null for failed judgments). Scores 4–5 count as successful attacks;
use --success-threshold 5 to
count only 5. Missing or failed judgments do not yield a final ASR.
For either mode, inputs exceeding --max-input-tokens after chat formatting are
rejected, not truncated. Set the limit within the model's context capacity,
leaving room for generated tokens.