
Noisegate: a differential privacy gateway that lets an untrusted LLM agent query sensitive data over MCP (Model Context Protocol), with a formal guarantee no individual's record can leak even if the agent is adversarial - enforcement lives in trusted code below the model, validated by a runnable attack gallery.
An AI agent that can study sensitive data without being able to single anyone out. The limit is mathematical, it's enforced in code the agent can't reach, and the repo ships the attacks that try to break it.
A recorded Claude Desktop session (replies trimmed; the chart cards are the session's own). An AI agent breaks 20 patients down by diagnosis, and the ±12 noise swamps every bin. Reminded that it cannot turn the noise off, it drains a three-answer budget until the gate returns a refusal instead of a quieter answer. On the 32,561-row census, a too-narrow slice is rejected at the trust boundary, while a full education breakdown comes back clean at scale. The refusal and the rejection are the live gateway's real enforcement, reproduced by python scripts/render_demo_gif.py.
You ask questions about a sensitive dataset in plain English. An LLM compiles each question into a small, constrained query. A differential-privacy engine executes it under a tracked privacy budget and returns a deliberately noisy answer with a stated confidence interval. Like its audio namesake, the gateway keeps every signal below a set threshold under the noise floor: any one individual's contribution is drowned out, while signal at the scale of the whole dataset passes through nearly untouched.
The interesting part isn't that an LLM can write queries. It's that the privacy guarantee does not depend on the LLM being trustworthy. The model is a convenience that proposes a query. It enforces nothing. Every privacy property is enforced downstream, by components that would behave the same way if a human typed the query by hand. This is the trust-boundary discipline you'd apply to any untrusted input in a production system, applied here to an AI agent.
The attack gallery runs in-process against the real DP engine:
pip install -e .
python -m attacks.patients_alice # re-identify Alice with privacy off, then watch
# the guard, the noise, and the budget defeat it
The gateway runs as an MCP stdio server for Claude Desktop. The agent becomes the untrusted query author, and it gets only structured tools (count, sum, average, histogram, get_budget) whose argument schemas are generated from the dataset policy. No API key is needed anywhere, because the connecting agent is the intelligence.
A local single-tenant demo. An API key is needed only for the untrusted NL→query compiler:
export ANTHROPIC_API_KEY=... # used only by the untrusted NL→query compiler
docker compose up # brings up the engine, API, and UI
# open http://localhost:8501
That HTTP + Streamlit surface is a local, single-tenant demo. Identity comes from a spoofable X-Identity header, so it is meant for one trusted operator on their own machine, not a public deployment (see what these surfaces are, and are not). For local non-Docker setup, running the tests, and configuration knobs, see SETUP.md.
Anyone can claim privacy. This repository ships the exploits that would break the claim, runs them against its own engine, and pins the outcomes in CI. The fastest way to understand what the gateway guarantees is to watch it defeat three classic attacks that break naive "query a database" systems.
A differencing attack isolates one person by asking two aggregate questions that differ by exactly that person.
Query A: "Total income of all 100 people in department X." → $7,240,000
Query B: "Total income of all people in department X except Alice." → $7,135,000
Attacker computes: A − B = $105,000 ← Alice's exact salary, leaked.
Both queries are "just aggregates." Neither names a single row. Yet together they expose an individual. The gallery (attacks/differencing.py) shows this attack succeeding with privacy disabled: the target's private value is recovered exactly. (The salary sketch above is illustrative; on the real UCI Adult data, "Alice" is the unique holder of her group's maximum capital gain.) It then shows the same attack defeated once DP is on: the calibrated noise on each answer makes the subtraction useless, and the budget accountant charges for the information released across both queries rather than treating them as independent.
A membership-inference attack determines whether a specific individual is in the dataset at all. For many datasets (a medical study, a list of defaulters), that fact is itself sensitive. An attacker with only query access tries to decide: "is this exact person in the data?"
The gallery runs this attack across a sweep of privacy budgets (ε, the dial that trades answer accuracy for privacy) and plots the result:
