Skip to content
KitploitKITPLOIT
ToolsExploitsBlog
Log in
Submit
ToolsExploitsBlog
Submit

Hacking, PenTest, and Cybersecurity Tools for Your Security Arsenal!

Kitploit is a directory of hacking, cybersecurity, and pentesting tools. Discover the latest project updates to find vulnerabilities, analyze systems, automate testing, and strengthen your security.

··Feeds·Contact·Privacy·© 2026 Kitploit

Tool Directory

Categories

View all categories
Loading categories
llm-differential-privacy-gateway — Noisegate: a differential privacy gateway that lets an untrusted LLM agent query sensitive data over MCP (Model Context Protocol), with a formal guarantee no individual's record can leak even if the agent is adversarial - enforcement lives in trusted code below the model, validated by a runnable attack gallery. | Kitploit
Tools/GitHubGitHub/yashmahajan10/llm-differential-privacy-gateway
Defensive ToolsPrivacyLearning & EducationAI SecurityLabs & Practice
GitHubyashmahajan10/llm-differential-privacy-gateway

llm-differential-privacy-gateway

View Repository

Most Popular

View all →

Discover the most used tools by our community.

Explore all tools

Browse our collection of tools

View all tools →
253161 month agoNot yet reviewed

About

Noisegate: a differential privacy gateway that lets an untrusted LLM agent query sensitive data over MCP (Model Context Protocol), with a formal guarantee no individual's record can leak even if the agent is adversarial - enforcement lives in trusted code below the model, validated by a runnable attack gallery.

Share

Noisegate: a differential-privacy gateway for untrusted AI agents

CI Python 3.13+ License: Apache-2.0

An AI agent that can study sensitive data without being able to single anyone out. The limit is mathematical, it's enforced in code the agent can't reach, and the repo ships the attacks that try to break it.

A Claude Desktop chat against the gateway: an AI agent gets honest noise-injected charts, is told the model itself cannot disable the noise, drains a tiny privacy budget until the gate refuses, has a too-narrow census query rejected at the trust boundary, and ends on a clean 16-bar histogram usable at scale

A recorded Claude Desktop session (replies trimmed; the chart cards are the session's own). An AI agent breaks 20 patients down by diagnosis, and the ±12 noise swamps every bin. Reminded that it cannot turn the noise off, it drains a three-answer budget until the gate returns a refusal instead of a quieter answer. On the 32,561-row census, a too-narrow slice is rejected at the trust boundary, while a full education breakdown comes back clean at scale. The refusal and the rejection are the live gateway's real enforcement, reproduced by python scripts/render_demo_gif.py.

At a glance

  • Runnable attacks, pinned in CI. Three classic privacy attacks run against the system's own engine: differencing, membership inference, and singling out by re-identification. Each one is shown succeeding with privacy off, defeated with privacy on, and checked on every build so the defense cannot quietly rot.
  • Independently verified math. The noise mechanism is built from scratch, and it matches OpenDP, the industry reference implementation, in all 35 noise-scale checks to within 1e-9.
  • More questions from tighter accounting. At the deployment's per-query ε, hybrid zCDP composition admits 308 queries against the same budget, against 268 under advanced composition and 100 under naive sum-of-ε. The guarantee is (ε, δ)-DP rather than pure ε.
  • Built for AI agents. Runs as an MCP server for Claude Desktop. The connecting agent is untrusted by design, and every privacy property is enforced below it.
  • Stack: Python · DuckDB · FastAPI · Streamlit · MCP SDK · Docker · GitHub Actions, with a 250+ test suite in CI.

You ask questions about a sensitive dataset in plain English. An LLM compiles each question into a small, constrained query. A differential-privacy engine executes it under a tracked privacy budget and returns a deliberately noisy answer with a stated confidence interval. Like its audio namesake, the gateway keeps every signal below a set threshold under the noise floor: any one individual's contribution is drowned out, while signal at the scale of the whole dataset passes through nearly untouched.

The interesting part isn't that an LLM can write queries. It's that the privacy guarantee does not depend on the LLM being trustworthy. The model is a convenience that proposes a query. It enforces nothing. Every privacy property is enforced downstream, by components that would behave the same way if a human typed the query by hand. This is the trust-boundary discipline you'd apply to any untrusted input in a production system, applied here to an AI agent.


Quickstart

1. Run the attacks: no API key, no data fetch, no server

The attack gallery runs in-process against the real DP engine:

pip install -e .
python -m attacks.patients_alice   # re-identify Alice with privacy off, then watch
                                   # the guard, the noise, and the budget defeat it

2. Connect an AI agent

The gateway runs as an MCP stdio server for Claude Desktop. The agent becomes the untrusted query author, and it gets only structured tools (count, sum, average, histogram, get_budget) whose argument schemas are generated from the dataset policy. No API key is needed anywhere, because the connecting agent is the intelligence.

Setup and full walkthrough →

3. The full natural-language UI

A local single-tenant demo. An API key is needed only for the untrusted NL→query compiler:

export ANTHROPIC_API_KEY=...   # used only by the untrusted NL→query compiler
docker compose up              # brings up the engine, API, and UI
# open http://localhost:8501

That HTTP + Streamlit surface is a local, single-tenant demo. Identity comes from a spoofable X-Identity header, so it is meant for one trusted operator on their own machine, not a public deployment (see what these surfaces are, and are not). For local non-Docker setup, running the tests, and configuration knobs, see SETUP.md.


The headline: an attack gallery

Anyone can claim privacy. This repository ships the exploits that would break the claim, runs them against its own engine, and pins the outcomes in CI. The fastest way to understand what the gateway guarantees is to watch it defeat three classic attacks that break naive "query a database" systems.

Attack 1: The differencing attack

A differencing attack isolates one person by asking two aggregate questions that differ by exactly that person.

Query A: "Total income of all 100 people in department X."        → $7,240,000
Query B: "Total income of all people in department X except Alice." → $7,135,000
Attacker computes: A − B = $105,000  ← Alice's exact salary, leaked.

Both queries are "just aggregates." Neither names a single row. Yet together they expose an individual. The gallery (attacks/differencing.py) shows this attack succeeding with privacy disabled: the target's private value is recovered exactly. (The salary sketch above is illustrative; on the real UCI Adult data, "Alice" is the unique holder of her group's maximum capital gain.) It then shows the same attack defeated once DP is on: the calibrated noise on each answer makes the subtraction useless, and the budget accountant charges for the information released across both queries rather than treating them as independent.

Attack 2: Membership inference

A membership-inference attack determines whether a specific individual is in the dataset at all. For many datasets (a medical study, a list of defaulters), that fact is itself sensitive. An attacker with only query access tries to decide: "is this exact person in the data?"

The gallery runs this attack across a sweep of privacy budgets (ε, the dial that trades answer accuracy for privacy) and plots the result:

Membership-inference success rate vs ε, with a utility curve overlaid

Download Tool