Skip to content
KitploitKITPLOIT
उपकरणएक्सप्लॉइटब्लॉग
Log in
जमा करें
उपकरणएक्सप्लॉइटब्लॉग
जमा करें

हैकिंग, पेनटेस्ट और साइबर सुरक्षा उपकरण आपके सुरक्षा शस्त्रागार के लिए!

Kitploit हैकिंग, साइबर सुरक्षा और पेंटेस्टिंग टूल्स की एक निर्देशिका है। कमजोरियों को खोजने, सिस्टम का विश्लेषण करने, परीक्षण को स्वचालित करने और अपनी सुरक्षा को मजबूत करने के लिए नवीनतम प्रोजेक्ट अपडेट खोजें।

फ़ीडसंपर्कगोपनीयता© 2026 Kitploit

टूल निर्देशिका

श्रेणियाँ

सभी श्रेणियाँ देखें
Loading categories
lm-fingerprint — Fingerprint OpenAI-compatible LLMs from tokenizer and behavior signals. | Kitploit
उपकरण/GitLabGitLab/wattocyber/lm-fingerprint
ReconnaissanceInformation GatheringPenetration TestingRed TeamingAPI SecurityAI Security
GitLabwattocyber/lm-fingerprint

lm-fingerprint

Fingerprint OpenAI-compatible LLMs from tokenizer and behavior signals.

रिपॉजिटरी देखें
261 महीना पहलेअभी तक समीक्षित नहीं

सबसे लोकप्रिय

सभी देखें →

हमारे समुदाय द्वारा सबसे अधिक उपयोग किए जाने वाले उपकरण खोजें।

सभी उपकरण खोजें

हमारे उपकरणों का संग्रह ब्राउज़ करें

सभी उपकरण देखें →
साझा करें
वेबसाइट
अनुरोधित भाषा में सामग्री उपलब्ध नहीं है। अंग्रेज़ी संस्करण दिखाया जा रहा है।

LM-Fingerprint

LM-Fingerprint

Fingerprints the serving stack behind an OpenAI-compatible chat endpoint.

It records tokenizer accounting, hidden template offset, validation prose, JSON/SSE shape, role acceptance, named limits, and gateway headers. It does not load weights, store API keys, or ask the model who it is.

Importlm_fingerprint
CLIlm-fingerprint
WhatVersioned infrastructure probes against /chat/completions. Offline synthetic library. Local Hugging Face tokenizer verify (no weights).
What it is notNot a gradient attack. Not a jailbreak. Not encode-map identity (tokprint). Not “which binary is on this port” (Julius).

python license repro

Authorized use only: endpoints you own, labs you control, in-scope bounties, written pentests. See SECURITY.md. --api-key-env names an environment variable. The key is never logged.

Contents

  1. Problem
  2. What a fingerprint is
  3. Architecture
  4. Threat model
  5. Install and the repro gate
  6. First run
  7. Against a live endpoint
  8. Local Hugging Face tokenizers
  9. CLI
  10. Library
  11. Reading a result JSON
  12. Measured vs ModelPrint
  13. Honesty rules
  14. Nearby tools
  15. License

Problem

OpenAI-compatible chat is a common wire. vLLM, SGLang, a lab gateway, and a reseller can all speak /chat/completions. The model field is a label. The host can swap the backend, wrap the tokenizer, or sit a router in front. Personality probes move when the system prompt moves.

What stays put is plumbing: how the server counts tokens, how many hidden template tokens it adds, the prose its validator wrote, the SSE dialect, which roles it accepts, the number it names when max_tokens is absurd, the Server / provider headers.

That is the question this tool answers: what stack is handling this chat path? Not which checkpoint, not which process is bound to the port, not who wrote a completion.

ModelPrint asked the same question in a browser with nine probes. This package is the CLI and library form: versioned probes, enrollable library, pairwise evidence, drift, and a local tokenizer path that never calls a hosted Inference API.

What a fingerprint is

A fingerprint is a versioned JSON record of probe values. Each probe must return the same value when the same stack is hit twice. Timestamps, request ids, latency, and keys are stripped.

Tokenizer counts reuse ModelPrint 1.0.0 texts (MIT, unclecode/modelprint), byte-exact, minus a one-character "a" baseline so a hidden template cancels. Those four numbers are a projection of the encode map (T). Equal counts do not mean (T_1 \equiv T_2). GPT-2 and OPT collide on them; the rust pipeline hash does not.

hf_identity (local HF path only) is SHA-256 of that pipeline commitment plus the chat-template hash. Same (T) and the same template match (tiny-gpt2 / gpt2). That is family + template, not a checkpoint.

Architecture

Layout:

src/lm_fingerprint/
 cli.py argparse. --api-key-env only. No --api-key.
 client.py HttpTransport, InProcessTransport, redact_secrets
 probes.py 38 ProbeSpec rows (tokenizer / errors / shape / roles / limits / proxy / surface)
 engine.py run suite → Fingerprint
 schema.py schema_version=1, probe_suite_version=lm-fingerprint-probes-v2
 similarity.py weighted exact-match, coverage, confidence, library rank, drift
 align.py constant-shift MAE + residual on the 15-script suite
 catalog.py offline supported_parameters twins
 library.py enroll / match. Packaged rows are synthetic fixtures
 synthetic.py in-process stacks for tests and the shipped library
 hf.py local AutoTokenizer + named validation profile. No weights
 localmap.py encode-pipeline-v2 commitment (rust tokenizer.json)
 proof.py measured ModelPrint 9-probe identity
 report.py text reports
 data/library.json

Two paths into the same probe suite:

live endpoint local Hub tokenizer (optional [hf])
 | |
 HttpTransport.chat HfChatTransport
 POST /chat/completions encode / chat_template for counts
 | synthetic profile for errors/shape
 +------------------+----------------------+
 |
 engine.run_probes
 |
 Fingerprint (JSON, keyless)
 |
 compare | match | drift | enroll

Probe authoring: docs/PROBES.md. Module map: docs/ARCHITECTURE.md.

Threat model

You holdWhat you can observeWhat this tool does
An authorized OpenAI-compat base URL + keyChat HTTP: usage, errors, headers, SSEfingerprint / compare / match / drift
Local tokenizer filesencode and chat_templatefingerprint-hf / verify-hf. Weights stay on disk
Nothing but a host:portBanner / /health / /api/tagsOut of scope. Use Julius

A match is shared infrastructure. One lab can serve two checkpoints on one stack. A router can answer with its own error wrapper. Count agreement is not identity of (T).

Install and the repro gate

Python 3.10+. Core runtime is the standard library.

git clone https://gitlab.com/WattoCyber/lm-fingerprint.git
cd lm-fingerprint
pip install -e ".[dev]"
PYTHONPATH=src python3 scripts/repro.py

Expect REPRO_OK. That gate is offline: no GPU, no production APIs. Local HF tests skip if the tokenizer is not cached.

pip install -e ".[hf]" # transformers, for fingerprint-hf / verify-hf

First run

lm-fingerprint list-probes
lm-fingerprint list

list prints the packaged stacks: openai-chat-strict, glm-compat, router-openai, permissive-compat. Those are fixtures, not scraped providers. Compare two of them with no network:

from lm_fingerprint.client import InProcessTransport
from lm_fingerprint.engine import fingerprint_endpoint
from lm_fingerprint.similarity import compare_fingerprints
from lm_fingerprint.synthetic import SYNTHETIC_STACKS, handler_for

def fp(stack_id):
 t = InProcessTransport(handler_for(stack_id), headers=SYNTHETIC_STACKS[stack_id].headers)
 return fingerprint_endpoint(t, model=f"synthetic/{stack_id}", base_url=f"inprocess://{stack_id}")

print(compare_fingerprints(fp("openai-chat-strict"), fp("router-openai"))["note"])

Same word tokenizer, different errors and headers. Score lands around 0.62. ModelPrint's four counts match; this tool does not call that a shared stack.

Against a live endpoint

lm-fingerprint fingerprint \
 --base-url https://your-lab.example/v1 \
 --model the-label \
 --api-key-env OPENAI_API_KEY \
 --out lab.json

lm-fingerprint match --fingerprint lab.json
lm-fingerprint enroll --fingerprint lab.json --stack-id my-lab --library my-lib.json

Do not pass the key on the command line. Do not scrape production APIs to fill a public library.

Local Hugging Face tokenizers

lm-fingerprint fingerprint-hf --model sshleifer/tiny-gpt2 --local-files-only
lm-fingerprint verify-hf --local-files-only
टूल डाउनलोड करें