Skip to content
KitploitKITPLOIT
OutilsExploitsBlog
Log in
Soumettre
OutilsExploitsBlog
Soumettre

Outils de Hacking, PenTest et Cybersécurité pour votre Arsenal de Sécurité !

Kitploit est un répertoire d'outils de hacking, de cybersécurité et de pentesting. Découvrez les dernières mises à jour des projets pour trouver des vulnérabilités, analyser des systèmes, automatiser les tests et renforcer votre sécurité.

··Flux·Contact·Confidentialité·© 2026 Kitploit

Répertoire d'outils

Catégories

Voir toutes les catégories
Loading categories
claude-awm — Evades LLM text watermarks by injecting Unicode variation selectors; includes the SynthID generator, mean-g detector, normalization defenses, and entropy experiments. | Kitploit
Outils/GitHubGitHub/aloshdenny/claude-awm
SteganographyPrivacyMachine LearningAI SecurityAdversarial Attack
GitHubaloshdenny/claude-awm

claude-awm

Evades LLM text watermarks by injecting Unicode variation selectors; includes the SynthID generator, mean-g detector, normalization defenses, and entropy experiments.

Voir le dépôt
61627il y a 26 joursPas encore vérifié

Populaires

Voir tout →

Découvrez les outils les plus utilisés par notre communauté.

Explorer tous les outils

Parcourez notre collection d'outils

Voir tous les outils →
Partager
Site web
Contenu non disponible dans la langue demandée. Affichage de la version anglaise.

claude-awm: can you scrub a SynthID text watermark by editing the text?

the dream, allegedly

Yes, but only one family of attack works, and it isn't the one everyone assumes.

Unicode variation selectors (category Mn, U+FE00 to U+FE0F and U+E0100 to U+E01EF) drive the detector below threshold and stay there. Every other invisible-character attack I tried gets fully reverted by one line of input normalization. Variation selectors don't, because they're meaningful codepoints (emoji presentation, CJK variants) that NFKC will not and should not fold away.

Replicated across models and domains:

modeldomainbaseline zafter vs16_30edit rate
gpt-oss-20bprose45.030.7257%
gpt-oss-20bcode37.240.6858%
Qwen3.8-27Bprose35.50-0.6757%

Threshold is z = 2.33. All three land under it after the attack, and stay under after normalization (0.09, 0.45, -0.78 respectively). Text renders identically to a human reader.

The second real finding needs no attack at all: low-entropy text is barely watermarked to begin with. With thinking disabled, Qwen3.8-27B's pure code output scores z = 4.31 unattacked, close to the 2.33 threshold with nothing done to it. Turn thinking back on and the same model, same domain, same length reaches z = 25.53, because the reasoning preamble is ordinary prose and carries the mark normally.

models tested

Everything in this repo was measured on these, nothing else:

modelparamswherewhat was run
Qwen3.5-0.8B0.8BMac (MPS)full attack ladder, 1k-8k
Qwen3.5-4B4BMac / 4090full attack ladder + the entropy experiment + stego raw/normalized
DeepSeek-R1-Distill-Qwen-14B14BMac (MLX, 4-bit)watermark + detect, reasoning traces
gpt-oss-20b20B3090 / 4090 / RunPodfull attack ladder to 32k, prose + code, variation selectors, insertion-rate scaling law to 131k
Qwen3.8-27B27BRunPod H100 / Modal H100prose + code, variation selectors, thinking-mode contrast
gpt-oss-120b120BModal H100 (MXFP4) / RunPod H100prose + code + reasoning at 2k / 8k / 32k, partial insertion-rate scaling to 65k
DeepSeek-V4-Flash284B MoERunPod 2x H200 (FP8)prose + code at 2k / 8k

Not every experiment ran on every model, and the tables below say which ran where. Larger models were added as the interesting questions narrowed, so the attack ladder is broad on the small ones and the length and domain work is concentrated on the large ones.

what this is

Anthropic (and Google DeepMind before them, the SynthID-Text paper) watermark generated text by biasing token sampling with a keyed tournament. The signal lives in which tokens got picked, not in any hidden character. I wanted to know how fragile that is to a motivated person with a find-and-replace, so I built the generator plus an untrained mean-g detector and threw attacks at it.

Detector threshold is z = 2.33 (1% false positive rate). Above that = watermark detected. I generated watermarked text, attacked it, re-tokenized, and scored. Every results table carries a roundtrip control (unattacked watermarked text) so you can see the detector is actually working in that row. Twice this study a bug made every z collapse to zero, and the control is the only thing that caught it.

scope of the attack: what actually works

The attack surface splits cleanly by Unicode category, which turned out to be the whole story.

Works (category Mn, nonspacing marks):

attackwhat it doesedit ratesurvives normalization?
vs16_30variation selector after ~30% of chars57%yes
vs16variation selector after ~10% of chars23%yes (z 3.46)
vs_suppsupplementary-plane selectors (U+E0100+)24%yes (z 3.40)
homoglyphCyrillic а for Latin a (category Ll)9%yes, but weak effect

Doesn't work (category Cf, format chars):

attackraw znormalized zverdict
zwsp_30-0.0935.68fully reverted
combo0.9235.68fully reverted
bidi24.3735.68fully reverted
nbsp41.0446.51barely moves it

Cf characters are pure formatting with no semantic content, so a normalizer can strip them wholesale and lose nothing. Mn variation selectors carry real meaning in emoji and CJK text, so blanket-stripping them would corrupt legitimate documents. That asymmetry is why one family is defensible and the other isn't.

The practical scope: this needs roughly 20% or more of characters touched to push a strong baseline under threshold, but the inserted characters are invisible, so the text reads normally at any edit rate. It is verified visually identical, see src/fidelity_check.py. Four families insert only characters that render nothing; homoglyphs are the honest exception (font-dependent, not guaranteed invisible).

what doesn't work: the obvious stuff

My starting fear was the boring one: that em-dashes and "and" to "&" and stripping markdown would just work, and the watermark would turn out to be security theater. A lot of people online assume exactly that.

It doesn't work. The whole surface-edit ladder on gpt-oss-20b, out to 32k tokens:

attackedit ratez @ 1kz @ 32k
roundtrip (control)0%25.7104.3
em-dash to hyphen~0%26.4113.7
strip all markdown13.6%27.2103.3
AmE to BrE + abbreviations1.3%~28~100
delete 40% of every word38%4.925.4

Across the three models this ladder ran on (0.8B, 4B, 20B = 294 cells), exactly one cell crossed threshold: deleting 39% of every word at the shortest length on the 4B. That's not an attack, that's shredding the document.

Total measured cells across the whole study: 440 (294 from the attack ladder above, plus 146 from the scaling-law sweep below extending prose to 131072 tokens).

Two things surprised me:

  • Edit count doesn't predict damage, edit geometry does. Stripping all markdown (13.6% of tokens) did nothing; it even scored slightly above baseline. Injecting stray spaces at 1.6% did 25x more damage per edit. Markdown markers cluster, so their corruption windows overlap and the long prose runs between them keep replaying the watermark seed intact. Scattered edits that desync the tokenizer hit fresh windows every time.
  • Length helps the detector, not the attacker. z grows like sqrt(tokens). "Fool it over a long context" is backwards, 32k is the hardest case to attack, not the easiest.

Full mechanism and per-attack tables in docs/FINDINGS.md.

the detector wins by waiting

watermark strength vs context

Watermark confidence grows with context while per-token signal stays flat, so a longer document is harder to attack, not easier. Prose carries the strongest mark at every length; DeepSeek-V4-Flash (dashed) reproduces the same climb on a 284B MoE.

attack cost vs context

And the attack has to keep up. Each cell is how many of 8 random insertion seeds beat the detector. 10% insertion clears 1k tokens but fails completely by 4k. The required insertion rate rises with context length, which is why the attack is not context-agnostic.

Télécharger l’outil