
Embeds and detects keyed, style-based watermarks in LLM-generated Python, Java, and C++ code via CST rewriting, preserving functional correctness and resisting code-editing attacks.
| Feature | Token-level watermarks (KGW, SWEET, Unigram, STONE, STA-1) | Post-hoc watermarks (ACW, SrcMarker, RoSeMary) | ✨ SEW |
|---|---|---|---|
| Access | Decoding-time (biases token selection) | Post-hoc (rewrites finished code) | Post-hoc (rewrites finished code, model-agnostic) |
| Functional correctness | Changes the program (detectability–correctness trade-off) | Can break programs (Java/C++ pass@1 ≈ 12% for neural rewriting) | Preserved (pass@1 equal to unwatermarked code) |
| Predictability | — | Recurring patterns (recovered from 10 watermarked programs) | Key- and context-dependent (choices vary with each program's structure) |
| Detection (TPR@FPR5%) | 14–60% | 31–98% | ⚡ 98.7–99.5% |
Watermarking LLM-generated code supports provenance tracking. Watermarks that modify token selection during generation trade detectability against functional correctness, and they require control over the generating model. Post-hoc methods instead watermark completed code with predefined transformations or trained neural models, but their recurring patterns make the watermark predictable across programs, and patterns that are already common in unwatermarked code are counted as watermark evidence, which causes false detections. SEW asks: can the code style of an already generated program carry a watermark that is correct by construction, hard to predict, and calibrated against human-written code?
SEW (Style-Encoded Watermarking) embeds and detects watermarks in already generated code through three components:
x += 1 / x = x + 1, range(n) / range(0, n), if (c) s; / if (c) { s; }) are matched on the concrete syntax tree (CST). Which variant a site takes is decided by a secret key and the site's structural context.Detection needs only the suspect code and the key — not the generating model, the original code, or any record of the embedding.
✅ Model-agnostic and correct by construction — SEW only rewrites style sites where both variants have the same semantics, so it works on the output of any model and keeps pass@1 equal to that of the unwatermarked code.
✅ Calibrated evidence — A Poisson-binomial test with style probabilities estimated from human-written code (LeetCode solutions, disjoint from the evaluation data) keeps false detections on human code low.
✅ Robust and hard to infer — SEW keeps its detection under formatting, linting, comment removal and variable renaming, and an adversary who observes watermarked programs cannot recover its style choices the way it recovers those of the post-hoc baselines.
Main result — detection on CodeContests (TPR@FPR5% / AUROC, %), averaged over three LLMs (Qwen3.5-9B, gemma-4-12B-it, gpt-oss-20b).
| Type | Method | Python | Java | C++ |
|---|---|---|---|---|
| Token-level | KGW | 48.33 / 80.56 | 43.30 / 72.31 | 59.31 / 84.54 |
| SWEET | 59.92 / 85.18 | 41.19 / 76.56 | 60.01 / 87.24 | |
| Unigram | 57.45 / 88.97 | 36.54 / 71.37 | 47.05 / 71.13 | |
| STONE | 28.09 / 62.33 | 14.38 / 62.95 | 23.65 / 64.79 | |
| STA-1 | 35.26 / 66.26 | 20.93 / 61.48 | 44.71 / 73.72 | |
| Post-hoc | ACW | 90.10 / 95.05 | – | – |
| SrcMarker | 90.21 / 97.86 | 71.52 / 93.58 | 69.88 / 81.55 | |
| RoSeMary | 97.86 / 97.32 | 31.00 / 95.46 | 81.26 / 89.44 | |
| SEW | 99.49 / 99.64 | 98.70 / 98.99 | 99.22 / 99.38 |
Functional correctness (pass@1 of the watermarked code, %; unwatermarked code: 57.63 / 52.05 / 53.74).
| Method | Python | Java | C++ |
|---|---|---|---|
| ACW | 56.00 | – | – |
| SrcMarker | 56.78 | 11.87 | 11.82 |
| RoSeMary | 56.39 | 11.71 | 11.52 |
| SEW | 57.63 | 52.05 | 53.74 |
Robustness to code-editing attacks (TPR@FPR5%, %, averaged over three LLMs and three languages; ACW: Python only).
| Method | No attack | Formatting | Linting | Comment removal | Renaming |
|---|---|---|---|---|---|
| KGW | 50.31 | 43.31 | 49.99 | 24.48 | 43.26 |
| SWEET | 53.71 | 49.79 | 53.00 | 19.34 | 48.83 |
| ACW | 90.10 | 0.51 | 93.56 | 89.35 | 3.17 |
| SrcMarker | 77.20 | 77.23 | 76.42 | 77.20 | 26.43 |
| RoSeMary | 70.04 | 67.00 | 69.42 | 70.04 | 22.26 |
| SEW | 99.14 | 95.57 | 98.99 | 99.14 | 99.14 |
Our experiments on CodeContests with three LLMs and three programming languages show: