
Un framework Pythonic pour la modélisation des menaces
La modélisation traditionnelle des menaces arrive trop souvent en retard, ou parfois pas du tout. De plus, la création manuelle de flux de données et de rapports peut prendre énormément de temps. L'objectif de pytm est de déplacer la modélisation des menaces vers la gauche, rendant la modélisation des menaces plus automatisée et axée sur les développeurs.
En fonction de vos entrées et de la définition de la conception architecturale, pytm peut générer automatiquement les éléments suivants :
tm.py est un exemple de modèle. Vous pouvez l'exécuter pour générer le rapport et les fichiers d'images de diagramme auxquels il fait référence :```
mkdir -p tm
./tm.py --report docs/basic_template.md | pandoc -f markdown -t html > tm/report.html
./tm.py --dfd | dot -Tpng -o tm/dfd.png
./tm.py --seq | java -Djava.awt.headless=true -jar $PLANTUML_PATH -tpng -pipe > tm/seq.png
Il existe aussi un exemple de `Makefile` qui regroupe tout cela en cibles faciles à partager pour plusieurs modèles. Si vous avez [GNU make](https://www.gnu.org/software/make/) installé (disponible par défaut sur les distributions Linux mais pas sur OSX), exécutez simplement :```
make MODEL=the_name_of_your_model_minus_.py
Vous devriez soit avoir plantuml.jar dans le même répertoire que votre modèle, soit définir PLANTUML_PATH.
Pour éviter d'installer toutes les dépendances, comme pandoc ou Java, le script peut être exécuté dans un conteneur :```
export USE_DOCKER=true make image
make
### Pour commencer - Variante Devbox
Pour simplifier l'utilisation de `pytm`, les dépendances hôtes peuvent être complètement isolées à l'aide de [`Devbox`](https://github.com/jetify-com/devbox). C'est généralement une alternative plus légère et plus pratique que l'approche par conteneur OCI.
- Installez Devbox sur Linux/MacOS : `curl -fsSL https://get.jetify.com/devbox | bash`
- Installez Devbox sur [Windows/WSL](https://www.jetify.com/docs/devbox/installing-devbox/index#installing-wsl2)
- Mettez à jour vers la dernière version de devbox : `devbox version update`
- Définissez votre jeton d'accès GitHub dans le fichier `~/.config/nix/nix.conf` : `access-tokens = github.com=YOUR_TOKEN_HERE`
- Créez un nouvel environnement shell isolé qui inclut tous les outils et paquets spécifiés dans le fichier `devbox.json` du projet : `devbox shell`
- Affichez le chemin complet de l'exécutable Python qui sera utilisé lorsque vous tapez simplement `python` dans votre terminal à l'aide de la commande which python. La sortie doit être le chemin suivant : `.devbox/nix/profile/default/bin/python`
- Testez en exécutant la commande suivante, qui doit générer un DFD sous forme de fichier PNG nommé `sample.png` : `./tm.py --dfd | dot -Tpng -o sample.png`
- Quittez l'environnement shell Devbox : `exit`
## Utilisation
Tous les arguments disponibles :```text
usage: tm.py [-h] [--debug] [--dfd] [--report REPORT] [--exclude EXCLUDE]
[--seq] [--list] [--colormap] [--describe DESCRIBE]
[--list-elements] [--json JSON] [--levels LEVELS [LEVELS ...]]
[--stale_days STALE_DAYS]
options:
-h, --help show this help message and exit
--debug print debug messages
--dfd output DFD
--report REPORT output report using the named template file (sample
template file is under docs/template.md)
--exclude EXCLUDE specify threat IDs to be ignored
--seq output sequential diagram
--list list all available threats
--colormap color the risk in the diagram
--describe DESCRIBE describe the properties available for a given element
--list-elements list all elements which can be part of a threat model
--json JSON output a JSON file
--levels LEVELS [LEVELS ...]
Select levels to be drawn in the threat model (int
separated by comma).
--stale_days STALE_DAYS
checks if the delta between the TM script and the code
described by it is bigger than the specified value in
days
L'argument stale_days tente de déterminer l'écart en jours entre le script du modèle (que vous écrivez) et le code qui implémente le système modélisé. Idéalement, ils devraient être assez proches dans la plupart des cas d'un système activement développé. Vous pouvez exécuter cette vérification périodiquement pour mesurer le pouls de votre projet et la « fraîcheur » de votre modèle de menace.
Les éléments actuellement disponibles sont : TM, Element, Server, ExternalEntity, Datastore, Actor, Process, SetOfProcesses, Dataflow, Boundary, Lambda, LLM et Agent.
Les propriétés disponibles d'un élément peuvent être listées en utilisant --describe suivi du nom d'un élément :```text
$ ./tm.py --describe Server
Server class attributes:
OS Operating system
default: ''
assumptions Assumptions about the element. These optionally allow to exclude threats with the given SIDs
default factory: list
controls Security controls for this element
default factory: Controls
data pytm.Data object(s) in incoming data flows
default factory: DataSet
description Description of the element
default: ''
findings Threats that apply to this element
default factory: list
handlesResources Does this asset handle resources?
default: False
inBoundary Trust boundary this element exists in
default: None
inScope Is the element in scope of the threat model
default: True
inputs incoming Dataflows
default factory: list
is_drawn default: False
levels List of levels (0, 1, 2, ...) to be drawn in the model
default factory:
maxClassification Maximum data classification this element can handle
default: <Classification.UNKNOWN: 0>
minTLSVersion Minimum TLS version required
default: <TLSVersion.NONE: 0>
name Name of the element
required
onAWS Is this asset on AWS?
default: False
outputs outgoing Dataflows
default factory: list
overrides Overrides to findings, allowing to set a custom response, CVSS score or override other attributes
default factory: list
port Default TCP port for incoming data flows
default: -1
protocol Default network protocol for incoming data flows
default: ''
severity Severity level of threats affecting this element
default: 0
sourceFiles Location of the source code that describes this element relative to the directory of the model script
default factory: list
usesCache Does this server use cache?
default: False
usesEnvironmentVariables Does this asset use environment variables?
default: False
usesSessionTokens Does this server use session tokens?
default: False
usesVPN Does this server use VPN?
default: False
usesXMLParser Does this server use XML parser?
default: False
uuid default factory:
L'argument *colormap*, utilisé avec *dfd*, produit un DFD codé en couleurs où les éléments sont peints en rouge, jaune ou vert selon leur niveau de risque (tel qu'identifié par l'exécution des règles).
## Utilisation - Variante Devbox
- `devbox shell`
- utilisation de `pytm` comme d'habitude
- `exit`
## Création d'un modèle de menace
Ce qui suit est un exemple de fichier `tm.py` qui décrit une application simple où un utilisateur se connecte à l'application
et publie des commentaires sur l'application. Le serveur de l'application stocke ces commentaires dans la base de données. Il y a une AWS Lambda
qui nettoie périodiquement la base de données.```python
#!/usr/bin/env python3
from pytm import TM, Server, Datastore, Dataflow, Boundary, Actor, Lambda, LLM, Data, Classification, DatastoreType
tm = TM("my test tm")
tm.description = "another test tm"
tm.isOrdered = True
User_Web = Boundary("User/Web")
Web_DB = Boundary("Web/DB")
user = Actor("User")
user.inBoundary = User_Web
web = Server("Web Server")
web.OS = "CloudOS"
web.controls.isHardened = True
web.sourceFiles = ["server/web.cc"]
db = Datastore("SQL Database (*)")
db.OS = "CentOS"
db.controls.isHardened = False
db.inBoundary = Web_DB
db.type = DatastoreType.SQL
db.inScope = False
db.sourceFiles = ["model/schema.sql"]
comments = Data(
name="Comments",
description="Comments in HTML or Markdown",
classification=Classification.PUBLIC,
isPII=False,
isCredentials=False,
# credentialsLife=Lifetime.LONG,
isStored=True,
isSourceEncryptedAtRest=False,
isDestEncryptedAtRest=True
)
results = Data(
name="results",
description="Results of insert op",
classification=Classification.SENSITIVE,
isPII=False,
isCredentials=False,
# credentialsLife=Lifetime.LONG,
isStored=True,
isSourceEncryptedAtRest=False,
isDestEncryptedAtRest=True
)
my_lambda = Lambda("cleanDBevery6hours")
my_lambda.controls.hasAccessControl = True
my_lambda.inBoundary = Web_DB
llm_api = LLM("AI Writing Assistant")
llm_api.isThirdParty = True
llm_api.processesPersonalData = True
llm_api.hasContentFiltering = False
llm_api.hasSystemPrompt = True
llm_api.processesUntrustedInput = True
my_lambda_to_db = Dataflow(my_lambda, db, "(λ)Periodically cleans DB")
my_lambda_to_db.protocol = "SQL"
my_lambda_to_db.dstPort = 3306
user_to_web = Dataflow(user, web, "User enters comments (*)")
user_to_web.protocol = "HTTP"
user_to_web.dstPort = 80
user_to_web.data = comments
web_to_user = Dataflow(web, user, "Comments saved (*)")
web_to_user.protocol = "HTTP"
web_to_db = Dataflow(web, db, "Insert query with comments")
web_to_db.protocol = "MySQL"
web_to_db.dstPort = 3306
db_to_web = Dataflow(db, web, "Comments contents")
db_to_web.protocol = "MySQL"
db_to_web.data = results
web_to_llm = Dataflow(web, llm_api, "Chat completion request")
web_to_llm.protocol = "HTTPS"
web_to_llm.dstPort = 443
tm.process()
Vous avez également la possibilité d'utiliser pytmGPT pour créer vos modèles à partir de texte en prose !
Les diagrammes sont générés au format Dot et PlantUML.
Lorsque l'argument --dfd est passé au fichier tm.py ci-dessus, il génère une sortie sur stdout, qui est ensuite transmise à Graphviz dot pour générer le diagramme de flux de données :```bash
tm.py --dfd | dot -Tpng -o sample.png
Génère ce diagramme :
dfd.png
L'ajout d'attributs ".levels = [1,2]" à un élément fera en sorte qu'il (et ses Dataflows associés si les deux extrémités de flux sont dans le même niveau DFD) s'affiche (ou non) selon l'argument de commande "--levels 1 2".
La commande suivante génère un diagramme de séquence.```bash
tm.py --seq | java -Djava.awt.headless=true -jar plantuml.jar -tpng -pipe > seq.png
Génère ce diagramme :
seq.png
Les diagrammes et les résultats peuvent être inclus dans le modèle pour créer un rapport final :```bash
tm.py --report docs/basic_template.md | pandoc -f markdown -t html > report.html
Le format de templating utilisé dans le modèle de rapport est très simple :```text
# Threat Model Sample
***
## System Description
{tm.description}
## Dataflow Diagram

## Dataflows
Name|From|To |Data|Protocol|Port
----|----|---|----|--------|----
{dataflows:repeat:{{item.name}}|{{item.source.name}}|{{item.sink.name}}|{{item.data}}|{{item.protocol}}|{{item.dstPort}}
}
## Findings
{findings:repeat:* {{item.description}} on element "{{item.target}}"
}
Pour regrouper les résultats par éléments, utilisez une boucle imbriquée plus avancée :```text
{elements🔁{{item.findings:if:
{{item.findings🔁 Threat: {{{{item.id}}}} - {{{{item.description}}}}
Severity: {{{{item.severity}}}}
Mitigations: {{{{item.mitigations}}}}
References: {{{{item.references}}}}
}}}}}
Tous les éléments d'une boucle doivent être échappés, en doublant les accolades, donc `{item.name}` devient `{{item.name}}`.
L'exemple ci-dessus utilise deux boucles imbriquées, donc les éléments de la boucle interne doivent être échappés deux fois, c'est pourquoi ils utilisent quatre accolades.
### Surcharges
Vous pouvez remplacer les attributs des findings (menaces correspondant aux assets et/ou dataflows du modèle), par exemple pour définir un score CVSS personnalisé et/ou un texte de réponse :```python
user_to_web = Dataflow(user, web, "User enters comments (*)", protocol="HTTP", dstPort="80")
user_to_web.overrides = [
Finding(
# Overflow Buffers
threat_id="INP02",
cvss="9.3",
response="""**To Mitigate**: run a memory sanitizer to validate the binary""",
severity="Very High",
)
]
If you are adding a Finding, make sure to add a severity: "Very High", "High", "Medium", "Low", "Very Low".
Pour le praticien en sécurité, vous pouvez fournir votre propre fichier de menaces en définissant TM.threatsFile. Il doit contenir des entrées comme :```json
{
"SID":"INP01",
"target": ["Lambda","Process"],
"description": "Buffer Overflow via Environment Variables",
"details": "This attack pattern involves causing a buffer overflow through manipulation of environment variables. Once the attacker finds that they can modify an environment variable, they may try to overflow associated buffers. This attack leverages implicit trust often placed in environment variables.",
"Likelihood Of Attack": "High",
"severity": "High",
"condition": "target.usesEnvironmentVariables is True and target.controls.sanitizesInput is False and target.controls.checksInputBounds is False",
"prerequisites": "The application uses environment variables.An environment variable exposed to the user is vulnerable to a buffer overflow.The vulnerable environment variable uses untrusted data.Tainted data used in the environment variables is not properly validated. For instance boundary checking is not done before copying the input data to a buffer.",
"mitigations": "Do not expose environment variable to the user.Do not use untrusted data in your environment variables. Use a language or compiler that performs automatic bounds checking. There are tools such as Sharefuzz [R.10.3] which is an environment variable fuzzer for Unix that support loading a shared library. You can use Sharefuzz to determine if you are exposing an environment variable vulnerable to buffer overflow.",
"example": "Attack Example: Buffer Overflow in $HOME A buffer overflow in sccw allows local users to gain root access via the $HOME environmental variable. Attack Example: Buffer Overflow in TERM A buffer overflow in the rlogin program involves its consumption of the TERM environmental variable.",
"references": "https://capec.mitre.org/data/definitions/10.html, CVE-1999-0906, CVE-1999-0046, http://cwe.mitre.org/data/definitions/120.html, http://cwe.mitre.org/data/definitions/119.html, http://cwe.mitre.org/data/definitions/680.html"
}
Le champ `target` répertorie les classes d'éléments de modèle à comparer avec cette menace.
Il peut s'agir d'actifs, comme : Actor, Datastore, Server, Process, SetOfProcesses, ExternalEntity,
Lambda, LLM, Agent ou Element, qui est la classe de base et correspond à n'importe quel élément. Cela peut également être un Dataflow qui relie deux actifs.
Tous les autres champs (sauf `condition`) sont disponibles pour l'affichage et peuvent être utilisés dans le modèle
pour lister les constats dans le [rapport](#report) final.
> **AVERTISSEMENT**
>
> Le fichier `threats.json` contient des chaînes qui passent par `eval()`. Assurez-vous que le fichier dispose des permissions correctes
> ou vous risquez qu'un attaquant modifie les chaînes et vous fasse exécuter du code en son nom.
La logique se trouve dans `condition`, où les membres de `target` peuvent être évalués logiquement.
Retourner vrai signifie que la règle génère un constat ; sinon, elle n'en génère pas.
La condition peut comparer des attributs de `target` et/ou des attributs de contrôle de 'target.control' et également appeler l'une de ces méthodes :
* `target.oneOf(class, ...)` où `class` est un ou plusieurs : Actor, Datastore, Server, Process, SetOfProcesses, ExternalEntity, Lambda, LLM, Agent ou Dataflow,
* `target.crosses(Boundary)`,
* `target.enters(Boundary)`,
* `target.exits(Boundary)`,
* `target.inside(Boundary)`.
Si `target` est un Dataflow, rappelez-vous que vous pouvez accéder à `target.source` et/ou `target.sink`, ainsi qu'à d'autres attributs.
Les conditions sur les actifs peuvent analyser tous les Dataflows entrants et sortants en inspectant
les attributs `target.input` et `target.output`. Par exemple, pour ne faire correspondre une menace qu'aux
serveurs avec du trafic entrant, utilisez `any(target.inputs)`. Un exemple plus avancé,
qui fait correspondre les éléments se connectant à des datastores SQL, serait `any(f.sink.oneOf(Datastore) and f.sink.type == DatastoreType.SQL for f in target.outputs)`.
## Importation depuis JSON
Avec un peu de code Python, il est possible d'importer un modèle de menace depuis JSON (notez le format spécial dans l'exemple trouvé dans `tests/input.json`). L'exemple suivant importe l'exemple `input.json` trouvé dans les tests. Enregistrez le code suivant sous le nom `tm2.py`.```python
#!/usr/bin/env python3
# Example tm2.py contents
# Run: python tm2.py --dfd | dot -Tpng -o sample_json.png
from pytm import (
TM,
Actor,
Boundary,
Classification,
Data,
Dataflow,
Datastore,
Lambda,
Server,
DatastoreType,
Assumption,
load,
)
json_file_string = './tests/input.json'
with open(json_file_string) as input_json:
TM.reset()
tm = load(input_json)
tm.process()
Nous pouvons appeler tm2.py de la même manière qu'avant, ici avec --dfd et ensuite rediriger la sortie vers Graphviz (dot) :```bash
python tm2.py --dfd | dot -Tpng -o sample_json.png
## Making slides!
Une fois le modèle de menace terminé et prêt, l'étape redoutée de la présentation arrive - et pytm peut maintenant vous y aider aussi, grâce à un modèle qui exprime votre modèle de menace sous forme de diapositives, en utilisant la puissance de (RevealMD)[https://github.com/webpro/reveal-md] ! Il suffit d'utiliser le modèle docs/revealjs.md et vous obtiendrez de jolies diapositives, entièrement configurables, que vous pourrez présenter et partager depuis votre navigateur.
https://github.com/izar/pytm/assets/368769/30218241-c7cc-4085-91e9-bbec2843f838
## Menaces actuellement prises en charge```text
INP01 - Buffer Overflow via Environment Variables
INP02 - Overflow Buffers
INP03 - Server Side Include (SSI) Injection
CR01 - Session Sidejacking
INP04 - HTTP Request Splitting
CR02 - Cross Site Tracing
INP05 - Command Line Execution through SQL Injection
INP06 - SQL Injection through SOAP Parameter Tampering
SC01 - JSON Hijacking (aka JavaScript Hijacking)
LB01 - API Manipulation
AA01 - Authentication Abuse/ByPass
DS01 - Excavation
DE01 - Interception
DE02 - Double Encoding
API01 - Exploit Test APIs
AC01 - Privilege Abuse
INP07 - Buffer Manipulation
AC02 - Shared Data Manipulation
DO01 - Flooding
HA01 - Path Traversal
AC03 - Subverting Environment Variable Values
DO02 - Excessive Allocation
DS02 - Try All Common Switches
INP08 - Format String Injection
INP09 - LDAP Injection
INP10 - Parameter Injection
INP11 - Relative Path Traversal
INP12 - Client-side Injection-induced Buffer Overflow
AC04 - XML Schema Poisoning
DO03 - XML Ping of the Death
AC05 - Content Spoofing
INP13 - Command Delimiters
INP14 - Input Data Manipulation
DE03 - Sniffing Attacks
CR03 - Dictionary-based Password Attack
API02 - Exploit Script-Based APIs
HA02 - White Box Reverse Engineering
DS03 - Footprinting
AC06 - Using Malicious Files
HA03 - Web Application Fingerprinting
SC02 - XSS Targeting Non-Script Elements
AC07 - Exploiting Incorrectly Configured Access Control Security Levels
INP15 - IMAP/SMTP Command Injection
HA04 - Reverse Engineering
SC03 - Embedding Scripts within Scripts
INP16 - PHP Remote File Inclusion
AA02 - Principal Spoof
CR04 - Session Credential Falsification through Forging
DO04 - XML Entity Expansion
DS04 - XSS Targeting Error Pages
SC04 - XSS Using Alternate Syntax
CR05 - Encryption Brute Forcing
AC08 - Manipulate Registry Information
DS05 - Lifting Sensitive Data Embedded in Cache
SC05 - Removing Important Client Functionality
INP17 - XSS Using MIME Type Mismatch
AA03 - Exploitation of Trusted Credentials
AC09 - Functionality Misuse
INP18 - Fuzzing and observing application log data/errors for application mapping
CR06 - Communication Channel Manipulation
AC10 - Exploiting Incorrectly Configured SSL
CR07 - XML Routing Detour Attacks
AA04 - Exploiting Trust in Client
CR08 - Client-Server Protocol Manipulation
INP19 - XML External Entities Blowup
INP20 - iFrame Overlay
AC11 - Session Credential Falsification through Manipulation
INP21 - DTD Injection
INP22 - XML Attribute Blowup
INP23 - File Content Injection
DO05 - XML Nested Payloads
AC12 - Privilege Escalation
AC13 - Hijacking a privileged process
AC14 - Catching exception throw/signal from privileged block
INP24 - Filter Failure through Buffer Overflow
INP25 - Resource Injection
INP26 - Code Injection
INP27 - XSS Targeting HTML Attributes
INP28 - XSS Targeting URI Placeholders
INP29 - XSS Using Doubled Characters
INP30 - XSS Using Invalid Characters
INP31 - Command Injection
INP32 - XML Injection
INP33 - Remote Code Inclusion
INP34 - SOAP Array Overflow
INP35 - Leverage Alternate Encoding
DE04 - Audit Log Manipulation
AC15 - Schema Poisoning
INP36 - HTTP Response Smuggling
INP37 - HTTP Request Smuggling
INP38 - DOM-Based XSS
AC16 - Session Credential Falsification through Prediction
INP39 - Reflected XSS
INP40 - Stored XSS
AC17 - Session Hijacking - ServerSide
AC18 - Session Hijacking - ClientSide
INP41 - Argument Injection
AC19 - Reusing Session IDs (aka Session Replay) - ServerSide
AC20 - Reusing Session IDs (aka Session Replay) - ClientSide
AC21 - Cross Site Request Forgery
DS06 - Data Leak
DR01 - Unprotected Sensitive Data
AC22 - Credentials Aging (deprecated)
AC23 - Credentials Disclosure
AC24 - Use of hardcoded credentials
LLM01 - Direct Prompt Injection
LLM02 - Indirect Prompt Injection via Retrieved Content
LLM03 - Sensitive Data Leakage to Third-Party Provider
LLM04 - Training Data Poisoning
LLM05 - Excessive Agency via Unauthorized Tool Use
LLM06 - Arbitrary Code Execution via LLM Agent
LLM07 - Jailbreaking and Safety Bypass
LLM08 - Sensitive Information Disclosure Through Output
LLM09 - Untrusted Tool Launch Configuration