
従来の脅威モデリングは、しばしば遅れて登場するか、まったく行われないことがあります。さらに、手動でデータフローやレポートを作成するのは非常に時間がかかります。pytmの目標は、脅威モデリングを左にシフトし、より自動化され、開発者中心のものにすることです。
あなたの入力とアーキテクチャ設計の定義に基づいて、pytmは以下の項目を自動生成できます:
tm.py はサンプルモデルです。これを実行すると、参照されているレポートと図の画像ファイルを生成できます。```
mkdir -p tm
./tm.py --report docs/basic_template.md | pandoc -f markdown -t html > tm/report.html
./tm.py --dfd | dot -Tpng -o tm/dfd.png
./tm.py --seq | java -Djava.awt.headless=true -jar $PLANTUML_PATH -tpng -pipe > tm/seq.png
また、これらすべてを複数のモデルで簡単に共有できるターゲットにまとめた `Makefile` の例もあります。[GNU make](https://www.gnu.org/software/make/) がインストールされている場合(Linuxディストリビューションではデフォルトで利用可能ですが、OSXでは利用できません)、以下のように実行してください:```
make MODEL=the_name_of_your_model_minus_.py
モデルと同じディレクトリに plantuml.jar を置くか、PLANTUML_PATH を設定する必要があります。
すべての依存関係(pandoc や Java など)をインストールする手間を省くため、スクリプトはコンテナ内で実行できます:```
export USE_DOCKER=true make image
make
### はじめよう - Devbox バリアント
`pytm` のホスト依存関係を完全に分離するために、[`Devbox`](https://github.com/jetify-com/devbox) を使用すると簡略化できます。これは通常、OCI コンテナアプローチよりもオーバーヘッドが低く、より便利な代替手段です。
- Linux/MacOS に Devbox をインストール: `curl -fsSL https://get.jetify.com/devbox | bash`
- [Windows/WSL](https://www.jetify.com/docs/devbox/installing-devbox/index#installing-wsl2) に Devbox をインストール
- devbox を最新バージョンに更新: `devbox version update`
- `~/.config/nix/nix.conf` ファイルに GitHub アクセストークンを設定: `access-tokens = github.com=YOUR_TOKEN_HERE`
- プロジェクトの `devbox.json` ファイルに指定されたすべてのツールとパッケージを含む、新しい独立したシェル環境を作成: `devbox shell`
- ターミナルで単に `python` と入力したときに使用される Python 実行ファイルのフルパスを `which python` コマンドを使用して表示します。出力は次のパスになるはずです: `.devbox/nix/profile/default/bin/python`
- 次のコマンドを実行してテストします。これにより、`sample.png` という名前の PNG ファイルとして DFD が生成されます: `./tm.py --dfd | dot -Tpng -o sample.png`
- Devbox シェル環境を終了: `exit`
## 使い方
利用可能なすべての引数:```text
usage: tm.py [-h] [--debug] [--dfd] [--report REPORT]
[--exclude EXCLUDE] [--seq] [--list] [--describe DESCRIBE]
[--list-elements] [--json JSON] [--levels LEVELS [LEVELS ...]]
[--stale_days STALE_DAYS]
optional arguments:
-h, --help show this help message and exit
--debug print debug messages
--dfd output DFD
--report REPORT output report using the named template file (sample
template file is under docs/template.md)
--exclude EXCLUDE specify threat IDs to be ignored
--seq output sequential diagram
--list list all available threats
--colormap color the risk in the diagram
--describe DESCRIBE describe the properties available for a given element
--list-elements list all elements which can be part of a threat model
--json JSON output a JSON file
--levels LEVELS [LEVELS ...]
Select levels to be drawn in the threat model (int
separated by comma).
--stale_days STALE_DAYS
checks if the delta between the TM script and the code
described by it is bigger than the specified value in
days
stale_days 引数は、モデルスクリプト(作成中のもの)と、モデル化対象のシステムを実装するコードとの間の日数差を判断しようとします。理想的には、活発に開発されているシステムのほとんどの場合、これらはかなり近い値であるべきです。これを定期的に実行することで、プロジェクトの状態と脅威モデルの「鮮度」を測定できます。
現在利用可能な要素は次のとおりです: TM、Element、Server、ExternalEntity、Datastore、Actor、Process、SetOfProcesses、Dataflow、Boundary、Lambda、LLM、Agent。
要素の利用可能なプロパティは、--describe に続けて要素の名前を指定することで一覧表示できます:```text
(pytm) ➜ pytm git:(master) ✗ ./tm.py --describe Element Element class attributes: OS definesConnectionTimeout default: False description handlesResources default: False implementsAuthenticationScheme default: False implementsNonce default: False inBoundary inScope Is the element in scope of the threat model, default: True isAdmin default: False isHardened default: False name required onAWS default: False
*colormap*引数は、*dfd*と一緒に使用すると、リスクレベル(ルールの実行によって特定される)に応じて要素が赤、黄、緑に着色された色分けされたDFDを出力します。
## Usage - Devbox Variant
- `devbox shell`
- `pytm` usage as usual
- `exit`
## 脅威モデルの作成
以下はサンプル`tm.py`ファイルで、ユーザーがアプリケーションにログインし、アプリにコメントを投稿するシンプルなアプリケーションを記述しています。アプリサーバーはそれらのコメントをデータベースに保存します。データベースを定期的にクリーンアップするAWS Lambdaがあります。```python
#!/usr/bin/env python3
from pytm import TM, Server, Datastore, Dataflow, Boundary, Actor, Lambda, LLM, Data, Classification
tm = TM("my test tm")
tm.description = "another test tm"
tm.isOrdered = True
User_Web = Boundary("User/Web")
Web_DB = Boundary("Web/DB")
user = Actor("User")
user.inBoundary = User_Web
web = Server("Web Server")
web.OS = "CloudOS"
web.isHardened = True
web.sourceCode = "server/web.cc"
db = Datastore("SQL Database (*)")
db.OS = "CentOS"
db.isHardened = False
db.inBoundary = Web_DB
db.isSql = True
db.inScope = False
db.sourceCode = "model/schema.sql"
comments = Data(
name="Comments",
description="Comments in HTML or Markdown",
classification=Classification.PUBLIC,
isPII=False,
isCredentials=False,
# credentialsLife=Lifetime.LONG,
isStored=True,
isSourceEncryptedAtRest=False,
isDestEncryptedAtRest=True
)
results = Data(
name="results",
description="Results of insert op",
classification=Classification.SENSITIVE,
isPII=False,
isCredentials=False,
# credentialsLife=Lifetime.LONG,
isStored=True,
isSourceEncryptedAtRest=False,
isDestEncryptedAtRest=True
)
my_lambda = Lambda("cleanDBevery6hours")
my_lambda.hasAccessControl = True
my_lambda.inBoundary = Web_DB
llm_api = LLM("AI Writing Assistant")
llm_api.isThirdParty = True
llm_api.processesPersonalData = True
llm_api.hasContentFiltering = False
llm_api.hasSystemPrompt = True
llm_api.processesUntrustedInput = True
my_lambda_to_db = Dataflow(my_lambda, db, "(λ)Periodically cleans DB")
my_lambda_to_db.protocol = "SQL"
my_lambda_to_db.dstPort = 3306
user_to_web = Dataflow(user, web, "User enters comments (*)")
user_to_web.protocol = "HTTP"
user_to_web.dstPort = 80
user_to_web.data = comments
web_to_user = Dataflow(web, user, "Comments saved (*)")
web_to_user.protocol = "HTTP"
web_to_db = Dataflow(web, db, "Insert query with comments")
web_to_db.protocol = "MySQL"
web_to_db.dstPort = 3306
db_to_web = Dataflow(db, web, "Comments contents")
db_to_web.protocol = "MySQL"
db_to_web.data = results
web_to_llm = Dataflow(web, llm_api, "Chat completion request")
web_to_llm.protocol = "HTTPS"
web_to_llm.dstPort = 443
tm.process()
You also have the option of using pytmGPT to create your models from prose!
上記の tm.py ファイルに --dfd 引数を渡すと、標準出力に結果が生成され、それが Graphviz の dot に渡されてデータフロー図が作成されます。```bash
tm.py --dfd | dot -Tpng -o sample.png
このダイアグラムを生成します:
dfd.png
要素に ".levels = [1,2]" 属性を追加すると、その要素(および両方のフローの終端が同じDFDレベルにある場合は関連するデータフローも)が、コマンド引数 "--levels 1 2" に応じて表示/非表示になります。
次のコマンドはシーケンス図を生成します。```bash
tm.py --seq | java -Djava.awt.headless=true -jar plantuml.jar -tpng -pipe > seq.png
この図を生成します:
seq.png
図と発見事項はテンプレートに含めて最終レポートを作成できます:```bash
tm.py --report docs/basic_template.md | pandoc -f markdown -t html > report.html
レポートテンプレートで使用されるテンプレート形式は非常にシンプルです:```text
# Threat Model Sample
***
## System Description
{tm.description}
## Dataflow Diagram

## Dataflows
Name|From|To |Data|Protocol|Port
----|----|---|----|--------|----
{dataflows:repeat:{{item.name}}|{{item.source.name}}|{{item.sink.name}}|{{item.data}}|{{item.protocol}}|{{item.dstPort}}
}
## Findings
{findings:repeat:* {{item.description}} on element "{{item.target}}"
}
要素ごとに結果をグループ化するには、より高度なネストされたループを使用します。```text
{elements🔁{{item.findings:if:
{{item.findings🔁 Threat: {{{{item.id}}}} - {{{{item.description}}}}
Severity: {{{{item.severity}}}}
Mitigations: {{{{item.mitigations}}}}
References: {{{{item.references}}}}
}}}}}
ループ内のすべてのアイテムはエスケープする必要があり、中括弧を二重にします。つまり、`{item.name}` は `{{item.name}}` になります。
上記の例では2つのネストされたループを使用しているため、内側のループのアイテムは2回エスケープする必要があり、そのため4つの中括弧を使用しています。
### オーバーライド
発見(モデルアセットやデータフローに一致する脅威)の属性をオーバーライドできます。例えば、カスタムCVSSスコアや応答テキストを設定するために:```python
user_to_web = Dataflow(user, web, "User enters comments (*)", protocol="HTTP", dstPort="80")
user_to_web.overrides = [
Finding(
# Overflow Buffers
threat_id="INP02",
cvss="9.3",
response="""**To Mitigate**: run a memory sanitizer to validate the binary""",
severity="Very High",
)
]
If you are adding a Finding, make sure to add a severity: "Very High", "High", "Medium", "Low", "Very Low".
セキュリティ実務者は、TM.threatsFile を設定することで、独自の脅威ファイルを提供できます。エントリは次のような形式である必要があります:```json
{
"SID":"INP01",
"target": ["Lambda","Process"],
"description": "Buffer Overflow via Environment Variables",
"details": "This attack pattern involves causing a buffer overflow through manipulation of environment variables. Once the attacker finds that they can modify an environment variable, they may try to overflow associated buffers. This attack leverages implicit trust often placed in environment variables.",
"Likelihood Of Attack": "High",
"severity": "High",
"condition": "target.usesEnvironmentVariables is True and target.controls.sanitizesInput is False and target.controls.checksInputBounds is False",
"prerequisites": "The application uses environment variables.An environment variable exposed to the user is vulnerable to a buffer overflow.The vulnerable environment variable uses untrusted data.Tainted data used in the environment variables is not properly validated. For instance boundary checking is not done before copying the input data to a buffer.",
"mitigations": "Do not expose environment variable to the user.Do not use untrusted data in your environment variables. Use a language or compiler that performs automatic bounds checking. There are tools such as Sharefuzz [R.10.3] which is an environment variable fuzzer for Unix that support loading a shared library. You can use Sharefuzz to determine if you are exposing an environment variable vulnerable to buffer overflow.",
"example": "Attack Example: Buffer Overflow in $HOME A buffer overflow in sccw allows local users to gain root access via the $HOME environmental variable. Attack Example: Buffer Overflow in TERM A buffer overflow in the rlogin program involves its consumption of the TERM environmental variable.",
"references": "https://capec.mitre.org/data/definitions/10.html, CVE-1999-0906, CVE-1999-0046, http://cwe.mitre.org/data/definitions/120.html, http://cwe.mitre.org/data/definitions/119.html, http://cwe.mitre.org/data/definitions/680.html"
}
`target`フィールドは、この脅威と照合するモデル要素のクラスをリストします。
これらは、Actor、Datastore、Server、Process、SetOfProcesses、ExternalEntity、Lambda、LLM、Agent、または任意の要素にマッチする基底クラスであるElementなどの資産です。また、2つの資産を接続するDataflowでもかまいません。
他のフィールド(`condition`を除く)は表示に利用可能で、最終的な[レポート](#report)で検出結果をリストするためにテンプレート内で使用できます。
> **警告**
>
> `threats.json`ファイルには、`eval()`を通過する文字列が含まれています。ファイルのパーミッションが正しく設定されていることを確認するか、攻撃者が文字列を変更して自分の代わりにコードを実行させるリスクを回避してください。
ロジックは`condition`にあり、`target`のメンバーを論理的に評価できます。
trueを返すとルールが検出結果を生成し、それ以外の場合は検出結果になりません。
Conditionは`target`の属性や、`target.control`の制御属性を比較することもでき、以下のメソッドのいずれかを呼び出すこともできます。
* `target.oneOf(class, ...)` ここで`class`は、Actor、Datastore、Server、Process、SetOfProcesses、ExternalEntity、Lambda、LLM、Agent、Dataflowのうち1つ以上です。
* `target.crosses(Boundary)`
* `target.enters(Boundary)`
* `target.exits(Boundary)`
* `target.inside(Boundary)`
`target`がDataflowの場合、`target.source`や`target.sink`、その他の属性にアクセスできることを忘れないでください。
資産に対する条件は、`target.input`および`target.output`属性を検査することで、すべての着信および発信Dataflowを分析できます。例えば、受信トラフィックがあるサーバーに対してのみ脅威をマッチさせるには、`any(target.inputs)`を使用します。より高度な例として、SQLデータストアに接続する要素をマッチさせるには、`any(f.sink.oneOf(Datastore) and f.sink.isSQL for f in target.outputs)`となります。
## JSONからのインポート
少しのPythonコードを使用することで、JSONから脅威モデルをインポートすることが可能です(`tests/input.json`にある例の特別な形式に注意してください)。次の例は、テストにある`input.json`の例をインポートします。以下のコードを`tm2.py`として保存してください。```python
#!/usr/bin/env python3
# Example tm2.py contents
# Run: python tm2.py --dfd | dot -Tpng -o sample_json.png
from pytm import (
TM,
Actor,
Boundary,
Classification,
Data,
Dataflow,
Datastore,
Lambda,
Server,
DatastoreType,
Assumption,
load,
)
json_file_string = './tests/input.json'
with open(json_file_string) as input_json:
TM.reset()
tm = load(input_json)
tm.process()
以前と同じように tm2.py を呼び出すことができます。ここでは --dfd を指定し、出力を Graphviz (dot) にリダイレクトします:```bash
python tm2.py --dfd | dot -Tpng -o sample_json.png
## スライド作成!
脅威モデルが完成して準備が整うと、恐ろしいプレゼンテーションの段階がやってきます。しかし今では、pytmがその場面でも役立ちます。(RevealMD)[https://github.com/webpro/reveal-md] の力を借りて、脅威モデルをスライドで表現するテンプレートが用意されています。テンプレートの docs/revealjs.md を使用するだけで、ブラウザから発表・共有できる、完全にカスタマイズ可能な美しいスライドが生成されます。
https://github.com/izar/pytm/assets/368769/30218241-c7cc-4085-91e9-bbec2843f838
## 現在対応している脅威```text
INP01 - Buffer Overflow via Environment Variables
INP02 - Overflow Buffers
INP03 - Server Side Include (SSI) Injection
CR01 - Session Sidejacking
INP04 - HTTP Request Splitting
CR02 - Cross Site Tracing
INP05 - Command Line Execution through SQL Injection
INP06 - SQL Injection through SOAP Parameter Tampering
SC01 - JSON Hijacking (aka JavaScript Hijacking)
LB01 - API Manipulation
AA01 - Authentication Abuse/ByPass
DS01 - Excavation
DE01 - Interception
DE02 - Double Encoding
API01 - Exploit Test APIs
AC01 - Privilege Abuse
INP07 - Buffer Manipulation
AC02 - Shared Data Manipulation
DO01 - Flooding
HA01 - Path Traversal
AC03 - Subverting Environment Variable Values
DO02 - Excessive Allocation
DS02 - Try All Common Switches
INP08 - Format String Injection
INP09 - LDAP Injection
INP10 - Parameter Injection
INP11 - Relative Path Traversal
INP12 - Client-side Injection-induced Buffer Overflow
AC04 - XML Schema Poisoning
DO03 - XML Ping of the Death
AC05 - Content Spoofing
INP13 - Command Delimiters
INP14 - Input Data Manipulation
DE03 - Sniffing Attacks
CR03 - Dictionary-based Password Attack
API02 - Exploit Script-Based APIs
HA02 - White Box Reverse Engineering
DS03 - Footprinting
AC06 - Using Malicious Files
HA03 - Web Application Fingerprinting
SC02 - XSS Targeting Non-Script Elements
AC07 - Exploiting Incorrectly Configured Access Control Security Levels
INP15 - IMAP/SMTP Command Injection
HA04 - Reverse Engineering
SC03 - Embedding Scripts within Scripts
INP16 - PHP Remote File Inclusion
AA02 - Principal Spoof
CR04 - Session Credential Falsification through Forging
DO04 - XML Entity Expansion
DS04 - XSS Targeting Error Pages
SC04 - XSS Using Alternate Syntax
CR05 - Encryption Brute Forcing
AC08 - Manipulate Registry Information
DS05 - Lifting Sensitive Data Embedded in Cache
SC05 - Removing Important Client Functionality
INP17 - XSS Using MIME Type Mismatch
AA03 - Exploitation of Trusted Credentials
AC09 - Functionality Misuse
INP18 - Fuzzing and observing application log data/errors for application mapping
CR06 - Communication Channel Manipulation
AC10 - Exploiting Incorrectly Configured SSL
CR07 - XML Routing Detour Attacks
AA04 - Exploiting Trust in Client
CR08 - Client-Server Protocol Manipulation
INP19 - XML External Entities Blowup
INP20 - iFrame Overlay
AC11 - Session Credential Falsification through Manipulation
INP21 - DTD Injection
INP22 - XML Attribute Blowup
INP23 - File Content Injection
DO05 - XML Nested Payloads
AC12 - Privilege Escalation
AC13 - Hijacking a privileged process
AC14 - Catching exception throw/signal from privileged block
INP24 - Filter Failure through Buffer Overflow
INP25 - Resource Injection
INP26 - Code Injection
INP27 - XSS Targeting HTML Attributes
INP28 - XSS Targeting URI Placeholders
INP29 - XSS Using Doubled Characters
INP30 - XSS Using Invalid Characters
INP31 - Command Injection
INP32 - XML Injection
INP33 - Remote Code Inclusion
INP34 - SOAP Array Overflow
INP35 - Leverage Alternate Encoding
DE04 - Audit Log Manipulation
AC15 - Schema Poisoning
INP36 - HTTP Response Smuggling
INP37 - HTTP Request Smuggling
INP38 - DOM-Based XSS
AC16 - Session Credential Falsification through Prediction
INP39 - Reflected XSS
INP40 - Stored XSS
AC17 - Session Hijacking - ServerSide
AC18 - Session Hijacking - ClientSide
INP41 - Argument Injection
AC19 - Reusing Session IDs (aka Session Replay) - ServerSide
AC20 - Reusing Session IDs (aka Session Replay) - ClientSide
AC21 - Cross Site Request Forgery
DS06 - Data Leak
DR01 - Unprotected Sensitive Data
AC22 - Credentials Aging (deprecated)
AC23 - Credentials Disclosure
AC24 - Use of hardcoded credentials
LLM01 - Direct Prompt Injection
LLM02 - Indirect Prompt Injection via Retrieved Content
LLM03 - Sensitive Data Leakage to Third-Party Provider
LLM04 - Training Data Poisoning
LLM05 - Excessive Agency via Unauthorized Tool Use
LLM06 - Arbitrary Code Execution via LLM Agent
LLM07 - Jailbreaking and Safety Bypass
LLM08 - Sensitive Information Disclosure Through Output
LLM09 - Untrusted Tool Launch Configuration