Skip to content
KitploitKITPLOIT
ToolsExploitsBlog
Log in
Submit
ToolsExploitsBlog
Submit

Hacking, PenTest, and Cybersecurity Tools for Your Security Arsenal!

Kitploit is a directory of hacking, cybersecurity, and pentesting tools. Discover the latest project updates to find vulnerabilities, analyze systems, automate testing, and strengthen your security.

··Feeds·Contact·Privacy·© 2026 Kitploit

Tool Directory

Categories

View all categories
Loading categories
standalone_tdo — Extracts structured Cyber Threat Intelligence (CTI) from PDF, DOCX, and TXT documents using LLMs. Generates MITRE ATT&CK attack flows, detection opportunities, and comprehensive entity-relationship graphs. | Kitploit
Tools/GitHubGitHub/blevene/standalone_tdo
Vulnerability AnalysisInformation GatheringThreat IntelligenceMachine LearningPapers & ResearchLearning & Education
GitHubblevene/standalone_tdo

standalone_tdo

Extracts structured Cyber Threat Intelligence (CTI) from PDF, DOCX, and TXT documents using LLMs. Generates MITRE ATT&CK attack flows, detection opportunities, and comprehensive entity-relationship graphs.

View Repository
4159 months agoNot yet reviewed

Most Popular

View all →

Discover the most used tools by our community.

Explore all tools

Browse our collection of tools

View all tools →
Share

TDO Standalone Extractor

A powerful, standalone command-line tool for extracting Cyber Threat Intelligence (CTI) from documents using Large Language Models with advanced structured output capabilities.

🚀 Features

  • Multi-format Document Support: PDF, DOCX, TXT, and Markdown files
  • Advanced LLM Integration: Google Gemini models with structured output and automatic fallback parsing
  • Comprehensive CTI Schema: 12 entity types and 24 relationship types with rich properties
  • Detection Opportunity Generation: Creates actionable threat detection rules with evidence backing
  • Attack Flow Synthesis: Generates evidence-based MITRE ATT&CK flow diagrams
  • Structured Output Reliability: Pydantic schemas with 100% JSON parsing success rate
  • Parallel Processing: Process multiple files concurrently with progress tracking
  • Portable & Self-contained: Minimal dependencies with environment-based configuration

📋 Quick Start

1. Installation

# Clone or download the standalone-tdo folder
cd standalone-tdo

# Create virtual environment (recommended)
python -m venv .venv
source .venv/bin/activate  # On Windows: .venv\Scripts\activate

# Install dependencies
pip install -r requirements.txt

2. Configuration

The tool uses environment variables loaded from a .env file for configuration:

# Copy the example configuration
cp env.example .env

# Edit .env with your settings
# Required variables:
GEMINI_API_KEY=your-google-ai-api-key-here
GEMINI_MODEL=gemini-2.5-flash

Getting a Gemini API Key:

  1. Visit Google AI Studio
  2. Create a new API key
  3. Copy the key to your .env file

Available Models:

  • gemini-2.5-flash-preview-05-20 (recommended - fast and cost-effective)
  • gemini-2.5-pro-preview-06-05 (more powerful, slower)

3. Basic Usage

# Process a single file with all features
python tdo_bulk.py report.pdf --flow --opps

# Process multiple files
python tdo_bulk.py file1.pdf file2.docx file3.txt

# Process all files in a directory with parallel workers
python tdo_bulk.py reports/ -j 4

# Generate comprehensive analysis with evaluation
python tdo_bulk.py report.pdf --flow --opps --eval-opps

🎯 Command Line Reference

usage: tdo_bulk.py [options] FILE [FILE ...]

Positional Arguments:
  FILE                  One or more files or directories to process

Core Options:
  -o, --output DIR      Output directory (default: ./extracted_data)
  --csv FILE            Export summary to CSV file
  -j, --jobs N          Number of parallel workers (default: 1)
  --retries N           LLM retry count (default: 2)
  --backoff SEC         Exponential back-off base (default: 1.5)

Analysis Features:
  --flow                Generate Attack-Flow JSON
  --opps                Generate Threat Detection Opportunities
  --eval-opps           Evaluate detection opportunities quality
  --debug-opps          Show debug information for opportunity generation

Output Control:
  -q, --quiet           Minimal console output
  -v, --verbose         Debug-level logging
  -h, --help            Show help message and exit

Example Commands

# Basic CTI extraction
python tdo_bulk.py threat_report.pdf

# Full analysis with all features
python tdo_bulk.py apt_report.pdf --flow --opps --eval-opps

# Batch processing with parallel workers
python tdo_bulk.py reports_folder/ -j 8 --flow --opps

# Export results to CSV
python tdo_bulk.py *.pdf --csv results.csv

# Quiet mode for automation
python tdo_bulk.py reports/ -q -o /var/soc/cti --opps

📊 Output Files

For each processed file, the tool generates:

Primary Outputs

  • {filename}_extracted.json: Structured CTI data with comprehensive schema
  • {filename}_{timestamp}.md: Human-readable markdown report with all analysis

Optional Outputs (with flags)

  • Attack Flow JSON: Evidence-based MITRE ATT&CK flow structure (with --flow)
  • Detection Opportunities: Actionable detection rules with evidence (with --opps)

🔍 CTI Schema Overview

Entity Types (12 Types)

Entity TypeDescriptionKey Properties
ThreatActorCyber threat groupsaliases, primary_motivation, first_seen
ToolLegitimate softwarefamily, capabilities, kill_chain_phases
MalwareMalicious softwarefamily, capabilities, first_seen, last_seen
TechniqueMITRE ATT&CK techniquesid (T1234), description, kill_chain_phases
TacticMITRE ATT&CK tacticsid (TA0001), description
InfrastructureIPs, domains, URLstype, tags, first_seen, last_seen
IndicatorFile hashes, patternsvalue, pattern, valid_from, valid_until
VulnerabilityCVE entriesid, cvss_score, affected_software
CampaignNamed attack campaignsobjective, status, first_seen
IdentityTarget organizationstype, description
CourseOfActionMitigations, patchestype, description
SourceDocument provenancefilename, document_title, file_size_mb

Relationship Types (24 Types)

Attribution & Actor Relationships:

  • USES_TOOL, USES_TECHNIQUE, CONDUCTS, ATTRIBUTED_TO, TARGETS

Technical Relationships:

  • TOOL_IMPLEMENTS_TECHNIQUE, HOSTS_ON, COMMUNICATES_WITH, VARIANT_OF
  • DROPS, DOWNLOADS, INDICATES, OBSERVED_ON, EXPLOITS

Advanced Relationships:

  • MITIGATES, DETECTS, IS_SUBTECHNIQUE_OF, FLOW_CONTAINS_STEP
  • FLOW_USED_BY_ACTOR, OBSERVES, SOURCED_FROM

All relationships support properties like confidence, source, first_seen, last_seen.

💡 Detection Opportunities

The tool generates evidence-based detection opportunities with:

Core Features

  • Technique Mapping: Links to specific MITRE ATT&CK techniques
  • Observable Artefacts: Specific indicators to detect
  • Behavioral Patterns: Sequences and patterns to monitor
  • Evidence Citations: Direct quotes from source reports
  • Confidence Scores: Reliability assessment (0.0-1.0)
  • Quality Evaluation: Automated scoring with criteria breakdown

Example Output

{
  "id": "opp-001",
  "name": "Detect PowerShell Process Injection",
  "technique_id": "T1055",
  "artefacts": ["powershell.exe with WriteProcessMemory calls"],
  "behaviours": ["Process injection into legitimate processes"],
  "rationale": "APT groups commonly use PowerShell for process injection",
  "confidence": 0.8,
  "source": "PowerShell used for process injection (T1055)",
  "evidence": [
    "• PowerShell executes WriteProcessMemory calls (line 45)",
    "• Relationship: ThreatActor USES_TECHNIQUE T1055"
  ]
}

🔗 Attack Flow Synthesis

Enhanced attack flows provide:

  • Evidence-based Ordering: Steps backed by temporal and causal evidence
  • Entity Referencing: Each step references extracted CTI entities
  • Explicit Reasoning: Justification for step ordering with citations
  • Comprehensive Coverage: Up to 25 steps covering full attack lifecycle

Example Attack Flow

{
  "flow": {
    "label": "AttackFlow",
    "pk": "attack-flow--uuid",
    "properties": {
      "name": "APT29 Multi-stage Attack",
      "description": "Sophisticated spear-phishing to data exfiltration flow"
    }
  },
  "steps": [
    {
      "order": 1,
      "entity": {"label": "Technique", "pk": "T1566.001"},
      "description": "Spear-phishing attachment delivery",
      "reason": "Initial access method cited in report section 2.1"
    }
  ]
}

🏗️ Architecture Overview

Core Components

🎯 Main CLI (tdo_bulk.py)

  • Entry point with argument parsing and environment validation
  • Loads .env configuration using python-dotenv
  • Orchestrates bulk processing with progress tracking
Download Tool