AST-based Static Code Analyzer with Agentic LLM-Powered Relationship Mapping to discover Python RCE paths and deep deserialization chains on AI, LLM, Robotics, Data Science, Machine Learning and Deep Learning (but not limited).

Far beyond a traditional scanner, it provides a generic, high-performance framework for auditing over 120 libraries and formats—including YAML, Msgpack, CBOR, and custom JSON hooks—where traditional trust in "safe" serialization hides critical logic-based RCE vectors. By resolving imports, aliases, and complex dotted attributes, the tool serves as a high-fidelity signal amplifier that prioritizes dangerous code paths in modern distributed architectures and AI/ML repositories.
The project operates through a modular, multi-phased workflow that transitions from raw detection to deep technical audit. Following the initial high-velocity SAST scan, the ecosystem leverages specialized relationship mappers to trace execution flows and result processors to generate detailed security reports. This systematic approach ensures that every finding is contextualized within the application's broader architecture, transforming high-volume telemetry into actionable research assets and structured milestones that simplify the mapping of infrastructure-level attack surfaces.
At its most advanced tier, Deserializer integrates an autonomous AI Security Agent (Phase 4), explicitly designed to navigate "self-command-injection" limitations and synthesize functional reproduction guides, based on HuggingFace inference API, Local LLM (like llama.cpp), or OpenAI API compatible.
As the project evolves, it continues to define the frontier of automated vulnerability research by bridging the gap between abstract syntax tree static analysis, relationship mapping, reporting, and functional exploit development based on documented research.
Deserializer directly supports security research by locating RCE and Insecure Deserialization paths across various large-scale AI, Robotics, and Data Science projects and environments, like Genesis World (v0.2.1), MuJoCo (v3.7.0), LeRobot (v0.5.1), Brax (v0.14.2), TensorFlow (v2.21.0), LangGraph (v1.1.6), VibeVoice (v0.0.1), Hugging Face Hub (v1.11.0), PyGlove (v0.4.5), and many others.
Its capabilities have directly powered the discovery of critical vulnerabilities in industry-leading frameworks, proving its efficacy in auditing complex MLOps and agentic AI environments.
The project is structured into four distinct phases, each designed to shift the analysis from high-volume automated telemetry to deep, functional security research:
| Phase | Title | Tooling / Engine | Objective |
|---|---|---|---|
| 1 | High-Velocity Detection | deserializer.py (Triple-Pass) | Perform massive-scale SAST to identify potential deserialization sinks. |
| 2 | Relationship Mapping | Result Processors / Mappers | Contextualize findings by tracing execution flows and component interdependencies. |
| 3 | Technical Synthesis | Research Documentation | Formalize findings into technical writeups, mapping infrastructure-level attack surfaces. |
| 4 | Autonomous AI Agent | AI Security Agent | Automate 0-day discovery and generate functional reproduction guides/exploits using HuggingFace inference API, Local LLM (like llama.cpp), or OpenAI API compatible. |
A high-level map of the repository's organization and the technical purpose of each specialized directory:
agent/docs/exploit_development/modules/reports/research/templates/This scanner features a high-performance parallel execution engine built on Python's concurrent.futures.ProcessPoolExecutor. It is designed to scale across all available CPU cores (controllable via the -j or --concurrency flag), making it capable of scanning tens of thousands of files in seconds.
SetConsoleCtrlHandler via ctypes to ensure that Ctrl+C is 100% responsive, even during heavy processing.ctypes for native ANSI color support in modern CMD and PowerShell environments.[!WARNING] Performance Warning: When analyzing extremely large or complex files (e.g., over 1MB, 2MB, or 3MB in size), the tool may experience significant slowdowns or appear "stuck" while parsing deep AST trees. If you encounter such bottlenecks, consider using the
--timeout(to skip slow files) and--max-size(to skip huge files) flags to maintain scan velocity.