AV Chaos Monkey
Distributed chaos engineering platform for load testing video conferencing systems. Simulates 1500+ WebRTC participants with H.264/Opus streams and injects network chaos spikes to validate system resilience under degraded conditions
Architecture
-
Media Processing Pipeline:
- FFmpeg converts input video to H.264 Annex-B and Ogg/Opus at startup
- NAL Reader parses H.264 stream (SPS/PPS/IDR/Slices)
- Opus Reader extracts 20ms audio frames from Ogg container
- Frames cached in memory, shared across all participants (zero-copy)
- Reduces CPU by ~90% vs per-participant encoding
-
Control Plane:
- HTTP Server (:8080) manages test lifecycle via REST API
- Spike Scheduler distributes chaos events (even/random/front/back/legacy)
- Network Degrader applies chaos: packet loss (1-25%), jitter (10-50ms), bitrate reduction (30-80%), frame drops (10-60%)
- Loaded chaos configuration applied to participant pool
-
Participant Pool:
- Auto-partitioned across pods using:
participant_id % total_partitions = partition_id
- Each participant generates RTP streams (PT=96 video, PT=111 audio)
- Participant ID embedded in RTP extension header (ID=1)
- Pool size: 1-100 (local), 100-500 (Docker), 500-1500 (Kubernetes)
-
Kubernetes Auto-Configuration:
- Pods auto-detect partition ID from pod name:
orchestrator-3 → PARTITION_ID=3
- Port allocation:
base_port + (partition_id × 10000) + participant_index
- Example: Partition 0 uses 5000-14999, Partition 1 uses 15000-24999
- StatefulSet with 10 replicas, each handling ~150 participants
- Resources: 1-4 CPU, 2-4Gi memory per pod
- Auto-configures based on host machine specs
-
UDP Relay Chain (Kubernetes only):
Orchestrator Pods (10×) → UDP :5000 → udp-relay Pod (Python)
→ Length-Prefixed TCP :5001 → kubectl port-forward 15001:5001
→ tools/udp-relay (Go) → UDP :5002 → Your Receiver
- Why: kubectl port-forward only supports TCP, not UDP
- In-cluster relay: Python script aggregates UDP from all pods, streams as TCP with 2-byte length prefix
- Local relay: Go tool converts TCP stream back to UDP packets
- Aggregates 1500 participant streams into single connection
-
WebRTC Infrastructure:
- Coturn StatefulSet: 3 initial replicas, HPA scales 1-10 based on load (~500 participants/replica)
- coturn-lb Service: Load balances TURN traffic across replicas
- webrtc-connector: Optional proxy layer (Deployment + HPA 2-10 replicas), handles SDP signaling
- Docker Mode: Single Coturn container for local testing
- Ports: 3478 (TURN), 49152-65535 (relay range)
- Credentials: webrtc/webrtc123
-
Client Integration:
- UDP Receiver: Receives aggregated RTP stream from all participants via relay chain
- WebRTC Receiver: Establishes 1:1 WebRTC connections via SDP exchange through TURN servers
- Both forward to your video call system under test (SFU/MCU/Mesh)
-
Observability Stack (Optional):
- Prometheus: Scrapes
/metrics endpoint from all orchestrator pods every 5s
- Grafana: Visualizes metrics via pre-configured dashboard (admin/admin)
- Metrics exposed: participant count, packets sent, bytes sent, active spikes, packet loss %, jitter, MOS score
- Access: Prometheus on :30090, Grafana on :30030 (NodePort)
- Orchestrator pods annotated for auto-discovery:
prometheus.io/scrape: "true"
Core Concepts
Participant Simulation
Each virtual participant generates real media streams:
- Video: H.264 NAL units from actual video files, packetized per RFC 6184
- Audio: Opus frames from Ogg containers, packetized per RFC 7587
- RTP: Standards-compliant headers with participant ID extensions
- Timing: Frame-accurate timing (30fps video, 20ms audio packets)
Chaos Injection
Five spike types simulate real-world network conditions:
- Packet Loss: Drops RTP packets at application layer (1-100%)
- Network Jitter: Adds latency variation (base + gaussian jitter)
- Bitrate Reduction: Throttles video encoding (30-80% reduction)
- Frame Drops: Skips video frames (10-60% drop rate)
- Bandwidth Limiting: Caps total throughput
Distribution Strategies
Spikes are distributed across test duration using configurable strategies:
- Even: Uniform spacing with jitter (predictable load)
- Random: Unpredictable timing (realistic chaos)
- Front-loaded: Dense spikes early (recovery testing)
- Back-loaded: Baseline then chaos (comparison testing)
- Legacy: Fixed interval ticker (runtime injection)
Partitioning
Kubernetes deployments use participant partitioning for horizontal scaling:
- Each pod handles
participant_id % total_partitions == partition_id
- Port allocation:
base_port + (partition_id * 10000) + participant_index
- Automatic load distribution across 1-10 pods
- Scales to 1500+ participants (150 per pod)
Running the System
1. Local Development (Native Go)
Best for: Development, debugging, small-scale tests (1-100 participants)
# Start orchestrator
go run cmd/main.go
# In another terminal: Start UDP receiver
go run examples/go/udp_receiver.go 5002
# Edit config/config.json to set num_participants: 10
# Run chaos test
go run tools/chaos-test/main.go -config config/config.json
What happens:
- Single orchestrator process on
:8080
- Participants send UDP to
127.0.0.1:5002
- Chaos spikes injected via HTTP API
- Real-time metrics displayed every 2s
Configuration (config/config.json):
{
"base_url": "http://localhost:8080",
"media_path": "public/rick-roll.mp4",
"num_participants": 10,
"duration_seconds": 300,
"spikes": {
"count": 20,
"interval_seconds": 5,
"types": { "rtp_packet_loss": {...}, "network_jitter": {...} }
},
"spike_distribution": {
"strategy": "random",
"min_spacing_seconds": 5,
"jitter_percent": 15
}
}
2. Docker Compose (Containerized)
Best for: Isolated testing, CI/CD, medium-scale tests (100-500 participants)
Prerequisites:
- Docker Desktop with 8-16GB memory allocation
docker-compose installed
# Build and start orchestrator container
./scripts/start_everything.sh build
# In another terminal: Start UDP receiver
go run examples/go/udp_receiver.go 5002
# Edit config/config.json to set num_participants: 100
# Run chaos test (targets container)
go run tools/chaos-test/main.go -config config/config.json
Resource Limits (edit docker-compose.yaml):
services:
orchestrator:
deploy:
resources:
limits:
cpus: "14.0"
memory: 6G # Increase for more participants
Scaling Guide:
| Docker Memory | Max Participants | CPU Cores |
|---|
| 8 GB | ~100 | 4 |
| 16 GB | ~250 | 8 |
| 24 GB | ~400 | 12 |
| 32 GB | ~500 | 14 |
3. Kubernetes with Nix (Production Scale)
Best for: Large-scale tests (500-1500 participants), horizontal scaling, production validation
Prerequisites:
- Nix with flakes enabled
- Docker Desktop or kind cluster
- kubectl configured
Step 1: Enter Nix Environment
# Nix provides: Go, Docker, kubectl, kind, ffmpeg
nix develop
# Or use direnv for auto-activation
echo "use flake" > .envrc
direnv allow
Step 2: Deploy to Kubernetes
# Auto-deploy with optimal settings (detects system resources)
./scripts/start_everything.sh run -config config/config.json
# Or specify custom media files
./scripts/start_everything.sh run --media=path/to/video.mp4 -config config/config.json