
CPRA is a high-performance infrastructure monitoring system designed for platform teams managing large-scale microservice architectures. Built on Entity-Component-System (ECS) architecture and queueing theory principles, CPRA handles 1,000,000+ concurrent health checks with automatic worker pool scaling to meet SLO targets.
Continuous Pulse and Recovery Agent
Checks services, sends alerts, and runs the recovery actions you configure.
CPRa is a self-hosted monitoring and recovery agent written in Go. It runs
health checks against your services on a schedule, opens and closes incidents
against configurable thresholds, sends notifications, and executes a recovery
action — restart a container, call a webhook, restart or scale a Kubernetes
workload, reboot an EC2 instance, restart a systemd unit — when a service
fails. It ships as a single static binary with an embedded read-only dashboard,
an HTTP API, and the cpractl command-line client. It is MIT-licensed.
Documentation: ziad-hsn.github.io/cpra — quickstart · monitor configuration · drivers · HTTP API · deployment · FAQ
The complete documentation source is now kept on main and
synchronized to the published site.
Versions and availability identifies the code behind each guide:
410fbfb.The candidate guides describe source that is separate from this branch. Their publication does not add those APIs, storage or installation commands to main. See the latest-change review for the corrections and source boundaries behind the documentation refresh.
Requires Go 1.25 or later and Make. The repository already contains the built dashboard assets, so a Go toolchain is enough.
make
cp examples/monitors.yaml monitors.yaml
# Set the service address in monitors.yaml.
./bin/cpra -yaml monitors.yaml
Open http://localhost:8060 or run ./bin/cpractl get monitors. Missing or malformed configuration stops startup. Empty configurations require -allow-empty.
The example checks an HTTP endpoint and writes incident transitions to alerts.jsonl. Each monitor can specify a check interval, timeout, failure threshold, recovery threshold, notification destinations, and a recovery action. Maintenance windows suppress alerts and recovery while checks continue; they use five-field cron expressions, a duration, and an IANA timezone.
| Function | Default build | Optional build tags |
|---|---|---|
| Checks | HTTP, TCP, ICMP, DNS, UDP, TLS, Docker, gRPC port reachability | redis postgres mysql mongo rabbitmq kafka |
| Recovery | Docker, HTTP webhook | kubernetes aws systemd |
| Alerts | Log, Slack, PagerDuty, email, webhook, Telegram, Discord, Opsgenie, Mattermost, VictorOps, Pushover, Datadog | teams twilio |
make BUILD_TAGS='redis postgres kubernetes'
The grpc check tests the TCP port; it does not call the gRPC health service. UDP checks require a payload and a reply. PagerDuty requires an Events API v2 routing key. Email uses an SMTP relay with STARTTLS; SMTP username/password authentication is not implemented.
TLS warn_days produces a yellow alert and a degraded monitor status without starting recovery; critical_days fails the check and follows the normal recovery policy. Pushover emergency priority accepts retry and expire in seconds, defaulting to 60 and 1800. Docker recovery preserves the daemon's stop grace when its timeout is omitted.
MongoDB checks require a direct mongodb:// URI. The selected driver cannot bound initial mongodb+srv:// discovery by the check deadline, so CPRa rejects that mode. Kubernetes recovery supports token, certificate, and in-cluster credentials; kubeconfig exec credential plugins are rejected because they can outlive the recovery deadline.
The server listens on loopback by default. Other bind addresses require an authentication token. Keep the token in a file readable by the service account:
./bin/cpra -yaml monitors.yaml -web.addr 0.0.0.0:8060 -web.auth-file /run/secrets/cpra-token
./bin/cpractl --server https://monitor.example.com --token-file /run/secrets/cpra-token get monitors
Browser login uses username cpra and the token as the password. The API accepts a Bearer token. The server and CLI also read CPRA_AUTH_TOKEN and CPRA_AUTH_TOKEN_FILE. Remote access needs an HTTPS reverse proxy; the built-in listener serves HTTP.
Manifests authorize checks, notification destinations, and recovery actions with the process account's permissions. Use trusted configuration. -ssrf-protect blocks non-public HTTP destinations at connection time; it does not restrict other protocols or local actions. Profiling is opt-in. Keep token files and manifests containing credentials outside version control.