
Like Envoy xDS, but for eBPF filters
Like Envoy xDS, but for eBPF filters.
Netfence runs as a daemon on your VM/container hosts and automatically injects eBPF filter programs into cgroups and network interfaces, with a built-in DNS server that resolves allowed domains and populates the IP allowlist.
Netfence daemons can be driven through their local Unix-socket API alone, or connect to a central control plane that you implement via gRPC to synchronize allowlists/denylists with your backend.
Your control plane pushes network rules like ALLOW *.pypi.org or ALLOW 10.0.0.0/16 to attached interfaces/cgroups. When a VM/container queries DNS, Netfence resolves it, adds the IPs to the eBPF filter, and drops traffic to unknown IPs before it leaves the host with warmed-path overhead that is effectively indistinguishable from a normal socket connect in current benchmarks.
In allowlist mode, IPv4 link-local (169.254.0.0/16) is no longer auto-allowed
by default — so the cloud metadata service (169.254.169.254) is blocked unless
explicitly allowlisted. This is deliberate: the metadata service is a
credential-theft target, and sandboxed workloads must not be able to reach it
implicitly. Localhost (127.0.0.0/8, ::1) and IPv6 neighbor discovery
(fe80::/10, ff02::/16) remain allowed by default so basic connectivity and NDP
keep working. To permit the metadata service for a workload, allowlist
169.254.169.254/32 (a per-attachment carve-out override via the control plane
is a planned follow-up).
IPv4 broadcast (255.255.255.255) and multicast (224.0.0.0/4) have no carve-out and are subject to policy, so under TC allowlist mode traffic like DHCP-renewal broadcasts is blocked unless explicitly allowlisted. Carve-out checks run before the denylist, so a carved range can only be blocked by turning its carve-out off — and because IPv4 link-local is now off by default, denylist mode can block the metadata service too.
A few major benefits to this solution that other options don't usually support:
secretdata.someattacker.com exfiltration)To my knowledge, no other solutions offers all of these features together.
Known limitation: cgroup attachments filter at the socket layer (connect/sendmsg
hooks), so a process with CAP_NET_RAW can craft raw packets that bypass them. Use
a TC (interface) attachment, which filters at the device layer, for workloads that
may hold CAP_NET_RAW.
However, this does have a bit more overhead than something like httpjail.
These numbers were measured in the privileged Docker Linux gate on linux/arm64
using make bench-docker. Values are medians of five samples.
The warm socket benchmark uses connected UDP sockets to isolate the
cgroup/connect4 eBPF hook cost from TCP handshake latency. In this path DNS has
already resolved the domain, the IP is still within TTL, and the IP/CIDR is
already present in the eBPF map.
| Path | Median latency |
|---|---|
| Normal socket connect, no eBPF | ~2.647 us |
| Warm allowlist, protected LPM hit | ~2.691 us |
| Warm allowlist, DNS exact-host hit | ~2.741 us |
| Allowlist miss, local block | ~1.652 us |
The measured spread between the normal, protected-LPM, and DNS exact-host connect paths is within sample noise.
There is no "kernel miss asks parent process" path today. A cgroup allowlist miss is decided locally by eBPF and is blocked immediately.
These numbers measure the DNS server path, not the warmed socket connect path.
| Path | Median latency |
|---|---|
| Proxy query cold, in-process policy function | ~31.336 us |
| Proxy query warm | ~27.964 us |
| Allowlist query cold with local upstream | ~53.510 us |
| Allowlist query warm with local upstream | ~53.432 us |
Cold rows synchronize through the real attachment mutation barrier and clear
the benchmark ownership graph and fake exact-map snapshot between queries.
The timer runs continuously to preserve UDP scheduler locality, while ns/op
subtracts the separately reported fixture-reset-ns/op wall time (including
any tail of the prior handler after the client received its packet) and thus
measures the current client Exchange. raw-total-ns/op reports both together.
The reset preserves configured policy domains and backing storage, and the
benchmark asserts one physical exact-map add per query. Warm rows prime
ownership once and assert one physical add across the run.
The internal ownership microbenchmarks below are scalability diagnostics, not end-to-end DNS query-path acceptance rows. The cached helper is retained only for tests and benchmarks; it wraps one record at a time and repeats domain validation. Both it and normal resolver traffic traverse the attachment mutation barrier, while normal resolver traffic admits each complete response as one transaction.
| Internal scalability diagnostic | Current median | Memory / allocations |
|---|---|---|
| Cold new-key admission, empty ownership graph | ~370.3 ns | 232 B, 5 allocs/op |
| Cold new-key admission, 4,095 unrelated entries | ~451.2 ns | 232 B, 5 allocs/op |
| Physical-capacity pressure and LRU replacement | ~3.820 ms | ~4.23 MB (4,226,243 B), 4,336 allocs/op |
| Exhausted physical-budget precheck | ~611.9 ns | 344 B, 9 allocs/op |
| Maximum-edge pressure, 64-address response | ~6.849 ms | ~7.66 MB (7,658,774 B), 2,233 allocs/op |
| Maximum-graph work guard, permitted full plan | ~2.763 ms | ~4.26 MB (4,264,386 B), 3,074 allocs/op |
| Maximum-graph work guard, exhausted pre-projection rejection | ~10.935 us | 8.76 KB (8,760 B), 14 allocs/op |
| Churn-budget operation near the numeric ceiling | ~18.98 ns | 0 B, 0 allocs/op |
| Coherent ownership-stats snapshot | ~2.094 ns | 0 B, 0 allocs/op |
| No-op expiry scan across 4,095 entries | ~74.849 us/scan | 0 B, 0 allocs/op |
+------------------+ +-------------------------+
| Your Control |<------->| Daemon (per host) |
| Plane (gRPC) | stream | |
+------------------+ | +-------------------+ |
| | DNS Server | |
| | (per-attachment) | |
| +-------------------+ |
+-------------------------+
|
+------+------+
| |
TC Filter Cgroup Filter
(veth, eth) (containers)
Each attachment gets a unique DNS address (port) provisioned by the daemon. Containers/VMs must be configured to use their assigned DNS address; filtering ordinary workload DNS traffic does not transparently redirect it.