
A lightweight, multi-layer Linux sandbox combining namespaces, pivot_root, seccomp-bpf, capability dropping, and an evidence-based verdict engine (Truthimatics Public Version) for secure, auditable code execution.
Multi-layer sandbox for native code execution on Linux.
Seven independent defence layers — no external dependencies, ~73 KiB PIE binary.
┌──────────────────────────────────────────────────────┐
│ Z-Jail │
├──────────────────────────────────────────────────────┤
│ Truthimatics PV (evidence-based verdict engine) │
│ Namespaces (mount, pid, net, ipc, uts) │
│ pivot_root (chroot on steroids) │
│ Capabilities (drop all, lock securebits) │
│ NO_NEW_PRIVS (no privilege escalation) │
│ seccomp-BPF (whitelist: 15 syscalls only) │
│ Audit (JSON logging + BLAKE2b hashing) │
└──────────────────────────────────────────────────────┘
git clone https://github.com/Division-36/Z-Jail.git
cd Z-Jail
make
sudo ./z_jail --root=/path/to/rootfs --seccomp-enforce -- /bin/ls
The --root directory should contain a minimal filesystem with the target binary and its dependencies (for static binaries, just the binary is enough).
Existing sandboxing solutions make trade-offs:
| Z-Jail | Firecracker | gVisor | bwrap | nsjail | |
|---|---|---|---|---|---|
| External deps | zero | libc, seccomp | Go runtime | libc | libc, protobuf |
| Binary size | ~73 KiB | 20+ MiB | 40+ MiB | ~70 KiB | ~1 MiB |
| VM isolation | no | yes (microVM) | no (sandbox) | no | no |
| seccomp whitelist | yes | no | yes | optional | yes |
| Content hashing | yes | no | no | no | no |
| Audit JSON | yes | no | yes | no | partial |
| Build complexity | one make | complex | complex | trivial | moderate |
Z-Jail fills the niche between bwrap (minimal, no seccomp-by-default) and nsjail (featureful, heavy deps). It is designed for CI pipelines, CTF jail challenges, and lightweight code evaluation where you need defence-in-depth without pulling in a container runtime.
flowchart LR
CLI[CLI args] --> P[parse_args]
P --> C{clone namespaces}
C -->|child| CR[child_run]
C -->|parent| W[waitpid]
CR --> RL[setrlimit]
RL --> FD[close fds >= 3]
FD --> DUMP[PR_SET_DUMPABLE=0]
DUMP --> PV[pivot_root]
PV --> NNP[PR_SET_NO_NEW_PRIVS]
NNP --> CAP[drop capabilities]
CAP --> SC[seccomp-BPF]
SC --> SIG[signal parent]
SIG --> EX[execve target]
W --> A[audit JSON]
A --> EXIT[exit]
Each layer is ordered so that a later layer can't be undone by an earlier one:
capset escalation after this pointsequenceDiagram
participant P as Parent
participant C as Child
P->>C: clone (NEWNS|NEWPID|NEWNET|NEWIPC|NEWUTS)
Note over C: setrlimit(CPU, AS, NOFILE, NPROC)
Note over C: close(all fds > 2)
Note over C: PR_SET_DUMPABLE=0
Note over C: pivot_root → chdir("/") → umount -l
Note over C: PR_SET_NO_NEW_PRIVS
Note over C: capset(all zero) + securebits
Note over C: seccomp(SECCOMP_MODE_FILTER, whitelist)
C->>P: write(pipe, ready=1)
Note over C: execve(target)
P->>P: waitpid
P->>P: write audit JSON
Evidence-based verdict engine. Collects weighted observations about the executed binary and determines a final verdict (DETERMINISTIC, REJECT, or UNCERTAIN). Each observation carries a weight; any single observation with weight >50% of total decides the verdict.
Five namespaces are created via clone():
| Namespace | Flag | Purpose |
|---|---|---|
| Mount | CLONE_NEWNS | Isolated filesystem tree |
| PID | CLONE_NEWPID | Process ID space (child is pid 1) |
| Net | CLONE_NEWNET | No network interfaces |
| IPC | CLONE_NEWIPC | No shared memory / semaphores |
| UTS | CLONE_NEWUTS | Separate hostname |
Requires CAP_SYS_ADMIN in the initial namespace.
Replaces the mount namespace root with the --root directory:
MS_BIND|MS_REC)pivot_root(new_root, put_old) — swap the mount treechdir("/") — move into the new rootumount2("/.pivot_old", MNT_DETACH) — detach old rootrmdir("/.pivot_old") — clean upThis is strictly stronger than chroot(2) — there is no way for the sandboxed process to escape back to the host root, even with CLONE_NEWNS from inside the sandbox (which is already blocked by seccomp).
All capabilities are dropped via:
capset(hdr, data) // data = {0, 0, 0}
prctl(SECBIT_KEEP_CAPS_LOCKED | SECBIT_NO_SETUID_FIXUP | ...)
The process drops setuid/setgid before capset so the uid change takes effect while CAP_SETUID is still held. After capset, all caps are gone and the securebits are locked — no re-enablement is possible.
prctl(PR_SET_NO_NEW_PRIVS, 1, 0, 0, 0);
Prevents the process or its children from gaining new privileges via setuid binaries, file capabilities, or LSM transitions. Irreversible.
Allow-list of 15 syscalls — anything not on the list gets SECCOMP_RET_KILL:
| Syscall | Number | Notes |
|---|---|---|
read | 0 | stdin |
write | 1 | stdout/stderr + report pipe |
openat | 257 | file access (not open) |
close | 3 | — |
lseek | 8 | — |
brk | 12 | heap management |
mmap | 9 | arg-restricted: flags & 4 == 0 (no MAP_SHARED), flags == 0x22 (MAP_PRIVATE|MAP_ANONYMOUS) |
munmap | 11 | — |
execve | 59 | single exec at startup |
exit_group | 231 | clean process exit |
rt_sigaction | 13 | signal handlers |
rt_sigprocmask | 14 | signal masking |
getrandom | 318 | random number source |
clock_gettime | 228 | timing |
fstat | 5 | file metadata |
The BPF filter is generated dynamically: for each whitelist entry, a jump chain is emitted that either allows (if syscall matches) or falls through to KILL. Architecture is checked first (AUDIT_ARCH_X86_64).