All articles

Namespaces and cgroups: The Two Boundaries Behind Containers

10 minutes read


Linux publication

Share this article

𝕏✉

Illustration for Namespaces and cgroups: The Two Boundaries Behind Containers

A container is not one kernel object. A runtime combines Linux mechanisms with a filesystem, process setup, networking, security policy, and lifecycle contract described by specifications such as OCI’s runtime contract 6. Two mechanisms carry much of the mental model:

  • namespaces give a process a particular view of resources;
  • cgroups organize processes for accounting, limits, and control.

Namespaces answer “what can this process see or identify?” Cgroups answer “how much can this group consume, and how is it tracked?” Neither alone is a complete security boundary.

Inspect namespace identities

ls -l /proc/self/ns
readlink /proc/self/ns/mnt
readlink /proc/self/ns/pid
readlink /proc/self/ns/net
readlink /proc/self/ns/user

The symlink identifiers let you compare two processes. Matching identifiers mean the processes refer to the same namespace of that type; different identifiers mean different views.

Linux documents mount, PID, network, IPC, UTS, user, cgroup, and time namespaces, among others 1. Their effects differ:

Namespace Isolates a view of
mount mount points and filesystem topology
PID process IDs and parent relationships
network interfaces, routes, ports, firewall state
UTS hostname and domain name
IPC System V IPC and POSIX message queues
user user/group IDs and capabilities relative to the namespace
cgroup cgroup path exposure
time selected boot and monotonic clock offsets

A process can share some namespaces with the host and differ in others. “Inside a container” is therefore not a sufficient troubleshooting fact.

Inspect cgroup v2

stat -fc '%T' /sys/fs/cgroup
cat /proc/self/cgroup

The verified lab printed:

cgroup2fs
0::/user.slice/user-1000.slice/user@1000.service/app.slice/…scope

cgroup2fs and the single 0:: hierarchy identify unified cgroup v2 on this host. The kernel’s cgroup v2 documentation defines a hierarchical organization: resource distribution occurs between a cgroup and its children, with controller files such as cpu.max, memory.max, and pids.max where delegated 2.

Read limits only where the process has access:

group=$(awk -F: '$1 == "0" { print $3 }' /proc/self/cgroup)
for file in memory.max memory.current cpu.max pids.max pids.current; do
  path="/sys/fs/cgroup${group}/${file}"
  if [ -r "$path" ]; then printf '%-16s %s\n' "$file" "$(cat "$path")"; fi
done

max commonly means no limit at that level. A parent may still impose a ceiling, and CPU weight is not the same as a hard CPU quota. Read the controller’s exact semantics before translating one file into a percentage.

Verify the model in a pinned container

The second lab used Docker Engine 29.7.1 and the official digest-pinned debian:13.6@sha256:34cd9e9fd437c0a095ec39cb2e73422c9f30821b0d0848ed74fd0d43bae4d958 image. It added no host mount, disabled networking, dropped every capability, enabled no-new-privileges, made the root read-only, and set explicit memory and process limits.

Caution: This command creates and automatically removes a disposable container. It still asks the Docker daemon to start a process; run it only on a lab host where that daemon and image are trusted.

docker run --rm --pull=never --read-only --network none \
  --cap-drop ALL --security-opt no-new-privileges \
  --memory=64m --pids-limit=64 \
  debian:13.6@sha256:34cd9e9fd437c0a095ec39cb2e73422c9f30821b0d0848ed74fd0d43bae4d958 \
  sh -c '
    readlink /proc/self/ns/mnt
    readlink /proc/self/ns/pid
    readlink /proc/self/ns/net
    cat /proc/self/cgroup
    cat /sys/fs/cgroup/memory.max
    cat /sys/fs/cgroup/pids.max
    grep -E "^(CapEff|CapBnd|NoNewPrivs):" /proc/self/status
  '

The observed controller values were 67108864 bytes and 64; CapEff and CapBnd were all zero, and NoNewPrivs was 1. Mount, PID, and network namespace identifiers differed from the host, while the user namespace was shared. The root mount reported overlay with ro. Docker documents that memory and CPU controls depend on kernel support and configuration 5, so verify the emitted cgroup files instead of assuming a flag took effect.

User namespaces do not grant host root

A user namespace can map an unprivileged host user to UID 0 inside that namespace. Capabilities are namespaced and checked relative to the governing user namespace; this is not UID 0 in the initial host namespace 4.

unshare --user --map-root-user sh -c '
  printf "uid="; id -u
  readlink /proc/self/ns/user
'

On the lab host, that attempt returned:

unshare: unshare failed: Operation not permitted

That denial is evidence, not a reason to loosen the machine. A kernel build, sysctl, LSM, container policy, or outer virtualization boundary can disable unprivileged namespace creation. Test in an approved disposable environment if the lesson requires it.

Capabilities divide privilege, but do not finish isolation

Linux capabilities split traditional root powers into units such as CAP_NET_ADMIN, CAP_SYS_ADMIN, and CAP_CHOWN 3. Inspect one process with:

grep -E '^(Uid|Gid|Cap(Inh|Prm|Eff|Bnd|Amb)|NoNewPrivs):' /proc/self/status

Dropping capabilities reduces authority. A strong container profile also uses filesystem permissions, seccomp syscall filtering, an LSM such as SELinux or AppArmor, read-only mounts, careful device exposure, and a narrow network path. Mounting the host Docker socket or broad device nodes can bypass the isolation story regardless of a tidy namespace list.

Read a container from the process outward

When a workload surprises you, capture:

printf 'process\n'; ps -o pid,ppid,user,stat,args -p 1,$$
printf 'namespaces\n'; ls -l /proc/self/ns
printf 'cgroup\n'; cat /proc/self/cgroup
printf 'mounts\n'; findmnt -o TARGET,SOURCE,FSTYPE,OPTIONS | sed -n '1,25p'
printf 'security\n'; grep -E '^(CapEff|CapBnd|NoNewPrivs):' /proc/self/status

Inside a container, PID 1 has special lifecycle responsibilities. The mount view may show an image plus writable overlay. Network interfaces can be namespaced while DNS configuration is injected as a file. The cgroup path connects the process to runtime resource controls.

The durable model

Do not ask only “is it containerized?” Ask:

  1. Which namespaces differ from the host?
  2. Which cgroup contains the process, and what do its ancestors limit?
  3. Which capabilities and devices remain available?
  4. Which host paths, sockets, and networks cross the boundary?
  5. Which process owns shutdown, children, and logs?

Namespaces shape visibility. Cgroups shape resource governance. Capabilities and security controls shape authority. A container runtime assembles them; the kernel does not turn that assembly into a tiny virtual machine.

Sources and further reading
  1. Linux man-pages — namespaces(7)
  2. Linux kernel documentation — cgroup v2
  3. Linux man-pages — capabilities(7)
  4. Linux man-pages — user_namespaces(7)
  5. Docker Engine — container resource constraints
  6. Open Container Initiative runtime specification