Docker 3 🐳 Linux Kernel Isolation Mechanics: Namespaces (PID, NET, MNT, IPC, UTS, USER)
Linux namespaces are the kernel feature that makes containers possible. A namespace wraps a global system resource in an abstraction that makes it appear to processes within the namespace that they have their own isolated instance of that resource . Changes made by one process inside a namespace are visible to other processes that are members of the same namespace, but are invisible to processes outside it. Docker creates a new set of namespaces for every container, so each container sees only its own processes, its own network stack, its own filesystem mounts, and its own hostname. The host kernel remains shared, but the view each container has of the system is partitioned.
This chapter examines the six namespace types most relevant to containers: PID, NET, MNT, IPC, UTS, and USER. Each isolates a different category of system resource, and together they construct the illusion of a self-contained system. Understanding what each namespace isolates, how namespaces nest, and what happens when the namespace’s init process terminates is essential for reasoning about container behavior and security.
Key point: Namespaces partition the view of global system resources — processes, network, mounts, IPC, hostname, and user IDs — so that each container sees only its own isolated instance of those resources while sharing the host kernel.
Why namespaces exist
The shared kernel problem. Containers share the host kernel. Without some mechanism to partition the kernel’s global resources, every process on the host would see every other process, every network interface, every mount point, and every hostname would be the same. Namespaces solve this by wrapping global resources in isolated views. The kernel maintains these partitions natively, so the isolation has no runtime cost beyond the bookkeeping the kernel already performs.
The PID collision problem. Process IDs are global in a traditional Linux system. Two containers both running an nginx master process would collide if they shared the same PID space. The PID namespace gives each container its own process ID number space, starting at 1, so the nginx master in one container can be PID 1 while the nginx master in another container is also PID 1 . The host kernel tracks the real PIDs, but each container sees only its own numbering.
The network stack problem. A container needs its own IP address, its own routing table, and its own port numbers. If containers shared the host network namespace, two containers could not both bind to port 80. The network namespace isolates network devices, IPv4 and IPv6 protocol stacks, IP routing tables, firewall rules, and port numbers . Each container gets its own network stack, with virtual ethernet (veth) pairs connecting it to the host or to other containers.
The filesystem isolation problem. A container must see its own root filesystem, not the host’s. The mount namespace isolates mount points, so a container can have its own /, its own /proc, its own /sys, and its own /tmp without conflicting with the host or other containers . When a container is created, the image layers are assembled into a root filesystem and the mount namespace makes that the container’s view of the filesystem hierarchy.
The IPC collision problem. System V IPC objects and POSIX message queues are identified by keys that are global in the traditional kernel. Two containers using the same key for a shared memory segment or message queue would collide. The IPC namespace gives each container its own set of System V IPC identifiers and its own POSIX message queue filesystem . Objects created in one IPC namespace are invisible to processes in other IPC namespaces.
The identity problem. Every Linux system has a hostname and a NIS domain name. Containers often need different hostnames — a web server container might want web-1 while a database container wants db-1. The UTS namespace isolates the hostname and domain name, so each container can set its own without affecting the host or other containers .
The privilege problem. A container running as root should not have the same power as root on the host. The user namespace maps a range of user and group IDs inside the container to a different range on the host, so root inside the container is an unprivileged user outside it . This is a security mechanism: even if an attacker escapes the container’s process isolation, they lack host-level privileges.
a. PID namespace: process ID isolation
The PID namespace isolates the process ID number space . Processes in different PID namespaces can have the same PID. When a new PID namespace is created, the first process in it has PID 1 and is the “init” process for that namespace. This process becomes the parent of any orphaned child processes within the namespace .
The init process has special significance. If the init process of a PID namespace terminates, the kernel terminates all other processes in that namespace via SIGKILL . This is because the init process is essential for the namespace’s operation — without it, orphaned processes would have no parent to reap them. After init terminates, it is impossible to create new processes in that namespace; fork() calls fail with ENOMEM .
Signal delivery is restricted inside a PID namespace. Only signals for which the init process has established a signal handler can be sent to it by other members of the namespace. This applies even to privileged processes and prevents a container process from accidentally killing PID 1. SIGKILL and SIGSTOP are exceptions when sent from an ancestor namespace: they are forcibly delivered and cannot be caught .
PID namespaces can be nested. Each namespace has a parent, except the initial root namespace, and the nesting depth is limited to 32 levels since Linux 3.7 . A process is visible to processes in its own namespace and in each direct ancestor namespace going back to the root. Processes in a child namespace cannot see processes in parent or ancestor namespaces . A process has one PID in each level of the hierarchy in which it is visible.
b. Network namespace: network stack isolation
The network namespace isolates all networking-related resources: network devices, IPv4 and IPv6 protocol stacks, IP routing tables, firewall rules, the /proc/net directory, /sys/class/net, files under /proc/sys/net, and port numbers . It also isolates the UNIX domain abstract socket namespace .
A physical network device can live in exactly one network namespace . When a network namespace is freed (when the last process in it terminates), its physical network devices are moved back to the initial network namespace, not to the namespace of the parent process . This behavior ensures that physical devices are not lost when a container exits.
Virtual ethernet (veth) pairs provide the mechanism for connecting network namespaces. A veth pair is a pipe-like abstraction with two ends, one in each namespace. Packets written to one end appear at the other. This is how Docker connects a container’s network namespace to the host’s network namespace, often through a bridge device . When a namespace is freed, the veth devices it contains are destroyed .
c. Mount namespace: filesystem view isolation
The mount namespace isolates mount points, giving each namespace its own view of the filesystem hierarchy . When a mount namespace is created, the new namespace receives a copy of the parent’s mount list. Subsequent mount and unmount operations in one namespace do not affect other namespaces unless propagation is explicitly configured.
Mount propagation types control how mount and unmount events spread between namespaces . MS_SHARED mounts share events with members of a peer group, so a mount under one propagates to all others in the group. MS_PRIVATE mounts do not propagate events in either direction. MS_SLAVE mounts receive events from a master peer group but do not send events back. MS_UNBINDABLE mounts are private and cannot be bind-mounted. Docker typically creates private mounts for containers so that container mounts do not leak into the host or other containers.
The mount namespace is what allows a container to have its own /proc, /sys, and /dev without conflicting with the host. The container’s root filesystem is assembled from image layers and made the root of the mount namespace. From inside the container, the host’s filesystem is invisible unless a specific mount is configured to expose it.
d. IPC namespace: inter-process communication isolation
The IPC namespace isolates System V IPC objects and POSIX message queues . System V IPC includes shared memory segments, semaphores, and message queues. POSIX message queues are a separate mechanism with a different API. These IPC objects are identified by keys or names rather than filesystem paths, which is why they need a separate namespace from the mount namespace.
Objects created in an IPC namespace are visible to all processes that are members of that namespace but are not visible to processes in other IPC namespaces . The /proc interfaces that report IPC limits and statistics are also distinct per IPC namespace: /proc/sys/fs/mqueue for POSIX message queues and /proc/sys/kernel and /proc/sysvipc for System V IPC .
When an IPC namespace is destroyed — when the last process in it terminates — all IPC objects in the namespace are automatically destroyed . This prevents orphaned shared memory segments and message queues from consuming resources after their creators have exited.
e. UTS namespace: hostname and domain isolation
The UTS namespace isolates two system identifiers: the hostname and the NIS domain name . These are set with sethostname() and setdomainname(), and retrieved with uname(), gethostname(), and getdomainname(). Changes made by a process in one UTS namespace are visible to all other processes in the same namespace but not to processes in other UTS namespaces .
When a new UTS namespace is created, the hostname and domain name are copied from the caller’s UTS namespace . A container can then change its hostname without affecting the host or other containers. This is why hostname inside a container returns the container’s name rather than the host’s name, even though the same kernel is running underneath.
The UTS namespace is one of the simpler namespaces, but it is important for applications that use the hostname for identification, logging, or service discovery. A container that reports the host’s hostname in its logs would be indistinguishable from the host itself.
f. User namespace: user and group ID mapping
The user namespace isolates user and group IDs . Processes in different user namespaces can have different mappings between the IDs they see and the IDs the kernel uses. The most common configuration maps root (UID 0) inside the container to an unprivileged UID (like 100000) on the host . Inside the container, the process believes it is root. Outside, the kernel sees an ordinary user.
Capabilities are also scoped to user namespaces. A process that has CAP_SYS_ADMIN inside a new user namespace does not have that capability in the host’s initial user namespace . This means a container can have capabilities that are normally restricted to privileged processes without granting those privileges on the host.
User namespaces are unusual in that creating one does not require the CAP_SYS_ADMIN capability. Since Linux 3.8, any user can create a user namespace . This allows unprivileged users to run containers, a feature that underpins rootless container runtimes. The other namespace types require CAP_SYS_ADMIN because creating them grants the creator the power to affect global resources visible to other processes.
Complete Example Session
# ============================================
# PART 1: VIEW NAMESPACES OF A PROCESS
# ============================================
# Every process has a /proc/PID/ns/ directory
# showing its namespace membership.
ls -l /proc/$$/ns
# Output shows symbolic links for each namespace:
# cgroup, ipc, mnt, net, pid, pid_for_children,
# time, user, uts
# ============================================
# PART 2: START A CONTAINER
# ============================================
# Run a container and find its host PID.
docker run -d --name ns-demo nginx
docker inspect -f '{{.State.Pid}}' ns-demo
# e.g., 12345
# ============================================
# PART 3: COMPARE HOST AND CONTAINER PID VIEW
# ============================================
# On the host, the container's init is PID 12345.
# Inside the container, it is PID 1.
ps -p 12345 -o pid,cmd
# Output: 12345 nginx: master process
docker exec ns-demo ps aux
# Output: PID 1 nginx: master process
# ============================================
# PART 4: INSPECT PID NAMESPACE LINKS
# ============================================
# The namespace symlink shows the inode number.
sudo ls -l /proc/12345/ns/pid
# Output: pid -> 'pid:[4026532300]'
docker exec ns-demo ls -l /proc/1/ns/pid
# Output: pid -> 'pid:[4026532300]'
# Same inode = same namespace.
# ============================================
# PART 5: INSPECT NETWORK NAMESPACE
# ============================================
# The container has its own network interfaces.
docker exec ns-demo ip addr
# Shows eth0 with its own IP, not the host's.
# On the host:
ip addr
# Does NOT show the container's eth0.
# ============================================
# PART 6: INSPECT MOUNT NAMESPACE
# ============================================
# The container sees its own root filesystem.
docker exec ns-demo ls /
# Shows container filesystem (bin, etc, usr, var...)
ls /
# On the host, shows the host's filesystem.
# ============================================
# PART 7: INSPECT IPC NAMESPACE
# ============================================
# IPC objects are isolated.
docker exec ns-demo ipcs
# Shows IPC objects in the container's namespace.
ipcs
# On the host, shows the host's IPC objects.
# ============================================
# PART 8: INSPECT UTS NAMESPACE
# ============================================
# The hostname is isolated.
docker exec ns-demo hostname
# Output: the container ID (short form)
hostname
# Output: the host's hostname
# ============================================
# PART 9: INSPECT USER NAMESPACE
# ============================================
# Check user namespace mapping (if enabled).
cat /proc/12345/uid_map
# Shows how container UIDs map to host UIDs.
docker exec ns-demo id
# Output: uid=0(root) gid=0(root)
# But on the host, the process runs as a
# different user if user namespaces are enabled.
# ============================================
# PART 10: COMPARE NAMESPACE INODES
# ============================================
# Same inode = same namespace.
sudo ls -l /proc/12345/ns/ | awk '{print $9, $10, $11}'
docker exec ns-demo ls -l /proc/1/ns/ | awk '{print $9, $10, $11}'
# Namespaces with the same inode are shared.
# Namespaces with different inodes are isolated.
These ten parts move from inspecting a process’s namespace links, through starting a container and comparing views, to examining each namespace type individually. The final comparison shows that the container’s namespaces have different inode numbers than the host’s, proving that Docker created new namespaces for the container.
Quick Reference
Six Namespaces and What They Isolate
| Namespace | Flag | Isolates | Kernel Config |
|---|---|---|---|
| PID | CLONE_NEWPID | Process IDs, process visibility | CONFIG_PID_NS |
| Network | CLONE_NEWNET | Network devices, stacks, ports, routes | CONFIG_NET_NS |
| Mount | CLONE_NEWNS | Mount points, filesystem view | (built-in) |
| IPC | CLONE_NEWIPC | System V IPC, POSIX message queues | CONFIG_IPC_NS |
| UTS | CLONE_NEWUTS | Hostname, NIS domain name | CONFIG_UTS_NS |
| User | CLONE_NEWUSER | User and group IDs, capabilities | (built-in) |
Namespace API System Calls
| Call | Purpose | Requires CAP_SYS_ADMIN |
|---|---|---|
clone(2) | Create process in new namespaces | Yes (except USER) |
unshare(2) | Move calling process to new namespaces | Yes (except USER) |
setns(2) | Join an existing namespace | Yes (except USER) |
Container View vs Host View
| Resource | Host View | Container View |
|---|---|---|
| PID | Container init is e.g. 12345 | Container init is 1 |
| Network | Host’s interfaces and routes | Container’s eth0 and routes |
| Filesystem | Host’s root, all mounts | Container’s root from image |
| Hostname | Host’s hostname | Container’s hostname |
| IPC | Host’s IPC objects | Container’s IPC objects |
| User | Real UIDs (e.g., 100000) | Mapped UIDs (e.g., root) |
Best Practices
✅ Do This:
docker run --pid=host nginx # Share host PID namespace (debug)
docker run --network=host nginx # Share host network namespace
docker run --ipc=host nginx # Share host IPC namespace
docker run --uts=host nginx # Share host UTS namespace
docker run --userns=host nginx # Disable user namespace mapping
docker run --rm alpine hostname # Verify UTS isolation
docker exec container ls /proc/1/ns/ # Inspect namespace inodes
❌ Don’t Do This:
docker run --privileged nginx # ❌ Disables all namespace isolation
docker run --pid=host nginx # ❌ For non-debug workloads (reduces security)
docker run --network=host nginx # ❌ When isolation is required
docker run --ipc=container:other nginx # ❌ Unnecessary namespace sharing
docker run --uts=host nginx # ❌ When container hostname matters
Common Pitfalls
| Pitfall | Why It Happens | Fix |
|---|---|---|
| Container PID 1 exits, all processes die | PID namespace init termination kills namespace | Ensure CMD runs a long-lived process |
| Cannot bind to port in container | Network namespace isolates ports | Publish port with -p host:container |
| Container cannot see host files | Mount namespace isolates filesystem | Use -v to mount host directories |
| Hostname changes not visible | UTS namespace isolates hostname | Set --hostname flag or accept container default |
| IPC shared memory not visible across containers | IPC namespace isolates objects | Use --ipc=container:<id> to share |
| User namespace not available | Kernel or runtime not configured | Check CONFIG_USER_NS and runtime support |
fork: Resource temporarily unavailable | PID namespace nesting limit reached | Reduce nesting depth (max 32) |
Real-World Examples
1. Share Network Namespace Between Containers
docker run -d --name web --network=container:proxy nginx
# web uses proxy's network namespace
2. Share IPC Namespace for Shared Memory
docker run -d --name producer --ipc=shared_mem producer
docker run -d --name consumer --ipc=container:producer consumer
3. Inspect Namespace Inode from Host
sudo ls -l /proc/$(docker inspect -f '{{.State.Pid}}' mycontainer)/ns/pid
4. Run Container with Custom Hostname
docker run --hostname web-01 nginx hostname
# Output: web-01
5. Enable User Namespace Mapping
dockerd --userns-remap=default
# Maps container root to an unprivileged host UID
6. Use Host Network for Performance
docker run --network=host nginx
# Container shares host network stack; no NAT overhead
7. Verify PID Namespace Isolation
docker exec mycontainer ps aux | head
# PID 1 is the container's init, not host init
8. Check Namespace Membership
readlink /proc/self/ns/net
# net:[4026531992]
9. List All Namespaces for a Container
sudo ls -l /proc/$(docker inspect -f '{{.State.Pid}}' mycontainer)/ns/
10. Compare Container and Host UTS
docker exec mycontainer hostname
hostname
# Different values prove UTS isolation
Visual
The Six Namespaces and Their Resources
┌──────────────────────────────────────────────────────────────┐
│ NAMESPACE TYPES AND ISOLATED RESOURCES │
│ │
│ PID (CLONE_NEWPID) │
│ └── Process IDs, process visibility │
│ │
│ NET (CLONE_NEWNET) │
│ └── Network devices, stacks, routes, ports, firewall │
│ │
│ MNT (CLONE_NEWNS) │
│ └── Mount points, filesystem hierarchy │
│ │
│ IPC (CLONE_NEWIPC) │
│ └── System V IPC, POSIX message queues │
│ │
│ UTS (CLONE_NEWUTS) │
│ └── Hostname, NIS domain name │
│ │
│ USER (CLONE_NEWUSER) │
│ └── User IDs, group IDs, capabilities │
│ │
│ Each namespace wraps a different category of global │
│ kernel resource in an isolated view. │
└──────────────────────────────────────────────────────────────┘
PID Namespace Hierarchy and Visibility
┌──────────────────────────────────────────────────────────────┐
│ PID NAMESPACE TREE AND PROCESS VISIBILITY │
│ │
│ ROOT PID NAMESPACE │
│ ┌────────────────────────────────────────────────────────┐ │
│ │ PID 100 (dockerd) │ │
│ │ PID 12345 (container init, real PID) │ │
│ │ PID 12346 (container process, real PID) │ │
│ └────────────────────────────────────────────────────────┘ │
│ │ │
│ │ clone(CLONE_NEWPID) │
│ ▼ │
│ CHILD PID NAMESPACE │
│ ┌────────────────────────────────────────────────────────┐ │
│ │ PID 1 (container init) ← sees only this namespace │ │
│ │ PID 2 (container process) │ │
│ │ (Host PID 100 is invisible to this namespace) │ │
│ └────────────────────────────────────────────────────────┘ │
│ │
│ Visibility rule: │
│ - A process can see processes in its own namespace │
│ and in all ancestor namespaces. │
│ - A process cannot see processes in descendant │
│ namespaces. │
│ - The root namespace can see all processes. │
└──────────────────────────────────────────────────────────────┘
Network Namespace and veth Pair
┌──────────────────────────────────────────────────────────────┐
│ CONNECTING NETWORK NAMESPACES WITH veth PAIRS │
│ │
│ HOST NETWORK NAMESPACE │
│ ┌────────────────────────────────────────────────────────┐ │
│ │ eth0 (physical) │ │
│ │ docker0 (bridge) │ │
│ │ veth1234 (one end of pair) │ │
│ └────────────────────────────────────────────────────────┘ │
│ │ │
│ veth pair (virtual cable) │
│ │ │
│ CONTAINER NETWORK NAMESPACE │
│ ┌────────────────────────────────────────────────────────┐ │
│ │ eth0 (other end of pair) │ │
│ │ IP: 172.17.0.2 │ │
│ │ Routes: default via 172.17.0.1 │ │
│ └────────────────────────────────────────────────────────┘ │
│ │
│ Packets written to one end appear at the other. │
│ Docker uses veth pairs to connect container network │
│ namespaces to the host bridge. │
└──────────────────────────────────────────────────────────────┘
User Namespace Mapping
┌──────────────────────────────────────────────────────────────┐
│ USER NAMESPACE: MAPPING CONTAINER ROOT TO HOST USER │
│ │
│ CONTAINER VIEW: │
│ ┌────────────────────────────────────────────────────────┐ │
│ │ uid=0 (root) │ │
│ │ Can do privileged operations within container │ │
│ └────────────────────────────────────────────────────────┘ │
│ │ │
│ │ uid_map: 0 100000 65536 │
│ ▼ │
│ HOST VIEW: │
│ ┌────────────────────────────────────────────────────────┐ │
│ │ uid=100000 (unprivileged) │ │
│ │ Cannot affect host-level resources │ │
│ └────────────────────────────────────────────────────────┘ │
│ │
│ The container process believes it is root. │
│ The host kernel sees an unprivileged user. │
│ Capabilities inside the namespace do not apply outside. │
└──────────────────────────────────────────────────────────────┘
Summary
| Item | Value |
|---|---|
| Namespace definition | Kernel feature that wraps global resource in isolated view |
| PID namespace | Isolates process IDs; init at PID 1; nesting depth 32 |
| Network namespace | Isolates network devices, stacks, routes, ports, firewall |
| Mount namespace | Isolates mount points; propagation types control sharing |
| IPC namespace | Isolates System V IPC and POSIX message queues |
| UTS namespace | Isolates hostname and NIS domain name |
| User namespace | Isolates UID/GID mapping; no privilege required to create |
| Creation API | clone(), unshare(), setns() |
| Privilege required | CAP_SYS_ADMIN for all except user namespace |
| Visibility rule | Process sees own namespace and ancestors, not descendants |
Key takeaways:
- Namespaces partition kernel resources into isolated views. Each container gets its own process ID space, network stack, filesystem mounts, IPC objects, hostname, and user mapping.
- The PID namespace makes containers feel like standalone systems. Process IDs start at 1, and init termination kills the entire namespace .
- The network namespace gives each container its own network identity. Virtual ethernet pairs connect container namespaces to the host or to each other .
- The mount namespace controls what filesystem the container sees. Docker assembles the container’s root from image layers and makes it the root of the mount namespace.
- The IPC namespace prevents collisions between containers. Shared memory segments and message queues are isolated per namespace .
- The UTS namespace allows each container its own hostname. Without it, all containers and the host would share the same identity .
- The user namespace maps container root to an unprivileged host user. This is the primary security mechanism for running containers as “root” safely .
- Namespaces are the first half of container isolation. Cgroups provide the second half by limiting how much of each resource a container can use.
Remember: Namespaces are Linux’s native mechanism for partitioning global resources. Docker creates a new set of namespaces for every container, so each container sees its own processes, network, filesystem, IPC objects, hostname, and user mapping. The host kernel is still shared, but the view each container has is isolated. The PID namespace makes containers feel like independent systems, with init at PID 1 and its termination signaling the end of the namespace. The network namespace gives each container its own IP stack, connected to the host through veth pairs. The mount namespace provides the container’s filesystem view, assembled from image layers. The IPC, UTS, and user namespaces complete the isolation by partitioning inter-process communication, hostname identity, and user privileges. Together with cgroups, which limit resource usage, namespaces form the foundation on which Docker builds its container abstraction.
Stop using slow, ad-bloated tool sites! 🤮
🔎 Search “KandZ Tools” on Google to use many professional utilities for free.
KandZ.me is the ultimate minimalist hub for:
✅ Finance (Mortgage, Interest, Inflation)
✅ Tech (Base64, JSON, Dev Suite, IP)
✅ Health (BMI, BMR, TDEE)
✅ Productivity (Timer, Workspace, QR)
⚡️ Fast & Private
🔒 No data leaves your device
💎 100% Free
🔗 Use it now: https://tools.kandz.me
🔖 Bookmark it—you’ll need it later!