| |

Docker 4 🐳 Resource Control and Limits: Linux Control Groups (cgroups v1 vs cgroups v2)

Namespaces isolate what a container can see. Control groups determine how much of each resource it can use. These two kernel features are complementary and together they form the complete foundation of containerization. Without cgroups, a single container running a runaway process could consume all available memory on the host, saturate every CPU core, or exhaust the disk I/O bandwidth, starving other containers and the host system itself. Cgroups provide the enforcement mechanism that makes resource guarantees possible in multi-tenant environments.

The control group subsystem has undergone a major revision. cgroups v1, the original implementation, organized resource controllers into separate hierarchies, one per resource type. This design worked but created inconsistencies — most notably, buffered disk writes were not properly attributed to the container that generated them, allowing a single workload to saturate the host’s storage without being throttled. cgroups v2 addresses these problems with a unified hierarchy where every process sits in exactly one cgroup for all controllers, enabling correct cross-resource accounting and more sophisticated control policies.

This chapter covers the architecture of both cgroups versions, the major differences in their interfaces and behavior, how Docker uses cgroups to enforce resource limits, and the current state of migration as modern distributions move to v2 by default. We will examine the practical commands Docker exposes for setting limits and the underlying files that carry those limits to the kernel.

Key point: cgroups v2 replaces v1’s separate per-controller hierarchies with a single unified tree, fixing buffered I/O accounting and adding graduated memory controls, pressure metrics, and safe delegation to unprivileged users.


Why cgroups v2 exists

The buffered I/O accounting problem. In cgroups v1, block I/O limits were enforced by the blkio controller. But the kernel’s page cache writeback happens asynchronously by kernel threads, not by the process that generated the writes. When a container wrote data that was buffered in the page cache, those writes were issued later by kernel threads that were not associated with the container’s cgroup. The I/O limits were bypassed entirely. A container could dirty gigabytes of page cache and saturate the disk, and the blkio controller would see none of it. cgroups v2 fixes this by charging writeback to the cgroup that dirtied the page cache, so io.max and io.weight limits actually apply to buffered writes.

The hierarchy fragmentation problem. In cgroups v1, each resource controller had its own independent tree under /sys/fs/cgroup/<controller>/. A process could be in one cgroup for CPU, a different cgroup for memory, and yet another for I/O. Controllers could not cooperate, and there was no unified view of a process’s resource constraints. cgroups v2 replaces this with a single tree at /sys/fs/cgroup, where every process belongs to exactly one cgroup and all controllers operate on that same membership. This simplifies management and eliminates the possibility of inconsistent placements.

The memory control granularity problem. cgroups v1 offered a hard memory limit (memory.limit_in_bytes) and a weak soft limit that was not reliably enforced. When a container hit the hard limit, the OOM killer triggered immediately. There was no effective intermediate state where the kernel could aggressively reclaim memory to keep the container alive while slowing its allocation rate. cgroups v2 introduces memory.high, a throttling threshold that triggers aggressive reclaim and slows the process before the hard limit is reached. This graduated model — memory.min, memory.low, memory.high, memory.max — provides much finer control over memory behavior.

The pressure visibility problem. In cgroups v1, there was no way to know how much a container’s processes were stalled waiting for resources. You could see that CPU was being consumed or that memory was full, but not that a process was blocked and unable to make progress. cgroups v2 adds Pressure Stall Information (PSI), exposing per-cgroup metrics for CPU, memory, and I/O pressure. These metrics report the percentage of time that at least one task (or all tasks) were stalled waiting for the resource, with averages over 10-second, 60-second, and 300-second windows. PSI enables monitoring systems to detect contention before it causes failures.

The delegation problem. In cgroups v1, safely delegating control of a cgroup subtree to an unprivileged user was difficult and error-prone. This blocked rootless containers, where a non-root user runs a container runtime and manages its own resource limits. cgroups v2 was designed from the start with safe delegation in mind, using cgroup.subtree_control to explicitly grant control of specific controllers to child cgroups. Rootless Docker and Podman rely on this capability to enforce resource limits without requiring root privileges.


a. cgroups v1 architecture: separate hierarchies

In cgroups v1, each resource controller has its own independent directory hierarchy. The standard mount point is /sys/fs/cgroup, and under it you find subdirectories for each controller: cpu/, memory/, blkio/, pids/, cpuset/, and others. A process can be a member of different cgroups in different hierarchies. For example, a container process might be in /sys/fs/cgroup/cpu/docker/abc123 for CPU limits and /sys/fs/cgroup/memory/docker/abc123 for memory limits.

The interface files in each controller reflect the v1 design. For CPU limits, you write to cpu.cfs_quota_us and cpu.cfs_period_us. The quota is the amount of CPU time (in microseconds) the cgroup can consume in each period. A quota of 50000 with a period of 100000 means 50% of one CPU core. For relative CPU priority, you set cpu.shares, where the default is 1024 and higher values mean more CPU time relative to other cgroups.

Memory limits in v1 use memory.limit_in_bytes for the hard limit and memory.memsw.limit_in_bytes for memory plus swap. There is no effective soft limit, and OOM behavior kills individual processes without group coordination. The tasks file allows moving individual threads between cgroups, a feature that caused problems and was removed in v2.

b. cgroups v2 architecture: unified hierarchy

cgroups v2 uses a single hierarchy at /sys/fs/cgroup. Every process belongs to exactly one cgroup, and all controllers operate on that membership. Controllers are enabled per subtree by writing controller names to cgroup.subtree_control. For example, writing +memory +cpu enables those controllers for all child cgroups.

The “no internal processes” rule is a key v2 constraint. A non-root cgroup cannot both have member processes and distribute resources to child cgroups. If you want to create a hierarchy of cgroups with resource limits, the parent cgroup must have no processes directly in it; all processes go into leaf cgroups. This rule exists to prevent the ambiguous situation where a cgroup’s own processes compete with its children for the resources it is supposed to distribute.

The CPU controller in v2 uses cpu.weight instead of cpu.shares. The range is 1 to 10000 with a default of 100, which is a more intuitive scale than v1’s 2 to 262144. For hard CPU limits, v2 uses cpu.max, a single file containing both quota and period separated by a space. Writing 50000 100000 to cpu.max means 50% of one CPU core. The value max in the quota position means no limit.

c. Memory controls in cgroups v2

The memory controller in v2 provides a graduated set of controls rather than the single hard limit of v1. memory.min is a hard protection: memory below this threshold is never reclaimed, even under severe pressure. memory.low is a soft protection: memory below this threshold is protected from reclaim as long as reclaimable memory exists elsewhere. memory.high is a throttling threshold: when usage exceeds it, the kernel aggressively reclaims memory and slows the allocating process, but does not kill it. memory.max is the hard limit: reaching it triggers the OOM killer.

This graduated model allows workloads to be slowed rather than killed when they exceed their expected memory usage. A container with memory.high set to 80% of memory.max will experience reclaim pressure and reduced allocation speed as it approaches the throttling threshold, potentially avoiding the OOM kill entirely. Monitoring systems should watch memory.high throttling events and PSI memory pressure rather than only OOM counts, because a workload that is constantly throttled is still degraded even if it never gets killed.

cgroups v2 also adds memory.oom.group, which allows the entire cgroup to be killed together when an OOM occurs, rather than killing individual processes. This is important for containers where killing one process might leave the container in an inconsistent state.

d. Docker resource limits with cgroups

Docker exposes resource limits through command-line flags that translate into cgroup settings. The --memory flag sets the hard memory limit, which Docker writes to memory.max on cgroups v2 systems. The --cpus flag sets a CPU quota; --cpus=0.5 becomes 50000 100000 in cpu.max, meaning 50% of one core. The --cpu-shares flag sets relative CPU weight, using v1’s range and values but being converted to v2’s cpu.weight by the container runtime.

Docker creates a cgroup scope for each container. On a systemd-based system using the systemd cgroup driver, this appears as /sys/fs/cgroup/system.slice/docker-<container-id>.scope/. You can verify the applied limits by reading the files in that directory. The --pids-limit flag sets pids.max, preventing fork bombs and runaway process creation. The --blkio-weight flag sets I/O priority in the range 10 to 1000.

For CPU sets, --cpuset-cpus restricts the container to specific host CPU cores. When both --cpus and --cpuset-cpus are used, the effective CPU availability is the minimum of the two limits. Docker reports the available CPU count based on the tightest constraint.


Complete Example Session

# ============================================
# PART 1: CHECK CGROUP VERSION
# ============================================
# The filesystem type reveals which version is active.

stat -fc %T /sys/fs/cgroup/
# cgroup2fs → v2
# tmpfs → v1

mount | grep cgroup
# v2: one line: cgroup2 on /sys/fs/cgroup type cgroup2
# v1: multiple lines, one per controller
# ============================================
# PART 2: RUN A CONTAINER WITH MEMORY LIMIT
# ============================================
# Limit the container to 256 MB of memory.

docker run -d --memory=256m --name memtest nginx:1.25-alpine
# ============================================
# PART 3: VERIFY MEMORY LIMIT IN CGROUPS
# ============================================
# Find the container's cgroup and read memory.max.

CID=$(docker inspect --format '{{.Id}}' memtest)
cat /sys/fs/cgroup/system.slice/docker-${CID}.scope/memory.max
# Output: 268435456  (256 MB in bytes)
# ============================================
# PART 4: RUN A CONTAINER WITH CPU LIMIT
# ============================================
# Limit the container to 50% of one CPU core.

docker run -d --cpus=0.5 --name cputest nginx:1.25-alpine
# ============================================
# PART 5: VERIFY CPU LIMIT IN CGROUPS
# ============================================
# Read cpu.max to see the quota and period.

CID=$(docker inspect --format '{{.Id}}' cputest)
cat /sys/fs/cgroup/system.slice/docker-${CID}.scope/cpu.max
# Output: 50000 100000
# 50000 / 100000 = 0.5 = 50% of one core
# ============================================
# PART 6: RUN A CONTAINER WITH PID LIMIT
# ============================================
# Prevent fork bombs by limiting process count.

docker run -d --pids-limit=100 --name pidtest nginx:1.25-alpine
# ============================================
# PART 7: VERIFY PID LIMIT IN CGROUPS
# ============================================
# Read pids.max to see the process limit.

CID=$(docker inspect --format '{{.Id}}' pidtest)
cat /sys/fs/cgroup/system.slice/docker-${CID}.scope/pids.max
# Output: 100
# ============================================
# PART 8: CHECK CPU WEIGHT (v2)
# ============================================
# --cpu-shares values are converted to cpu.weight.

docker run -d --cpu-shares=512 --name weighttest nginx:1.25-alpine
CID=$(docker inspect --format '{{.Id}}' weighttest)
cat /sys/fs/cgroup/system.slice/docker-${CID}.scope/cpu.weight
# Output: a value in the range 1-10000
# ============================================
# PART 9: MONITOR PSI METRICS (v2)
# ============================================
# Pressure Stall Information shows resource contention.

cat /sys/fs/cgroup/system.slice/docker-${CID}.scope/memory.pressure
# Output: some avg10=0.00 avg60=0.00 avg300=0.00 total=0
# ============================================
# PART 10: CLEAN UP
# ============================================
# Remove the test containers.

docker rm -f memtest cputest pidtest weighttest

These ten parts move from identifying the cgroup version through applying and verifying memory, CPU, PID, and CPU weight limits, to reading PSI metrics and cleaning up. The commands demonstrate that Docker flags translate directly into cgroup filesystem settings that you can inspect and verify.


Quick Reference

cgroups v1 vs v2 Comparison

Aspectcgroups v1cgroups v2
HierarchyOne tree per controllerSingle unified tree
Mount point/sys/fs/cgroup/<controller>//sys/fs/cgroup/
Buffered I/ONot attributed to writerCharged to owning cgroup
CPU weightcpu.shares (2–262144, default 1024)cpu.weight (1–10000, default 100)
CPU quotacpu.cfs_quota_us + cpu.cfs_period_uscpu.max (single file)
Memory hard limitmemory.limit_in_bytesmemory.max
Memory soft limitWeak, unreliablememory.high (throttle)
OOM group killNot supportedmemory.oom.group
Pressure metricsNonePSI (cpu/memory/io.pressure)
DelegationUnsafe for unprivileged usersDesigned for safe delegation

Docker Resource Flags

FlagResourcecgroups v2 File
--memoryMemory hard limitmemory.max
--memory-swapMemory + swapmemory.swap.max
--cpusCPU quotacpu.max
--cpu-sharesCPU weightcpu.weight (converted)
--cpuset-cpusCPU affinitycpuset.cpus
--pids-limitProcess countpids.max
--blkio-weightI/O priorityio.weight

PSI Metrics

FileReports
cpu.pressureCPU contention
memory.pressureMemory contention
io.pressureI/O contention

Best Practices

✅ Do This:

docker run --memory=512m --cpus=1.0 nginx           # Set explicit limits
docker run --pids-limit=100 nginx                    # Prevent fork bombs
docker run --memory=512m --memory-swap=512m nginx   # Disable swap
stat -fc %T /sys/fs/cgroup/                         # Check cgroup version
cat /sys/fs/cgroup/.../memory.pressure              # Monitor PSI

❌ Don’t Do This:

docker run nginx                                    # ❌ No limits; container can starve host
docker run --memory=0 nginx                         # ❌ Zero means unlimited, not "no memory"
docker run --cpus=0 nginx                           # ❌ Invalid; use fractional or integer
docker run --pids-limit=-1 nginx                    # ❌ Unlimited; fork bomb risk

Common Pitfalls

PitfallWhy It HappensFix
Container OOM killed despite limitSwap not limitedUse --memory-swap equal to --memory
CPU limit not enforcedcgroups v1 with old runtimeUpdate to cgroups v2 or newer Docker
Memory throttling before OOMmemory.high configuredAdjust memory.high or monitor PSI
Buffered writes bypass limitsRunning on cgroups v1Migrate to cgroups v2
Old JVM sees host memoryJVM not cgroup-awareUse JDK 11.0.16+, 17+, or set -XX:MaxRAMPercentage
pids.max not enforcedcgroups v1 pids controller not mountedEnable pids controller or use v2

Real-World Examples

1. Web Server with Memory and CPU Limits

docker run -d --memory=1g --cpus=2 -p 80:80 nginx

2. Database with Swap Disabled

docker run -d --memory=4g --memory-swap=4g postgres:16

3. Build Container with PID Limit

docker run --rm --pids-limit=500 -v $(pwd):/src golang:1.22 go build

4. CPU Pinning for Latency-Sensitive Workload

docker run -d --cpuset-cpus="0-3" --cpus=2 redis:7

5. Check Memory Throttling

cat /sys/fs/cgroup/system.slice/docker-<id>.scope/memory.events

6. Monitor CPU Pressure

cat /sys/fs/cgroup/system.slice/docker-<id>.scope/cpu.pressure

7. Verify OOM Group Configuration

docker run -d --memory=256m --oom-kill-disable nginx  # Not recommended

8. cgroups v1 Memory Limit

cat /sys/fs/cgroup/memory/docker/<id>/memory.limit_in_bytes

9. Check cgroup Driver

docker info | grep -i cgroup

10. Enable cgroups v2 on Ubuntu

# /etc/default/grub: GRUB_CMDLINE_LINUX="systemd.unified_cgroup_hierarchy=1"
sudo update-grub && reboot

Visual

cgroups v1 vs v2 Hierarchy

┌──────────────────────────────────────────────────────────────┐
│  cgroups v1: SEPARATE TREES PER CONTROLLER                   │
│                                                              │
│  /sys/fs/cgroup/                                             │
│  ├── cpu/                                                    │
│  │   └── docker/                                             │
│  │       └── <container-id>/   (cpu.shares, cpu.cfs_quota)   │
│  ├── memory/                                                 │
│  │   └── docker/                                             │
│  │       └── <container-id>/   (memory.limit_in_bytes)       │
│  └── blkio/                                                  │
│      └── docker/                                             │
│          └── <container-id>/   (blkio.throttle.*)            │
│                                                              │
│  A process can be in different cgroups in each tree.         │
│  Controllers cannot cooperate.                               │
│                                                              │
│  ─────────────────────────────────────────                   │
│                                                              │
│  cgroups v2: SINGLE UNIFIED TREE                             │
│                                                              │
│  /sys/fs/cgroup/                                             │
│  └── system.slice/                                           │
│      └── docker-<container-id>.scope/                        │
│          ├── cpu.max                                         │
│          ├── cpu.weight                                      │
│          ├── memory.max                                      │
│          ├── memory.high                                     │
│          ├── memory.pressure                                 │
│          ├── pids.max                                        │
│          └── io.max                                          │
│                                                              │
│  One cgroup per container, all controllers together.         │
│  Cross-resource accounting is correct.                       │
└──────────────────────────────────────────────────────────────┘

Memory Control Graduation in cgroups v2

┌──────────────────────────────────────────────────────────────┐
│  MEMORY CONTROLS: FROM PROTECTION TO KILL                    │
│                                                              │
│  memory.min ───────────────────────────────────────────────  │
│  │  Hard protection: never reclaimed                         │
│  │                                                           │
│  memory.low ───────────────────────────────────────────────  │
│  │  Soft protection: reclaimed only if no other memory       │
│  │                                                           │
│  memory.high ──────────────────────────────────────────────  │
│  │  Throttling: aggressive reclaim, slows allocation         │
│  │  Watch this threshold; throttling means degradation        │
│  │                                                           │
│  memory.max ───────────────────────────────────────────────  │
│  │  Hard limit: OOM killer triggers                          │
│  │                                                           │
│  ─────────────────────────────────────────                   │
│                                                              │
│  v1 only had a hard limit (memory.limit_in_bytes).           │
│  v2 provides graduated control for graceful degradation.     │
└──────────────────────────────────────────────────────────────┘

CPU Quota and Weight

┌──────────────────────────────────────────────────────────────┐
│  CPU LIMITS IN cgroups v2                                    │
│                                                              │
│  cpu.max = "50000 100000"                                    │
│  ├── quota: 50000 microseconds                               │
│  ├── period: 100000 microseconds                             │
│  └── 50000/100000 = 0.5 = 50% of one CPU core                │
│                                                              │
│  cpu.weight = 100 (default)                                  │
│  ├── range: 1 to 10000                                       │
│  ├── higher = more CPU time when contended                   │
│  └── only matters when CPU is saturated                      │
│                                                              │
│  Relationship:                                               │
│  - cpu.max enforces a hard ceiling                           │
│  - cpu.weight distributes available CPU among competing      │
│    cgroups when the ceiling is not reached                   │
│                                                              │
│  Docker --cpus=0.5 → cpu.max = 50000 100000                  │
│  Docker --cpu-shares=1024 → cpu.weight = 100                 │
└──────────────────────────────────────────────────────────────┘

Docker Flag to cgroup File Mapping

┌──────────────────────────────────────────────────────────────┐
│  HOW DOCKER FLAGS BECOME CGROUP SETTINGS                     │
│                                                              │
│  docker run --memory=256m                                    │
│       │                                                      │
│       ▼                                                      │
│  /sys/fs/cgroup/system.slice/docker-<id>.scope/memory.max    │
│  Value: 268435456                                            │
│                                                              │
│  docker run --cpus=0.5                                       │
│       │                                                      │
│       ▼                                                      │
│  /sys/fs/cgroup/system.slice/docker-<id>.scope/cpu.max       │
│  Value: 50000 100000                                         │
│                                                              │
│  docker run --pids-limit=100                                 │
│       │                                                      │
│       ▼                                                      │
│  /sys/fs/cgroup/system.slice/docker-<id>.scope/pids.max      │
│  Value: 100                                                  │
│                                                              │
│  You can verify any limit by reading the corresponding       │
│  file in the container's cgroup directory.                   │
└──────────────────────────────────────────────────────────────┘

Summary

ItemValue
cgroups definitionKernel feature for limiting, accounting, and isolating resource usage
cgroups v1Separate hierarchy per controller; buffered I/O bypassed limits
cgroups v2Unified single hierarchy; correct I/O accounting
Memory controls v2memory.min, low, high, max (graduated)
CPU controls v2cpu.max (quota/period), cpu.weight (1–10000)
PSIPer-cgroup pressure metrics for CPU, memory, I/O
DelegationSafe in v2; enables rootless containers
Docker memory flag--memory → memory.max
Docker CPU flag--cpus → cpu.max
Docker PID flag--pids-limit → pids.max
Modern defaultsRHEL 9+, Ubuntu 21.10+, Debian 11+, Fedora 31+ use v2

Key takeaways:

  • cgroups control how much, namespaces control what. Together they provide complete container isolation. Cgroups enforce CPU, memory, I/O, and PID limits that prevent one container from starving others.
  • cgroups v2 fixes the buffered I/O accounting bug. In v1, page cache writeback was attributed to kernel threads, not the container, so I/O limits were bypassed. V2 charges writeback to the cgroup that dirtied the page.
  • The unified hierarchy simplifies management. One cgroup per container, all controllers together, eliminates the inconsistency of separate trees.
  • Memory controls are graduated in v2. memory.high throttles before memory.max kills, allowing graceful degradation rather than abrupt OOM.
  • PSI provides visibility into contention. Pressure metrics reveal when processes are stalled waiting for resources, enabling proactive monitoring.
  • Docker flags map directly to cgroup files. You can verify any limit by reading the corresponding file in the container’s cgroup directory.
  • cgroups v2 is the modern default. RHEL 9+, Ubuntu 21.10+, Debian 11+, and Fedora 31+ ship with v2 enabled.
  • Rootless containers depend on v2 delegation. The safe delegation model in v2 enables unprivileged users to run containers with enforced resource limits.

Remember: Namespaces and cgroups are the two pillars of containerization. Namespaces give each container its own isolated view of processes, network, filesystem, and identity. Cgroups enforce how much CPU, memory, disk I/O, and process count that container can consume. The transition from cgroups v1 to v2 is not a cosmetic change — it fixes fundamental accounting problems, introduces graduated memory controls, adds pressure metrics, and enables safe delegation. When you set --memory=256m or --cpus=0.5 on a Docker container, you are writing values to files in /sys/fs/cgroup that the kernel reads to enforce limits. Understanding this mapping lets you verify limits directly, diagnose resource contention, and reason about why a container behaves the way it does.



Stop using slow, ad-bloated tool sites! 🤮

🔎 Search “KandZ Tools” on Google to use many professional utilities for free.

KandZ.me is the ultimate minimalist hub for:
✅ Finance (Mortgage, Interest, Inflation)
✅ Tech (Base64, JSON, Dev Suite, IP)
✅ Health (BMI, BMR, TDEE)
✅ Productivity (Timer, Workspace, QR)

⚡️ Fast & Private
🔒 No data leaves your device
💎 100% Free

🔗 Use it now: https://tools.kandz.me
🔖 Bookmark it—you’ll need it later!