| | |

Docker 10 🐳 Understanding Container File Systems: Read-Only Layers, Copy-on-Write (CoW), and Ephemeral Storage

A container’s filesystem is not a single block of disk space. It is a stack of read-only image layers with a thin writable layer placed on top when the container starts. Every file the image contains lives in the read-only layers, and every change the running container makes—writing a log, modifying a config, creating a temporary file—is stored in the writable layer through a mechanism called copy-on-write. This architecture is what makes containers lightweight and fast to start, and it is also what makes container storage ephemeral by default: when the container is removed, the writable layer is removed with it, and any data that was not written to a volume or bind mount is lost .

Understanding this model is the difference between using containers effectively and losing data unexpectedly. The read-only layers are shared between every container based on the same image, which saves disk space and memory. The writable layer is unique to each container, which means two containers from the same image can each write to their own /app/config.json without interfering with each other. But the writable layer is tied to the container’s lifecycle, and that lifecycle is short. This chapter covers the layer stack, the copy-on-write operation, the difference between the writable layer and volumes, and the patterns for deciding where data should live.

Key point: A container’s filesystem is a stack of read-only image layers plus a thin writable layer. Modifications trigger a copy-on-write operation: the file is copied from the read-only layer to the writable layer, then modified. The writable layer is unique per container and is destroyed with the container. Data that must persist belongs in a volume or bind mount, not in the writable layer.


Why the layered filesystem exists

The image-sharing problem. If every container from an image required a full copy of the image’s filesystem, running ten containers would consume ten times the disk space of the image. The layered model eliminates this duplication: all containers from the same image share the read-only layers, and each container only adds the small writable layer for its own changes . This is the same principle that makes virtual machine disk images inefficient by comparison—VMs often copy entire disk images per instance .

The immutability problem. An image should not change after it is built. The read-only layers enforce this: the container cannot modify them, so the image’s base filesystem remains consistent across every container that uses it. The writable layer is the only place where changes occur, and it is isolated per container .

The fast-startup problem. Starting a container does not require copying the image’s files. Docker only needs to create the thin writable layer and mount the stack. This is why containers start in milliseconds while virtual machines take seconds or minutes .

The write-performance problem. The copy-on-write mechanism adds overhead the first time a file is modified, because the entire file must be copied from the read-only layer to the writable layer before the change can be written . This overhead is a trade-off for the space savings, and it is why write-heavy workloads should use volumes instead of relying on the writable layer .

The data-persistence problem. The writable layer is destroyed when the container is removed. This is by design: containers are meant to be disposable, and their state should live in volumes or external storage. Data written to the writable layer is ephemeral, and understanding this prevents the common mistake of storing database files or user uploads in the container’s filesystem .


a. The layer stack

A Docker image is composed of a series of layers. Each layer represents a set of filesystem changes relative to its parent layer—files added, modified, or deleted . When you build an image with a Dockerfile, each instruction (FROM, RUN, COPY, ADD) creates a new layer, except for metadata-only instructions like CMD or ENV .

When a container starts, Docker adds a writable layer on top of the image’s read-only layers. The stack looks like this:

┌─────────────────────────────────────────┐
│  Writable Layer (unique per container)  │  ← Container writes here
├─────────────────────────────────────────┤
│  Layer N: RUN npm install               │  ← Read-only
├─────────────────────────────────────────┤
│  Layer 2: COPY app /app                 │  ← Read-only
├─────────────────────────────────────────┤
│  Layer 1: FROM node:20-alpine           │  ← Read-only
└─────────────────────────────────────────┘

All containers based on the same image share layers 1 through N. Only the writable layer is unique . If you run five containers from the same image and none of them write to disk, the additional storage used is negligible—just the writable layer’s metadata .

The storage driver implements this stack. On modern Linux systems, overlay2 is the default driver. It uses OverlayFS to combine the layers into a single unified view called the merged directory, which is the container’s mount point .


b. Copy-on-write: reading and modifying files

Copy-on-write is the strategy that makes the layered filesystem efficient. The behavior depends on whether the container is reading a file or modifying it.

Reading a file that exists only in the read-only layers. The container reads directly from the image layer. No copy is made, and the overhead is minimal .

Reading a file that exists in both the read-only and writable layers. The writable layer’s version takes precedence. The file in the read-only layer is obscured .

Modifying a file for the first time. This triggers the copy_up operation. The storage driver locates the file in the read-only layers, copies it into the writable layer, and then applies the modification. From that point on, the container sees the writable layer’s copy, and the read-only layer’s version is invisible .

The copy_up operation works at the file level, not the block level. Even if only a single byte of a large file is changed, the entire file is copied . This is why write-heavy workloads can consume significant space in the writable layer: every modification to a large file duplicates it.

Deleting a file. The read-only layer cannot be modified, so deletion creates a whiteout file in the writable layer. The whiteout marks the file as deleted in the unified view, effectively hiding it from the container without removing it from the image .

Creating a new file. New files are created directly in the writable layer. No copy-up is needed because the file does not exist in the read-only layers.


c. The writable layer is ephemeral

The writable layer is tied to the container’s lifecycle. It is created when the container is created and destroyed when the container is removed. This is the defining characteristic of container storage .

Stopping a container does not destroy the writable layer. The container can be restarted, and its changes are still present. A stopped container retains its writable layer until it is removed .

Removing a container (docker rm) destroys the writable layer. Any data written only to the writable layer is gone. There is no recovery mechanism—the layer is deleted, not archived .

Recreating a container from the same image produces a new writable layer. The old layer is not reused. A new container with the same name does not inherit the previous container’s writable layer .

This is why the writable layer is often called “ephemeral storage.” It exists for the container’s lifetime and no longer. The Docker documentation is explicit: “Data written to the container layer doesn’t persist when the container is destroyed” .


d. Volumes and bind mounts: where persistent data lives

Data that must survive the container’s removal belongs in a volume or a bind mount, not in the writable layer.

Named volumes are managed by Docker and stored in /var/lib/docker/volumes/ on Linux. They are independent of the container’s lifecycle: removing the container does not remove the volume. A new container can mount the same volume and see the data .

docker run -d --name db -v pgdata:/var/lib/postgresql/data postgres:16

The volume pgdata persists even after docker rm db. It must be explicitly removed with docker volume rm pgdata .

Bind mounts map a host directory into the container. The data lives on the host filesystem, outside Docker’s management. Removing the container does not affect the host directory .

docker run -d --name web -v /srv/html:/usr/share/nginx/html nginx

The host’s /srv/html directory is the source of truth. The container sees it and can write to it, but the data’s lifecycle is the host’s, not the container’s .

Anonymous volumes are created automatically when an image declares a VOLUME instruction and the user does not specify a mount. They persist after container removal but have no name—only a random ID—making them difficult to manage and easy to orphan .

The distinction between the writable layer and volumes is not about performance or size. It is about lifecycle ownership. The writable layer is owned by the container. A volume is owned by Docker. A bind mount is owned by the host .


e. Performance implications

The copy-on-write mechanism has performance costs that vary by storage driver.

overlay2 performs a file-level copy-up. Large files and deep directory trees make the first write to a file more expensive. Subsequent writes to the same file are fast because the file is already in the writable layer .

vfs does not support copy-on-write at all. Each layer is a full deep copy of its parent, which means every container from the same image consumes the full size of the image. It is stable and works everywhere but is the least efficient driver .

Write-heavy workloads suffer the most from copy-on-write. A database that constantly modifies data files will trigger copy-up for each file it touches, duplicating them in the writable layer. This is why databases should always use volumes, not the writable layer . Volumes bypass the storage driver entirely and write directly to the host filesystem .

Read-heavy workloads are not affected. Reading from the read-only layers incurs minimal overhead, and page cache sharing means multiple containers reading the same file share a single cached copy in memory .


f. Practical patterns for where data lives

The decision of where to store data is a decision about lifecycle. Four options exist, each with a different owner and persistence guarantee.

Storage LocationOwnerPersists After docker rm
Writable layerContainerNo
Named volumeDockerYes
Bind mountHostYes
tmpfsMemoryNo

Use the writable layer for data that is truly disposable: logs that are also shipped elsewhere, package manager caches, temporary build artifacts, and anything that can be regenerated on startup. The writable layer is the right place for nothing important .

Use a named volume for data that must survive container removal and is managed by Docker: database files, application state, uploaded content that should persist. Volumes are Docker’s recommended solution for persistent data .

Use a bind mount for data that the host must control directly: source code during development, configuration files managed by an external system, logs that a host process reads. Bind mounts give the host full control over the data’s lifecycle .

Use tmpfs for sensitive data that should never touch disk: session tokens, temporary encryption keys, scratch space for memory-intensive operations. tmpfs is stored in memory and disappears when the container stops .

The Docker documentation’s recommendation is direct: “Use volumes for write-heavy applications. Don’t store the data in the container for write-heavy applications” . This is not just about persistence—it is also about performance, because volumes avoid the copy-on-write overhead entirely .


Complete Example Session

# ============================================
# PART 1: RUN A CONTAINER WITH NO VOLUME
# ============================================
# All writes go to the writable layer.

docker run -d --name ephemeral alpine sleep 3600
docker exec ephemeral sh -c "echo 'data' > /tmp/test.txt"
# ============================================
# PART 2: VERIFY DATA IN WRITABLE LAYER
# ============================================
docker exec ephemeral cat /tmp/test.txt
# data
# ============================================
# PART 3: REMOVE AND RECREATE
# ============================================
# The data is gone because the writable layer
# was destroyed with the container.

docker rm -f ephemeral
docker run -d --name ephemeral alpine sleep 3600
docker exec ephemeral cat /tmp/test.txt
# cat: /tmp/test.txt: No such file or directory
# ============================================
# PART 4: RUN WITH A NAMED VOLUME
# ============================================
# Data written to the volume survives removal.

docker run -d --name persisted -v mydata:/data alpine sleep 3600
docker exec persisted sh -c "echo 'persistent' > /data/test.txt"
# ============================================
# PART 5: REMOVE AND RECREATE WITH VOLUME
# ============================================
docker rm -f persisted
docker run -d --name persisted2 -v mydata:/data alpine sleep 3600
docker exec persisted2 cat /data/test.txt
# persistent
# ============================================
# PART 6: INSPECT THE WRITABLE LAYER
# ============================================
# Show the container's storage details.

docker inspect --format '{{.GraphDriver.Data.UpperDir}}' persisted2
# ============================================
# PART 7: INSPECT THE VOLUME
# ============================================
docker volume inspect mydata
# ============================================
# PART 8: COMPARE SIZES
# ============================================
# The container with a volume has a smaller
# writable layer because data went to the volume.

docker ps -s
# ============================================
# PART 9: USE A BIND MOUNT
# ============================================
mkdir -p /tmp/html
echo "hello" > /tmp/html/index.html
docker run -d --name bind -v /tmp/html:/usr/share/nginx/html nginx
docker exec bind cat /usr/share/nginx/html/index.html
# hello
# ============================================
# PART 10: CLEANUP
# ============================================
docker rm -f persisted2 bind
docker volume rm mydata

These ten parts demonstrate the ephemeral nature of the writable layer, the persistence of volumes, the difference in storage location, and the use of bind mounts for host-controlled data.


Quick Reference

The Layer Stack

LayerRead/WriteShared
Image layersRead-onlyYes, across containers
Writable layerRead-writeNo, unique per container

Copy-on-Write Operations

OperationBehavior
Read from read-only layerDirect read, no copy
First write to a fileCopy file to writable layer, then modify
Subsequent writesWrite to the copy in the writable layer
Delete a fileCreate whiteout in writable layer
Create a new fileWrite directly to writable layer

Storage Location Comparison

LocationPersists After docker rmManaged By
Writable layerNoContainer
Named volumeYesDocker
Bind mountYesHost
tmpfsNoMemory

Storage Drivers

DriverCoWNotes
overlay2YesDefault on modern Linux, file-level copy
vfsNoFull deep copy per layer, least efficient
btrfs, zfsYesBlock-level, require specific filesystems

Best Practices

✅ Do This:

docker run -d -v pgdata:/var/lib/postgresql/data postgres   # Volume for DB
docker run -d -v /srv/html:/usr/share/nginx/html nginx       # Bind for host data
docker run --rm -v /tmp:/tmp alpine                          # Disposable with --rm
docker volume ls                                             # Track volumes
docker volume inspect mydata                                 # Inspect volume

❌ Don’t Do This:

docker run -d postgres                                      # ❌ No volume; data in writable layer
docker exec container sh -c "echo data > /app/data"         # ❌ Data lost on rm
docker run -v /host/path:/container/path image              # ❌ -v creates dir if missing
docker rm -f $(docker ps -aq)                               # ❌ Does not remove volumes

Common Pitfalls

PitfallWhy It HappensFix
Data lost after docker rmData was in writable layerUse a volume or bind mount
Writable layer grows largeWrite-heavy workload in containerMove writes to a volume
Copy-up performance hitModifying large files in read-only layersUse volumes for write-heavy files
Anonymous volumes orphanedImage declares VOLUME, no explicit mountUse named volumes
Cannot migrate dataData only in one container’s writable layerUse volumes for portability
Disk full from containerLogs or temp files accumulatingUse volume with size limit or log rotation

Real-World Examples

1. Database with Named Volume

docker run -d --name db -v pgdata:/var/lib/postgresql/data postgres:16

2. Web Server with Bind Mount

docker run -d -v $(pwd)/html:/usr/share/nginx/html nginx

3. Disposable Test Container

docker run --rm -v /tmp/test:/data alpine sh -c "echo test > /data/file"

4. Inspect Writable Layer Size

docker ps -s --format "table {{.Names}}\t{{.Size}}"

5. Find Anonymous Volumes

docker volume ls -f dangling=true

6. Backup a Volume

docker run --rm -v pgdata:/data -v $(pwd):/backup alpine tar czf /backup/data.tar.gz -C /data .

7. Restore a Volume

docker run --rm -v pgdata:/data -v $(pwd):/backup alpine tar xzf /backup/data.tar.gz -C /data

8. tmpfs for Sensitive Data

docker run -d --tmpfs /run/secrets:size=16m myapp

9. Check Storage Driver

docker info | grep "Storage Driver"

10. Compare Container Sizes

docker ps -s --format "{{.Names}}: {{.Size}}"

Visual

The Layer Stack

┌─────────────────────────────────────────────┐
│  Writable Layer (unique per container)      │  ← Writes go here
├─────────────────────────────────────────────┤
│  Image Layer N: RUN npm install             │  ← Read-only
├─────────────────────────────────────────────┤
│  Image Layer 2: COPY app /app               │  ← Read-only
├─────────────────────────────────────────────┤
│  Image Layer 1: FROM node:20-alpine         │  ← Read-only
└─────────────────────────────────────────────┘
         │
         ▼
    Shared by all containers
    based on this image

Copy-on-Write Operation

┌─────────────────────────────────────────────┐
│  Before write:                              │
│  ┌───────────────────────────────────────┐  │
│  │  Writable Layer (empty)               │  │
│  ├───────────────────────────────────────┤  │
│  │  Read-only Layer: /app/config.json    │  │
│  └───────────────────────────────────────┘  │
└─────────────────────────────────────────────┘
                    │
                    │ container writes to /app/config.json
                    ▼
┌─────────────────────────────────────────────┐
│  After write:                               │
│  ┌───────────────────────────────────────┐  │
│  │  Writable Layer: config.json (copy)   │  │  ← Modified version
│  ├───────────────────────────────────────┤  │
│  │  Read-only Layer: config.json         │  │  ← Original, obscured
│  └───────────────────────────────────────┘  │
└─────────────────────────────────────────────┘

Storage Location and Lifecycle

┌─────────────────────────────────────────────────────────────┐
│  WRITABLE LAYER              │  VOLUME                     │
│  ┌───────────────────────┐   │  ┌───────────────────────┐  │
│  │  Owned by: Container  │   │  │  Owned by: Docker     │  │
│  │  Lifecycle: Container │   │  │  Lifecycle: Docker    │  │
│  │  Persists: No         │   │  │  Persists: Yes        │  │
│  └───────────────────────┘   │  └───────────────────────┘  │
│                               │                              │
│  BIND MOUNT                   │  TMPFS                       │
│  ┌───────────────────────┐   │  ┌───────────────────────┐  │
│  │  Owned by: Host       │   │  │  Owned by: Memory     │  │
│  │  Lifecycle: Host      │   │  │  Lifecycle: Container │  │
│  │  Persists: Yes        │   │  │  Persists: No         │  │
│  └───────────────────────┘   │  └───────────────────────┘  │
└─────────────────────────────────────────────────────────────┘

When to Use Each Storage Location

┌─────────────────────────────────────────────────────────────┐
│  DATA TYPE                    │  LOCATION                    │
│  ─────────────────────────────┼───────────────────────────── │
│  Database files               │  Named volume                │
│  User uploads                 │  Named volume                │
│  Source code (dev)            │  Bind mount                  │
│  Config files (host-managed)  │  Bind mount                  │
│  Logs (also shipped)          │  Writable layer / volume     │
│  Temp files                   │  tmpfs                       │
│  Session tokens               │  tmpfs                       │
│  Package caches               │  Writable layer              │
└─────────────────────────────────────────────────────────────┘

Summary

ItemValue
Layer stackRead-only image layers + writable container layer
Shared layersImage layers shared across containers
Unique layerWritable layer unique per container
Copy-on-writeFile copied to writable layer on first modification
Copy-up granularityFile-level, entire file copied
Delete operationWhiteout file in writable layer
Writable layer lifecycleDestroyed with container
Named volumeManaged by Docker, persists after removal
Bind mountManaged by host, persists after removal
tmpfsMemory-only, destroyed with container
Default driveroverlay2 on modern Linux
Write-heavy workloadsUse volumes, not writable layer

Key takeaways:

  • A container’s filesystem is a stack of read-only image layers plus a writable layer. All containers from the same image share the read-only layers, which saves disk space and memory .
  • Copy-on-write copies a file to the writable layer on first modification. The read-only layers are never modified. The copy-up is file-level, so modifying a small part of a large file duplicates the entire file .
  • The writable layer is ephemeral. It is destroyed when the container is removed. Stopping and restarting a container preserves the layer, but docker rm does not .
  • Volumes and bind mounts are where persistent data belongs. Named volumes are managed by Docker and survive container removal. Bind mounts are managed by the host and give the host direct control over the data .
  • Copy-on-write has a performance cost. The first write to a file is slower because the file must be copied. Write-heavy workloads should use volumes, which bypass the storage driver entirely .
  • Anonymous volumes are created automatically but are hard to manage. They persist after container removal but have random names, making them easy to orphan. Named volumes are the managed alternative .
  • The choice of storage location is a lifecycle decision. The writable layer belongs to the container, volumes belong to Docker, and bind mounts belong to the host. Choose based on who should own the data’s lifecycle .

Remember: A container’s filesystem is not a single disk; it is a stack of layers with a thin writable layer on top. The read-only layers are the image, shared by every container that uses it. The writable layer is the container’s private space, and it exists only as long as the container does. Copy-on-write makes this efficient: files are only copied to the writable layer when they are modified, and the copy is what the container sees from that point on. But the writable layer is not storage. It is scratch space. Data that matters—database files, user uploads, configuration that must survive—belongs in a volume or bind mount, where the lifecycle is independent of the container. The distinction between the writable layer and volumes is not about performance or size; it is about who owns the data and how long it should live. Understanding this distinction is understanding how to use containers without losing data.



Stop using slow, ad-bloated tool sites! 🤮

🔎 Search “KandZ Tools” on Google to use many professional utilities for free.

KandZ.me is the ultimate minimalist hub for:
✅ Finance (Mortgage, Interest, Inflation)
✅ Tech (Base64, JSON, Dev Suite, IP)
✅ Health (BMI, BMR, TDEE)
✅ Productivity (Timer, Workspace, QR)

⚡️ Fast & Private
🔒 No data leaves your device
💎 100% Free

🔗 Use it now: https://tools.kandz.me
🔖 Bookmark it—you’ll need it later!