LFCA 66 ๐ง Monitoring Disk and Network
The previous chapter covered CPU and memory monitoring โ the resources that determine whether the system is under computational pressure. This chapter covers the other two pillars: disk and network. A system can have idle CPU and plenty of RAM and still be unusable because the disk is full or the network is saturated. These are the failures that users notice first, and they are the ones that the LFCA exam expects you to diagnose.
The LFCA exam places system monitoring under System Administration Fundamentals, which carries 20%โ30% of the total weight . The study plan explicitly lists disk usage (df, du) and I/O statistics (iostat) as monitoring topics . Network monitoring is covered through the tools that display interface state, connection tables, and real-time throughput. Knowing which tool answers which question โ and knowing the difference between “disk is full” and “disk is slow” โ is the skill this chapter builds.
Key point: Disk monitoring has two distinct concerns. Space โ how much room is left โ is answered by df and du. Performance โ how fast the disk responds โ is answered by iostat and iotop. Network monitoring follows the same split. Interface state and connections are answered by ip and ss. Throughput and traffic are answered by iftop, nload, and sar -n. Confusing these categories leads to the wrong tool for the problem.
Why disk and network monitoring matter
CPU and memory get the attention because they are the resources that slow down under load. Disk and network get the attention when they fail completely. A full disk stops services. A saturated network makes every interaction feel broken.
The disk-full problem. When a filesystem reaches 100% capacity, writes fail. Databases stop accepting transactions. Logs stop being written. The application crashes, and the cause โ a disk that filled up over weeks โ is buried in the history. df -h reveals the problem in one command, but only if you look .
The disk-slow problem. A disk that is not full can still be the bottleneck. If the utilization (%util) is at 100% and the average wait time (await) is climbing, the disk is saturated. The application is waiting on I/O, not on CPU. iostat -x shows this. Without it, you might mistakenly attribute the slowdown to the application .
The network-saturation problem. A network interface that is transmitting at its maximum capacity will drop packets, increase latency, and cause timeouts. iftop shows which connections are consuming the bandwidth. Without it, you know the network is slow but not who is making it slow .
The connection-leak problem. A service that opens connections and never closes them will eventually exhaust the connection table. ss -s shows the total number of established and time-wait connections. A climbing time-wait count or an established count that matches a connection limit is the signature of a leak .
The trade-off. Each of these tools measures one specific thing. df does not tell you if the disk is slow. iostat does not tell you if the disk is full. iftop does not tell you if the interface is down. The discipline is knowing which category the problem belongs to โ space or speed, interface or traffic โ and reaching for the right tool.
a. Disk Space: df and du
The df command (disk free) reports the space usage of mounted filesystems. The -h flag makes the output human-readable, and -T adds the filesystem type. This is the first command to run when a service fails unexpectedly .
# Show all mounted filesystems with human-readable sizes
df -h
# Include the filesystem type
df -hT
# Check a specific path
df -h /var/log
The output has columns for total size, used space, available space, and use percentage. The Use% column is the one to watch. A filesystem at 100% is a problem. A filesystem at 95% is a warning โ it will hit 100% soon unless the growth stops .
The du command (disk usage) reports the space consumed by files and directories. While df answers “how much is left on the filesystem,” du answers “what is using the space.” The -s flag summarizes a directory, and -h makes the output readable .
# Show the total size of a directory
du -sh /var/log
# Show the size of each subdirectory, sorted
du -sh /var/log/* | sort -hr | head -10
# Find the largest directories in the home folder
du -sh ~/* | sort -hr | head -10
The combination of df and du is the standard disk-full workflow. df identifies which filesystem is full. du identifies which directory within that filesystem is consuming the space. The sort -hr pipeline orders the results by size, largest first, so the culprit is at the top .
One important distinction: df reports the filesystem’s view of free space, which accounts for reserved blocks and inodes. du reports the sum of file sizes, which may be smaller than the used space reported by df. The difference is the space consumed by deleted files still held open by processes, filesystem metadata, and reserved blocks. When df shows a full disk but du sums to much less, the cause is usually deleted files that are still open. The lsof +L1 command reveals them.
b. Disk Performance: iostat and iotop
The iostat command reports disk I/O statistics. It is part of the sysstat package and is the primary tool for diagnosing disk performance problems . The -x flag shows extended statistics, and the 1 specifies a one-second interval.
# Extended statistics every second
iostat -x 1
# Five samples at two-second intervals
iostat -x 2 5
# Per-device breakdown, ignoring the CPU report
iostat -dx 1
The columns that matter for diagnosis are %util, await, r/s, and w/s . %util is the percentage of time the device had I/O requests in flight. A value near 100% means the disk is saturated โ it cannot process requests faster than they arrive. await is the average time in milliseconds for an I/O request to complete, including time spent waiting in the queue. A climbing await with a saturated %util is the signature of an I/O bottleneck. r/s and w/s are the read and write operations per second. A sudden spike in these values indicates an application performing an I/O storm .
The iotop command is the process-level counterpart to iostat. It shows which processes are performing disk I/O, sorted by usage. While iostat tells you the disk is busy, iotop tells you who is making it busy. It requires root privileges and the iotop package.
# Show processes performing disk I/O
sudo iotop
# Show only processes doing I/O, not threads
sudo iotop -o
The -o flag filters the output to processes actively performing I/O, which is useful on busy systems where most processes are idle. The iotop display updates in real time, like top, and shows the read and write rates per process.
c. Network State and Throughput
Network monitoring splits into two categories: the state of the interfaces and connections, and the throughput of the traffic itself.
The ip command is the modern replacement for ifconfig. ip addr show displays all network interfaces with their IP addresses and state (UP or DOWN). ip link show displays the interface state and statistics. ip route show displays the routing table .
# Show interfaces with addresses
ip addr show
# Show interface state and statistics
ip link show
# Show routing table
ip route show
The ss command (socket statistics) is the modern replacement for netstat. It displays network connections, listening sockets, and statistics. ss -tuln shows all listening TCP and UDP ports. ss -an shows all connections. ss -s shows a summary of socket usage by state .
# Show listening TCP and UDP ports
ss -tuln
# Show all connections
ss -an
# Show summary of socket states
ss -s
The ss -s output is particularly useful for detecting connection leaks. It shows the count of sockets in each state: estab (established), timewait (waiting to close), close-wait, and others. A large timewait count or a climbing estab count is a signal that something is not closing connections properly .
For throughput monitoring, iftop is the standard tool. It shows the bandwidth usage of connections in real time, sorted by throughput. It requires root privileges and the iftop package. sudo iftop -i eth0 monitors a specific interface and shows the top bandwidth consumers with their transfer rates in both directions .
# Monitor bandwidth on a specific interface
sudo iftop -i eth0
The nload tool provides a simpler alternative: a live graph of incoming and outgoing traffic on one or more interfaces. It is less detailed than iftop but easier to read at a glance. The vnstat tool records historical network traffic and reports it over time โ daily, monthly, or yearly โ which is useful for capacity planning .
The sar -n DEV command provides historical network statistics. It shows packets and bytes transmitted and received per second for each interface. Because sar records data every 10 minutes by default, it can answer “what was the network traffic at 3 AM last Tuesday” โ a question that real-time tools cannot answer .
# Live network stats every 2 seconds
sar -n DEV 2 5
# Historical network stats from a saved file
sar -n DEV -f /var/log/sa/sa10
Complete Example Session
This session demonstrates the disk-full workflow, the disk-performance workflow, and the network-state and throughput workflows.
# ============================================
# PART 1: THE DISK-FULL WORKFLOW
# ============================================
# Check all filesystems
df -h
# Filesystem Size Used Avail Use% Mounted on
# /dev/sda1 50G 47G 3.0G 95% /
# tmpfs 3.9G 0 3.9G 0% /dev/shm
# The root filesystem is 95% full.
# Find what is using the space:
du -sh /var/* | sort -hr | head -5
# 32G /var/log
# 8G /var/lib
# 2G /var/cache
# Drill into /var/log:
du -sh /var/log/* | sort -hr | head -5
# 28G /var/log/journal
# 3G /var/log/syslog
# 500M /var/log/auth.log
# The journal is consuming 28G.
# ============================================
# PART 2: THE JOURNAL CLEANUP
# ============================================
# Check current journal usage
journalctl --disk-usage
# Archived and active journals take up 28.0G in the file system.
# Vacuum to 2G
sudo journalctl --vacuum-size=2G
# Deleted archived journal files.
# Verify
df -h / | tail -1
# /dev/sda1 50G 21G 27G 44% /
# ============================================
# PART 3: THE DISK-PERFORMANCE WORKFLOW
# ============================================
# Check disk I/O statistics
iostat -x 1 3
# Device r/s w/s rkB/s wkB/s await %util
# sda 150.2 320.5 4801.6 10241.6 45.2 98.3
# %util is 98.3 โ the disk is saturated.
# await is 45.2ms โ requests are waiting in the queue.
# Find the process doing the I/O
sudo iotop -o
# PID USER DISK READ DISK WRITE COMMAND
# 1234 postgres 0.00 B/s 45.2 M/s postgres: wal writer
# The PostgreSQL WAL writer is the source.
# ============================================
# PART 4: THE NETWORK-STATE CHECK
# ============================================
# Check interface state
ip link show
# 2: eth0: <BROADCAST,MULTICAST,UP,LOWER_UP> mtu 1500
# link/ether 00:11:22:33:44:55 brd ff:ff:ff:ff:ff:ff
# Interface is UP.
# Check IP addresses
ip addr show eth0
# inet 192.168.1.100/24 brd 192.168.1.255 scope global eth0
# Check listening ports
ss -tuln
# Netid State Local Address:Port
# tcp LISTEN 0.0.0.0:22
# tcp LISTEN 0.0.0.0:80
# tcp LISTEN 0.0.0.0:443
# ============================================
# PART 5: THE CONNECTION LEAK CHECK
# ============================================
# Check socket summary
ss -s
# Total: 847
# TCP: 823 (estab 512, closed 0, orphaned 0, timewait 311)
# 512 established connections and 311 in time-wait.
# That is high for a single application.
# Find which process is holding them
ss -tnp | grep ESTAB | awk '{print $6}' | cut -d: -f1 | sort | uniq -c | sort -nr
# 487 192.168.1.100:5432
# 25 192.168.1.100:6379
# The application is holding 487 connections to PostgreSQL.
# ============================================
# PART 6: THE NETWORK-THROUGHPUT CHECK
# ============================================
# Monitor interface bandwidth
sudo iftop -i eth0
# 12.5Mb 25.0Mb 37.5Mb
# โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
# โ 192.168.1.50:443 => 10.0.0.25:52341 2.5Mb
# โ 192.168.1.50:5432 => 10.0.0.25:52342 8.2Mb
# โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
# 10.0.0.25 is downloading 8.2 Mb/s from the database.
# ============================================
# PART 7: THE HISTORICAL NETWORK CHECK
# ============================================
# Check network stats from last week
sar -n DEV -f /var/log/sa/sa$(date -d "7 days ago" +%d)
# IFACE rxpck/s txpck/s rxkB/s txkB/s
# eth0 1250.3 980.2 4500.1 3200.5
# Compare to today
sar -n DEV 1 3
# IFACE rxpck/s txpck/s rxkB/s txkB/s
# eth0 2500.6 1900.4 9000.2 6400.1
# Traffic has doubled since last week.
# ============================================
# PART 8: THE COMBINED DIAGNOSIS
# ============================================
# The application is slow.
# Check disk: df shows 44% used โ not full.
# Check disk: iostat shows %util 98% โ disk is busy.
# Check network: iftop shows 8.2 Mb/s on the DB connection.
# Check connections: ss -s shows 512 established.
# The application is holding too many DB connections,
# causing the DB to perform excessive I/O.
# The fix is connection pooling.
# ============================================
# PART 9: THE QUICK REFERENCE CARD
# ============================================
# Disk space:
df -h # filesystem usage
du -sh /path # directory size
# Disk performance:
iostat -x 1 # device I/O stats
iotop -o # per-process I/O
# Network state:
ip addr show # interfaces
ss -tuln # listening ports
ss -s # socket summary
# Network throughput:
iftop -i eth0 # real-time bandwidth
nload # simple traffic graph
sar -n DEV # historical network
# ============================================
# PART 10: THE MONITORING PYRAMID
# ============================================
# Summary: df, ip, ss -s
# โ Is there a problem?
#
# Detail: du, iostat, iotop, iftop
# โ What is causing it?
#
# History: sar -d, sar -n DEV
# โ What was normal?
The ten parts cover the disk-full workflow, journal cleanup, disk-performance workflow, network-state check, connection-leak check, throughput check, historical network check, combined diagnosis, quick reference card, and the monitoring pyramid.
Quick Reference
The Disk Space Commands
| Command | Purpose |
|---|---|
df -h | Filesystem usage, human-readable |
df -hT | Include filesystem type |
du -sh /path | Total size of a directory |
du -sh /path/* | sort -hr | Directory sizes, sorted |
The Disk Performance Commands
| Command | Purpose |
|---|---|
iostat -x 1 | Extended device I/O stats |
iostat -dx 1 | Per-device, no CPU report |
iotop -o | Per-process I/O (active only) |
The Network State Commands
| Command | Purpose |
|---|---|
ip addr show | Interfaces with addresses |
ip link show | Interface state and stats |
ss -tuln | Listening TCP/UDP ports |
ss -an | All connections |
ss -s | Socket state summary |
The Network Throughput Commands
| Command | Purpose |
|---|---|
iftop -i eth0 | Real-time bandwidth by connection |
nload | Simple traffic graph |
vnstat | Historical traffic statistics |
sar -n DEV | Historical network stats |
The Key Metrics
| Metric | Source | What It Means |
|---|---|---|
| Use% | df | Filesystem fullness |
%util | iostat | Device saturation |
await | iostat | I/O latency |
estab | ss -s | Active connections |
timewait | ss -s | Closing connections |
Best Practices
โ Do This:
# Start with df when a service fails
df -h # โ
# Use du to find the directory consuming space
du -sh /var/* | sort -hr | head -10 # โ
# Check iostat before blaming the application
iostat -x 1 3 # โ
# Use ss -s for a quick connection summary
ss -s # โ
# Use iftop to find bandwidth consumers
sudo iftop -i eth0 # โ
โ Don’t Do This:
# Don't assume a full disk is the only disk problem
# Check %util for saturation. # โ
# Don't use ifconfig in new scripts
ifconfig # โ use ip
# Don't use netstat in new scripts
netstat -tuln # โ use ss
# Don't ignore time-wait count
ss -s | grep timewait # โ investigate if high
Common Pitfalls
| Pitfall | Why It Happens | Fix |
|---|---|---|
| df shows full, du shows less | Deleted files still open | lsof +L1 |
| iostat %util high but no load | I/O wait, not CPU | Check await, identify process |
| iftop shows no traffic | Wrong interface | Check ip link for the right one |
| ss -s shows many time-wait | Connections not closing | Check application connection handling |
| sar returns no data | sysstat not enabled | systemctl enable sysstat |
Real-World Examples
1. Check Filesystem Usage
df -h
2. Find Largest Directories
du -sh /var/* | sort -hr | head -10
3. Check Disk I/O
iostat -x 1 3
4. Find I/O Process
sudo iotop -o
5. Check Interface State
ip link show
6. List Listening Ports
ss -tuln
7. Check Socket Summary
ss -s
8. Monitor Bandwidth
sudo iftop -i eth0
9. Historical Network Stats
sar -n DEV -f /var/log/sa/sa10
10. Check Historical Disk I/O
sar -d -f /var/log/sa/sa10
Visual
The Disk Monitoring Split
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ DISK MONITORING โ
โ โ
โ SPACE (how much is left?) โ
โ df -h โ filesystem usage โ
โ du -sh โ directory usage โ
โ โ
โ PERFORMANCE (how fast is it?) โ
โ iostat -x โ device I/O stats โ
โ iotop -o โ per-process I/O โ
โ โ
โ A full disk is a different problem โ
โ than a slow disk. โ
โ โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
The Network Monitoring Split
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ NETWORK MONITORING โ
โ โ
โ STATE (is it up? who is connected?) โ
โ ip addr show โ interfaces โ
โ ss -tuln โ listening ports โ
โ ss -s โ socket summary โ
โ โ
โ THROUGHPUT (how much traffic?) โ
โ iftop โ real-time by connection โ
โ nload โ simple traffic graph โ
โ sar -n DEV โ historical stats โ
โ โ
โ An interface can be UP and still be โ
โ saturated. These are different questions. โ
โ โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
The Disk-Full Workflow
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ DISK-FULL WORKFLOW โ
โ โ
โ 1. df -h โ
โ โโ Which filesystem is full? โ
โ โ
โ 2. du -sh /path/* | sort -hr โ
โ โโ Which directory is consuming space? โ
โ โ
โ 3. Drill down โ
โ โโ du -sh /path/subdir/* โ
โ โ
โ 4. Clean up โ
โ โโ Delete old logs, vacuum journal, etc. โ
โ โ
โ 5. df -h โ
โ โโ Verify the space is freed. โ
โ โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
The Connection State Flow
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ CONNECTION STATES (ss -s) โ
โ โ
โ estab โ active connections โ
โ timewait โ closed, waiting for timeout โ
โ close-wait โ remote closed, local not โ
โ listen โ listening sockets โ
โ โ
โ High timewait: many short-lived connections โ
โ โโ Normal for web servers โ
โ โ
โ High estab: many open connections โ
โ โโ Check for connection leaks โ
โ โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
Summary
| Item | Value |
|---|---|
| Disk space | df -h, du -sh |
| Disk performance | iostat -x, iotop -o |
| Interface state | ip addr show, ip link show |
| Listening ports | ss -tuln |
| Connection summary | ss -s |
| Real-time bandwidth | iftop -i eth0 |
| Historical network | sar -n DEV |
| Historical disk I/O | sar -d |
| Disk saturation metric | %util from iostat |
| Disk latency metric | await from iostat |
Key takeaways:
- Disk monitoring splits into space and performance.
dfandduanswer “how much is left.”iostatandiotopanswer “how fast is it.” A disk can be 40% full and still be the bottleneck. Check both . - The disk-full workflow is a sequence.
df -hidentifies the full filesystem.du -sh /path/* | sort -hridentifies the directory consuming the space. Drill down until you find the files. Clean up. Verify withdf -hagain . %utilandawaitare the two metrics that matter for disk performance.%utilnear 100% means the disk is saturated. Climbingawaitmeans requests are waiting in the queue. Together they identify an I/O bottleneck .iotopattributes disk I/O to processes.iostattells you the disk is busy.iotop -otells you which process is making it busy. It requires root .ipandssreplaceifconfigandnetstat. Theipcommand shows interface state and addresses.ssshows connections, listening ports, and socket summaries. Both are the modern tools and are installed by default on current distributions .ss -sis the connection-leak detector. It shows the count of sockets in each state. A climbingestabcount or a largetimewaitcount is a signal that something is not closing connections properly .iftopshows who is consuming bandwidth. It ranks connections by throughput and shows the transfer rate in both directions. It is the network equivalent oftop.sar -n DEVprovides historical network data. Likesar -ufor CPU, it can answer “what was the network traffic at 3 AM last Tuesday.” It requiressysstatto be enabled .
Remember: Disk and network monitoring are the two halves of the resource story that CPU and memory do not tell. A system with idle CPU and free memory can still be unusable because the disk is full or the network is saturated. The discipline is the same as for CPU and memory: start with the summary, drill into the detail, check the history. df for space, iostat for speed, ip and ss for network state, iftop for throughput. Know which question you are asking, and the tool is obvious.
Stop using slow, ad-bloated tool sites! ๐คฎ
๐ Search “KandZ Tools” on Google to use many professional utilities for free.
KandZ.me is the ultimate minimalist hub for:
โ
Finance (Mortgage, Interest, Inflation)
โ
Tech (Base64, JSON, Dev Suite, IP)
โ
Health (BMI, BMR, TDEE)
โ
Productivity (Timer, Workspace, QR)
โก๏ธ Fast & Private
๐ No data leaves your device
๐ 100% Free
๐ Use it now: https://tools.kandz.me
๐ Bookmark itโyouโll need it later!