| | |

LFCA 65 🐧 Monitoring CPU and Memory

A Linux system under load tells you nothing unless you know where to look. The CPU might be pegged at 100% while memory sits idle, or memory might be exhausted while the CPU waits for disk I/O that never comes. Monitoring is the practice of reading the right signals at the right time, and the LFCA exam expects you to know which commands reveal which story.

The LFCA exam places system monitoring under System Administration Fundamentals, which carries 30% of the total weight in the updated competency list . The specific competencies include troubleshooting and best practices — and both depend on knowing how to observe a system before it fails. Monitoring is not about watching numbers scroll. It is about establishing a baseline, spotting deviations, and finding the process responsible.

Key point: CPU and memory monitoring on Linux works in layers. The first layer is the summary — uptime, free, vmstat — which tells you whether there is a problem at all. The second layer is the process — top, ps, pidstat — which tells you which process is responsible. The third layer is the historical record — sar — which tells you what the system looked like before you arrived. Every investigation moves through these layers in order.


Why CPU and memory monitoring exist

A Linux system does not announce that it is struggling. It does not slow down gracefully or display a warning. It simply stops responding to requests, or crashes, or becomes so slow that users abandon it. By the time the problem is visible to a user, the root cause may already be gone — the process that consumed all the memory has been killed, or the CPU spike has subsided.

The invisibility problem. CPU and memory are shared resources. A single runaway process can consume 100% of a CPU core or exhaust all available memory without any warning. Without monitoring, the first symptom is the crash, and the cause is lost. Monitoring tools make the consumption visible before it becomes fatal.

The baseline problem. A system using 80% of its memory is not necessarily in trouble. Linux uses free memory for disk caching, and the kernel releases those pages when applications need them . The free column showing low numbers is normal on a healthy system; the available column is what matters . Without a baseline of what “normal” looks like, you cannot distinguish healthy cache usage from genuine memory pressure.

The attribution problem. Knowing that CPU is high is not useful. Knowing that the mysql process is consuming 90% of a core is actionable. Monitoring tools progress from the summary to the process level, from “the system is busy” to “this specific process is the cause.” The pidstat command exists precisely for this: it attributes CPU, memory, and I/O activity to individual tasks .

The history problem. By the time you log in to investigate, the spike may be over. sar records system activity in the background — CPU, memory, I/O, network — and stores it in /var/log/sa/ files that persist across the event . Without sar running before the problem, the investigation is limited to what is visible in the moment.

The trade-off. Monitoring tools consume resources themselves. top and vmstat are lightweight, but sar running every 10 minutes accumulates data over time. The sysstat package must be installed and enabled, and its data directory must be managed. For a small system, the overhead is negligible. For a large fleet, centralized monitoring is the only practical approach.


a. The Summary Commands: uptime, free, vmstat

The first question in any investigation is whether the system is actually under pressure. The summary commands answer that question in a single line or two.

uptime displays how long the system has been running and the load average over the last 1, 5, and 15 minutes . The load average is the number of processes waiting for CPU time or in uninterruptible sleep. On a system with 4 CPU cores, a load average of 4.0 means the CPUs are fully utilized. A load average of 8.0 means processes are waiting. A load average that climbs steadily across the 1, 5, and 15-minute windows indicates sustained pressure; a spike that appears only in the 1-minute average indicates a transient event.

free displays memory usage: total, used, free, shared, buff/cache, and available . The critical column is available — an estimate of how much memory is available for new applications without swapping . The free column will often be low on a healthy system because the kernel uses unused RAM for disk caching. The buff/cache column shows that cache, and the available column shows how much of it can be reclaimed . The -h flag makes the output human-readable .

vmstat provides a single-line overview of CPU, memory, paging, block I/O, interrupts, and context switches . Running vmstat 1 5 shows five samples at one-second intervals, revealing trends rather than a snapshot. The key columns for CPU are us (user), sy (system), id (idle), and wa (wait for I/O). A high wa indicates disk I/O is the bottleneck, not CPU. The key columns for memory are swpd (swap used), free, buff, and cache, plus the si (swap in) and so (swap out) columns under the swap section. Sustained swap-in and swap-out activity indicates memory pressure .

# Load average
uptime

# Memory usage
free -h

# CPU, memory, I/O trends
vmstat 1 5

b. The Process Commands: top, ps, pidstat

Once you know the system is under pressure, the next step is finding the process responsible.

top is the interactive process monitor. It displays a live list of processes sorted by CPU usage by default . The header shows load average, task counts, CPU states, and memory usage — a complete summary in one screen. Pressing 1 toggles per-CPU display, showing whether a single core is pegged or the load is spread across all cores . Pressing M sorts by memory usage instead of CPU. Pressing P returns to CPU sort. The top output updates every few seconds, and htop provides a more colorful, navigable alternative when installed .

ps is the non-interactive process lister. The command ps -eo pcpu,pmem,pid,user,args | sort -r -k1 | head -6 lists the top five processes by CPU usage . Replacing -k1 with -k2 sorts by memory. The ps command is useful for scripting and for capturing a snapshot when you cannot run an interactive monitor.

pidstat is the per-process statistics tool, part of the sysstat package. It monitors individual tasks and can report CPU, memory, and I/O usage per process . pidstat -u 1 5 reports CPU usage for all processes at one-second intervals, five times. pidstat -r 1 5 reports memory usage . The -w option reports context switches, useful for diagnosing processes that are causing excessive task switching .

# Interactive process monitor
top

# Top 5 CPU processes
ps -eo pcpu,pmem,pid,user,args | sort -r -k1 | head -6

# Per-process memory usage
pidstat -r 1 5

c. The Historical Record: sar

sar is the system activity reporter, part of the sysstat package. It collects data in the background — by default every 10 minutes — and stores it in /var/log/sa/ files named saDD where DD is the day of the month . This is the only tool that answers the question “what was the system doing before the problem?”

The sysstat package must be installed and the sysstat service must be enabled for data collection to occur. Without it, sar has no history to report . Once enabled, sar accumulates CPU, memory, I/O, and network statistics that can be queried retroactively.

Common sar queries:

# CPU usage for all processors (live)
sar -u -P ALL 1 3

# CPU usage from a historical file
sar -u -f /var/log/sa/sa10

# Memory usage
sar -r

# Load average
sar -q

The -f option reads from a specific historical file, allowing you to examine a day that has already passed. The -s and -e options filter by start and end time within that file. For example, sar -u -s 14:00:00 -e 15:00:00 -f /var/log/sa/sa10 shows CPU usage between 2 PM and 3 PM on the 10th of the month .

The value of sar is in establishing a baseline. When a performance issue occurs, you compare the sar data from the problem period against the sar data from a normal period. Without the baseline, you cannot quantify how much the system has degraded .


Complete Example Session

This session demonstrates a complete investigation: spotting pressure with summary commands, finding the process with process commands, and checking history with sar.

# ============================================
# PART 1: CHECK LOAD AVERAGE
# ============================================

uptime
# 14:32:01 up 12 days,  3:45,  2 users,  load average: 4.12, 3.87, 2.91

# The 1-minute average is higher than the 15-minute average.
# The system is under increasing load.

# ============================================
# PART 2: CHECK MEMORY
# ============================================

free -h
#                total        used        free      shared  buff/cache   available
# Mem:            31Gi        19Gi       575Mi       3.3Gi        12Gi       8.8Gi
# Swap:          8.0Gi       6.6Gi       1.4Gi

# 6.6 GiB of swap is in use. This is a sign of memory pressure.

# ============================================
# PART 3: CHECK SWAP ACTIVITY
# ============================================

vmstat 1 5
# procs -----------memory---------- ---swap-- -----io---- -system-- ------cpu-----
#  r  b   swpd   free   buff  cache   si   so    bi    bo   in   cs us sy id wa st
#  2  0 7102404 1392528     36 12335148   8   21   130   724 2851   19 15  7 77  0  0
#  0  0 7102404 1392560     36 12335188   0    0     0     0 5779 7246 14...

# si (swap in) and so (swap out) show activity.
# The system is paging memory to and from disk.

# ============================================
# PART 4: FIND THE TOP CPU PROCESSES
# ============================================

ps -eo pcpu,pmem,pid,user,args | sort -r -k1 | head -6
# %CPU %MEM   PID USER     COMMAND
# 95.2  2.1  1234 mysql    /usr/sbin/mysqld
# 12.4  0.5  5678 apache2  /usr/sbin/apache2 -k start
#  8.1  0.3  9012 postgres /usr/lib/postgresql/14/bin/postgres

# mysql is consuming 95% of a CPU core.

# ============================================
# PART 5: INTERACTIVE MONITORING WITH TOP
# ============================================

top
# Press "1" to see per-CPU usage.
# Press "M" to sort by memory.
# Press "P" to sort by CPU.

# ============================================
# PART 6: PER-PROCESS MEMORY USAGE
# ============================================

pidstat -r 1 3
# Average:      UID       PID  minflt/s  majflt/s     VSZ     RSS   %MEM  Command
# Average:      106      1234      0.00      0.00 1234567  654321   2.10  mysqld

# mysql is using 2.10% of memory, but the high CPU usage
# and swap activity suggest memory is being used inefficiently.

# ============================================
# PART 7: CHECK HISTORICAL CPU USAGE
# ============================================

sar -u -f /var/log/sa/sa$(date +%d)
# 14:00:01     CPU     %user     %nice   %system   %iowait    %steal     %idle
# 14:10:01     all      15.23      0.00      7.45      0.00      0.00     77.32
# 14:20:01     all      14.87      0.00      7.12      0.00      0.00     78.01
# 14:30:01     all      42.11      0.00     28.34      0.00      0.00     29.55

# CPU usage jumped between 14:20 and 14:30.

# ============================================
# PART 8: CHECK HISTORICAL MEMORY USAGE
# ============================================

sar -r -f /var/log/sa/sa$(date +%d)
# 14:20:01 kbmemfree kbavail kbmemused  %memused kbbuffers  kbcached
# 14:20:01   1392560 9227464   18786044     60.23      36     12335188

# Memory usage has been climbing.

# ============================================
# PART 9: CHECK CONTEXT SWITCHES
# ============================================

pidstat -w 1 3
# Average:      UID       PID   cswch/s nvcswch/s  Command
# Average:      106      1234   1234.56     12.34  mysqld

# mysql is causing a high number of voluntary context switches.

# ============================================
# PART 10: TAKE ACTION
# ============================================

# If mysql is the problem, investigate the query log.
# If it is a runaway process, consider restarting the service.
# If it is a memory leak, the process may need to be restarted.
# Use sar data to quantify the problem before and after the fix.

The ten parts cover load average, memory, swap activity, top CPU processes, interactive top, per-process memory, historical CPU usage, historical memory usage, context switches, and taking action.


Quick Reference

The Summary Commands

CommandShowsKey Flags
uptimeLoad average (1, 5, 15 min)—
freeMemory and swap-h human-readable, -s repeat
vmstatCPU, memory, I/O, swap[delay] [count]

The Process Commands

CommandShowsKey Flags
topInteractive process list1 per-CPU, M memory sort, P CPU sort
psProcess snapshot-eo pcpu,pmem,pid,user,args
pidstatPer-process stats-u CPU, -r memory, -w switches

The Historical Command

CommandShowsKey Flags
sarHistorical system activity-u CPU, -r memory, -f file

The Key Metrics

MetricSourceWhat It Means
Load averageuptimeProcesses waiting for CPU
Available memoryfreeMemory usable without swapping
Swap in/outvmstatMemory pressure
%wavmstatI/O wait (not CPU-bound)
%uservmstatUser-space CPU usage
%systemvmstatKernel CPU usage

The sysstat Package

ComponentPurpose
sarHistorical activity reporter
pidstatPer-process statistics
mpstatPer-CPU statistics
iostatPer-device I/O statistics

Best Practices

✅ Do This:

# Use human-readable output
free -h                                                        # ✅
# Check available memory, not free memory
free -h | grep Mem                                             # ✅
# Use vmstat with an interval to see trends
vmstat 1 5                                                     # ✅
# Sort ps output by CPU or memory
ps -eo pcpu,pmem,pid,user,args | sort -r -k1 | head -6        # ✅
# Enable sysstat for historical data
sudo systemctl enable sysstat                                  # ✅
# Check sar data for the problem period
sar -u -f /var/log/sa/sa10                                     # ✅

❌ Don’t Do This:

# Don't panic at low free memory
free                                                           # ❌ available matters
# Don't rely on a single snapshot
top -n 1                                                       # ❌ trends matter
# Don't forget to enable sysstat
sar -u                                                         # ❌ no data if disabled
# Don't ignore iowait
# High %wa means disk is the bottleneck, not CPU               # ❌

Common Pitfalls

PitfallWhy It HappensFix
Low free memory looks alarmingLinux uses RAM for cacheCheck available column
High load average but low CPUI/O wait or uninterruptible sleepCheck vmstat wa column
sar returns no datasysstat not enabledsystemctl enable sysstat
pidstat not foundsysstat not installedInstall sysstat package
CPU usage exceeds 100%Multithreaded process across coresNormal, check per-CPU with top 1
Swap usage ignoredOnly looking at RAMCheck free swap line and vmstat si/so

Real-World Examples

1. Check Load Average

uptime

2. Check Memory Human-Readable

free -h

3. Watch Memory Every Second

free -h -s 1

4. Watch CPU and I/O Trends

vmstat 1 5

5. Interactive Process Monitor

top

6. Top 5 CPU Processes

ps -eo pcpu,pmem,pid,user,args | sort -r -k1 | head -6

7. Per-Process Memory

pidstat -r 1 5

8. Historical CPU Usage

sar -u -f /var/log/sa/sa10

9. Historical Memory Usage

sar -r -f /var/log/sa/sa10

10. Check Swap Devices

swapon -s

Visual

The Monitoring Layers

┌──────────────────────────────────────────────┐
│  LAYER 1: SUMMARY                            │
│                                              │
│  uptime      → load average                  │
│  free        → memory and swap               │
│  vmstat      → CPU, memory, I/O trends       │
│                                              │
│  Question: Is there a problem?               │
│                                              │
├──────────────────────────────────────────────┤
│  LAYER 2: PROCESS                            │
│                                              │
│  top         → interactive process list      │
│  ps          → process snapshot              │
│  pidstat     → per-process stats             │
│                                              │
│  Question: Which process is responsible?     │
│                                              │
├──────────────────────────────────────────────┤
│  LAYER 3: HISTORY                            │
│                                              │
│  sar         → historical activity           │
│  /var/log/sa/ → stored data files            │
│                                              │
│  Question: What was the baseline?            │
│                                              │
└──────────────────────────────────────────────┘

The free Output

┌──────────────────────────────────────────────┐
│  free -h                                     │
│                                              │
│               total  used  free  buff/cache  │
│  Mem:          31Gi  19Gi  575Mi       12Gi  │
│  Swap:        8.0Gi 6.6Gi 1.4Gi              │
│                                              │
│  available:  8.8Gi                           │
│                                              │
│  free = idle and unused                      │
│  available = usable without swapping         │
│  buff/cache = disk cache (reclaimable)       │
│                                              │
│  Low free is normal. Low available is not.   │
│                                              │
└──────────────────────────────────────────────┘

vmstat Columns

┌──────────────────────────────────────────────┐
│  vmstat 1 5                                  │
│                                              │
│  procs:  r (runnable)  b (blocked)           │
│                                              │
│  memory: swpd  free  buff  cache             │
│                                              │
│  swap:   si (in)  so (out)                   │
│  ─── high si/so = memory pressure ───        │
│                                              │
│  io:     bi (blocks in)  bo (blocks out)     │
│                                              │
│  system: in (interrupts)  cs (context sw)    │
│                                              │
│  cpu:    us  sy  id  wa  st                  │
│  ─── high wa = I/O bound, not CPU bound ─── │
│                                              │
└──────────────────────────────────────────────┘

The sar Data Flow

┌──────────────────────────────────────────────┐
│  SAR COLLECTION AND QUERY                    │
│                                              │
│  sysstat service (every 10 min)              │
│       │                                      │
│       ▼                                      │
│  /var/log/sa/saDD                            │
│       │                                      │
│       ▼                                      │
│  sar -u -f /var/log/sa/sa10                  │
│       │                                      │
│       ▼                                      │
│  Historical CPU usage for day 10             │
│                                              │
│  Compare problem period to baseline.         │
│                                              │
└──────────────────────────────────────────────┘

Summary

ItemValue
Load averageuptime
Memory summaryfree -h
CPU/memory/I/O trendsvmstat 1 5
Interactive process monitortop
Process snapshotps -eo pcpu,pmem,pid,user,args
Per-process statspidstat -u, pidstat -r, pidstat -w
Historical activitysar
Data files/var/log/sa/saDD
sysstat packageProvides sar, pidstat, mpstat, iostat
Critical memory metricavailable, not free
Critical CPU metric%wa for I/O wait

Key takeaways:

  • Monitoring works in layers: summary, process, history. Start with uptime, free, and vmstat to determine whether there is a problem. Move to top, ps, and pidstat to find the process responsible. Use sar to compare against a baseline .
  • Load average measures waiting processes, not just CPU. A load average of 4.0 on a 4-core system means full utilization. A load average that is high but CPU usage is low usually indicates I/O wait or uninterruptible sleep .
  • free memory is not the metric to watch. Linux uses unused RAM for disk caching. The available column estimates how much memory can be used without swapping . Low free with high available is normal. Low available with rising swap is a problem .
  • vmstat reveals the type of pressure. High si and so indicate memory pressure. High wa indicates disk I/O is the bottleneck, not CPU. High us indicates user-space CPU consumption .
  • pidstat attributes usage to processes. pidstat -u for CPU, pidstat -r for memory, pidstat -w for context switches. It is part of the sysstat package and reports at configurable intervals .
  • sar is the historical record. It collects data every 10 minutes by default and stores it in /var/log/sa/. Without sar running before a problem, the investigation is limited to the present .
  • The sysstat package provides the monitoring toolkit. sar, pidstat, mpstat, and iostat are all part of it. It must be installed and the service enabled for historical data collection .

Remember: CPU and memory monitoring is not about watching numbers. It is about answering three questions in order: Is there a problem? What is causing it? What was normal before? uptime, free, and vmstat answer the first. top, ps, and pidstat answer the second. sar answers the third. Learn the layers, learn which metric matters in each layer, and the investigation becomes a sequence rather than a guess. The LFCA exam tests whether you know the tools and whether you know what they mean.


Stop using slow, ad-bloated tool sites! 🤮

🔎 Search “KandZ Tools” on Google to use many professional utilities for free.

KandZ.me is the ultimate minimalist hub for:
✅ Finance (Mortgage, Interest, Inflation)
✅ Tech (Base64, JSON, Dev Suite, IP)
✅ Health (BMI, BMR, TDEE)
✅ Productivity (Timer, Workspace, QR)

⚡️ Fast & Private
🔒 No data leaves your device
💎 100% Free

🔗 Use it now: https://tools.kandz.me
🔖 Bookmark it—you’ll need it later!