LFCA 65 🐧 Monitoring CPU and Memory
A Linux system under load tells you nothing unless you know where to look. The CPU might be pegged at 100% while memory sits idle, or memory might be exhausted while the CPU waits for disk I/O that never comes. Monitoring is the practice of reading the right signals at the right time, and the LFCA exam expects you to know which commands reveal which story.
The LFCA exam places system monitoring under System Administration Fundamentals, which carries 30% of the total weight in the updated competency list . The specific competencies include troubleshooting and best practices — and both depend on knowing how to observe a system before it fails. Monitoring is not about watching numbers scroll. It is about establishing a baseline, spotting deviations, and finding the process responsible.
Key point: CPU and memory monitoring on Linux works in layers. The first layer is the summary — uptime, free, vmstat — which tells you whether there is a problem at all. The second layer is the process — top, ps, pidstat — which tells you which process is responsible. The third layer is the historical record — sar — which tells you what the system looked like before you arrived. Every investigation moves through these layers in order.
Why CPU and memory monitoring exist
A Linux system does not announce that it is struggling. It does not slow down gracefully or display a warning. It simply stops responding to requests, or crashes, or becomes so slow that users abandon it. By the time the problem is visible to a user, the root cause may already be gone — the process that consumed all the memory has been killed, or the CPU spike has subsided.
The invisibility problem. CPU and memory are shared resources. A single runaway process can consume 100% of a CPU core or exhaust all available memory without any warning. Without monitoring, the first symptom is the crash, and the cause is lost. Monitoring tools make the consumption visible before it becomes fatal.
The baseline problem. A system using 80% of its memory is not necessarily in trouble. Linux uses free memory for disk caching, and the kernel releases those pages when applications need them . The free column showing low numbers is normal on a healthy system; the available column is what matters . Without a baseline of what “normal” looks like, you cannot distinguish healthy cache usage from genuine memory pressure.
The attribution problem. Knowing that CPU is high is not useful. Knowing that the mysql process is consuming 90% of a core is actionable. Monitoring tools progress from the summary to the process level, from “the system is busy” to “this specific process is the cause.” The pidstat command exists precisely for this: it attributes CPU, memory, and I/O activity to individual tasks .
The history problem. By the time you log in to investigate, the spike may be over. sar records system activity in the background — CPU, memory, I/O, network — and stores it in /var/log/sa/ files that persist across the event . Without sar running before the problem, the investigation is limited to what is visible in the moment.
The trade-off. Monitoring tools consume resources themselves. top and vmstat are lightweight, but sar running every 10 minutes accumulates data over time. The sysstat package must be installed and enabled, and its data directory must be managed. For a small system, the overhead is negligible. For a large fleet, centralized monitoring is the only practical approach.
a. The Summary Commands: uptime, free, vmstat
The first question in any investigation is whether the system is actually under pressure. The summary commands answer that question in a single line or two.
uptime displays how long the system has been running and the load average over the last 1, 5, and 15 minutes . The load average is the number of processes waiting for CPU time or in uninterruptible sleep. On a system with 4 CPU cores, a load average of 4.0 means the CPUs are fully utilized. A load average of 8.0 means processes are waiting. A load average that climbs steadily across the 1, 5, and 15-minute windows indicates sustained pressure; a spike that appears only in the 1-minute average indicates a transient event.
free displays memory usage: total, used, free, shared, buff/cache, and available . The critical column is available — an estimate of how much memory is available for new applications without swapping . The free column will often be low on a healthy system because the kernel uses unused RAM for disk caching. The buff/cache column shows that cache, and the available column shows how much of it can be reclaimed . The -h flag makes the output human-readable .
vmstat provides a single-line overview of CPU, memory, paging, block I/O, interrupts, and context switches . Running vmstat 1 5 shows five samples at one-second intervals, revealing trends rather than a snapshot. The key columns for CPU are us (user), sy (system), id (idle), and wa (wait for I/O). A high wa indicates disk I/O is the bottleneck, not CPU. The key columns for memory are swpd (swap used), free, buff, and cache, plus the si (swap in) and so (swap out) columns under the swap section. Sustained swap-in and swap-out activity indicates memory pressure .
# Load average
uptime
# Memory usage
free -h
# CPU, memory, I/O trends
vmstat 1 5
b. The Process Commands: top, ps, pidstat
Once you know the system is under pressure, the next step is finding the process responsible.
top is the interactive process monitor. It displays a live list of processes sorted by CPU usage by default . The header shows load average, task counts, CPU states, and memory usage — a complete summary in one screen. Pressing 1 toggles per-CPU display, showing whether a single core is pegged or the load is spread across all cores . Pressing M sorts by memory usage instead of CPU. Pressing P returns to CPU sort. The top output updates every few seconds, and htop provides a more colorful, navigable alternative when installed .
ps is the non-interactive process lister. The command ps -eo pcpu,pmem,pid,user,args | sort -r -k1 | head -6 lists the top five processes by CPU usage . Replacing -k1 with -k2 sorts by memory. The ps command is useful for scripting and for capturing a snapshot when you cannot run an interactive monitor.
pidstat is the per-process statistics tool, part of the sysstat package. It monitors individual tasks and can report CPU, memory, and I/O usage per process . pidstat -u 1 5 reports CPU usage for all processes at one-second intervals, five times. pidstat -r 1 5 reports memory usage . The -w option reports context switches, useful for diagnosing processes that are causing excessive task switching .
# Interactive process monitor
top
# Top 5 CPU processes
ps -eo pcpu,pmem,pid,user,args | sort -r -k1 | head -6
# Per-process memory usage
pidstat -r 1 5
c. The Historical Record: sar
sar is the system activity reporter, part of the sysstat package. It collects data in the background — by default every 10 minutes — and stores it in /var/log/sa/ files named saDD where DD is the day of the month . This is the only tool that answers the question “what was the system doing before the problem?”
The sysstat package must be installed and the sysstat service must be enabled for data collection to occur. Without it, sar has no history to report . Once enabled, sar accumulates CPU, memory, I/O, and network statistics that can be queried retroactively.
Common sar queries:
# CPU usage for all processors (live)
sar -u -P ALL 1 3
# CPU usage from a historical file
sar -u -f /var/log/sa/sa10
# Memory usage
sar -r
# Load average
sar -q
The -f option reads from a specific historical file, allowing you to examine a day that has already passed. The -s and -e options filter by start and end time within that file. For example, sar -u -s 14:00:00 -e 15:00:00 -f /var/log/sa/sa10 shows CPU usage between 2 PM and 3 PM on the 10th of the month .
The value of sar is in establishing a baseline. When a performance issue occurs, you compare the sar data from the problem period against the sar data from a normal period. Without the baseline, you cannot quantify how much the system has degraded .
Complete Example Session
This session demonstrates a complete investigation: spotting pressure with summary commands, finding the process with process commands, and checking history with sar.
# ============================================
# PART 1: CHECK LOAD AVERAGE
# ============================================
uptime
# 14:32:01 up 12 days, 3:45, 2 users, load average: 4.12, 3.87, 2.91
# The 1-minute average is higher than the 15-minute average.
# The system is under increasing load.
# ============================================
# PART 2: CHECK MEMORY
# ============================================
free -h
# total used free shared buff/cache available
# Mem: 31Gi 19Gi 575Mi 3.3Gi 12Gi 8.8Gi
# Swap: 8.0Gi 6.6Gi 1.4Gi
# 6.6 GiB of swap is in use. This is a sign of memory pressure.
# ============================================
# PART 3: CHECK SWAP ACTIVITY
# ============================================
vmstat 1 5
# procs -----------memory---------- ---swap-- -----io---- -system-- ------cpu-----
# r b swpd free buff cache si so bi bo in cs us sy id wa st
# 2 0 7102404 1392528 36 12335148 8 21 130 724 2851 19 15 7 77 0 0
# 0 0 7102404 1392560 36 12335188 0 0 0 0 5779 7246 14...
# si (swap in) and so (swap out) show activity.
# The system is paging memory to and from disk.
# ============================================
# PART 4: FIND THE TOP CPU PROCESSES
# ============================================
ps -eo pcpu,pmem,pid,user,args | sort -r -k1 | head -6
# %CPU %MEM PID USER COMMAND
# 95.2 2.1 1234 mysql /usr/sbin/mysqld
# 12.4 0.5 5678 apache2 /usr/sbin/apache2 -k start
# 8.1 0.3 9012 postgres /usr/lib/postgresql/14/bin/postgres
# mysql is consuming 95% of a CPU core.
# ============================================
# PART 5: INTERACTIVE MONITORING WITH TOP
# ============================================
top
# Press "1" to see per-CPU usage.
# Press "M" to sort by memory.
# Press "P" to sort by CPU.
# ============================================
# PART 6: PER-PROCESS MEMORY USAGE
# ============================================
pidstat -r 1 3
# Average: UID PID minflt/s majflt/s VSZ RSS %MEM Command
# Average: 106 1234 0.00 0.00 1234567 654321 2.10 mysqld
# mysql is using 2.10% of memory, but the high CPU usage
# and swap activity suggest memory is being used inefficiently.
# ============================================
# PART 7: CHECK HISTORICAL CPU USAGE
# ============================================
sar -u -f /var/log/sa/sa$(date +%d)
# 14:00:01 CPU %user %nice %system %iowait %steal %idle
# 14:10:01 all 15.23 0.00 7.45 0.00 0.00 77.32
# 14:20:01 all 14.87 0.00 7.12 0.00 0.00 78.01
# 14:30:01 all 42.11 0.00 28.34 0.00 0.00 29.55
# CPU usage jumped between 14:20 and 14:30.
# ============================================
# PART 8: CHECK HISTORICAL MEMORY USAGE
# ============================================
sar -r -f /var/log/sa/sa$(date +%d)
# 14:20:01 kbmemfree kbavail kbmemused %memused kbbuffers kbcached
# 14:20:01 1392560 9227464 18786044 60.23 36 12335188
# Memory usage has been climbing.
# ============================================
# PART 9: CHECK CONTEXT SWITCHES
# ============================================
pidstat -w 1 3
# Average: UID PID cswch/s nvcswch/s Command
# Average: 106 1234 1234.56 12.34 mysqld
# mysql is causing a high number of voluntary context switches.
# ============================================
# PART 10: TAKE ACTION
# ============================================
# If mysql is the problem, investigate the query log.
# If it is a runaway process, consider restarting the service.
# If it is a memory leak, the process may need to be restarted.
# Use sar data to quantify the problem before and after the fix.
The ten parts cover load average, memory, swap activity, top CPU processes, interactive top, per-process memory, historical CPU usage, historical memory usage, context switches, and taking action.
Quick Reference
The Summary Commands
| Command | Shows | Key Flags |
|---|---|---|
uptime | Load average (1, 5, 15 min) | — |
free | Memory and swap | -h human-readable, -s repeat |
vmstat | CPU, memory, I/O, swap | [delay] [count] |
The Process Commands
| Command | Shows | Key Flags |
|---|---|---|
top | Interactive process list | 1 per-CPU, M memory sort, P CPU sort |
ps | Process snapshot | -eo pcpu,pmem,pid,user,args |
pidstat | Per-process stats | -u CPU, -r memory, -w switches |
The Historical Command
| Command | Shows | Key Flags |
|---|---|---|
sar | Historical system activity | -u CPU, -r memory, -f file |
The Key Metrics
| Metric | Source | What It Means |
|---|---|---|
| Load average | uptime | Processes waiting for CPU |
| Available memory | free | Memory usable without swapping |
| Swap in/out | vmstat | Memory pressure |
%wa | vmstat | I/O wait (not CPU-bound) |
%user | vmstat | User-space CPU usage |
%system | vmstat | Kernel CPU usage |
The sysstat Package
| Component | Purpose |
|---|---|
sar | Historical activity reporter |
pidstat | Per-process statistics |
mpstat | Per-CPU statistics |
iostat | Per-device I/O statistics |
Best Practices
✅ Do This:
# Use human-readable output
free -h # ✅
# Check available memory, not free memory
free -h | grep Mem # ✅
# Use vmstat with an interval to see trends
vmstat 1 5 # ✅
# Sort ps output by CPU or memory
ps -eo pcpu,pmem,pid,user,args | sort -r -k1 | head -6 # ✅
# Enable sysstat for historical data
sudo systemctl enable sysstat # ✅
# Check sar data for the problem period
sar -u -f /var/log/sa/sa10 # ✅
❌ Don’t Do This:
# Don't panic at low free memory
free # ❌ available matters
# Don't rely on a single snapshot
top -n 1 # ❌ trends matter
# Don't forget to enable sysstat
sar -u # ❌ no data if disabled
# Don't ignore iowait
# High %wa means disk is the bottleneck, not CPU # ❌
Common Pitfalls
| Pitfall | Why It Happens | Fix |
|---|---|---|
| Low free memory looks alarming | Linux uses RAM for cache | Check available column |
| High load average but low CPU | I/O wait or uninterruptible sleep | Check vmstat wa column |
sar returns no data | sysstat not enabled | systemctl enable sysstat |
pidstat not found | sysstat not installed | Install sysstat package |
| CPU usage exceeds 100% | Multithreaded process across cores | Normal, check per-CPU with top 1 |
| Swap usage ignored | Only looking at RAM | Check free swap line and vmstat si/so |
Real-World Examples
1. Check Load Average
uptime
2. Check Memory Human-Readable
free -h
3. Watch Memory Every Second
free -h -s 1
4. Watch CPU and I/O Trends
vmstat 1 5
5. Interactive Process Monitor
top
6. Top 5 CPU Processes
ps -eo pcpu,pmem,pid,user,args | sort -r -k1 | head -6
7. Per-Process Memory
pidstat -r 1 5
8. Historical CPU Usage
sar -u -f /var/log/sa/sa10
9. Historical Memory Usage
sar -r -f /var/log/sa/sa10
10. Check Swap Devices
swapon -s
Visual
The Monitoring Layers
┌──────────────────────────────────────────────┐
│ LAYER 1: SUMMARY │
│ │
│ uptime → load average │
│ free → memory and swap │
│ vmstat → CPU, memory, I/O trends │
│ │
│ Question: Is there a problem? │
│ │
├──────────────────────────────────────────────┤
│ LAYER 2: PROCESS │
│ │
│ top → interactive process list │
│ ps → process snapshot │
│ pidstat → per-process stats │
│ │
│ Question: Which process is responsible? │
│ │
├──────────────────────────────────────────────┤
│ LAYER 3: HISTORY │
│ │
│ sar → historical activity │
│ /var/log/sa/ → stored data files │
│ │
│ Question: What was the baseline? │
│ │
└──────────────────────────────────────────────┘
The free Output
┌──────────────────────────────────────────────┐
│ free -h │
│ │
│ total used free buff/cache │
│ Mem: 31Gi 19Gi 575Mi 12Gi │
│ Swap: 8.0Gi 6.6Gi 1.4Gi │
│ │
│ available: 8.8Gi │
│ │
│ free = idle and unused │
│ available = usable without swapping │
│ buff/cache = disk cache (reclaimable) │
│ │
│ Low free is normal. Low available is not. │
│ │
└──────────────────────────────────────────────┘
vmstat Columns
┌──────────────────────────────────────────────┐
│ vmstat 1 5 │
│ │
│ procs: r (runnable) b (blocked) │
│ │
│ memory: swpd free buff cache │
│ │
│ swap: si (in) so (out) │
│ ─── high si/so = memory pressure ─── │
│ │
│ io: bi (blocks in) bo (blocks out) │
│ │
│ system: in (interrupts) cs (context sw) │
│ │
│ cpu: us sy id wa st │
│ ─── high wa = I/O bound, not CPU bound ─── │
│ │
└──────────────────────────────────────────────┘
The sar Data Flow
┌──────────────────────────────────────────────┐
│ SAR COLLECTION AND QUERY │
│ │
│ sysstat service (every 10 min) │
│ │ │
│ ▼ │
│ /var/log/sa/saDD │
│ │ │
│ ▼ │
│ sar -u -f /var/log/sa/sa10 │
│ │ │
│ ▼ │
│ Historical CPU usage for day 10 │
│ │
│ Compare problem period to baseline. │
│ │
└──────────────────────────────────────────────┘
Summary
| Item | Value |
|---|---|
| Load average | uptime |
| Memory summary | free -h |
| CPU/memory/I/O trends | vmstat 1 5 |
| Interactive process monitor | top |
| Process snapshot | ps -eo pcpu,pmem,pid,user,args |
| Per-process stats | pidstat -u, pidstat -r, pidstat -w |
| Historical activity | sar |
| Data files | /var/log/sa/saDD |
| sysstat package | Provides sar, pidstat, mpstat, iostat |
| Critical memory metric | available, not free |
| Critical CPU metric | %wa for I/O wait |
Key takeaways:
- Monitoring works in layers: summary, process, history. Start with
uptime,free, andvmstatto determine whether there is a problem. Move totop,ps, andpidstatto find the process responsible. Usesarto compare against a baseline . - Load average measures waiting processes, not just CPU. A load average of 4.0 on a 4-core system means full utilization. A load average that is high but CPU usage is low usually indicates I/O wait or uninterruptible sleep .
freememory is not the metric to watch. Linux uses unused RAM for disk caching. Theavailablecolumn estimates how much memory can be used without swapping . Lowfreewith highavailableis normal. Lowavailablewith rising swap is a problem .vmstatreveals the type of pressure. Highsiandsoindicate memory pressure. Highwaindicates disk I/O is the bottleneck, not CPU. Highusindicates user-space CPU consumption .pidstatattributes usage to processes.pidstat -ufor CPU,pidstat -rfor memory,pidstat -wfor context switches. It is part of thesysstatpackage and reports at configurable intervals .saris the historical record. It collects data every 10 minutes by default and stores it in/var/log/sa/. Withoutsarrunning before a problem, the investigation is limited to the present .- The
sysstatpackage provides the monitoring toolkit.sar,pidstat,mpstat, andiostatare all part of it. It must be installed and the service enabled for historical data collection .
Remember: CPU and memory monitoring is not about watching numbers. It is about answering three questions in order: Is there a problem? What is causing it? What was normal before? uptime, free, and vmstat answer the first. top, ps, and pidstat answer the second. sar answers the third. Learn the layers, learn which metric matters in each layer, and the investigation becomes a sequence rather than a guess. The LFCA exam tests whether you know the tools and whether you know what they mean.
Stop using slow, ad-bloated tool sites! 🤮
🔎 Search “KandZ Tools” on Google to use many professional utilities for free.
KandZ.me is the ultimate minimalist hub for:
✅ Finance (Mortgage, Interest, Inflation)
✅ Tech (Base64, JSON, Dev Suite, IP)
✅ Health (BMI, BMR, TDEE)
✅ Productivity (Timer, Workspace, QR)
⚡️ Fast & Private
🔒 No data leaves your device
💎 100% Free
🔗 Use it now: https://tools.kandz.me
🔖 Bookmark it—you’ll need it later!