LFCA 117 🐧 Common Troubleshooting Scenarios
Knowing the method—assessment, resource check, logs, network, resolution—is only half the battle. The other half is recognizing which scenario you are facing and applying the method efficiently. A service that refuses to start is not the same problem as a system that crawls under load. A network that drops intermittently requires different tools than a disk that fills up silently. Each scenario has characteristic symptoms and a natural first place to look.
This chapter covers the most common troubleshooting scenarios you will encounter as a system administrator. For each scenario, you will see the symptom, the likely causes, and the specific commands that move the investigation forward. The goal is to build pattern recognition: when you see a service fail, your fingers should already be reaching for journalctl -u. When a system is slow, htop and iostat should be your first instinct.
By the end, you will have a mental catalog of scenarios mapped to diagnostic paths. You will also understand why certain commands belong to certain scenarios—not as memorized trivia but as logical consequences of what each scenario actually involves.
Key point: Most troubleshooting scenarios fall into a small number of categories: service failures, resource exhaustion, network problems, and configuration errors. Recognizing the category tells you where to start, and starting in the right place saves the most time.
Why scenario recognition matters
The speed problem. A methodical approach is thorough, but thoroughness alone is not enough under pressure. When a production service is down, every minute of investigation has a cost. Scenario recognition shortcuts the “where do I start” question. You do not need to re-derive the entire troubleshooting process from first principles for every incident. You recognize the pattern—service failed, system slow, network unreachable—and you start with the tools that are most likely to reveal the cause.
The LFCA exam problem. The LFCA certification tests troubleshooting as scenario-based reasoning. You will be given a description of a problem and asked which command or diagnostic step is appropriate . The exam does not expect you to know every Linux command. It expects you to know which class of command applies to which class of problem. Recognizing that “service won’t start” maps to systemctl status and journalctl -u is exactly the kind of reasoning the exam evaluates.
The common failure patterns. Linux systems fail in predictable ways. Services crash because of configuration errors, missing dependencies, or permission problems. Systems slow down because of CPU hogs, memory leaks, or disk I/O saturation. Networks fail because of interface issues, DNS problems, or firewall rules. These patterns repeat across distributions, across applications, and across environments. Learning to recognize them is more valuable than memorizing every flag of every command.
The escalation problem. When you escalate an issue to a senior administrator or a vendor, the first question is always the same: “What have you already tried?” A structured investigation—even one that has not yet found the answer—demonstrates that you have ruled out the obvious causes. Scenario recognition lets you rule out entire categories quickly, so the escalation includes useful negative results as well as the positive symptoms.
a. Service failures
A service that fails to start is the most common troubleshooting scenario in system administration. The symptom is clear: you run systemctl start service, and it does not start. Or it starts and immediately dies. Or it starts but refuses connections.
The first command is always the same: check the service status.
systemctl status myservice
The status output tells you several critical things: whether the service is active, inactive, or failed; the exit code or signal that terminated the process; and the last few log lines from the service . The exit code is often the most diagnostic piece of information.
journalctl -u myservice -n 100 --no-pager
The journalctl command with -u filters the systemd journal to a specific unit . The -n 100 shows the last hundred lines. The --no-pager option ensures the output goes directly to the terminal rather than opening in a pager, which is useful for scripting or for copying output into a ticket.
Common failure patterns emerge from the exit code. Exit code 203/EXEC means systemd could not execute the binary specified in ExecStart—usually a typo in the path or a missing package . Exit code 217/USER means the User= directive in the unit file specifies a user that does not exist . Exit code 200/CHDIR means the WorkingDirectory= points to a directory that does not exist . Exit codes 1-199 typically indicate application-level failures—the binary ran but crashed, and you need the application’s own logs, not just the systemd journal .
A subtle class of service failures comes from systemd sandboxing directives. Modern unit files often include ProtectSystem=, PrivateTmp=, or ReadOnlyPaths= to harden services . These directives can cause a service to fail in ways that do not reproduce when you run the command manually as the service user. The symptom is characteristic: the command works in a shell but fails under systemctl start with a permission denied or read-only filesystem error . Checking the effective unit configuration with systemctl cat myservice reveals these directives.
b. High load and resource exhaustion
A system that is slow, unresponsive, or refusing connections may be experiencing resource exhaustion. The four fundamental resources are CPU, memory, disk space, and disk I/O. Each has a characteristic symptom and a characteristic diagnostic command.
The first command for any high-load investigation is htop or top.
htop
# or
top -bn1 | head -20
htop provides an interactive view of processes sorted by CPU or memory usage. Pressing F6 allows sorting by different metrics . This quickly identifies the process consuming the most resources. If no single process stands out, the problem may be I/O wait or memory pressure rather than raw CPU consumption.
Memory exhaustion has a specific signature: the system starts swapping to disk, which degrades performance dramatically. The free -h command shows memory and swap usage in human-readable format .
free -h
If swap is heavily used and ps aux --sort=-%mem | head shows a process consuming an unusual amount of memory, you may be looking at a memory leak . The kernel’s OOM killer may also terminate processes to free memory, which appears in dmesg as “Out of memory: Killed process” .
Disk space is the most straightforward resource to check.
df -h
df -i
The -h flag shows disk usage in human-readable units. The -i flag shows inode usage. A filesystem can be completely full in inodes while still reporting free space in bytes, a common cause of mysterious “disk full” errors when df -h shows plenty of space .
Disk I/O saturation is subtler. The iostat -x 1 5 command shows extended disk statistics over five one-second intervals . The %util column shows disk utilization as a percentage. Values close to 100% indicate a saturated disk. The await column shows average I/O wait time in milliseconds; values over 20ms indicate slow disk response .
c. Network problems
Network problems have their own characteristic diagnostic sequence. The method is layered: test the interface, then the gateway, then DNS, then the target service .
The first check is whether the interface is up and has an IP address.
ip addr show
If the interface is down or has no address, the problem is at the link or configuration layer . If the interface is up but you cannot reach the gateway, the problem may be a cable, switch port, or VLAN mismatch .
DNS problems are a separate class. The symptom is that IP addresses work but hostnames do not. The dig command tests DNS resolution directly.
dig example.com
cat /etc/resolv.conf
If DNS fails, the problem is usually in /etc/resolv.conf or the systemd-resolved configuration . A missing or misconfigured nameserver entry causes all hostname lookups to fail.
Firewall rules are a common cause of “connection refused” or “connection timed out” errors. The commands depend on which firewall system is in use.
# nftables
sudo nft list ruleset
# firewalld (RHEL/Fedora)
sudo firewall-cmd --list-all
# ufw (Ubuntu)
sudo ufw status verbose
If a firewall rule is blocking traffic, temporarily disabling the firewall can confirm the diagnosis . But disabling a firewall on a production system is a security risk; the rule should be adjusted rather than the firewall disabled.
Intermittent network drops are the most difficult network problem to diagnose because the symptom is absent when you look for it. The approach is to set up continuous monitoring that logs diagnostic information when a drop occurs .
ip -s link show eth0
This command shows interface error counters. Increasing error counts indicate physical layer problems—cable, NIC, or switch port issues . The ethtool -S eth0 command provides driver-level statistics that may reveal hardware errors .
Complete Example Session
# ============================================
# PART 1: SERVICE FAILURE — INITIAL STATUS
# ============================================
systemctl status nginx
# Output shows: Active: failed (Result: exit-code)
# Process: 1234 ExecStart=/usr/sbin/nginx (code=exited, status=1/FAILURE)
# ============================================
# PART 2: SERVICE FAILURE — JOURNAL LOGS
# ============================================
journalctl -u nginx -n 50 --no-pager
# Output shows: nginx: [emerg] bind() to 0.0.0.0:80 failed (98: Address already in use)
# ============================================
# PART 3: SERVICE FAILURE — PORT CONFLICT
# ============================================
ss -tulpn | grep :80
# Output shows another process using port 80
# ============================================
# PART 4: HIGH LOAD — IDENTIFY PROCESS
# ============================================
htop
# Or: ps aux --sort=-%cpu | head -10
# Shows process consuming 95% CPU
# ============================================
# PART 5: HIGH LOAD — MEMORY CHECK
# ============================================
free -h
# Output shows swap heavily used, little free RAM
# ============================================
# PART 6: HIGH LOAD — DISK I/O CHECK
# ============================================
iostat -x 1 5
# Output shows %util near 100% on /dev/sda
# ============================================
# PART 7: NETWORK — INTERFACE AND GATEWAY
# ============================================
ip addr show
ping -c 3 $(ip route | grep default | awk '{print $3}')
# ============================================
# PART 8: NETWORK — DNS RESOLUTION
# ============================================
dig example.com
cat /etc/resolv.conf
# ============================================
# PART 9: NETWORK — FIREWALL CHECK
# ============================================
sudo ufw status verbose
# or: sudo firewall-cmd --list-all
# ============================================
# PART 10: RESOLUTION AND VERIFICATION
# ============================================
systemctl restart nginx
systemctl status nginx
curl -I http://localhost
The ten parts moved through the three major scenario categories: service failure diagnosis through status and logs, high load investigation through CPU, memory, and I/O checks, and network troubleshooting through interface, gateway, DNS, and firewall layers.
Quick Reference
Service Failure Diagnosis
| Symptom | Command | What to Look For |
|---|---|---|
| Service won’t start | systemctl status svc | Exit code, last log lines |
| Need full logs | journalctl -u svc -n 100 | Error messages before failure |
| Config syntax | nginx -t or apachectl configtest | Syntax errors |
| Port conflict | ss -tulpn | grep :PORT | Another process on port |
| Permission issue | systemctl cat svc | Sandboxing directives |
Resource Exhaustion
| Resource | Command | Warning Sign |
|---|---|---|
| CPU | htop or top | Process at 90%+ |
| Memory | free -h | Swap heavily used |
| Disk space | df -h and df -i | 100% usage (bytes or inodes) |
| Disk I/O | iostat -x 1 5 | %util near 100%, await > 20ms |
| OOM kills | dmesg | grep -i "out of memory" | Killed process messages |
Network Diagnostics
| Layer | Command | Failure Indicates |
|---|---|---|
| Interface | ip addr show | No IP or interface down |
| Gateway | ping gateway_ip | Link or routing problem |
| DNS | dig domain | Nameserver misconfiguration |
| Target | curl -v http://target | Service or firewall issue |
| Firewall | ufw status or firewall-cmd --list-all | Blocking rules |
Best Practices
✅ Do This:
# Check service status before anything else
systemctl status myservice # ✅
# Read full logs before restarting
journalctl -u myservice -n 100 --no-pager # ✅
# Check all four resources
htop && free -h && df -h && iostat -x 1 5 # ✅
# Test network layers in order
ip addr && ping gateway && dig domain && curl target # ✅
# Check inodes when df -h shows free space
df -i # ✅
# Verify configuration syntax before restart
nginx -t && systemctl reload nginx # ✅
❌ Don’t Do This:
# Restart without reading logs
systemctl restart myservice # ❌
# Check only one resource
free -h # ❌
# Test DNS before interface
dig example.com # ❌ (check interface first)
# Disable firewall on production
sudo ufw disable # ❌
# Ignore exit codes
journalctl -u myservice | tail # ❌ (look for exit code)
Common Pitfalls
| Pitfall | Why It Happens | Fix |
|---|---|---|
| Restarting without logs | Impatience or habit | journalctl -u first |
| Missing inode exhaustion | Only checking df -h | Also run df -i |
| Wrong port in firewall | Service listens on different port | Verify with ss -tulpn |
| DNS misdiagnosed as network | Hostname fails, IP works | Test dig separately |
| Exit code ignored | Reading logs but not status | systemctl status shows code |
| Sandboxing failure missed | Works manually, fails in systemd | systemctl cat reveals directives |
| I/O wait mistaken for CPU | High load but low CPU usage | iostat reveals disk bottleneck |
Real-World Examples
1. Service Failed — Port Conflict
systemctl status nginx
# status=1/FAILURE
ss -tulpn | grep :80
2. Service Failed — Bad ExecStart Path
journalctl -u myservice -n 20
# ExecStart path not found
ls -l /path/to/binary
3. High Load — Runaway Process
ps aux --sort=-%cpu | head -5
kill -15 PID
4. High Load — Memory Leak
free -h
ps aux --sort=-%mem | head -5
5. Disk Full — Large Log Files
df -h
du -sh /var/log/* | sort -rh | head -10
6. Disk Full — Inode Exhaustion
df -i
find / -xdev -type f | wc -l
7. Network — Interface Down
ip addr show
ip link set eth0 up
8. Network — DNS Failure
dig example.com
cat /etc/resolv.conf
9. Network — Firewall Blocking
sudo ufw status
sudo ufw allow 443/tcp
10. System Unresponsive — OOM Kill
dmesg | grep -i "out of memory"
free -h
Visual
Scenario to Diagnostic Path
┌─────────────────────────────────────────────────────────────┐
│ SCENARIO RECOGNITION MAP │
│ │
│ Symptom: Service won't start │
│ └──▶ systemctl status svc │
│ └──▶ journalctl -u svc -n 50 │
│ └──▶ Check exit code, port conflict, config │
│ │
│ Symptom: System is slow │
│ └──▶ htop / top │
│ └──▶ free -h, df -h, iostat -x 1 5 │
│ └──▶ Identify CPU, memory, disk, or I/O │
│ │
│ Symptom: Network unreachable │
│ └──▶ ip addr show │
│ └──▶ ping gateway │
│ └──▶ dig domain │
│ └──▶ curl target │
│ └──▶ firewall check │
│ │
└─────────────────────────────────────────────────────────────┘
Service Failure Decision Tree
┌─────────────────────────────────────────────────────────────┐
│ SERVICE FAILED TO START │
│ │
│ systemctl status svc │
│ │ │
│ ├── Exit 203/EXEC ──▶ Binary path wrong or missing │
│ │ │
│ ├── Exit 217/USER ──▶ User= doesn't exist │
│ │ │
│ ├── Exit 200/CHDIR ──▶ WorkingDirectory missing │
│ │ │
│ ├── Exit 1-199 ────▶ Application crash (check app logs) │
│ │ │
│ └── Port conflict ─▶ ss -tulpn | grep :PORT │
│ │
└─────────────────────────────────────────────────────────────┘
Resource Exhaustion Symptom Map
┌─────────────────────────────────────────────────────────────┐
│ HIGH LOAD — WHICH RESOURCE? │
│ │
│ htop shows high CPU usage? │
│ └──▶ ps aux --sort=-%cpu | head │
│ │
│ free -h shows swap used? │
│ └──▶ ps aux --sort=-%mem | head │
│ └──▶ dmesg | grep -i "out of memory" │
│ │
│ df -h shows full? │
│ └──▶ du -sh /var/log/* | sort -rh │
│ │
│ iostat shows %util 100%? │
│ └──▶ Check disk health, large I/O processes │
│ │
└─────────────────────────────────────────────────────────────┘
Network Layered Diagnosis
┌─────────────────────────────────────────────────────────────┐
│ NETWORK TROUBLESHOOTING LAYERS │
│ │
│ Layer 1: Interface │
│ ip addr show → is the interface UP with an IP? │
│ │
│ Layer 2: Gateway │
│ ping gateway → can we reach the local router? │
│ │
│ Layer 3: DNS │
│ dig domain → does the name resolve? │
│ │
│ Layer 4: Service │
│ curl -v → is the service listening? │
│ │
│ Layer 5: Firewall │
│ ufw status → is traffic blocked? │
│ │
│ Each layer assumes the previous layer works. │
│ │
└─────────────────────────────────────────────────────────────┘
Summary
| Item | Value |
|---|---|
| Scenario category | Service failures |
| First command | systemctl status svc |
| Log command | journalctl -u svc -n 100 |
| Scenario category | High load / resource exhaustion |
| First command | htop or top |
| Resource commands | free -h, df -h, iostat -x 1 5 |
| Scenario category | Network problems |
| First command | ip addr show |
| Diagnostic order | Interface → gateway → DNS → service → firewall |
| Common hidden issue | Inode exhaustion (df -i) |
| Service exit code 203 | Binary path wrong or missing |
| Service exit code 217 | User= doesn’t exist |
Key takeaways:
- Service failures start with
systemctl statusandjournalctl -u. The status output gives the exit code; the journal gives the context. Together they identify the failure class. - Exit codes are diagnostic. 203/EXEC means bad path, 217/USER means bad user, 200/CHDIR means bad working directory. Exit codes 1-199 mean the application ran but crashed.
- High load is not always CPU. Check memory with
free -h, disk space withdf -h, and disk I/O withiostat -x. The bottleneck may be anywhere. - Inode exhaustion masquerades as disk space.
df -hcan show free space whiledf -ishows 100% inode usage. Both checks are required. - Network diagnostics are layered. Test the interface, then the gateway, then DNS, then the service, then the firewall. Each layer builds on the previous one.
- Sandboxing directives cause service failures that don’t reproduce manually. If a command works in a shell but fails under
systemctl, checksystemctl catforProtectSystem=,PrivateTmp=, and similar directives. - The LFCA exam tests scenario recognition. Knowing which command belongs to which scenario is the core skill the exam evaluates. Practice mapping symptoms to diagnostic paths.
Remember: Troubleshooting scenarios are patterns, not unique events. A service that fails to start follows the same diagnostic path whether it runs on Ubuntu or RHEL, whether it is nginx or PostgreSQL. A system under high load reveals its bottleneck through the same four resource checks regardless of what application is running. The value of scenario recognition is that it turns novel incidents into recognizable shapes. You have seen this before—not this exact service or this exact load spike, but this category of problem. The category tells you where to start, and starting in the right place is half the work.
Stop using slow, ad-bloated tool sites! 🤮
🔎 Search “KandZ Tools” on Google to use many professional utilities for free.
KandZ.me is the ultimate minimalist hub for:
✅ Finance (Mortgage, Interest, Inflation)
✅ Tech (Base64, JSON, Dev Suite, IP)
✅ Health (BMI, BMR, TDEE)
✅ Productivity (Timer, Workspace, QR)
⚡️ Fast & Private
🔒 No data leaves your device
💎 100% Free
🔗 Use it now: https://tools.kandz.me
🔖 Bookmark it—you’ll need it later!