| |

LFCA 116 ๐Ÿง Troubleshooting โ€” A Methodical Approach

Troubleshooting is not guessing. It is not randomly restarting services or rebooting the machine in the hope that something changes. A methodical approach to troubleshooting is a repeatable process: gather information, form a hypothesis, test it, and either confirm or discard it. The difference between a technician who solves problems and one who creates new ones is discipline.

This chapter covers troubleshooting as a system administration competency within the LFCA System Administration Fundamentals domain. You will learn the phases of a structured investigation: initial assessment, resource analysis, process investigation, log analysis, network diagnostics, service troubleshooting, and resolution. Each phase has specific commands and specific questions it answers. The goal is not to memorize every tool but to internalize the order in which questions should be asked.

By the end, you will understand why checking logs comes before restarting services, why documenting findings matters even when the fix seems obvious, and how to avoid the trap of fixing symptoms instead of causes.

Key point: Troubleshooting is a process of elimination, not a process of inspiration. Start broad, narrow systematically, and never change more than one thing at a time.


Why a methodical approach exists

The random-fix problem. When a service fails, the instinct is to try the first solution that comes to mind. Restart the service. Reboot the server. Reinstall the package. Sometimes this works. More often, it masks the problem temporarily, and when the failure recurs, the technician has no idea what actually changed or why the fix worked. A methodical approach eliminates this uncertainty by ensuring every action is based on evidence, not guesswork.

The complexity problem. A modern Linux system has hundreds of interacting components: kernel, systemd, network stack, filesystem, application services, security modules. A single symptom can have dozens of possible causes. Without a structured investigation, the search space is impossibly large. The methodical approach narrows that space layer by layer, from general system health to specific component behavior.

The documentation problem. The LFCA exam tests troubleshooting as scenario-based reasoning, not just command recognition . You will be given a symptom and asked to identify the correct diagnostic step. This requires understanding why certain checks come before others. Checking network connectivity before checking whether the network interface is up is backwards. Checking disk space after investigating application logs is inefficient. The order matters, and the exam tests that understanding.

The production problem. On a production system, every action has consequences. Restarting a service disrupts users. Rebooting a server takes down applications. Running untested commands as root can cause data loss. A methodical approach minimizes disruption by requiring evidence before action. You do not restart a service until you have identified the reason it failed, because restarting without understanding means the failure will recur.

The prevention problem. The final phase of troubleshooting is not “the problem is fixed.” It is “the problem is understood, documented, and prevented from recurring.” A methodical approach produces documentation that feeds into monitoring, alerting, and capacity planning. It turns a one-time incident into institutional knowledge.


a. Initial assessment and information gathering

The first phase is not about fixing anything. It is about understanding the problem. Before you run a single diagnostic command, you need answers to basic questions: What is the symptom? When did it start? What changed recently? Who is affected?

The commands in this phase are about orientation. uptime tells you how long the system has been running and what the load average looks like. hostnamectl identifies the operating system and kernel version. dmesg | tail -50 shows the most recent kernel messages, which often contain hardware errors, driver failures, or boot-time warnings .

uptime
hostnamectl
cat /etc/os-release
dmesg | tail -50

These commands do not diagnose the problem. They establish the baseline. A load average of 12 on a 4-core system is a signal. A kernel version that is unexpectedly old is a signal. A dmesg output full of ATA errors is a signal. You are collecting evidence before forming a hypothesis.

The output of these commands should be documented. If the problem requires escalation to a colleague or a support ticket, this is the information they will ask for first.


b. Resource analysis

Once you know what the system is and how long it has been running, the next question is whether the system has the resources it needs. CPU, memory, disk space, and I/O are the four fundamental resources. A deficiency in any one of them can manifest as application slowness, service failures, or unresponsive behavior .

top -bn1 | head -20
free -h
df -h
iostat -x 1 5

top -bn1 provides a snapshot of CPU usage and the most active processes. free -h shows memory availability in human-readable units. df -h shows disk space on all mounted filesystems. iostat -x 1 5 shows disk I/O statistics over a five-second interval .

The key insight is that resource problems often present as application problems. A web server that refuses connections might be out of file descriptors, not broken. A database that runs slowly might be waiting on disk I/O, not misconfigured. Checking resources before checking application logs prevents wasted effort in the wrong layer of the stack.

A common pitfall is checking only one resource. Memory exhaustion can cause the kernel to kill processes. Disk exhaustion can prevent services from writing logs or temporary files. High I/O wait can make a system appear unresponsive even when CPU utilization is low. The four checks together provide a complete picture.


c. Process investigation, log analysis, and network diagnostics

If resources are adequate but the problem persists, the investigation moves to processes, logs, and network. These three areas overlap, and the order depends on the symptom. A service that will not start suggests process investigation. An application that throws errors suggests log analysis. A connectivity failure suggests network diagnostics.

Process investigation begins with identifying what is running and what is consuming resources.

ps aux --sort=-%cpu | head -10
pstree -p
lsof -p PID
strace -p PID

ps aux --sort=-%cpu shows the top CPU consumers. pstree -p shows process parent-child relationships, which reveals whether a service is managed by systemd or spawned manually. lsof -p PID lists all files and network connections opened by a process. strace -p PID traces system calls, which is useful for debugging a process that hangs or fails silently .

Log analysis is the most information-dense phase of troubleshooting. Linux maintains logs in /var/log/ and through systemd’s journal . The most important files are /var/log/syslog or /var/log/messages (general system messages), /var/log/auth.log (authentication events), and /var/log/kern.log (kernel messages) . For systemd-managed services, journalctl -u service_name provides service-specific logs .

journalctl -xe
tail -f /var/log/syslog
grep -i error /var/log/*

journalctl -xe shows recent systemd journal entries with explanations. tail -f follows a log file in real time, which is useful for watching a problem reproduce. grep -i error searches across log files for error messages .

A critical principle in log analysis is correlation. A single error message in isolation may be a red herring. The same error repeated across multiple services at the same timestamp suggests a shared causeโ€”a disk failure, a network outage, or a resource exhaustion event. The timeline of events is often more informative than any individual message.

Network diagnostics follow a layer-by-layer approach . The first question is whether the network interface is up. The second is whether the system can reach its gateway. The third is whether DNS resolution works. The fourth is whether the target service is reachable.

ip addr show
ss -tulpn
curl -v http://target
dig domain

ip addr show displays interface configuration. ss -tulpn shows listening ports and active connections. curl -v tests HTTP connectivity with verbose output. dig tests DNS resolution .

The layering matters. Testing DNS before confirming the interface is up is wasted effort. Testing HTTP connectivity before confirming DNS works is wasted effort. Each test builds on the previous one.


Complete Example Session

# ============================================
# PART 1: INITIAL ASSESSMENT
# ============================================
uptime
hostnamectl
dmesg | tail -20
# ============================================
# PART 2: RESOURCE CHECK
# ============================================
free -h
df -h
top -bn1 | head -15
# ============================================
# PART 3: PROCESS INSPECTION
# ============================================
ps aux --sort=-%mem | head -10
systemctl status nginx
# ============================================
# PART 4: LOG EXAMINATION
# ============================================
journalctl -u nginx --since "1 hour ago"
tail -20 /var/log/nginx/error.log
# ============================================
# PART 5: NETWORK VERIFICATION
# ============================================
ip addr show
ss -tulpn | grep :80
curl -v http://localhost
# ============================================
# PART 6: CONFIGURATION CHECK
# ============================================
nginx -t
ls -la /etc/nginx/sites-enabled/
# ============================================
# PART 7: DEPENDENCY VERIFICATION
# ============================================
systemctl list-dependencies nginx
ldd $(which nginx)
# ============================================
# PART 8: SECURITY LAYER CHECK
# ============================================
getenforce
iptables -L -n
# ============================================
# PART 9: RESOLUTION AND VERIFICATION
# ============================================
systemctl restart nginx
systemctl status nginx
curl -I http://localhost
# ============================================
# PART 10: DOCUMENTATION
# ============================================
# Record in ticket:
# - Symptom: nginx not responding on port 80
# - Root cause: configuration syntax error in sites-enabled/default
# - Fix: corrected missing semicolon, ran nginx -t, restarted
# - Prevention: add nginx -t to deployment pipeline

The ten parts moved through the troubleshooting workflow: orientation, resource check, process inspection, log analysis, network verification, configuration validation, dependency check, security layer review, resolution, and documentation.


Quick Reference

Troubleshooting Phases

PhasePurposeKey Commands
AssessmentUnderstand the problemuptime, dmesg, hostnamectl
ResourcesCheck CPU, memory, disk, I/Otop, free, df, iostat
ProcessesIdentify running servicesps, systemctl status, lsof
LogsFind error messagesjournalctl, tail, grep
NetworkVerify connectivityip, ss, curl, dig
ResolutionFix and verifysystemctl restart, test

Log File Locations

FileContent
/var/log/syslogGeneral system messages
/var/log/auth.logAuthentication events
/var/log/kern.logKernel messages
/var/log/dmesgBoot-time kernel messages
/var/log/nginx/Nginx-specific logs

Diagnostic Commands

CommandPurpose
journalctl -u svcService logs from systemd
journalctl -xeRecent journal entries with details
`dmesgtail`
ss -tulpnListening ports and connections
systemctl status svcService state and recent logs

Best Practices

โœ… Do This:

# Check logs before restarting services
journalctl -u nginx --since "10 min ago"                    # โœ…

# Document findings as you investigate
echo "14:32 - nginx error: bind() failed" >> /tmp/troubleshoot.log  # โœ…

# Verify one layer at a time
ping gateway && dig domain && curl localhost                # โœ…

# Test configuration after changes
nginx -t && systemctl reload nginx                          # โœ…

โŒ Don’t Do This:

# Restart without understanding the failure
systemctl restart nginx                                     # โŒ

# Change multiple things at once
sed -i 's/80/8080/' config && systemctl restart nginx       # โŒ

# Ignore logs because the fix worked
# (the same problem will recur)                             # โŒ

# Skip documentation
# (no record of what was done or why)                       # โŒ

Common Pitfalls

PitfallWhy It HappensFix
Fixing symptom, not causeRestart masks underlying issueRead logs before restarting
Skipping resource checkAssumes application is at faultCheck CPU, memory, disk first
Ignoring timestampsIndividual errors misleadCorrelate events by time
Changing multiple variablesUnclear which change fixed itChange one thing at a time
Not documentingKnowledge lost after resolutionRecord findings immediately

Real-World Examples

1. Service Fails to Start

systemctl status myservice
journalctl -u myservice --no-pager | tail -30

2. Disk Full Prevents Logging

df -h
du -sh /var/log/* | sort -rh | head -10

3. High Load Average

uptime
top -bn1 | head -15

4. Network Unreachable

ip addr show
ping -c 3 8.8.8.8

5. DNS Resolution Failure

dig example.com
cat /etc/resolv.conf

6. Port Already in Use

ss -tulpn | grep :8080
lsof -i :8080

7. Permission Denied

ls -la /path/to/file
namei -l /path/to/file

8. SSH Connection Refused

systemctl status sshd
ss -tulpn | grep :22

9. Memory Exhaustion

free -h
ps aux --sort=-%mem | head -5

10. Configuration Syntax Error

nginx -t
apachectl configtest

Visual

The Troubleshooting Funnel

โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚  METHODICAL TROUBLESHOOTING FUNNEL                          โ”‚
โ”‚                                                             โ”‚
โ”‚  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”    โ”‚
โ”‚  โ”‚  INITIAL ASSESSMENT                                 โ”‚    โ”‚
โ”‚  โ”‚  What is the system? What changed?                  โ”‚    โ”‚
โ”‚  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜    โ”‚
โ”‚                        โ–ผ                                    โ”‚
โ”‚  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”    โ”‚
โ”‚  โ”‚  RESOURCE CHECK                                     โ”‚    โ”‚
โ”‚  โ”‚  CPU, memory, disk, I/O โ€” is the system healthy?    โ”‚    โ”‚
โ”‚  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜    โ”‚
โ”‚                        โ–ผ                                    โ”‚
โ”‚  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”    โ”‚
โ”‚  โ”‚  PROCESS & LOG ANALYSIS                             โ”‚    โ”‚
โ”‚  โ”‚  What is running? What does it say?                 โ”‚    โ”‚
โ”‚  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜    โ”‚
โ”‚                        โ–ผ                                    โ”‚
โ”‚  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”    โ”‚
โ”‚  โ”‚  NETWORK VERIFICATION                               โ”‚    โ”‚
โ”‚  โ”‚  Is the path to the service clear?                  โ”‚    โ”‚
โ”‚  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜    โ”‚
โ”‚                        โ–ผ                                    โ”‚
โ”‚  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”    โ”‚
โ”‚  โ”‚  RESOLUTION                                         โ”‚    โ”‚
โ”‚  โ”‚  Fix, verify, document.                             โ”‚    โ”‚
โ”‚  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜    โ”‚
โ”‚                                                             โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜

Layer-by-Layer Network Diagnosis

โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚  NETWORK TROUBLESHOOTING LAYERS                             โ”‚
โ”‚                                                             โ”‚
โ”‚  Layer 1: Interface                                         โ”‚
โ”‚  ip addr show โ†’ is the interface UP with an IP?             โ”‚
โ”‚                                                             โ”‚
โ”‚  Layer 2: Gateway                                           โ”‚
โ”‚  ping gateway โ†’ can we reach the local router?              โ”‚
โ”‚                                                             โ”‚
โ”‚  Layer 3: DNS                                               โ”‚
โ”‚  dig domain โ†’ does the name resolve?                        โ”‚
โ”‚                                                             โ”‚
โ”‚  Layer 4: Service                                           โ”‚
โ”‚  curl localhost โ†’ is the service listening?                 โ”‚
โ”‚                                                             โ”‚
โ”‚  Layer 5: Firewall                                          โ”‚
โ”‚  iptables -L โ†’ is traffic blocked?                          โ”‚
โ”‚                                                             โ”‚
โ”‚  Each layer assumes the previous layer is working.          โ”‚
โ”‚                                                             โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜

The Investigation Loop

โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚  HYPOTHESIS-DRIVEN TROUBLESHOOTING                          โ”‚
โ”‚                                                             โ”‚
โ”‚          โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”                               โ”‚
โ”‚          โ”‚   OBSERVE        โ”‚                               โ”‚
โ”‚          โ”‚   (symptoms)     โ”‚                               โ”‚
โ”‚          โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜                               โ”‚
โ”‚                   โ–ผ                                         โ”‚
โ”‚          โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”                               โ”‚
โ”‚          โ”‚   HYPOTHESIZE    โ”‚                               โ”‚
โ”‚          โ”‚   (likely cause) โ”‚                               โ”‚
โ”‚          โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜                               โ”‚
โ”‚                   โ–ผ                                         โ”‚
โ”‚          โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”                               โ”‚
โ”‚          โ”‚   TEST           โ”‚                               โ”‚
โ”‚          โ”‚   (one variable) โ”‚                               โ”‚
โ”‚          โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜                               โ”‚
โ”‚                   โ–ผ                                         โ”‚
โ”‚          โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”                               โ”‚
โ”‚          โ”‚   CONFIRM?       โ”‚                               โ”‚
โ”‚          โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜                               โ”‚
โ”‚                   โ”‚                                         โ”‚
โ”‚         โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”                               โ”‚
โ”‚         โ–ผ                   โ–ผ                               โ”‚
โ”‚    โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”         โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”                          โ”‚
โ”‚    โ”‚  FIX    โ”‚         โ”‚ DISCARD โ”‚                          โ”‚
โ”‚    โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜         โ””โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”˜                          โ”‚
โ”‚                             โ”‚                               โ”‚
โ”‚                             โ””โ”€โ”€โ”€โ”€โ”€โ”€โ–ถ back to HYPOTHESIZE    โ”‚
โ”‚                                                             โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜

Documentation Template

โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚  INCIDENT DOCUMENTATION                                     โ”‚
โ”‚                                                             โ”‚
โ”‚  Date/Time:                                                 โ”‚
โ”‚  Reported by:                                               โ”‚
โ”‚                                                             โ”‚
โ”‚  SYMPTOM:                                                   โ”‚
โ”‚  What was observed?                                         โ”‚
โ”‚                                                             โ”‚
โ”‚  INVESTIGATION:                                             โ”‚
โ”‚  Commands run, outputs observed.                            โ”‚
โ”‚                                                             โ”‚
โ”‚  ROOT CAUSE:                                                โ”‚
โ”‚  What was actually wrong?                                   โ”‚
โ”‚                                                             โ”‚
โ”‚  RESOLUTION:                                                โ”‚
โ”‚  What was changed to fix it?                                โ”‚
โ”‚                                                             โ”‚
โ”‚  PREVENTION:                                                โ”‚
โ”‚  What monitoring or process change prevents recurrence?     โ”‚
โ”‚                                                             โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜

Summary

ItemValue
DomainSystem Administration Fundamentals
First phaseInitial assessment and orientation
Second phaseResource analysis
Third phaseProcess and log investigation
Fourth phaseNetwork verification
Fifth phaseResolution and verification
Final phaseDocumentation and prevention
Key principleChange one thing at a time
Log commandjournalctl -xe for systemd journal
Network commandss -tulpn for listening ports
DocumentationRequired for every incident

Key takeaways:

  • Troubleshooting is a process, not an instinct. The methodical approach provides a repeatable sequence that narrows the problem space layer by layer.
  • Logs before restarts. Restarting a service without reading its logs destroys the evidence of why it failed. The failure will recur, and you will have learned nothing.
  • Check resources early. CPU, memory, disk, and I/O problems masquerade as application problems. Ruling out resource exhaustion prevents wasted time in the wrong layer.
  • Network diagnostics are layered. Test the interface, then the gateway, then DNS, then the service. Each layer assumes the previous one works.
  • One change at a time. Changing multiple variables makes it impossible to know which change fixed the problem. The methodical approach requires controlled experiments.
  • Documentation is part of the fix. An undocumented resolution is a temporary fix. The next occurrence will require the entire investigation to be repeated.

Remember: The LFCA exam tests troubleshooting as scenario-based reasoning. You will be given a symptom and asked which diagnostic step comes next. The answer is always the step that gathers evidence without assuming the cause. A methodical approach is not just good practice for real systemsโ€”it is the specific reasoning pattern the exam evaluates.



Stop using slow, ad-bloated tool sites! ๐Ÿคฎ

๐Ÿ”Ž Search “KandZ Tools” on Google to use many professional utilities for free.

KandZ.me is the ultimate minimalist hub for:
โœ… Finance (Mortgage, Interest, Inflation)
โœ… Tech (Base64, JSON, Dev Suite, IP)
โœ… Health (BMI, BMR, TDEE)
โœ… Productivity (Timer, Workspace, QR)

โšก๏ธ Fast & Private
๐Ÿ”’ No data leaves your device
๐Ÿ’Ž 100% Free

๐Ÿ”— Use it now: https://tools.kandz.me
๐Ÿ”– Bookmark itโ€”youโ€™ll need it later!