LFCA 114 🐧 SLI, SLO, and SLA
SLI, SLO, and SLA are three terms that describe how reliability is measured, targeted, and contracted. They are often confused because they sound similar and are used together, but they answer different questions. An SLI asks “what did the service actually do?” An SLO asks “what is the target for that measurement?” An SLA asks “what happens if the target is missed?” Confusing them leads to agreements that cannot be measured, targets that are not enforced, and consequences that surprise everyone when they are triggered.
The distinction matters because each term has a different audience and a different purpose. Engineers define and measure SLIs. Engineering and product together agree on SLOs. Business and legal negotiate SLAs with customers. When the three are aligned, the system has a shared definition of reliability that everyone understands. When they are not, the organization argues about whether the service is “good enough” without a common language to resolve the question.
Key point: An SLI is a measurement. An SLO is a target for that measurement. An SLA is a contract that references the SLI and specifies consequences for missing the target. The error budget is 100% minus the SLO. SLAs are typically looser than SLOs to provide a buffer.
Why SLI, SLO, and SLA exist
The measurement problem. “The service is slow” is not actionable. “The 95th percentile latency is 420 milliseconds, and the SLO is 200 milliseconds” is actionable. An SLI replaces subjective impressions with a number that can be tracked, alerted on, and improved.
The alignment problem. Developers, operators, and product managers have different intuitions about how reliable a service should be. Without an agreed target, the discussion is subjective. An SLO makes the target explicit and quantifies the trade-off between reliability and feature velocity.
The contract problem. Customers want a commitment. If the service is unavailable for an hour, what happens? An SLA answers this question with specific terms. It references an SLI so the commitment is measurable and specifies what the customer receives if the target is missed.
The buffer problem. An SLA is a public commitment with financial consequences. An SLO is an internal target that is stricter, so the team has room to act before the SLA is breached. If the SLA is 99.9% and the SLO is 99.95%, the team has a warning window before the customer-facing commitment is violated.
The prioritization problem. Not all services need the same reliability. A billing service might need 99.99% availability, while an internal reporting tool might be fine at 99%. SLIs and SLOs let the organization assign the appropriate level of reliability to each service instead of treating all services as equally critical.
a. The Service Level Indicator (SLI)
An SLI is a quantitative measure of some aspect of the service’s behavior that matters to users. It is always a ratio: the number of good events divided by the total number of events. This ratio form makes the SLI a percentage between 0% and 100%, which can be compared to a target.
The standard SLI categories are:
| Category | Question it answers | Example |
|---|---|---|
| Availability | Did the request succeed? | Successful requests / total requests |
| Latency | Was the response fast enough? | Requests under 200ms / total requests |
| Throughput | Did the system process the volume? | Records processed / records received |
| Correctness | Was the response accurate? | Correct results / total results |
| Freshness | Was the data up to date? | Reads with data younger than 5 minutes / total reads |
| Durability | Was the data preserved? | Records recoverable / records stored |
A well-formed SLI has three parts: the measurement (what is being counted), the numerator (what counts as a good event), and the denominator (the total population). For example:
sli:
name: api_availability
measurement: HTTP requests to /api/*
good: response status in 2xx and 3xx
total: all responses
value: good / total
The SLI must be measured from a place that reflects the user experience. Measuring availability inside the service misses requests that never reached it; measuring at the load balancer or from the client captures the full picture.
Raw measurements are not SLIs. A counter of 500 errors is a raw measurement. The ratio of errors to total requests is an SLI. The ratio is what can be compared to a target and what can be expressed as a percentage.
b. The Service Level Objective (SLO)
An SLO is a target value or range for an SLI, measured over a specified window. It is the reliability target that the team commits to internally.
slo:
sli: api_availability
target: 99.9%
window: 30 days
An SLO has three parts: the SLI it references, the target, and the window over which the target is evaluated. The window matters because a 99.9% target over 30 days is different from a 99.9% target over 7 days. The 30-day window is common because it matches a monthly reporting cycle.
SLOs can be expressed as a single value or a range:
slo:
sli: api_latency
target: 95% of requests under 200ms
window: 28 days
The target is not always the maximum. For latency, a common SLO is “95% of requests complete in under 200ms,” which allows the slowest 5% to be slower without violating the objective.
Choosing an SLO target. The target should reflect what users actually need, not what is theoretically achievable. If users cannot tell the difference between 99.9% and 99.99% availability for a given service, the extra reliability is wasted effort. The cost of each additional nine is roughly ten times the previous one, so the target should be the lowest value that satisfies users.
| SLO | Error Budget | Downtime per Month |
|---|---|---|
| 99% | 1% | ~7.2 hours |
| 99.5% | 0.5% | ~3.6 hours |
| 99.9% | 0.1% | ~43 minutes |
| 99.95% | 0.05% | ~21.5 minutes |
| 99.99% | 0.01% | ~4.3 minutes |
| 99.999% | 0.001% | ~26 seconds |
The error budget. The error budget is 100% minus the SLO. An SLO of 99.9% means the service may be unreliable for 0.1% of the time. The budget is the amount of unreliability the team is allowed before the SLO is violated.
The error budget is the mechanism that balances reliability against feature velocity. When the budget is healthy, the team ships features aggressively. When the budget is exhausted, the team freezes releases and focuses on reliability until the budget recovers.
c. The Service Level Agreement (SLA)
An SLA is a contract between a service provider and a customer. It defines the level of service the customer can expect and the consequences if the provider fails to meet it. The consequences are typically financial: service credits, refunds, or penalties.
An SLA references one or more SLIs so the commitment is measurable:
sla:
customer: Acme Corp
sli: api_availability
target: 99.5%
window: 30 days
consequence:
if_target_missed: 10% service credit
The SLA target is usually looser than the internal SLO. If the SLA is 99.5% and the SLO is 99.9%, the team has a buffer: the internal monitoring alerts before the contractual target is violated. The buffer protects against the financial consequences of a breach.
The SLA is a business document. It is negotiated with customers, reviewed by legal, and referenced in contracts. It is not an engineering artifact. Engineers should not be surprised by the consequences of an SLA breach, which is why the SLO should be stricter than the SLA and the error budget should be monitored continuously.
| Aspect | SLO | SLA |
|---|---|---|
| Audience | Internal (engineering, product) | External (customers, legal) |
| Target | Stricter | Looser |
| Consequences | Process (freeze releases) | Financial (credits, penalties) |
| Purpose | Balance reliability and velocity | Contractual commitment |
d. How the three relate
The three terms form a chain: the SLI is measured, the SLO sets a target for the SLI, and the SLA references the SLI with contractual consequences.
| Term | Question | Form | Audience |
|---|---|---|---|
| SLI | What did the service do? | Ratio | Engineering |
| SLO | What is the target? | Percentage over a window | Engineering and product |
| SLA | What happens if we miss? | Contract with consequences | Business and customers |
A worked example:
SLI: api_availability = successful_requests / total_requests
Measured over 30 days from the load balancer.
SLO: api_availability >= 99.9%
Internal target. Error budget = 0.1% = ~43 minutes/month.
SLA: api_availability >= 99.5%
Contractual commitment. If missed, customer receives a 10% credit.
If the service drops below 99.9%, the SLO is violated, the error budget is exhausted, and the team freezes feature releases. If it drops below 99.5%, the SLA is violated, and the customer receives a credit. The buffer between 99.9% and 99.5% is the team’s warning window.
e. Common mistakes
Defining SLIs from the wrong place. An SLI measured inside the service does not capture requests that never reached it. Measuring at the load balancer or from the client reflects the user experience more accurately.
Setting SLOs without a window. A target without a measurement window is ambiguous. The window determines the error budget and the alerting behavior.
Setting SLAs equal to SLOs. Without a buffer, the team has no warning before the contractual commitment is breached. The SLO should be stricter than the SLA.
Chasing 100%. A 100% target is impossible, wasteful, and discouraging. The cost of the last fraction of a percent is enormous, and users cannot perceive the difference.
Ignoring the error budget. The error budget is the mechanism that balances reliability and velocity. If it is not monitored, the balance is not enforced.
Using raw measurements as SLIs. A count of failures is not an SLI. The ratio of failures to total requests is.
Complete Example Session
# ============================================
# PART 1: DEFINING AN SLI
# ============================================
sli:
name: api_availability
measurement: HTTP requests to the public API
good: response status 2xx or 3xx
total: all responses
value: good / total
measured_at: load balancer
window: 30 days
# ============================================
# PART 2: SETTING AN SLO
# ============================================
slo:
sli: api_availability
target: 99.9%
window: 30 days
error_budget: 0.1%
# ============================================
# PART 3: CALCULATING THE ERROR BUDGET
# ============================================
error_budget:
formula: 100% - SLO
value: 0.1%
minutes_per_month: 43
seconds_per_month: 2592
# ============================================
# PART 4: DEFINING AN SLA
# ============================================
sla:
customer: Acme Corp
sli: api_availability
target: 99.5%
window: 30 days
consequence: 10% service credit
# ============================================
# PART 5: DEFINING A LATENCY SLI
# ============================================
sli:
name: api_latency
measurement: response time of API requests
good: duration < 200ms
total: all requests
value: good / total
# ============================================
# PART 6: SETTING A LATENCY SLO
# ============================================
slo:
sli: api_latency
target: 95%
window: 28 days
# ============================================
# PART 7: ERROR BUDGET POLICY
# ============================================
policy:
budget_remaining_50_percent: normal velocity
budget_remaining_25_percent: review releases
budget_exhausted: freeze features, focus on reliability
# ============================================
# PART 8: ALERTING ON BURN RATE
# ============================================
alert:
name: high_burn_rate
condition: error_budget_consumed_2_percent_in_1_hour
action: page on-call
rationale: fast burn will exhaust the budget before the window ends
# ============================================
# PART 9: COMPARING SLO AND SLA
# ============================================
comparison:
slo: 99.9% (internal target)
sla: 99.5% (contractual commitment)
buffer: 0.4% (warning window)
# ============================================
# PART 10: REPORTING
# ============================================
report:
sli_value: 99.94%
slo_status: met
error_budget_remaining: 0.04%
sla_status: met
These ten parts cover defining an SLI, setting an SLO, calculating the error budget, defining an SLA, defining a latency SLI, setting a latency SLO, the error budget policy, alerting on burn rate, comparing SLO and SLA, and reporting.
Quick Reference
The Three Terms
| Term | Question | Form | Audience |
|---|---|---|---|
| SLI | What did the service do? | Ratio | Engineering |
| SLO | What is the target? | Percentage over a window | Engineering and product |
| SLA | What if we miss? | Contract with consequences | Business and customers |
SLI Categories
| Category | Question | Example |
|---|---|---|
| Availability | Did it succeed? | successful / total |
| Latency | Was it fast? | under threshold / total |
| Throughput | Did it process? | processed / received |
| Correctness | Was it accurate? | correct / total |
| Freshness | Was it current? | fresh reads / total |
| Durability | Was it preserved? | recoverable / stored |
SLO Targets and Error Budgets
| SLO | Error Budget | Downtime per Month |
|---|---|---|
| 99% | 1% | ~7.2 hours |
| 99.5% | 0.5% | ~3.6 hours |
| 99.9% | 0.1% | ~43 minutes |
| 99.95% | 0.05% | ~21.5 minutes |
| 99.99% | 0.01% | ~4.3 minutes |
SLO vs SLA
| Aspect | SLO | SLA |
|---|---|---|
| Audience | Internal | External |
| Target | Stricter | Looser |
| Consequences | Process | Financial |
| Purpose | Balance reliability and velocity | Contractual commitment |
Error Budget Policy
| Budget State | Action |
|---|---|
| Healthy | Ship features normally |
| 25% remaining | Review release cadence |
| Exhausted | Freeze features, focus on reliability |
Burn Rate Alerting
| Condition | Action |
|---|---|
| 2% budget in 1 hour | Page on-call |
| 5% budget in 6 hours | Page on-call |
| 10% budget in 3 days | Ticket |
Best Practices
✅ Do This:
# Define SLIs as ratios
sli: successful_requests / total_requests
# Set SLOs with a window
slo:
target: 99.9%
window: 30 days
# Set SLAs looser than SLOs
slo: 99.9%
sla: 99.5%
# Monitor the error budget
error_budget: 100% - SLO
# Alert on burn rate
alert: 2% budget in 1 hour
❌ Don’t Do This:
# Use raw counts as SLIs
sli: 500_errors # ❌ use a ratio
# Set SLO without a window
slo: 99.9% # ❌ over what period?
# Set SLA equal to SLO
slo: 99.9% # ❌ no buffer
sla: 99.9%
# Chase 100%
slo: 100% # ❌ impossible and wasteful
Common Pitfalls
| Pitfall | Why It Happens | Fix |
|---|---|---|
| SLI measured in the wrong place | Internal measurement only | Measure at the load balancer or client |
| SLO without a window | Ambiguity | Specify the evaluation period |
| SLA equal to SLO | No buffer | Make the SLO stricter |
| Chasing 100% | Ambition without cost analysis | Set the lowest target that satisfies users |
| Error budget ignored | No policy | Define actions for budget states |
| No burn rate alerting | Only end-of-window check | Alert on fast consumption |
| Too many SLOs | Trying to measure everything | Focus on user-facing signals |
Real-World Examples
1. Availability SLI
sli: successful_requests / total_requests
2. Latency SLI
sli: requests_under_200ms / total_requests
3. Availability SLO
slo: 99.9% over 30 days
4. Latency SLO
slo: 95% under 200ms over 28 days
5. SLA
sla: 99.5% with 10% credit
6. Error Budget
error_budget: 0.1% = 43 minutes/month
7. Error Budget Policy
if_exhausted: freeze releases
8. Burn Rate Alert
alert: 2% in 1 hour
9. Throughput SLI
sli: records_processed / records_received
10. Freshness SLI
sli: reads_with_data_under_5min / total_reads
Visual
The SLI, SLO, SLA Chain
┌──────────────────────────────────────────────────────────────┐
│ SLI │
│ └── successful_requests / total_requests = 99.94% │
│ │ │
│ ▼ │
│ SLO │
│ └── target: 99.9% over 30 days │
│ error budget: 0.1% │
│ │ │
│ ▼ │
│ SLA │
│ └── contractual: 99.5% │
│ consequence: 10% credit │
└──────────────────────────────────────────────────────────────┘
SLO Targets and Error Budgets
┌──────────────────────────────────────────────────────────────┐
│ SLO ERROR BUDGET DOWNTIME/MONTH │
│ ───────────┼────────────────┼───────────────────────────── │
│ 99% │ 1% │ ~7.2 hours │
│ 99.5% │ 0.5% │ ~3.6 hours │
│ 99.9% │ 0.1% │ ~43 minutes │
│ 99.99% │ 0.01% │ ~4.3 minutes │
└──────────────────────────────────────────────────────────────┘
SLO vs SLA Buffer
┌──────────────────────────────────────────────────────────────┐
│ SLO (internal): 99.9% ──▶ freeze releases if missed │
│ SLA (contractual): 99.5% ──▶ customer credit if missed │
│ │
│ Buffer: 0.4% │
│ └── The team has a warning window before the SLA is │
│ breached. │
└──────────────────────────────────────────────────────────────┘
Error Budget Policy
┌──────────────────────────────────────────────────────────────┐
│ Budget healthy: ──▶ ship features │
│ Budget at 25%: ──▶ review release cadence │
│ Budget exhausted: ──▶ freeze features │
│ │
│ The budget is the mechanism that balances reliability │
│ against feature velocity. │
└──────────────────────────────────────────────────────────────┘
Summary
| Item | Value |
|---|---|
| SLI | Measurement of service behavior as a ratio |
| SLI categories | Availability, latency, throughput, correctness, freshness, durability |
| SLO | Target for an SLI over a window |
| Error budget | 100% minus the SLO |
| SLA | Contract with consequences |
| SLO vs SLA | SLO stricter; SLA looser with buffer |
| Error budget policy | Healthy: ship; exhausted: freeze |
| Burn rate alerting | Alert on fast budget consumption |
| Measurement location | Load balancer or client, not internal |
| 100% target | Impossible and wasteful |
Key takeaways:
- An SLI is a measurement, an SLO is a target, and an SLA is a contract. The SLI answers “what did the service do?” The SLO answers “what is the target?” The SLA answers “what happens if we miss?”
- SLIs are ratios, not raw counts. The ratio of good events to total events is what can be compared to a target and expressed as a percentage.
- SLOs are defined over a window. A target without a window is ambiguous. The window determines the error budget and the alerting behavior.
- The error budget is 100% minus the SLO. It is the amount of unreliability the team is allowed. When the budget is healthy, the team ships features. When it is exhausted, reliability work takes priority.
- SLAs are looser than SLOs. The buffer between them is the warning window. If the SLO is 99.9% and the SLA is 99.5%, the team has room to act before the customer-facing commitment is breached.
- Measure SLIs at the user’s boundary. An SLI measured inside the service misses requests that never reached it. Measuring at the load balancer or from the client reflects the user experience.
- Do not chase 100%. The cost of each additional nine is roughly ten times the previous one, and users cannot perceive the difference. Set the lowest target that satisfies users.
Remember: SLI, SLO, and SLA are the vocabulary of reliability. The SLI is the measurement, the SLO is the target, and the SLA is the contract. Together they turn a subjective discussion about whether a service is “good enough” into a quantitative one that engineering, product, and business can all participate in. The error budget is the mechanism that balances reliability against feature velocity, and the buffer between the SLO and the SLA protects the organization from financial consequences. Understanding the three terms and their relationships is the foundation for defining, measuring, and improving reliability.
Stop using slow, ad-bloated tool sites! 🤮
🔎 Search “KandZ Tools” on Google to use many professional utilities for free.
KandZ.me is the ultimate minimalist hub for:
✅ Finance (Mortgage, Interest, Inflation)
✅ Tech (Base64, JSON, Dev Suite, IP)
✅ Health (BMI, BMR, TDEE)
✅ Productivity (Timer, Workspace, QR)
⚡️ Fast & Private
🔒 No data leaves your device
💎 100% Free
🔗 Use it now: https://tools.kandz.me
🔖 Bookmark it—you’ll need it later!