| |

LFCA 114 🐧 SLI, SLO, and SLA

SLI, SLO, and SLA are three terms that describe how reliability is measured, targeted, and contracted. They are often confused because they sound similar and are used together, but they answer different questions. An SLI asks “what did the service actually do?” An SLO asks “what is the target for that measurement?” An SLA asks “what happens if the target is missed?” Confusing them leads to agreements that cannot be measured, targets that are not enforced, and consequences that surprise everyone when they are triggered.

The distinction matters because each term has a different audience and a different purpose. Engineers define and measure SLIs. Engineering and product together agree on SLOs. Business and legal negotiate SLAs with customers. When the three are aligned, the system has a shared definition of reliability that everyone understands. When they are not, the organization argues about whether the service is “good enough” without a common language to resolve the question.

Key point: An SLI is a measurement. An SLO is a target for that measurement. An SLA is a contract that references the SLI and specifies consequences for missing the target. The error budget is 100% minus the SLO. SLAs are typically looser than SLOs to provide a buffer.


Why SLI, SLO, and SLA exist

The measurement problem. “The service is slow” is not actionable. “The 95th percentile latency is 420 milliseconds, and the SLO is 200 milliseconds” is actionable. An SLI replaces subjective impressions with a number that can be tracked, alerted on, and improved.

The alignment problem. Developers, operators, and product managers have different intuitions about how reliable a service should be. Without an agreed target, the discussion is subjective. An SLO makes the target explicit and quantifies the trade-off between reliability and feature velocity.

The contract problem. Customers want a commitment. If the service is unavailable for an hour, what happens? An SLA answers this question with specific terms. It references an SLI so the commitment is measurable and specifies what the customer receives if the target is missed.

The buffer problem. An SLA is a public commitment with financial consequences. An SLO is an internal target that is stricter, so the team has room to act before the SLA is breached. If the SLA is 99.9% and the SLO is 99.95%, the team has a warning window before the customer-facing commitment is violated.

The prioritization problem. Not all services need the same reliability. A billing service might need 99.99% availability, while an internal reporting tool might be fine at 99%. SLIs and SLOs let the organization assign the appropriate level of reliability to each service instead of treating all services as equally critical.


a. The Service Level Indicator (SLI)

An SLI is a quantitative measure of some aspect of the service’s behavior that matters to users. It is always a ratio: the number of good events divided by the total number of events. This ratio form makes the SLI a percentage between 0% and 100%, which can be compared to a target.

The standard SLI categories are:

CategoryQuestion it answersExample
AvailabilityDid the request succeed?Successful requests / total requests
LatencyWas the response fast enough?Requests under 200ms / total requests
ThroughputDid the system process the volume?Records processed / records received
CorrectnessWas the response accurate?Correct results / total results
FreshnessWas the data up to date?Reads with data younger than 5 minutes / total reads
DurabilityWas the data preserved?Records recoverable / records stored

A well-formed SLI has three parts: the measurement (what is being counted), the numerator (what counts as a good event), and the denominator (the total population). For example:

sli:
  name: api_availability
  measurement: HTTP requests to /api/*
  good: response status in 2xx and 3xx
  total: all responses
  value: good / total

The SLI must be measured from a place that reflects the user experience. Measuring availability inside the service misses requests that never reached it; measuring at the load balancer or from the client captures the full picture.

Raw measurements are not SLIs. A counter of 500 errors is a raw measurement. The ratio of errors to total requests is an SLI. The ratio is what can be compared to a target and what can be expressed as a percentage.


b. The Service Level Objective (SLO)

An SLO is a target value or range for an SLI, measured over a specified window. It is the reliability target that the team commits to internally.

slo:
  sli: api_availability
  target: 99.9%
  window: 30 days

An SLO has three parts: the SLI it references, the target, and the window over which the target is evaluated. The window matters because a 99.9% target over 30 days is different from a 99.9% target over 7 days. The 30-day window is common because it matches a monthly reporting cycle.

SLOs can be expressed as a single value or a range:

slo:
  sli: api_latency
  target: 95% of requests under 200ms
  window: 28 days

The target is not always the maximum. For latency, a common SLO is “95% of requests complete in under 200ms,” which allows the slowest 5% to be slower without violating the objective.

Choosing an SLO target. The target should reflect what users actually need, not what is theoretically achievable. If users cannot tell the difference between 99.9% and 99.99% availability for a given service, the extra reliability is wasted effort. The cost of each additional nine is roughly ten times the previous one, so the target should be the lowest value that satisfies users.

SLOError BudgetDowntime per Month
99%1%~7.2 hours
99.5%0.5%~3.6 hours
99.9%0.1%~43 minutes
99.95%0.05%~21.5 minutes
99.99%0.01%~4.3 minutes
99.999%0.001%~26 seconds

The error budget. The error budget is 100% minus the SLO. An SLO of 99.9% means the service may be unreliable for 0.1% of the time. The budget is the amount of unreliability the team is allowed before the SLO is violated.

The error budget is the mechanism that balances reliability against feature velocity. When the budget is healthy, the team ships features aggressively. When the budget is exhausted, the team freezes releases and focuses on reliability until the budget recovers.


c. The Service Level Agreement (SLA)

An SLA is a contract between a service provider and a customer. It defines the level of service the customer can expect and the consequences if the provider fails to meet it. The consequences are typically financial: service credits, refunds, or penalties.

An SLA references one or more SLIs so the commitment is measurable:

sla:
  customer: Acme Corp
  sli: api_availability
  target: 99.5%
  window: 30 days
  consequence:
    if_target_missed: 10% service credit

The SLA target is usually looser than the internal SLO. If the SLA is 99.5% and the SLO is 99.9%, the team has a buffer: the internal monitoring alerts before the contractual target is violated. The buffer protects against the financial consequences of a breach.

The SLA is a business document. It is negotiated with customers, reviewed by legal, and referenced in contracts. It is not an engineering artifact. Engineers should not be surprised by the consequences of an SLA breach, which is why the SLO should be stricter than the SLA and the error budget should be monitored continuously.

AspectSLOSLA
AudienceInternal (engineering, product)External (customers, legal)
TargetStricterLooser
ConsequencesProcess (freeze releases)Financial (credits, penalties)
PurposeBalance reliability and velocityContractual commitment

d. How the three relate

The three terms form a chain: the SLI is measured, the SLO sets a target for the SLI, and the SLA references the SLI with contractual consequences.

TermQuestionFormAudience
SLIWhat did the service do?RatioEngineering
SLOWhat is the target?Percentage over a windowEngineering and product
SLAWhat happens if we miss?Contract with consequencesBusiness and customers

A worked example:

SLI:    api_availability = successful_requests / total_requests
        Measured over 30 days from the load balancer.

SLO:    api_availability >= 99.9%
        Internal target. Error budget = 0.1% = ~43 minutes/month.

SLA:    api_availability >= 99.5%
        Contractual commitment. If missed, customer receives a 10% credit.

If the service drops below 99.9%, the SLO is violated, the error budget is exhausted, and the team freezes feature releases. If it drops below 99.5%, the SLA is violated, and the customer receives a credit. The buffer between 99.9% and 99.5% is the team’s warning window.


e. Common mistakes

Defining SLIs from the wrong place. An SLI measured inside the service does not capture requests that never reached it. Measuring at the load balancer or from the client reflects the user experience more accurately.

Setting SLOs without a window. A target without a measurement window is ambiguous. The window determines the error budget and the alerting behavior.

Setting SLAs equal to SLOs. Without a buffer, the team has no warning before the contractual commitment is breached. The SLO should be stricter than the SLA.

Chasing 100%. A 100% target is impossible, wasteful, and discouraging. The cost of the last fraction of a percent is enormous, and users cannot perceive the difference.

Ignoring the error budget. The error budget is the mechanism that balances reliability and velocity. If it is not monitored, the balance is not enforced.

Using raw measurements as SLIs. A count of failures is not an SLI. The ratio of failures to total requests is.


Complete Example Session

# ============================================
# PART 1: DEFINING AN SLI
# ============================================
sli:
  name: api_availability
  measurement: HTTP requests to the public API
  good: response status 2xx or 3xx
  total: all responses
  value: good / total
  measured_at: load balancer
  window: 30 days
# ============================================
# PART 2: SETTING AN SLO
# ============================================
slo:
  sli: api_availability
  target: 99.9%
  window: 30 days
  error_budget: 0.1%
# ============================================
# PART 3: CALCULATING THE ERROR BUDGET
# ============================================
error_budget:
  formula: 100% - SLO
  value: 0.1%
  minutes_per_month: 43
  seconds_per_month: 2592
# ============================================
# PART 4: DEFINING AN SLA
# ============================================
sla:
  customer: Acme Corp
  sli: api_availability
  target: 99.5%
  window: 30 days
  consequence: 10% service credit
# ============================================
# PART 5: DEFINING A LATENCY SLI
# ============================================
sli:
  name: api_latency
  measurement: response time of API requests
  good: duration < 200ms
  total: all requests
  value: good / total
# ============================================
# PART 6: SETTING A LATENCY SLO
# ============================================
slo:
  sli: api_latency
  target: 95%
  window: 28 days
# ============================================
# PART 7: ERROR BUDGET POLICY
# ============================================
policy:
  budget_remaining_50_percent: normal velocity
  budget_remaining_25_percent: review releases
  budget_exhausted: freeze features, focus on reliability
# ============================================
# PART 8: ALERTING ON BURN RATE
# ============================================
alert:
  name: high_burn_rate
  condition: error_budget_consumed_2_percent_in_1_hour
  action: page on-call
  rationale: fast burn will exhaust the budget before the window ends
# ============================================
# PART 9: COMPARING SLO AND SLA
# ============================================
comparison:
  slo: 99.9% (internal target)
  sla: 99.5% (contractual commitment)
  buffer: 0.4% (warning window)
# ============================================
# PART 10: REPORTING
# ============================================
report:
  sli_value: 99.94%
  slo_status: met
  error_budget_remaining: 0.04%
  sla_status: met

These ten parts cover defining an SLI, setting an SLO, calculating the error budget, defining an SLA, defining a latency SLI, setting a latency SLO, the error budget policy, alerting on burn rate, comparing SLO and SLA, and reporting.


Quick Reference

The Three Terms

TermQuestionFormAudience
SLIWhat did the service do?RatioEngineering
SLOWhat is the target?Percentage over a windowEngineering and product
SLAWhat if we miss?Contract with consequencesBusiness and customers

SLI Categories

CategoryQuestionExample
AvailabilityDid it succeed?successful / total
LatencyWas it fast?under threshold / total
ThroughputDid it process?processed / received
CorrectnessWas it accurate?correct / total
FreshnessWas it current?fresh reads / total
DurabilityWas it preserved?recoverable / stored

SLO Targets and Error Budgets

SLOError BudgetDowntime per Month
99%1%~7.2 hours
99.5%0.5%~3.6 hours
99.9%0.1%~43 minutes
99.95%0.05%~21.5 minutes
99.99%0.01%~4.3 minutes

SLO vs SLA

AspectSLOSLA
AudienceInternalExternal
TargetStricterLooser
ConsequencesProcessFinancial
PurposeBalance reliability and velocityContractual commitment

Error Budget Policy

Budget StateAction
HealthyShip features normally
25% remainingReview release cadence
ExhaustedFreeze features, focus on reliability

Burn Rate Alerting

ConditionAction
2% budget in 1 hourPage on-call
5% budget in 6 hoursPage on-call
10% budget in 3 daysTicket

Best Practices

✅ Do This:

# Define SLIs as ratios
sli: successful_requests / total_requests
# Set SLOs with a window
slo:
  target: 99.9%
  window: 30 days
# Set SLAs looser than SLOs
slo: 99.9%
sla: 99.5%
# Monitor the error budget
error_budget: 100% - SLO
# Alert on burn rate
alert: 2% budget in 1 hour

❌ Don’t Do This:

# Use raw counts as SLIs
sli: 500_errors  # ❌ use a ratio
# Set SLO without a window
slo: 99.9%  # ❌ over what period?
# Set SLA equal to SLO
slo: 99.9%  # ❌ no buffer
sla: 99.9%
# Chase 100%
slo: 100%  # ❌ impossible and wasteful

Common Pitfalls

PitfallWhy It HappensFix
SLI measured in the wrong placeInternal measurement onlyMeasure at the load balancer or client
SLO without a windowAmbiguitySpecify the evaluation period
SLA equal to SLONo bufferMake the SLO stricter
Chasing 100%Ambition without cost analysisSet the lowest target that satisfies users
Error budget ignoredNo policyDefine actions for budget states
No burn rate alertingOnly end-of-window checkAlert on fast consumption
Too many SLOsTrying to measure everythingFocus on user-facing signals

Real-World Examples

1. Availability SLI

sli: successful_requests / total_requests

2. Latency SLI

sli: requests_under_200ms / total_requests

3. Availability SLO

slo: 99.9% over 30 days

4. Latency SLO

slo: 95% under 200ms over 28 days

5. SLA

sla: 99.5% with 10% credit

6. Error Budget

error_budget: 0.1% = 43 minutes/month

7. Error Budget Policy

if_exhausted: freeze releases

8. Burn Rate Alert

alert: 2% in 1 hour

9. Throughput SLI

sli: records_processed / records_received

10. Freshness SLI

sli: reads_with_data_under_5min / total_reads

Visual

The SLI, SLO, SLA Chain

┌──────────────────────────────────────────────────────────────┐
│  SLI                                                         │
│  └── successful_requests / total_requests = 99.94%           │
│       │                                                      │
│       ▼                                                      │
│  SLO                                                         │
│  └── target: 99.9% over 30 days                              │
│      error budget: 0.1%                                      │
│       │                                                      │
│       ▼                                                      │
│  SLA                                                         │
│  └── contractual: 99.5%                                      │
│      consequence: 10% credit                                 │
└──────────────────────────────────────────────────────────────┘

SLO Targets and Error Budgets

┌──────────────────────────────────────────────────────────────┐
│  SLO        ERROR BUDGET     DOWNTIME/MONTH                  │
│  ───────────┼────────────────┼─────────────────────────────  │
│  99%        │ 1%             │ ~7.2 hours                    │
│  99.5%      │ 0.5%           │ ~3.6 hours                    │
│  99.9%      │ 0.1%           │ ~43 minutes                   │
│  99.99%     │ 0.01%          │ ~4.3 minutes                  │
└──────────────────────────────────────────────────────────────┘

SLO vs SLA Buffer

┌──────────────────────────────────────────────────────────────┐
│  SLO (internal):    99.9%  ──▶ freeze releases if missed     │
│  SLA (contractual): 99.5%  ──▶ customer credit if missed     │
│                                                              │
│  Buffer: 0.4%                                                │
│  └── The team has a warning window before the SLA is         │
│      breached.                                               │
└──────────────────────────────────────────────────────────────┘

Error Budget Policy

┌──────────────────────────────────────────────────────────────┐
│  Budget healthy:        ──▶ ship features                    │
│  Budget at 25%:         ──▶ review release cadence           │
│  Budget exhausted:      ──▶ freeze features                  │
│                                                              │
│  The budget is the mechanism that balances reliability       │
│  against feature velocity.                                   │
└──────────────────────────────────────────────────────────────┘

Summary

ItemValue
SLIMeasurement of service behavior as a ratio
SLI categoriesAvailability, latency, throughput, correctness, freshness, durability
SLOTarget for an SLI over a window
Error budget100% minus the SLO
SLAContract with consequences
SLO vs SLASLO stricter; SLA looser with buffer
Error budget policyHealthy: ship; exhausted: freeze
Burn rate alertingAlert on fast budget consumption
Measurement locationLoad balancer or client, not internal
100% targetImpossible and wasteful

Key takeaways:

  • An SLI is a measurement, an SLO is a target, and an SLA is a contract. The SLI answers “what did the service do?” The SLO answers “what is the target?” The SLA answers “what happens if we miss?”
  • SLIs are ratios, not raw counts. The ratio of good events to total events is what can be compared to a target and expressed as a percentage.
  • SLOs are defined over a window. A target without a window is ambiguous. The window determines the error budget and the alerting behavior.
  • The error budget is 100% minus the SLO. It is the amount of unreliability the team is allowed. When the budget is healthy, the team ships features. When it is exhausted, reliability work takes priority.
  • SLAs are looser than SLOs. The buffer between them is the warning window. If the SLO is 99.9% and the SLA is 99.5%, the team has room to act before the customer-facing commitment is breached.
  • Measure SLIs at the user’s boundary. An SLI measured inside the service misses requests that never reached it. Measuring at the load balancer or from the client reflects the user experience.
  • Do not chase 100%. The cost of each additional nine is roughly ten times the previous one, and users cannot perceive the difference. Set the lowest target that satisfies users.

Remember: SLI, SLO, and SLA are the vocabulary of reliability. The SLI is the measurement, the SLO is the target, and the SLA is the contract. Together they turn a subjective discussion about whether a service is “good enough” into a quantitative one that engineering, product, and business can all participate in. The error budget is the mechanism that balances reliability against feature velocity, and the buffer between the SLO and the SLA protects the organization from financial consequences. Understanding the three terms and their relationships is the foundation for defining, measuring, and improving reliability.



Stop using slow, ad-bloated tool sites! 🤮

🔎 Search “KandZ Tools” on Google to use many professional utilities for free.

KandZ.me is the ultimate minimalist hub for:
✅ Finance (Mortgage, Interest, Inflation)
✅ Tech (Base64, JSON, Dev Suite, IP)
✅ Health (BMI, BMR, TDEE)
✅ Productivity (Timer, Workspace, QR)

⚡️ Fast & Private
🔒 No data leaves your device
💎 100% Free

🔗 Use it now: https://tools.kandz.me
🔖 Bookmark it—you’ll need it later!