| |

LFCA 115 🐧 Error Budgets and Toil

Error budgets and toil are the two operational mechanisms that make Site Reliability Engineering sustainable. The error budget is the amount of unreliability a service is allowed before the reliability target is violated, and it is the tool that resolves the tension between shipping features and keeping systems stable. Toil is the repetitive, manual work that consumes an operations team’s time without producing lasting value, and it is the enemy of the engineering work that improves systems over time. Together they define the boundary between what a team should accept and what it should engineer away.

The two concepts are connected. An error budget that is exhausted by preventable failures caused by toil is a signal that the system needs automation. A toil cap of 50% is what gives the team the time to build that automation. Without the error budget, reliability and velocity are argued about subjectively. Without the toil cap, the team is too busy firefighting to improve anything.

Key point: The error budget is 100% minus the SLO. It defines how much unreliability is acceptable. When the budget is exhausted, feature releases freeze and reliability work takes priority. Toil is manual, repetitive, automatable work that scales with system growth. SRE caps toil at 50% of the team’s time and uses the remaining 50% for engineering work that reduces future toil.


Why error budgets and toil limits exist

The reliability-velocity tension. Developers want to ship features; operators want to keep systems stable. Without a shared metric, the conversation is a negotiation of opinions. The error budget makes the trade-off explicit: as long as the service is within its budget, shipping is the priority; when the budget is spent, stability is.

The unsustainable on-call problem. If a team is woken every night, the engineers cannot sustain the work. They leave, and the team loses its institutional knowledge. The error budget and the toil cap are the mechanisms that make on-call humane: they force the organization to invest in reliability and automation instead of relying on heroics.

The toil creep problem. Toil accumulates quietly. A manual deployment here, a manual restart there, a manual certificate renewal somewhere else. Each task is small, but the aggregate consumes the team’s time, and the work that would reduce the toil never gets done. The 50% cap makes the accumulation visible and forces prioritization.

The measurement problem. Neither the error budget nor the toil percentage is useful without measurement. The error budget requires an SLI that is measured continuously. The toil percentage requires tracking what the team spends its time on. Both make the invisible visible.

The prioritization problem. When everything is on fire, nothing is a priority. The error budget policy defines what happens at each budget state, so the team does not have to decide in the moment. The toil cap defines what work is protected, so the engineering projects do not get crowded out.


a. The error budget

The error budget is 100% minus the SLO. If the SLO is 99.9% availability over 30 days, the error budget is 0.1%, which is approximately 43 minutes of downtime per month.

SLOError BudgetDowntime per Month
99%1%~7.2 hours
99.5%0.5%~3.6 hours
99.9%0.1%~43 minutes
99.95%0.05%~21.5 minutes
99.99%0.01%~4.3 minutes

The budget is not a target to spend. It is a limit. A team that consistently spends 100% of its error budget is a team that is consistently at the edge of violating its SLO. Healthy teams leave budget in reserve for unexpected events.

The error budget is tracked by measuring the SLI continuously and comparing the accumulated failures to the budget. If the SLI is availability, the budget is consumed by failed requests. If the SLI is latency, the budget is consumed by slow requests.

Burn rate. The rate at which the error budget is consumed is called the burn rate. A burn rate of 1 means the budget will be exhausted exactly at the end of the window. A burn rate of 10 means the budget will be exhausted in one-tenth of the window. Burn rate is the metric that drives alerting: a fast burn rate indicates a problem that will exhaust the budget before the window ends.

alert:
  name: fast_burn
  condition: burn_rate > 14.4 over 1 hour
  action: page on-call
  rationale: at this rate, the 30-day budget is consumed in 2 days

b. The error budget policy

The error budget policy defines what the team does at each budget state. It removes the debate from the moment of crisis by deciding in advance.

A common policy has three states:

Budget StateAction
Healthy (>50% remaining)Ship features at normal velocity
Warning (25–50% remaining)Review release cadence; prioritize reliability work
Exhausted (<25% remaining)Freeze feature releases; focus entirely on reliability

The policy is agreed upon by engineering, product, and leadership. It is not a technical decision; it is a business decision about how much risk is acceptable.

When the budget is exhausted, the freeze is not a punishment. It is a signal that the system is not meeting its reliability target, and the work that restores reliability takes priority over new features. The freeze ends when the budget recovers, which usually requires both fixing the underlying issues and waiting for the window to roll forward.


c. What toil is

Toil is the work that is:

PropertyMeaning
ManualDone by a person, not automated
RepetitiveDone over and over
AutomatableA machine could do it
TacticalReactive, not strategic
No lasting valueThe system is not better after it is done
Scales with growthMore of it as the system grows

Examples of toil include:

  • Manually restarting a crashed service
  • Manually provisioning a new environment
  • Manually rotating a certificate across many servers
  • Manually correlating logs during an incident
  • Manually updating a configuration file in many places
  • Manually approving a deployment that could be automated

The distinction between toil and engineering is not about the task’s difficulty. Restarting a service is not hard, but it is toil. Writing a script that automatically restarts the service when it fails is engineering, because the script reduces future toil.


d. The toil cap

Google’s internal benchmark recommends that SRE teams spend no more than 50% of their time on toil. The remaining 50% is dedicated to engineering work that improves the system.

The 50% cap is not arbitrary. If toil consumes more than half the time, the team cannot keep up with the engineering work that reduces future toil. The toil grows, the engineering shrinks, and the team enters a downward spiral. The cap breaks the spiral by protecting the engineering time.

Tracking the toil percentage requires measuring how the team spends its time. A simple approach is a weekly log: each engineer records the hours spent on toil and the hours spent on engineering. The aggregate percentage is reported and compared to the 50% target. When the percentage exceeds the cap, the team prioritizes automation projects that address the largest sources of toil.


e. Reducing toil through automation

The response to toil is automation. The goal is not to eliminate all toil, because some toil is unavoidable, but to reduce it below the cap.

ToilAutomation
Manual restartHealth check with auto-restart
Manual deploymentCI/CD pipeline
Manual provisioningInfrastructure as code
Manual certificate rotationAutomated renewal
Manual log correlationStructured logging with trace IDs
Manual scalingAutoscaling based on metrics

Each automation project has a cost and a return. The return is the toil hours saved per week or per month. Prioritizing projects by return on investment ensures that the team addresses the largest sources of toil first.

A useful practice is to estimate the toil hours saved by a proposed automation and compare it to the engineering hours required to build it. A project that saves 10 hours per week and takes 40 hours to build pays for itself in four weeks. A project that saves 1 hour per week and takes 40 hours takes ten months.


Complete Example Session

# ============================================
# PART 1: DEFINING THE SLO
# ============================================
slo:
  sli: api_availability
  target: 99.9%
  window: 30 days
# ============================================
# PART 2: CALCULATING THE ERROR BUDGET
# ============================================
error_budget:
  formula: 100% - SLO
  value: 0.1%
  minutes: 43
  seconds: 2592
# ============================================
# PART 3: MEASURING THE BUDGET
# ============================================
measurement:
  sli_value: 99.94%
  budget_consumed: 0.06%
  budget_remaining: 0.04%
  status: healthy
# ============================================
# PART 4: BURN RATE ALERTING
# ============================================
alerts:
  - name: fast_burn
    condition: 2% budget in 1 hour
    action: page on-call
  - name: slow_burn
    condition: 5% budget in 6 hours
    action: page on-call
  - name: ticket_burn
    condition: 10% budget in 3 days
    action: create ticket
# ============================================
# PART 5: ERROR BUDGET POLICY
# ============================================
policy:
  healthy:
    condition: budget_remaining > 50%
    action: ship features at normal velocity
  warning:
    condition: 25% < budget_remaining <= 50%
    action: review release cadence, prioritize reliability
  exhausted:
    condition: budget_remaining <= 25%
    action: freeze feature releases, focus on reliability
# ============================================
# PART 6: IDENTIFYING TOIL
# ============================================
toil:
  - manual restart of crashed service
  - manual certificate rotation
  - manual log correlation during incidents
  - manual deployment approval
  - manual environment provisioning
# ============================================
# PART 7: MEASURING THE TOIL PERCENTAGE
# ============================================
toil_measurement:
  week: 42
  engineer_hours:
    alice: { toil: 18, engineering: 22 }
    bob:   { toil: 24, engineering: 16 }
    carol: { toil: 14, engineering: 26 }
  total_toil: 56
  total_engineering: 64
  toil_percentage: 46.7%
  cap: 50%
  status: within_cap
# ============================================
# PART 8: AUTOMATION PROJECT
# ============================================
automation:
  project: auto-restart crashed service
  build_effort: 40 hours
  toil_saved: 10 hours per week
  payback_period: 4 weeks
  priority: high
# ============================================
# PART 9: TOIL REDUCTION PLAN
# ============================================
plan:
  quarter: Q4 2026
  target_toil_percentage: 30%
  projects:
    - auto-restart service
    - automated certificate renewal
    - CI/CD pipeline for all services
    - structured logging with trace IDs
# ============================================
# PART 10: REPORTING
# ============================================
report:
  slo_status: met
  error_budget_remaining: 0.04%
  toil_percentage: 46.7%
  toil_cap: 50%
  engineering_projects_active: 3

These ten parts cover defining the SLO, calculating the error budget, measuring the budget, burn rate alerting, the error budget policy, identifying toil, measuring the toil percentage, an automation project, a toil reduction plan, and reporting.


Quick Reference

Error Budget

SLOError BudgetDowntime per Month
99%1%~7.2 hours
99.5%0.5%~3.6 hours
99.9%0.1%~43 minutes
99.95%0.05%~21.5 minutes
99.99%0.01%~4.3 minutes

Error Budget Policy

Budget StateAction
Healthy (>50%)Ship features
Warning (25–50%)Review cadence, prioritize reliability
Exhausted (<25%)Freeze features, focus on reliability

Burn Rate Alerts

ConditionAction
2% in 1 hourPage on-call
5% in 6 hoursPage on-call
10% in 3 daysCreate ticket

Toil Properties

PropertyMeaning
ManualDone by a person
RepetitiveDone over and over
AutomatableA machine could do it
TacticalReactive
No lasting valueSystem not improved
Scales with growthMore as system grows

Toil Cap

AspectValue
Maximum toil50% of time
Engineering50% of time
MeasurementWeekly time log
ActionAutomate largest sources

Toil Reduction

ToilAutomation
Manual restartAuto-restart
Manual deployCI/CD
Manual provisionIaC
Manual rotationAuto-renewal
Manual correlationStructured logging

Best Practices

✅ Do This:

# Define the error budget from the SLO
error_budget: 100% - SLO
# Alert on burn rate
alert: 2% budget in 1 hour
# Define the budget policy in advance
if_exhausted: freeze feature releases
# Measure toil weekly
toil_percentage: 46.7%
# Prioritize automation by ROI
payback_period: 4 weeks

❌ Don’t Do This:

# Spend the entire budget every month
# ❌ no reserve for unexpected events
# Ignore the error budget policy
# ❌ the trade-off is not enforced
# Let toil exceed 50%
# ❌ no time for engineering work
# Automate the smallest toil first
# ❌ prioritize by impact, not ease

Common Pitfalls

PitfallWhy It HappensFix
Budget always exhaustedSLO too aggressiveReassess the target
Budget never consumedSLO too looseRaise the target
No burn rate alertingOnly end-of-window checkAlert on fast consumption
Policy not enforcedLeadership not committedAgree on the policy upfront
Toil unmeasuredNo trackingWeekly time log
Toil exceeds capNo automation investmentPrioritize automation by ROI
Automation of low-impact toilEase over impactEstimate hours saved

Real-World Examples

1. Error Budget Calculation

slo: 99.9%
error_budget: 0.1% = 43 minutes/month

2. Burn Rate Alert

alert: 2% in 1 hour → page on-call

3. Budget Policy

if_exhausted: freeze feature releases

4. Toil Identification

toil: manual certificate rotation

5. Toil Measurement

toil_percentage: 46.7%

6. Automation Project

project: auto-restart service
payback: 4 weeks

7. Toil Reduction Plan

target: 30% toil
projects: [auto-restart, CI/CD, structured logging]

8. Reporting

slo_status: met
toil_percentage: 46.7%

9. Budget Reserve

# Keep 50% of the budget in reserve

10. Engineering Time Protection

# 50% of time on engineering work

Visual

The Error Budget

┌──────────────────────────────────────────────────────────────┐
│  SLO: 99.9%                                                  │
│  └── Error budget: 0.1% (43 min/month)                       │
│                                                              │
│  Budget consumed by:                                         │
│  ├── Failed requests                                         │
│  ├── Slow responses                                          │
│  └── Other SLI violations                                    │
│                                                              │
│  Budget remaining = 0.1% - consumed                          │
└──────────────────────────────────────────────────────────────┘

Error Budget Policy

┌──────────────────────────────────────────────────────────────┐
│  HEALTHY (>50% remaining):                                   │
│  └── Ship features at normal velocity                        │
│                                                              │
│  WARNING (25-50% remaining):                                 │
│  └── Review release cadence, prioritize reliability          │
│                                                              │
│  EXHAUSTED (<25% remaining):                                 │
│  └── Freeze feature releases, focus on reliability           │
└──────────────────────────────────────────────────────────────┘

Burn Rate

┌──────────────────────────────────────────────────────────────┐
│  Burn rate = rate of budget consumption                      │
│                                                              │
│  Burn rate 1:   budget exhausted at end of window            │
│  Burn rate 10:  budget exhausted in 1/10 of window           │
│  Burn rate 14.4: 30-day budget exhausted in 2 days           │
│                                                              │
│  Alert on fast burn to catch problems early.                 │
└──────────────────────────────────────────────────────────────┘

Toil and Engineering Split

┌──────────────────────────────────────────────────────────────┐
│  SRE TEAM TIME                                               │
│  ┌────────────────────────┬─────────────────────────────────┐│
│  │  TOIL (≤50%)           │  ENGINEERING (≥50%)             ││
│  │  Manual, repetitive    │  Automation, improvement        ││
│  │  No lasting value      │  Reduces future toil            ││
│  └────────────────────────┴─────────────────────────────────┘│
│                                                              │
│  If toil > 50%, engineering work is starved and the          │
│  toil grows. The cap breaks the downward spiral.             │
└──────────────────────────────────────────────────────────────┘

Summary

ItemValue
Error budget100% minus the SLO
Budget purposeBalance reliability and velocity
Budget statesHealthy, warning, exhausted
Burn rateRate of budget consumption
Burn rate alertingCatch fast consumption early
ToilManual, repetitive, automatable work
Toil cap50% of team time
Engineering time50% of team time
Toil reductionAutomation
Automation prioritizationBy return on investment
Payback periodBuild effort / toil saved

Key takeaways:

  • The error budget is 100% minus the SLO. It is the amount of unreliability the team is allowed. It makes the trade-off between reliability and feature velocity explicit and measurable.
  • The error budget policy defines actions in advance. Healthy budget means ship features. Exhausted budget means freeze features and focus on reliability. The policy removes the debate from the moment of crisis.
  • Burn rate alerts catch problems before the budget is exhausted. A fast burn rate indicates a problem that will consume the budget before the window ends. Alerting on burn rate is more effective than alerting on the budget at the end of the window.
  • Toil is manual, repetitive, automatable work with no lasting value. It scales with system growth and consumes the team’s time. Examples include manual restarts, manual deployments, and manual certificate rotation.
  • The toil cap is 50% of the team’s time. The remaining 50% is dedicated to engineering work that reduces future toil. Without the cap, the toil grows and the engineering shrinks.
  • Automation is the response to toil. Each automation project has a cost and a return. Prioritizing by return on investment ensures the team addresses the largest sources of toil first.
  • The payback period is the build effort divided by the toil saved per period. A project that saves 10 hours per week and takes 40 hours to build pays for itself in four weeks.

Remember: Error budgets and toil limits are the operational mechanisms that make SRE sustainable. The error budget balances reliability against feature velocity, and the error budget policy defines what happens when the budget is exhausted. Toil is the repetitive work that consumes the team’s time, and the 50% cap protects the engineering time that reduces future toil. Together they answer two questions: how reliable should the service be, and how should the team spend its time? The answers are quantitative, agreed upon in advance, and enforced by policy. Without them, the organization relies on heroics, and heroics do not scale.



Stop using slow, ad-bloated tool sites! 🤮

🔎 Search “KandZ Tools” on Google to use many professional utilities for free.

KandZ.me is the ultimate minimalist hub for:
✅ Finance (Mortgage, Interest, Inflation)
✅ Tech (Base64, JSON, Dev Suite, IP)
✅ Health (BMI, BMR, TDEE)
✅ Productivity (Timer, Workspace, QR)

⚡️ Fast & Private
🔒 No data leaves your device
💎 100% Free

🔗 Use it now: https://tools.kandz.me
🔖 Bookmark it—you’ll need it later!