SRE-201 Intermediate ⏱ 6 weeks

SRE Fundamentals

Run reliable services: SLIs, SLOs and error budgets, monitoring with Prometheus and Grafana, and structured incident response.

Prerequisites: LNX-201 Linux System Administration

Syllabus

Module 1 — Reliability engineering

  • SRE vs DevOps
  • SLIs, SLOs, SLAs
  • Error budgets and toil

Module 2 — Monitoring

  • Prometheus architecture
  • PromQL
  • Grafana dashboards
  • Alertmanager

Module 3 — Incident management

  • On-call fundamentals
  • Incident response process
  • Blameless postmortems

Module 4 — Practical reliability

  • Health checks and runbooks
  • Capacity basics
  • Chaos engineering intro