Site Reliability Engineer

Ecomifly · Remote

Ecomifly
Full-timeRemote$140,000 - $175,000
Apply Now

About this role

Ensure our platform stays fast and reliable at scale. You will define SLOs, build observability tooling, and lead incident response.

Responsibilities

  • Define and track SLIs/SLOs across critical services
  • Build dashboards and alerting with Prometheus/Grafana
  • Lead incident response and drive blameless postmortems
  • Partner with product engineering on reliability reviews

Requirements

  • 5+ years in SRE or infrastructure engineering
  • Strong grasp of observability tooling (Prometheus, Grafana, OpenTelemetry)
  • Experience leading incident response for production systems
  • Comfortable writing automation in Python or Go
SREObservabilityPrometheusIncident Response