Site Reliability Engineer
Ecomifly · Remote

Full-timeRemote$140,000 - $175,000
Apply NowAbout this role
Ensure our platform stays fast and reliable at scale. You will define SLOs, build observability tooling, and lead incident response.
Responsibilities
- Define and track SLIs/SLOs across critical services
- Build dashboards and alerting with Prometheus/Grafana
- Lead incident response and drive blameless postmortems
- Partner with product engineering on reliability reviews
Requirements
- 5+ years in SRE or infrastructure engineering
- Strong grasp of observability tooling (Prometheus, Grafana, OpenTelemetry)
- Experience leading incident response for production systems
- Comfortable writing automation in Python or Go
SREObservabilityPrometheusIncident Response