Open Role
Site Reliability Engineer, Kubernetes
A 12-month contract embedded with a client platform team, keeping high-traffic Kubernetes workloads reliable. You'll own SLOs, incident response and the reliability roadmap through a major migration.
Engineering
Department
Remote — North America
Location
Contract
Employment type
Responsibilities
- Define and track SLOs, error budgets and reliability metrics with client teams
- Operate and tune production Kubernetes clusters across multiple environments
- Lead incident response and run blameless postmortems
- Automate toil out of the operational path
- Harden capacity, failover and disaster-recovery posture ahead of migration cutover
Requirements
- 5+ years in SRE or production platform operations
- Deep Kubernetes operational experience at scale
- Strong observability tooling skills (Prometheus, Grafana, OpenTelemetry)
- Automation depth in Go, Python or Bash
- Comfortable joining an existing client team and being productive quickly
Nice to Have
- Experience with multi-region or active-active architectures
- Chaos engineering or resilience testing background
- Availability for occasional on-call rotation coverage
Engineering Remote — North America
Interested in this role? Send your CV to careers@webneuron.com or get in touch below.