Webneuron
Open Role

Site Reliability Engineer, Kubernetes

A 12-month contract embedded with a client platform team, keeping high-traffic Kubernetes workloads reliable. You'll own SLOs, incident response and the reliability roadmap through a major migration.

Engineering

Department

Remote — North America

Location

Contract

Employment type

Responsibilities

  • Define and track SLOs, error budgets and reliability metrics with client teams
  • Operate and tune production Kubernetes clusters across multiple environments
  • Lead incident response and run blameless postmortems
  • Automate toil out of the operational path
  • Harden capacity, failover and disaster-recovery posture ahead of migration cutover

Requirements

  • 5+ years in SRE or production platform operations
  • Deep Kubernetes operational experience at scale
  • Strong observability tooling skills (Prometheus, Grafana, OpenTelemetry)
  • Automation depth in Go, Python or Bash
  • Comfortable joining an existing client team and being productive quickly

Nice to Have

  • Experience with multi-region or active-active architectures
  • Chaos engineering or resilience testing background
  • Availability for occasional on-call rotation coverage

Engineering Remote — North America

Interested in this role? Send your CV to careers@webneuron.com or get in touch below.

Get in Touch