About the role
You will join a follow-the-sun reliability team responsible for the availability of client environments across AWS, Azure and Google Cloud. This is an engineering role, not a monitoring role: you will write the automation, define the service level objectives and lead the reviews when we miss them.
What you'll do
- Define and maintain SLOs and error budgets across client environments
- Lead incident response and write blameless postmortems that change the system
- Build automation that removes recurring manual operational work
- Partner with migration teams to make new environments operable from day one
- Mentor engineers joining the on-call rotation
What we're looking for
- 5+ years operating production infrastructure at scale
- Depth in at least one major cloud provider and comfort in the others
- Strong Go, Python or equivalent for automation work
- Experience with Terraform, Kubernetes and modern observability stacks
- Track record of leading incidents in a regulated or high-stakes environment