← All roles
ContractMedia
Senior Site Reliability Engineer
Streaming media company, London
Keep a service reliable through live-event traffic spikes, and make the spikes boring through capacity work rather than heroics.

About the client
A consumer streaming service with traffic that swings 20x around live events.
We name the company on the intake call. It is not on this page because the search is not public — the current team does not know they are hiring for it yet.
What you would work on
- Own capacity planning and load testing ahead of scheduled live events
- Improve autoscaling behaviour that currently reacts too slowly to matter
- Define SLOs with product teams and hold the error budget conversations
- Run blameless incident reviews and see the follow-up actions through
What we are looking for
- Deep AWS and Kubernetes experience under real production load
- Fluency with observability tooling — you have built dashboards people trust
- Incident command experience on user-facing outages
- Able to join a follow-the-sun on-call rotation, roughly one week in five
Apply for this role
We read every application ourselves and reply within three working days, including when the answer is no. Nothing goes to the client without your say-so.