This role is with one of Dex's trusted partner companies. We work closely with their teams to truly understand their culture, goals, and what they're looking for, so we can match you with the right opportunity and give you context about the role before you commit to a process.
If you're interested sign up to Dex to apply.
Dex is an AI recruiter agent that helps you run your job search. Tell Dex your stack, seniority, and what you want to build. We will manage your applications and surface other opportunities that are a fit.
The role
This company is building a privacy-first perception engine, turning existing cameras and sensors into real-time operational intelligence without facial recognition or hardware swaps. They're tackling a petabyte-scale data infrastructure challenge: ingesting and processing real-time feeds from thousands of customer sites. This isn't a typical SaaS reliability surface; think financial market data aggregation, but for the physical world.
You'll be the founding engineer for a new Resilience function, owning the availability, performance, and scalability of this unique platform. This is a greenfield opportunity to apply a software engineering mindset to operations, designing systems from scratch. It's not about running a ticket queue; it's about building the infrastructure that makes petabyte-scale real-time data reliable.
The work
-
Design, build, and operate a centralized, petabyte-scale data ingestion pipeline, solving novel reliability challenges akin to financial market data.
-
Architect and implement production Kubernetes clusters and GitOps workflows from scratch, ensuring high availability and scalability.
-
Automate manual operations work across deployment, monitoring, and incident response, embedding a software engineering approach into infrastructure.
-
Define and manage SLOs/SLIs for critical systems, leading root cause analyses that result in tangible preventative changes.
-
Establish and champion documentation standards (runbooks, post-mortems, system docs) as a core engineering deliverable.
What You Bring
-
You've built at least one production Kubernetes cluster from scratch, not just operated an existing one.
-
You're fluent with GitOps workflows in production and have strong opinions on their optimal setup and operation.
-
You write code to automate operations work across deployment, monitoring, and incident response.
-
You've defined and managed SLOs/SLIs for production systems and led RCAs that resulted in preventative changes.
-
You are eligible to work in the UK and ideally based in London.
Why apply through Dex
This is a rare, high-impact role that won't be widely advertised. Apply through Dex to get properly briefed on the company and role, skip the cold application process, and get matched to other similar, unlisted opportunities. We help you cut through the noise and find the roles that truly fit.
If you're interested, sign up to Dex to apply -
https://jobs.meetdex.ai/jobs/3dc041c8-3c13-4de2-b049-f756ab2bfdb0
As part of the recruitment process at Dex, we process your personal data in accordance with our Privacy Notice for Job Applicants. This notice explains how and why your data is collected and used, and how you can contact us if you have any concerns.
We review applications on a rolling basis, so early applicants are encouraged.