Site Reliability Engineer, Infrastructure Platforms — UK (Intermediate to Senior Staff)

GitLab · Remote

RemoteWorkplace
1d agoPosted · Sep 8
GreenhouseSource
$142kdevops engineer median
Apply now Opens the original posting at GitLab. PivotHop does not host applications.

Skills in this posting

The posting

GitLab is the intelligent orchestration platform for DevSecOps. GitLab enables organizations to increase developer productivity, improve operational efficiency, reduce security and compliance risk, and accelerate digital transformation. More than 50 million registered users and more than 50% of the Fortune 100* trust GitLab to ship better, more secure software faster.

The same principles built into our products are reflected in how our team works: we embrace AI as a core productivity multiplier, with all team members expected to incorporate AI into their daily workflows to drive efficiency, innovation, and impact. GitLab is where careers accelerate, innovation flourishes, and every voice is valued.

Our high-performance culture is driven by our values and continuous knowledge exchange, enabling our team members to reach their full potential while collaborating with industry leaders to solve complex problems. Co-create the future with us as we build technology that transforms how the world develops software.

* Fortune 500® is a registered trademark of Fortune Media IP Limited, used under license. Claim based on GitLab data. Fortune 100 refers to the top 20% ranked companies in the 2025 Fortune 500 list, published in June 2025. Fortune and Fortune Media IP Limited are not affiliated with, and do not endorse products or services of GitLab.

An overview of this role

Site Reliability Engineers keep GitLab's user-facing services and production systems running reliably at scale. They combine software engineering with operational excellence, applying sound engineering principles, automation, and continuous improvement to build, operate, and evolve our production infrastructure.

This is a single application for Site Reliability Engineering opportunities across Infrastructure Platforms. Rather than asking you to choose the right team or level upfront, we evaluate your skills holistically and match you to the opportunity that best aligns with your experience and our hiring needs. We hire Site Reliability Engineers from Intermediate through Senior Staff across multiple Infrastructure Platforms teams.

We don't expect every candidate to have experience with every technology in our environment. We're looking for engineers with strong technical fundamentals, a growth mindset, and the ability to learn quickly. We'll support you in becoming successful with GitLab's tools, systems, and ways of working.

Please note: This position is open to candidates based in the United Kingdom only. Candidates based in the United States or Canada can apply to this posting: Site Reliability Engineer, Infrastructure Platforms — AMER (Intermediate to Senior Staff)

How our SRE hiring works

Because this is a single application for SRE roles across Infrastructure Platforms, our process is built to evaluate you once and match you well, rather than interviewing separately for every team.

Recruiter Screen: A conversation about your background, what you're looking for, and the level and teams that fit, so we can point your process in the right direction.

Core Technical: The shared assessment every SRE candidate takes, regardless of eventual team. A low-stress, collaborative discussion covering system architecture and incident review.

Hiring Manager Interview: A conversation about ownership, judgment, execution, collaboration, and growth, the non-technical signals that make an SRE effective at GitLab.

Peer Technical: Team-specific depth, run by SREs from the team you're most likely to join, focused on the problems that team actually works on.

Skip-Level Interview: A conversation with a senior leader on values alignment, and how you'll work across teams.

After your interviews, we consider your performance alongside our current hiring needs to confirm the level and team where you'll do your best work. Interview results are a major factor, and final placement also reflects our active hiring priorities at the time.

We’ll calibrate your level throughout the interview process based on the scope and impact of your experience.

Intermediate: You independently deliver meaningful reliability improvements within a defined area.

Senior: You own complex reliability work end to end and raise the effectiveness of your team.

Staff: You shape reliability across multiple teams, solving systemic problems and creating approaches others can reuse.

Senior Staff: You set technical direction across a broader Infrastructure area and influence reliability strategy at organizational scale.

What you'll do

Keep user-facing services and production systems reliable, scalable, and efficient

Build automation and tooling that reduces toil and replaces manual work with repeatable, infrastructure-as-code-driven workflows

Operate and troubleshoot production systems on Kubernetes, including deployments, rollouts, and scaling

Write and maintain infrastructure as code, and ship changes safely through CI/CD and GitOps

Participate in on-call, triage alerts, follow and improve runbooks, and escalate appropriately

Contribute to the observability stack, using metrics, logs, and SLOs to detect symptoms early rather than just outages

Take part in incident response and post-incident reviews, turning learnings into changes in automation and process

Document runbooks, architecture decisions, and reviews so your findings become repeatable practices

What you'll bring

Experience keeping production systems reliable, combining an operations mindset with real software engineering practice

Experience building net-new infrastructure tooling and automation, not just configuring existing tools. For example, Terraform modules, Kubernetes operators or controllers, or production automation and services written from scratch

The ability to read, debug, and reason about code. Most of our teams work in Go; some work in Ruby. You can discuss a piece of code's behavior, performance, and failure modes

Experience with infrastructure as code, and with Kubernetes and its ecosystem, at a depth appropriate to your level

Hands-on experience with at least one major cloud provider (GCP or AWS)

Familiarity with observability practices, including metrics, logging, alerting, and SLOs or SLIs, and using data to inform operational decisions

Comfort participating in on-call and incident response, with a structured approach to troubleshooting under pressure

Strong written communication and the ability to operate as a manager-of-one in an async, distributed environment

A track record of using automation, and increasingly AI, to reduce toil and improve how you and your team work

Alignment with GitLab's values and a commitment to working in accordance with them

About the team

Infrastructure Platforms is responsible for the availability, reliability, performance, and scalability of GitLab’s user-facing services, most notably GitLab.com . The organization spans teams across Production Engineering , GitLab Dedicated , GitLab Delivery , and Developer Experience , covering everything from the production fleet and networking platform to observability, incident response, deployment infrastructure, tenant scale, and our single-tenant Dedicated offering.

The PivotHop read

Where these skills also reach

More devops engineer roles

Backfilled listing, refreshed with the nightly scrape; the employer has not claimed it yet. Are you the employer? Claim this listing and it can be featured to the candidates whose skills already reach it, first month free.

© 2026 PivotHopReal data, real career moves