Principal Production Engineer

Canva · Adelaide,

RemoteWorkplace
2w agoPosted · Jul 20
RemoteOKSource
Apply now Opens the original posting at Canva. PivotHop does not host applications.

Skills in this posting

Extracted from the posting text by the instrument — the demand side, read literally.

The posting

Job Description

Join the team redefining how the world experiences design.

Hey, g'day, mabuhay, kia ora, 你好, hallo, vítejte!

Thanks for stopping by. We know job hunting can be a little time consuming and you're probably keen to find out what's on offer, so we'll get straight to the point.

What You'd Be Doing In This Role

The Production Engineering team sits at the intersection of software engineering and the hardest reliability problems in Canva's infrastructure. At 240M MAUs, the hardest problems aren't on any product roadmap. Production Engineering exists to find them first and fix them properly. Writing software. Changing how production behaves. When it works, every team ships with more confidence and Canva gets faster and more resilient for the people who use it every day.

The strategic bet is a different model entirely. Canva's own take on what production reliability looks like, built for how we work. Senior software engineers embedded long-term in the areas that carry the most technical risk, working shoulder to shoulder with product teams, close enough to the roadmap to shape how features land in production before the problems compound. Not operationalising systems. Not running alerts. Writing software that changes how production behaves.

The engineers who do this work well have gone deep in systems most people only operate. They can walk into a codebase they didn't write, understand what's actually happening at scale, win the technical respect of the team they're embedded with, and then bend the software to make it more reliable, more efficient, and more resilient.

At Principal level, you're also the person who defines what this practice looks like at Canva. The calibration anchor for every hire, every engagement, and every standard the team sets.

At the moment, this role is focused on

Defining the engagement model: How Production Engineering pairs with product and infrastructure teams, how engagements are scoped, and what handoff actually looks like. This model is new. You're shaping it.

Leading the hardest engagements: Taking personal ownership of the most technically complex areas, sharding, multi-region architecture, JVM performance at scale, while the team builds depth in adjacent domains.

Setting the technical bar: What it means to be a production engineer at Canva. The standard for technical credibility. The archetype that future hiring calibrates against.

Pairing strategy across the team: Deciding how staff and mid-level engineers are paired and what they should be learning from each engagement. Growing production engineering capability across the org, not just within the team.

Connecting to the product roadmap: Working with engineering leadership across Canva to ensure Production Engineering is upstream of problems, not downstream. Influencing how product teams think about production readiness before they ship.

Building the measurement story: Incident severity and duration trending down. Feature launches going to production cleanly. You define what the metrics are and how they're tracked.

Compounding at organisation scale: One well-placed Production Engineering engagement changes how engineers build for years. At Principal level, your leverage is the sum of all those engagements, plus the standard you set for how each one runs.

What success looks like: As a secondee, developing trusted relationship with your team. Guiding them towards, shipping at velocity, with more confidence and less toil.

You’re probably a match

We'd love to hear from you if you fit one or more of these. You don't need to meet all of them, but the more the better and if you join the team, we're invested in helping you grow.

Experience

Production at scale: You've owned reliability in large-scale distributed systems. When things went brake, you investigate how and shipped the solution that lasts forever.

Technical leadership in embedded models: You've led or helped shape a function where engineers work across team boundaries rather than within a single one. You know what makes that model work and what makes it fail.

Hands-on through seniority: You've stayed close to the code. At this level, you're the engineer others consult when the problem is genuinely hard.

Cross-org influence: You've shaped how teams outside your own make technical decisions because your technical judgement is trusted.

JVM or systems depth: You've built real things in Java, Go, Rust, C++, or a comparable systems language at production scale. Language matters less than depth.

Distributed systems in practice: You've navigated sharding, replication, failure modes, and consistency tradeoffs in production.

Technical knowledge

Linux internals: You can reason about process scheduling, memory, I/O, and the network stack when a system misbehaves.

Distributed systems: You've navigated sharding, replication, failure modes, and consistency tradeoffs in production. As well as consistent hashing, leader election, consensus, backpressure, circuit breakers

Observability tooling: You've built the tracing, dashboards, and alerting that tells you what's wrong.

Containerisation and orchestration: Kubernetes in production, at the scheduler level.

Performance analysis: You've profiled JVM applications or systems-level processes and fixed what you found.

Cloud infrastructure: AWS in production, across the failure modes that matter at scale.

Incident response: You've been on-call and have opinions about what good looks like.

Nice to have

Enterprise SaaS background: You've done this specific kind of work at an org that's done it well. You know what "production engineering" means when it's not just a job title.

JVM internals: You've tuned GC and profiled threads in production.

Multi-region or sharding experience: You've been involved in a data store migration or multi-region architecture where getting it wrong was not an option.

About The Group And Team

Join the Production Engineering Group at Canva, where our mission is to make every system that powers Canva fast, reliable, and ready for the next scale. Infra owns the infrastructure layer that every other team builds on: compute, storage, networking, developer experience, and reliability.

The Reliability Platform subgroup is where Canva thinks seriously about the technical risk that comes with operating at hundreds of millions of users. It's a group with broad scope, from the tooling that helps teams run incidents well, to the engineering work that stops incidents from happening in the first place.

Production Engineering sits within Reliability Platform. A small team of senior software engineers embedded in Canva's highest-risk technical areas, working alongside the product and infrastructure teams who own those systems for a long-term engagements. When it works, other teams ship more confidently and the incidents that do happen resolve faster and hurt less.

What’s in it for you?

Achieving our crazy big goals motivates us to work hard — and we do — but you'll experience lots of moments of magic, connectivity and fun woven throughout life at Canva, too. We also offer a range of benefits to set you up for every success in and outside of work.

Excerpt from the original listing. The full, current text lives at the source. Read and apply there →

The PivotHop read

Where these skills also reach

Adjacent occupations measured from the same postings — readiness is what an industrial engineer’s profile already covers.

More industrial engineer roles

Backfilled listing, refreshed with the nightly scrape; the employer has not claimed it yet. Are you the employer? Claim this listing and it can be featured to the candidates whose skills already reach it, first month free.

© 2026 PivotHopReal data, real career moves