Principal Site Reliability Engineer

Okta · Bengaluru, India

On-siteWorkplace
Jun 15Posted · Jun 15
GreenhouseSource
Apply now Opens the original posting at Okta. PivotHop does not host applications.

Skills in this posting

Extracted from the posting text by the instrument — the demand side, read literally.

The posting

Secure Every Identity, from AI to Human

Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence.

This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk.

Get to know Okta

Okta is The World’s Identity Company. We free everyone to safely use any technology—anywhere, on any device or app. Our Workforce and Customer Identity Clouds enable secure yet flexible access, authentication, and automation that transforms how people move through the digital world, putting Identity at the heart of business security and growth.

At Okta, we celebrate a variety of perspectives and experiences. We are not looking for someone who checks every single box, we’re looking for lifelong learners and people who can make us better with their unique experiences.

Join our team! We’re building a world where Identity belongs to you.

The Engineering Opportunity

We are seeking a Principal Site Reliability Engineer to serve as a technical leader for reliability engineering within Okta's Emerging Products Group (EPG).

This role extends beyond operating production systems. You will define technical strategy, influence platform architecture, establish reliability standards, and lead transformational initiatives that improve scalability, resilience, security, and operational excellence for one of Okta's fastest-growing product areas.

Initially, you will partner closely with the Spera / Identity Security Posture Management (ISPM) engineering organization to establish reliability strategy, operational excellence, and platform maturity. Over time, you will help drive broader reliability initiatives across EPG and contribute to the evolution of reliability engineering practices across multiple products including Workflows, IGA, PAM, and ISPM.

You will work closely with engineering leadership, product leadership, architects, and Staff engineers to shape the future of Okta's cloud infrastructure and reliability practices.

The ideal candidate combines deep technical expertise with strong organizational influence and has a proven track record of leading large-scale engineering initiatives that drive measurable business outcomes.

What You'll Be Doing

Reliability Strategy & Architecture

  • Define and drive the reliability strategy for critical product and platform services.
  • Establish standards for availability, resilience, observability, incident management, and operational readiness.
  • Lead architecture reviews for critical services and platform initiatives.
  • Partner with engineering leaders to ensure reliability objectives align with business priorities and customer expectations.
  • Create frameworks, standards, and operational guardrails that enable engineering teams to operate safely at scale.
  • Guide service architecture toward simplicity, scalability, resilience, and operational excellence.
  • Drive major initiatives that improve platform maturity and long-term sustainability.

Product & Platform Leadership

  • Own reliability architecture and operational excellence for the Spera / ISPM product area.
  • Collaborate closely with engineering leadership to establish reliability objectives and technical roadmaps.
  • Lead large-scale scalability, resiliency, and performance initiatives.
  • Partner with platform and product engineering teams to build self-service operational capabilities that improve developer productivity while strengthening reliability and security.
  • Influence technical direction through data-driven recommendations, engineering expertise, and collaborative leadership.
  • Support highly available, large-scale cloud environments as part of an on-call rotation.

Engineering & Automation

  • Design, build, and operate large-scale cloud infrastructure and production services.
  • Develop software, automation, and infrastructure using Go, Python, Terraform, and related technologies.
  • Eliminate operational toil through automation, tooling, and platform engineering.
  • Improve deployment safety, operational workflows, and platform consistency through GitOps and Infrastructure-as-Code practices.
  • Collaborate on modernizing existing workloads and aligning them with evolving platform capabilities.
  • Lead complex engineering initiatives from conception through production rollout and long-term operational ownership.

Technical Leadership

  • Mentor Staff and Senior engineers across multiple teams and organizations.
  • Lead technical reviews, design reviews, and operational readiness assessments.
  • Build engineering consensus across teams with differing priorities and objectives.
  • Help develop the next generation of technical leaders within Okta.
  • Drive adoption of reliability engineering best practices across EPG.
  • Share patterns, tooling, and operational practices across Workflows, Inbox, PAM, and ISPM teams.
  • Influence technical direction through expertise, collaboration, and execution rather than organizational authority.

AI & Agentic Operations

  • Lead the exploration and adoption of AI-assisted reliability engineering practices across EPG.
  • Design and champion agentic systems that accelerate troubleshooting, incident response, root-cause analysis, and operational decision-making.
  • Evaluate emerging AI technologies and identify practical opportunities to improve reliability engineering workflows.
  • Establish best practices for safe, effective, and measurable use of AI within production operations.
  • Drive initiatives that reduce operational toil and improve engineering productivity through intelligent automation.

Our Tech Stack

  • Infrastructure/Orchestration: Kubernetes (EKS/GKE), Terraform, Helm, Git, ArgoCD, Gitops
  • Programming: Golang, Python
  • Observability: Datadog, Splunk
  • Data Stores: PostgreSQL, Redis, OpenSearch

What We Are Looking For

Technical Excellence

  • Extensive experience designing and operating large-scale production systems in AWS and/or GCP.
  • Deep expertise with Kubernetes in production environments.
  • Experience designing reliability strategies for Kubernetes-based platforms.
  • Strong expertise troubleshooting Kubernetes networking, storage, scheduling, scaling, and workload lifecycle challenges.
  • Extensive experience with Infrastructure as Code technologies such as Terraform and Helm.
  • Strong software engineering skills in Golang and/or Python.
  • Experience building internal platforms, developer tooling, and operational automation.
  • Deep understanding of distributed systems architecture and cloud-native application design.
  • Strong understanding of cloud networking fundamentals including DNS, service discovery, ingress, load balancing, TLS, traffic management, and multi-region architectures.
  • Experience operating and troubleshooting distributed data platforms such as PostgreSQL, Redis, OpenSearch, MySQL, Cassandra, or similar technologies.

Excerpt from the original listing. The full, current text lives at the source. Read and apply there →

The PivotHop read

Where these skills also reach

Adjacent occupations measured from the same postings — readiness is what a devops engineer’s profile already covers.

More devops engineer roles

Backfilled listing, refreshed with the nightly scrape; the employer has not claimed it yet. Are you the employer? Claim this listing and it can be featured to the candidates whose skills already reach it, first month free.

© 2026 PivotHopReal data, real career moves