Senior Software Engineering Manager - FinOps Platform Services
ServiceNow · Pleasanton, US
Skills in this posting
Benefits
The posting
What you get to do in this role
Platform Ownership & Operations
Own the operational health and reliability of Trino, Lightdash, Coder, Jupyter, Redash, Hive Metastore, and Nessie across development and production environments.
Establish and maintain SLOs for platform availability, query performance, and workspace provisioning. Build the dashboards and alerting to track them.
Own Trino cluster operations end to end, including deployment, scaling, upgrades, performance tuning, resource group management, query optimization support, and user access controls.
Drive the platform upgrade and patching cadence, balancing stability with staying current on security fixes and feature releases across all services.
Build runbooks, on-call processes, and incident-response practices so the team can respond to and resolve production issues quickly and learn from them.
Ensure platform security across all services, including access controls, authentication (SSO/OIDC integration), secrets management, and audit logging.
Platform Evolution & Roadmap
Lead the migration from Hive Metastore to Nessie as the versioned Iceberg catalog, delivering Git-like branching semantics, safe multi-writer coordination, and auditable catalog history.
Drive Lightdash platform improvements including version upgrades, performance optimization, row-level security configuration, and the governed self-service analytics experience.
Evolve the Coder platform through workspace template lifecycle management, resource policies, idle-stop tuning, and onboarding new users and use cases including AI coding agents.
Own the Jupyter and Redash platforms, ensuring availability, scaling, integration with Trino and the lakehouse, and user lifecycle management.
Evaluate and adopt new open-source technologies where they raise the platform’s ceiling or reduce operational burden.
People Leadership
Manage, mentor, and grow a team of 3 to 5 platform engineers. Set clear expectations, provide regular feedback, and create career development paths.
Hire and build the team to match the platform’s growing scope and user base.
Foster a culture of operational excellence, automation over toil, and blameless incident retrospectives.
Set engineering standards for how the team builds, deploys, monitors, and documents platform services.
Collaboration & Stakeholder Management
Partner with the DevOps/infrastructure team on Kubernetes capacity, networking, storage, and CI/CD pipeline needs for your platform services.
Serve as the platform liaison to data engineers, analysts, and FinOps practitioners. Understand their workflows, gather feedback, and prioritize improvements that unblock them.
Collaborate with the Data Platform and Data Governance teams to ensure platform services align with enterprise standards for security, lineage, and access control.
Support the broader Cloudera-to-lakehouse migration by ensuring Trino, Nessie, and the catalog layer are production-ready for migrated workloads.
Apply AI/ML tooling where it accelerates platform operations, monitoring, or user support.
What success looks like
Platform services meet their SLOs consistently, and the team has the observability and processes to detect and resolve issues before users are affected.
Trino queries perform reliably at scale with well-managed resource groups and a clear upgrade cadence.
The Hive Metastore to Nessie migration is planned, sequenced, and executing without disruption to downstream users.
Lightdash and Coder are stable, current, and adopted broadly across the organization with minimal friction for new users.
The team is healthy, growing, and operating with clear ownership, automation, and documentation.
Internal users trust the platform and rarely lose productive time to platform instability.
To be successful in this role, you have
Experience leveraging or critically thinking about how to integrate AI into work processes, decision-making, or problem-solving.
12+ years of experience in software or platform engineering, with 5+ years in engineering management leading teams that own production platform services, with a Bachelor’s degree; or 10 years and a Master’s degree; or a PhD with 7 years of experience in Computer Science, Engineering, or a related technical field; or equivalent experience.
Proven track record managing teams that operate and scale open-source data infrastructure (query engines, BI platforms, developer environments, or similar) in production.
Hands-on experience operating distributed query engines (Trino, Presto, Spark, or similar) including cluster tuning, scaling, and performance optimization.
Strong knowledge of Kubernetes and containerized service deployment, enough to architect solutions and debug issues even if a separate team owns the clusters.
Demonstrated ability to establish SLOs, observability, and incident-response practices for platform services and to drive operational maturity over time.
Experience managing platform upgrades, migrations, and version lifecycle for open-source technologies in production without disrupting users.
Proven people leadership. Experience hiring, developing, and retaining strong platform engineers, and building team culture around operational excellence and automation.
Strong bias toward automation over manual toil, with experience building or directing the development of internal tooling and self-service workflows.
Excellent collaboration skills across engineering, data, DevOps, and business stakeholders.
Full professional proficiency in English.
Technical Expertise
Distributed query engines. Trino or Presto operations including deployment, scaling, resource group management, query optimization, connector configuration, and upgrades.
Data catalog and lakehouse. Hive Metastore operations and familiarity with modern catalog alternatives (Nessie, AWS Glue, Unity Catalog, Polaris). Apache Iceberg table format concepts.
BI and analytics platforms. Operating self-hosted BI tools such as Lightdash, Redash, Metabase, or Superset, including deployment, scaling, SSO integration, and user management.
Developer platforms. Coder, JupyterHub, or similar cloud development environment platforms, including workspace provisioning, template management, and resource policies.
Observability. Monitoring, alerting, and logging for platform services (Splunk, Prometheus, Grafana, CloudWatch, or similar). SLO design and tracking.
Security and access control. SSO/OIDC integration, RBAC, row-level security, secrets management, and audit logging across platform services.
Infrastructure familiarity. Kubernetes, Helm, Docker, Infrastructure as Code (Terraform, CDK), and CI/CD pipelines. Enough depth to partner effectively with infra teams and architect platform deployments.
Scripting and automation. Python, Bash, or Go for operational tooling, automation, and integration work.
Leadership & Communication
Proven ability to balance hands-on technical work with people leadership, knowing when to go deep and when to delegate.
Strong technical judgment with the ability to evaluate open-source technologies, make build-vs-buy decisions, and sequence a platform roadmap.
The PivotHop read
- What an engineering manager actually earnsmedian, seniority, by country
- Careers an engineering manager can move intoevery measured route out
- All open engineering manager rolesthe full board
Where these skills also reach
- 1742 open product manager roles56% readiness from engineering manager
- 7 open conversation designer roles54% readiness from engineering manager
- 1246 open solutions architect roles52% readiness from engineering manager
- 24 open ai product manager roles49% readiness from engineering manager
More engineering manager roles
Engineering Manager, Product Engineering at BoulevardUSA · Remote1d agoApply
Director of Engineering, Data at BoulevardUSA · Remote1d agoApply
Engineering Manager Senior at Factor ITColombia1d agoApply
Senior Engineering Manager - Enterprise Trust & Reliability at MultiverseLondon1d agoApply
Senior Engineering Manager, Brand Ads and Promotions at DeliverooLondon, England; London, United Kingdom - Deliveroo1d agoUnlock
Backfilled listing, refreshed with the nightly scrape; the employer has not claimed it yet. Are you the employer? Claim this listing and it can be featured to the candidates whose skills already reach it, first month free.