Linux Infrastructure Engineer (Bare Metal, Storage & AI Factory Infrastructure)

Uvation · Singapore

RemoteWorkplace
TodayPosted · Sep 8
HimalayasSource
$142kdevops engineer median
Apply now Opens the original posting at Uvation. PivotHop does not host applications.

Skills in this posting

The posting

Job Overview

We are seeking a highly experienced Senior Linux Infrastructure Engineer with deep expertise in Linux administration, bare metal infrastructure, enterprise storage, and next-generation AI Factory / GPU infrastructure platforms . This role is focused on designing, deploying, operating, and troubleshooting large-scale Linux-based infrastructure that powers both traditional enterprise workloads and modern AI/ML environments.

This is not a DevOps-focused role . We already have a dedicated DevOps team and are looking for an engineer with extensive hands-on experience in Bare Metal as a Service (BMaaS), GPU infrastructure, high-performance storage, data center operations, and enterprise Linux platforms .

The ideal candidate will have experience building and managing infrastructure from the hardware layer up, including servers, networking, storage, GPU clusters, and AI-ready platforms. They should be comfortable working with high-performance computing (HPC), AI Factory environments, and large-scale Linux deployments where performance, reliability, and operational excellence are critical.

Key Responsibilities & Required Skills

Linux & Bare Metal Infrastructure

Expert-level Linux administration (Ubuntu required; Red Hat and SUSE preferred)

Deep expertise in bare metal server deployment, architecture, provisioning, and lifecycle management

Experience operating Bare Metal as a Service (BMaaS) platforms and large-scale infrastructure environments

BIOS/UEFI

RAID controllers

Firmware management

iLO/iDRAC/IPMI

NICs and SmartNICs

HBA cards

Hardware diagnostics and troubleshooting

Experience designing, implementing, and supporting enterprise Linux infrastructure at scale

AI Factory & GPU Infrastructure

Experience deploying and managing GPU-accelerated infrastructure for AI/ML workloads

Understanding of NVIDIA GPU technologies including

A100, H100, H200, B200, or equivalent GPU platforms

NVIDIA DGX and OEM GPU servers

GPU provisioning and lifecycle management

GPU monitoring and performance optimization

Knowledge of AI Factory architecture and infrastructure requirements

Experience supporting GPU clusters, AI training environments, and high-performance computing (HPC) workloads

Understanding of

GPU resource allocation and scheduling

Multi-GPU systems

GPU networking requirements

High-bandwidth, low-latency infrastructure design

NCCL

GPUDirect Storage

NVIDIA Fabric Manager

NVIDIA Base Command (preferred)

Enterprise Storage & Data Platforms

Advanced Linux storage administration

LVM

XFS, EXT4

NFS

iSCSI

Fibre Channel SAN

Multipath I/O

Strong hands-on experience with Ceph , including

Cluster architecture

MON, OSD, MDS

RBD, CephFS, RGW

Capacity planning

Performance tuning

Failure recovery

The PivotHop read

Where these skills also reach

More devops engineer roles

Backfilled listing, refreshed with the nightly scrape; the employer has not claimed it yet. Are you the employer? Claim this listing and it can be featured to the candidates whose skills already reach it, first month free.

© 2026 PivotHopReal data, real career moves