Linux Infrastructure Engineer (Bare Metal, Storage & AI Factory Infrastructure)
Uvation · Singapore
Skills in this posting
The posting
Job Overview
We are seeking a highly experienced Senior Linux Infrastructure Engineer with deep expertise in Linux administration, bare metal infrastructure, enterprise storage, and next-generation AI Factory / GPU infrastructure platforms . This role is focused on designing, deploying, operating, and troubleshooting large-scale Linux-based infrastructure that powers both traditional enterprise workloads and modern AI/ML environments.
This is not a DevOps-focused role . We already have a dedicated DevOps team and are looking for an engineer with extensive hands-on experience in Bare Metal as a Service (BMaaS), GPU infrastructure, high-performance storage, data center operations, and enterprise Linux platforms .
The ideal candidate will have experience building and managing infrastructure from the hardware layer up, including servers, networking, storage, GPU clusters, and AI-ready platforms. They should be comfortable working with high-performance computing (HPC), AI Factory environments, and large-scale Linux deployments where performance, reliability, and operational excellence are critical.
Key Responsibilities & Required Skills
Linux & Bare Metal Infrastructure
Expert-level Linux administration (Ubuntu required; Red Hat and SUSE preferred)
Deep expertise in bare metal server deployment, architecture, provisioning, and lifecycle management
Experience operating Bare Metal as a Service (BMaaS) platforms and large-scale infrastructure environments
BIOS/UEFI
RAID controllers
Firmware management
iLO/iDRAC/IPMI
NICs and SmartNICs
HBA cards
Hardware diagnostics and troubleshooting
Experience designing, implementing, and supporting enterprise Linux infrastructure at scale
AI Factory & GPU Infrastructure
Experience deploying and managing GPU-accelerated infrastructure for AI/ML workloads
Understanding of NVIDIA GPU technologies including
A100, H100, H200, B200, or equivalent GPU platforms
NVIDIA DGX and OEM GPU servers
GPU provisioning and lifecycle management
GPU monitoring and performance optimization
Knowledge of AI Factory architecture and infrastructure requirements
Experience supporting GPU clusters, AI training environments, and high-performance computing (HPC) workloads
Understanding of
GPU resource allocation and scheduling
Multi-GPU systems
GPU networking requirements
High-bandwidth, low-latency infrastructure design
NCCL
GPUDirect Storage
NVIDIA Fabric Manager
NVIDIA Base Command (preferred)
Enterprise Storage & Data Platforms
Advanced Linux storage administration
LVM
XFS, EXT4
NFS
iSCSI
Fibre Channel SAN
Multipath I/O
Strong hands-on experience with Ceph , including
Cluster architecture
MON, OSD, MDS
RBD, CephFS, RGW
Capacity planning
Performance tuning
Failure recovery
The PivotHop read
- What a devops engineer actually earnsmedian, seniority, by country
- What devops engineers do insteadevery measured route out
- All open devops engineer rolesthe full board
Where these skills also reach
- 600 open systems administrator roles62% readiness from devops engineer
- 600 open solutions architect roles51% readiness from devops engineer
- 138 open network engineer roles51% readiness from devops engineer
- 600 open backend developer roles45% readiness from devops engineer
More devops engineer roles
AI Platform Engineer (f/m/d) at IdnowBerlin, Berlin; München, BavariaTodayApply
Senior Site Reliability Engineer / SRE – Kubernetes & Hybrid Cloud (m/f/d) at Fact FinderBerlinTodayApply- Platform engineer at SecclLondonTodayApply
(Senior) DevOps Engineer (F/M/*) at AmberAachenTodayApply
DevOps Engineer (Senior/Staff) at DualEntryRemote$100k–$165kTodayApply
Backfilled listing, refreshed with the nightly scrape; the employer has not claimed it yet. Are you the employer? Claim this listing and it can be featured to the candidates whose skills already reach it, first month free.