Machine Learning Engineer — Inference Optimization
Featherless AI · Canada; Germany; India; United Kingdom; United States
Skills in this posting
The posting
About the Role We’re looking for a Machine Learning Engineer to own and push the limits of model inference performance at scale . You’ll work at the intersection of research and production—turning cutting-edge models into fast, reliable, and cost-efficient systems that serve real users.
This role is ideal for someone who enjoys deep technical work, profiling systems down to the kernel/GPU level, and translating research ideas into production-grade performance gains.
What You’ll Do Optimize inference latency, throughput, and cost for large-scale ML models in production Profile and bottleneck GPU/CPU inference pipelines (memory, kernels, batching, IO) Implement and tune techniques such as: Quantization (fp16, bf16, int8, fp8) KV-cache optimization & reuse Speculative decoding, batching, and streaming Model pruning or architectural simplifications for inference Collaborate with research engineers to productionize new model architectures Build and maintain inference-serving systems (e.g.
Triton, custom runtimes, or bespoke stacks) Benchmark performance across hardware (NVIDIA / AMD GPUs, CPUs) and cloud setups Improve system reliability, observability, and cost efficiency under real workloads What We’re Looking For Strong experience in ML inference optimization or high-performance ML systems Solid understanding of deep learning internals (attention, memory layout, compute graphs) Hands-on experience with PyTorch (or similar) and model deployment Familiarity with GPU performance tuning (CUDA, ROCm, Triton, or kernel-level optimizations) Experience scaling inference for real users (not just research benchmarks) Comfortable working in fast-moving startup environments with ownership and ambiguity Nice to Have Experience with LLM or long-context model inference Knowledge of inference frameworks (TensorRT, ONNX Runtime, vLLM, Triton) Experience optimizing across different hardware vendors Open-source contributions in ML systems or inference tooling Background in distributed systems or low-latency services Why Join Us Real ownership over performance-critical systems Direct impact on product reliability and unit economics Close collaboration with research, infra, and product Competitive compensation + meaningful equity at Series A A team that cares about engineering quality, not hype Originally posted on Himalayas
The PivotHop read
- What a machine learning engineer actually earnsmedian, seniority, by country
- Careers a machine learning engineer can move intoevery measured route out
- Data Scientist → Machine Learning Engineer54% readiness
- MLOps Engineer → Machine Learning Engineer49% readiness
- All open machine learning engineer rolesthe full board
Where these skills also reach
- 501 open data scientist roles70% readiness from machine learning engineer
- 414 open ai engineer roles59% readiness from machine learning engineer
- 9 open conversation designer roles54% readiness from machine learning engineer
- 32 open mlops engineer roles48% readiness from machine learning engineer
More machine learning engineer roles
Senior Machine learning Engineer at BoschGroupbangalore, INTodayApply- AI/ML Engineer at XebiaceeBulgaria; Poland; Romania1d agoApply
AI/ML Engineer_MS at BoschGrouptelengana, IN1d agoApply- Senior Machine Learning Engineer at SeatGeekUSA · Remote1d agoApply
(Senior) Machine Learning Engineer – Pricing (m/f/d) at Auxmoney GmbhRemote / Düsseldorf1d agoApply
Backfilled listing, refreshed with the nightly scrape; the employer has not claimed it yet. Are you the employer? Claim this listing and it can be featured to the candidates whose skills already reach it, first month free.