Compute Infrastructure Lead
Enregistrez cette offre et organisez votre recherche
Créez un compte gratuit pour enregistrer des offres d'emploi, créer des alertes et revenir à cette liste depuis votre tableau de bord.
Your Mission
As Compute Infrastructure Lead, you will own and scale the compute backbone of UMA: the systems that provision, schedule, and run training, evaluation, and data-processing workloads — reliably, efficiently, and at scale — so our models can go from research to production without the cluster becoming the bottleneck.
This is a hands-on, high-impact role. We already train on a dedicated GPU cluster with a working training stack and a strong team behind it, so you won't be starting from zero — but you'll have the mandate to shape the architecture that takes us from a research cluster to a production-scale, multi-provider fleet, and to production-grade reliability as we start deploying POCs with industry partners. You'll take ownership of the compute platform end to end — multi-provider capacity, scheduling, distributed training / eval / processing, virtualized developer environments, observability, cost, and the tooling researchers and engineers actually use — and, if that's where you want to go, grow into leading the compute infrastructure team as it scales.
The technical problem is unusually rich for this stage. We are de-risking a stack built on pre-training and online RL, then industrializing it: heterogeneous hardware (training GPUs, cheaper eval and processing GPUs, CPU), a real-time learning loop, orchestrated data processing, interactive VMs on the cluster, and a fleet that will grow fast. Much of what we need — a true multi-provider compute fabric with elastic scheduling, dynamic checkpointing, unified observability, and jobs that resume themselves — does not exist off the shelf. Data infrastructure (datasets, storage, versioning) is owned by a sister role; this role is compute, including how processing jobs actually run on it.
Key responsibilities :
Own our compute platform end to end — from provisioning GPU capacity across cloud providers to keeping training, eval, and processing jobs running at high utilization, with reliability, cost, and researcher velocity as first-class goals
Build a multi-provider management layer so we can place, burst, and fail over workloads across GPU clouds, hyperscalers, and HPC without rewriting jobs
Design and operate the cloud scheduler — quotas, priority, preemption, topology-aware placement, and dynamic checkpointing so jobs survive node failure, preemption, and provider switches
Stand up a distributed compute framework for training and evaluation on heterogeneous hardware (e.g. Ray / similar), including the real-time / online-learning path
Orchestrate data-processing workloads at scale — CPU and cheaper GPUs, batch and streaming — so post-processing, dataset jobs, and training share one reliable compute fabric instead of ad-hoc scripts
Deliver virtualized GPU/CPU dev sessions (VMs on the cluster) so engineers iterate interactively on the same hardware and software they train on, without burning dedicated boxes
Build observability that works from any provider — system metrics (Prometheus, Grafana), job traces and logs, and model metrics (e.g. MLflow) — so a hung NCCL job, a silent GPU, or a broken training curve is diagnosable in minutes, not days
Own capacity, cost, and provider relationships as a technical lead: forecast demand, pick the right mix of hardware and contracts, and help negotiate pricing and terms. This is not a commercial role — but procurement is part of making the infra succeed
Help set production-grade practices (testing, reliability, fast iteration) as we move from R&D to partner POCs, and grow into leading the compute infra team if that's the path you want
What You Bring to the Table
8+ years in ML/compute infrastructure, HPC, or large-scale GPU platform engineering, at a senior, lead, or staff level
Proven track record building and operating infrastructure for large-scale AI model training — not inference-only. Multi-node GPU clusters, distributed training (PyTorch / NCCL or equivalent), and keeping long-running jobs healthy at scale
Deep, hands-on experience with GPU clouds and cluster operations: provisioning, Linux, high-performance networking (InfiniBand / RoCE), storage for training, utilization, and GPU/node failure modes
Built or owned schedulers and distributed frameworks (Slurm, Kubernetes, Ray, SkyPilot, or similar) — including checkpointing, elasticity, and preemption — so GPUs stay busy and jobs come back from failure
Treat observability, reliability, and cost as core engineering concerns, not afterthoughts
Strong Python and systems engineering, with the taste to build tooling that researchers actually want to use
Experience working with GPU providers on capacity and commercial terms — you can read a contract, push on price and availability, and still be the person who debugs the cluster at 2am
Ability to reason about systems end-to-end — performance, scalability, reliability, cost — and make and defend the right trade-offs
Thrive in a hands-on, fast-paced startup, building from a real (but small) cluster toward a production fleet: autonomous, rigorous, execution-driven, easy to work with, and broadly curious about AI and systems
Bonus : online / continuous RL, real-time training loops, or other always-on learning systems
Bonus : multi-cloud / multi-provider fabrics (SkyPilot, similar), Ray / Anyscale, HPC centers, VM-based GPU workstations / interactive cluster sessions, or standing up clusters from tens to hundreds of nodes
Bonus : robotics, autonomous vehicles, or other embodied/physical-AI training stacks — adjacent large-scale training (LLMs, multimodal, AV) counts strongly; robotics itself is not required
Bonus : public projects, open-source contributions, maintained tools, or technical writing
We value exceptional builders over perfect resumes. If you have a world-class track record building training infrastructure at scale and the drive to build the compute backbone that lets a robotics company scale, we strongly encourage you to apply — even if you don't tick every box. Robotics experience is a plus, not a requirement.