Member of Technical Staff — Infrastructure
Il y a 1 semaine
Paris, Ile-de-France
Remanence
Temps plein
Gratuit avec email ou Google
Enregistrez cette offre et organisez votre recherche
Créez un compte gratuit pour enregistrer des offres d'emploi, créer des alertes et revenir à cette liste depuis votre tableau de bord.
Gratuit avec email ou Google
En continuant, vous acceptez nos Conditions d’utilisation & Politique de confidentialité.
Member of Technical Staff — Infrastructure
Remanence is pioneering the next era of enterprise AI by building intelligent systems that learn continuously from real-world execution. We transform complex enterprise workflows and business context into dynamic, interactive environments where AI agents can safely learn, adapt, and improve. By combining high-fidelity simulation environments with state-of-the-art training loops, we build specialized models that solve long-horizon, complex tasks with unmatched reliability. We’re building the most talent-dense AI team in Europe to make this happen.
You’ll own the compute and execution platform supporting training, inference, task generation, and evaluation. Your work will make GPU resources productive and research workflows straightforward to operate and debug.
What you’ll work on
• Build and operate GPU clusters, job scheduling, networking, and storage for distributed ML workloads.
• Develop the execution platform for large numbers of concurrent environments, including sandbox isolation, resource limits, retries, and state recovery.
• Build reliable data and artifact pipelines for datasets, trajectories, model weights, and checkpoints.
• Own platform observability and recovery; improve capacity allocation, startup latency, and reliability through measurable changes.
• Give researchers reproducible environments and simple tools to launch, inspect, and debug experiments. Relevant technologies
• Compute and orchestration: Linux, Docker, Kubernetes, Slurm, and Terraform.
• Execution and storage: Python or Go, Ray, S3-compatible object storage, and PostgreSQL.
• Observability: Prometheus, Grafana, OpenTelemetry, and NVIDIA DCGM; familiarity with GPU networking and NCCL is mandatory. About you You have strong systems fundamentals and can reason carefully about concurrency, resource contention, and failure recovery. You enjoy making complex infrastructure understandable and dependable for the people using it. We welcome infrastructure specialists and exceptionally fast-learning generalists. Prior ML infrastructure experience is valuable; evidence of building and operating reliable systems matters.
What we offer
• Competitive compensation and equity.
• A fast-paced environment combining frontier research with impactful real-world applications.
• Visa sponsorship and relocation support for candidates joining us in Paris or London.
• A flexible hybrid setup, with a preference for working together in person.
• Build and operate GPU clusters, job scheduling, networking, and storage for distributed ML workloads.
• Develop the execution platform for large numbers of concurrent environments, including sandbox isolation, resource limits, retries, and state recovery.
• Build reliable data and artifact pipelines for datasets, trajectories, model weights, and checkpoints.
• Own platform observability and recovery; improve capacity allocation, startup latency, and reliability through measurable changes.
• Give researchers reproducible environments and simple tools to launch, inspect, and debug experiments. Relevant technologies
• Compute and orchestration: Linux, Docker, Kubernetes, Slurm, and Terraform.
• Execution and storage: Python or Go, Ray, S3-compatible object storage, and PostgreSQL.
• Observability: Prometheus, Grafana, OpenTelemetry, and NVIDIA DCGM; familiarity with GPU networking and NCCL is mandatory. About you You have strong systems fundamentals and can reason carefully about concurrency, resource contention, and failure recovery. You enjoy making complex infrastructure understandable and dependable for the people using it. We welcome infrastructure specialists and exceptionally fast-learning generalists. Prior ML infrastructure experience is valuable; evidence of building and operating reliable systems matters.
What we offer
• Competitive compensation and equity.
• A fast-paced environment combining frontier research with impactful real-world applications.
• Visa sponsorship and relocation support for candidates joining us in Paris or London.
• A flexible hybrid setup, with a preference for working together in person.