Member of Technical Staff — Platform Engineering
Il y a 6 jours
Paris, Ile-de-France
Remanence
Temps plein
Gratuit avec email ou Google
Enregistrez cette offre et organisez votre recherche
Créez un compte gratuit pour enregistrer des offres d'emploi, créer des alertes et revenir à cette liste depuis votre tableau de bord.
Gratuit avec email ou Google
En continuant, vous acceptez nos Conditions d’utilisation & Politique de confidentialité.
Member of Technical Staff — Platform Engineering
Remanence is pioneering the next era of enterprise AI by building intelligent systems that learn continuously from real-world execution. We transform complex enterprise workflows and business context into dynamic, interactive environments where AI agents can safely learn, adapt, and improve. By combining high-fidelity simulation environments with state-of-the-art training loops, we build specialized models that solve long-horizon, complex tasks with unmatched reliability. We're building the most talent-dense AI team in Europe to make this happen.
You’ll make the Remanence platform deployable wherever our customers need it to run: in their own cloud, in their data center, or eventually in fully air-gapped environments.
The work sits at the intersection of platform engineering, deployment infrastructure, and distributed systems. You’ll turn our environment, training, inference, and evaluation systems into a product that can be installed, upgraded, operated, and debugged reliably on infrastructure we don’t control.
What you'll work on
• Package the full Remanence stack for deployment on-premise and in customer-owned clouds (BYOC), across AWS, GCP, Azure, and bare-metal infrastructure. This includes our environment, training, inference, and evaluation systems.
• Build and maintain the Infrastructure-as-Code and deployment layer around the platform using tools such as Terraform, Helm, and Kubernetes operators. Aim for a reproducible path from an empty account or cluster to a fully working deployment.
• Design for constrained and security-sensitive environments: private registries, restricted networking, customer identity providers, secrets management, enterprise security and compliance requirements, and eventually fully air-gapped installations.
• Own CI/CD, release engineering, and upgrade paths. Make installations, upgrades, migrations, and rollbacks predictable, automated, and continuously tested across supported environments.
• Build observability, diagnostics, and support tooling that makes deployments operable without direct access to customer infrastructure, including health checks, support bundles, logs, metrics, and distributed tracing.
• Work closely with the Infrastructure and Research teams on the underlying platform: sandboxed execution, distributed job orchestration, GPU scheduling, checkpoint storage, and resource isolation.
• Occasionally work alongside customer infrastructure, platform, and security teams during deployments, then turn what you learn into abstractions and tooling that make the next deployment easier. Relevant technologies
• Deployment and IaC: Terraform, Helm, Kubernetes (EKS, GKE, AKS, OpenShift, RKE2/k3s), Argo CD or Flux.
• Infrastructure: Linux, Docker, networking (VPC, DNS, TLS, load balancing), IAM and secrets management (Vault, KMS).
• Platform: Python or Go, Ray, Slurm, S3-compatible object storage (e.g. MinIO), PostgreSQL, NVIDIA GPU Operator.
• CI/CD and observability: GitHub Actions, container registries (e.g. Harbor), Prometheus, Grafana, OpenTelemetry.
• Hands-on production experience with Kubernetes and Terraform is required. About you You have strong systems fundamentals and have shipped software that needs to run reliably outside infrastructure you fully control. You think naturally about reproducibility, failure modes, isolation, observability, backward compatibility, and upgrade paths. You’re comfortable debugging across the entire stack, whether that means Terraform state, networking, IAM, Kubernetes scheduling, storage, or a GPU workload that refuses to start. You’re equally comfortable building the platform and working directly with customer infrastructure or security teams when necessary. More importantly, you know how to turn what would otherwise become a collection of one-off deployment scripts into a clean, repeatable product. We welcome both infrastructure specialists and exceptionally fast-learning generalists. Experience deploying software into enterprise, regulated, or security-sensitive environments is valuable, but what matters most is evidence that you can build and operate reliable systems in production.
What we offer
• Competitive compensation and equity.
• A fast-paced environment combining frontier research with impactful real-world applications.
• Visa sponsorship and relocation support for candidates joining us in Paris or London.
• A flexible hybrid setup, with a preference for working together in person.
• Package the full Remanence stack for deployment on-premise and in customer-owned clouds (BYOC), across AWS, GCP, Azure, and bare-metal infrastructure. This includes our environment, training, inference, and evaluation systems.
• Build and maintain the Infrastructure-as-Code and deployment layer around the platform using tools such as Terraform, Helm, and Kubernetes operators. Aim for a reproducible path from an empty account or cluster to a fully working deployment.
• Design for constrained and security-sensitive environments: private registries, restricted networking, customer identity providers, secrets management, enterprise security and compliance requirements, and eventually fully air-gapped installations.
• Own CI/CD, release engineering, and upgrade paths. Make installations, upgrades, migrations, and rollbacks predictable, automated, and continuously tested across supported environments.
• Build observability, diagnostics, and support tooling that makes deployments operable without direct access to customer infrastructure, including health checks, support bundles, logs, metrics, and distributed tracing.
• Work closely with the Infrastructure and Research teams on the underlying platform: sandboxed execution, distributed job orchestration, GPU scheduling, checkpoint storage, and resource isolation.
• Occasionally work alongside customer infrastructure, platform, and security teams during deployments, then turn what you learn into abstractions and tooling that make the next deployment easier. Relevant technologies
• Deployment and IaC: Terraform, Helm, Kubernetes (EKS, GKE, AKS, OpenShift, RKE2/k3s), Argo CD or Flux.
• Infrastructure: Linux, Docker, networking (VPC, DNS, TLS, load balancing), IAM and secrets management (Vault, KMS).
• Platform: Python or Go, Ray, Slurm, S3-compatible object storage (e.g. MinIO), PostgreSQL, NVIDIA GPU Operator.
• CI/CD and observability: GitHub Actions, container registries (e.g. Harbor), Prometheus, Grafana, OpenTelemetry.
• Hands-on production experience with Kubernetes and Terraform is required. About you You have strong systems fundamentals and have shipped software that needs to run reliably outside infrastructure you fully control. You think naturally about reproducibility, failure modes, isolation, observability, backward compatibility, and upgrade paths. You’re comfortable debugging across the entire stack, whether that means Terraform state, networking, IAM, Kubernetes scheduling, storage, or a GPU workload that refuses to start. You’re equally comfortable building the platform and working directly with customer infrastructure or security teams when necessary. More importantly, you know how to turn what would otherwise become a collection of one-off deployment scripts into a clean, repeatable product. We welcome both infrastructure specialists and exceptionally fast-learning generalists. Experience deploying software into enterprise, regulated, or security-sensitive environments is valuable, but what matters most is evidence that you can build and operate reliable systems in production.
What we offer
• Competitive compensation and equity.
• A fast-paced environment combining frontier research with impactful real-world applications.
• Visa sponsorship and relocation support for candidates joining us in Paris or London.
• A flexible hybrid setup, with a preference for working together in person.