Senior Site Reliability Engineer — Token Factory

Il y a 2 jours

France, Auvergne-Rhône-Alpes Jobgether Temps plein 110 000 € - 135 000 € Contrat

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Site Reliability Engineer — Token Factory (Inference Platform) based in France.

This is a senior engineering role focused on the reliability, performance, and observability of a large-scale AI inference platform.

You will help operate infrastructure serving foundation models across text, vision, audio, and emerging multimodal workloads.

The role combines Kubernetes, infrastructure-as-code, observability, automation, and production incident management at significant scale.

You will optimize GPU-heavy workloads, strengthen resilience, and ensure high-throughput APIs meet demanding reliability and cost targets.

You will work closely with software engineers and infrastructure teams to build self-healing systems and robust operational processes.

The environment is fast-moving, highly technical, international, and focused on solving complex infrastructure challenges for the AI ecosystem.

This is an opportunity to have a direct impact on the infrastructure powering next-generation AI applications.

Accountabilities

  • Own the reliability, performance, and observability of the inference platform and its supporting infrastructure.
  • Design, implement, and continuously improve telemetry pipelines covering metrics, logs, and traces.
  • Build monitoring and observability solutions capable of processing large volumes of production signals and converting them into actionable insights.
  • Configure and optimize Kubernetes infrastructure for high availability, scalability, and efficient resource utilization.
  • Tune Kubernetes autoscaling mechanisms to improve the efficiency and utilization of GPU resources.
  • Develop and maintain Terraform modules and infrastructure-as-code patterns that embed resilience and reliability into new clusters and services.
  • Design and improve request-routing, retry, and failure-handling mechanisms to minimize the impact of transient infrastructure or service failures.
  • Develop automation and operational tooling to detect, isolate, and remediate incidents quickly.
  • Create, maintain, and improve runbooks for incident response and operational procedures.
  • Participate in production incident management, troubleshooting issues and restoring services within demanding reliability objectives.
  • Lead or contribute to post-mortem processes and implement corrective actions to prevent recurring incidents.
  • Define and improve reliability practices for high-throughput APIs, including alerting strategies and Service Level Objectives (SLOs).
  • Investigate distributed-system failures and performance issues across infrastructure and application layers.
  • Optimize systems from the kernel and infrastructure layer through to the application layer.
  • Support and improve the operation of GPU-intensive inference workloads and accelerator-based infrastructure.
  • Contribute to scaling the inference platform while balancing performance, reliability, and infrastructure costs.
  • Collaborate closely with software engineers to incorporate reliability and operational excellence into product and platform development.
  • Promote automation, self-healing capabilities, and engineering practices that reduce operational overhead and improve system resilience.

Requirements:

  • Significant experience in Site Reliability Engineering, Production Engineering, DevOps, or a closely related infrastructure discipline.
  • Deep practical knowledge of Kubernetes in production environments.
  • Strong experience with Prometheus and Grafana for monitoring, metrics, dashboards, and observability.
  • Advanced experience with Terraform and infrastructure-as-code practices.
  • Strong scripting and automation skills using Python and/or Bash.
  • Solid understanding of distributed systems and the ways production backends can fail under real-world conditions.
  • Experience designing effective alerts, monitoring strategies, and SLOs for high-throughput services or APIs.
  • Strong troubleshooting and debugging skills across infrastructure, networking, operating systems, and application layers.
  • Experience designing systems for high availability, resilience, scalability, and graceful failure recovery.
  • Hands-on experience with GPU-heavy workloads or accelerator-based infrastructure is highly valuable.
  • Familiarity with GPU inference technologies such as vLLM, Triton, Ray, or comparable accelerator and model-serving stacks.
  • Experience with MLOps, model hosting, AI infrastructure, or machine-learning platforms