Inference Systems Engineer: Scale

Il y a 5 jours

Paris, Île-de-France Mistral Temps plein 120 000 € - 180 000 €/an

Mistral is seeking an experienced engineer to own and evolve the inference foundation that powers production LLM serving and frontier model training. You will optimize the core stack, manage releases, and push upstream improvements to high-performance serving systems.

The role blends engine, platform development, and capacity engineering in a hybrid setup. You will work across CUDA, NCCL, and distributed architectures to deliver low-latency, scalable inference.