Member of Technical Staff — Post-training
Il y a 1 semaine
Paris, Ile-de-France
Remanence
Temps plein
Gratuit avec email ou Google
Enregistrez cette offre et organisez votre recherche
Créez un compte gratuit pour enregistrer des offres d'emploi, créer des alertes et revenir à cette liste depuis votre tableau de bord.
Gratuit avec email ou Google
En continuant, vous acceptez nos Conditions d’utilisation & Politique de confidentialité.
Member of Technical Staff — Post-training / RL
Remanence is pioneering the next era of enterprise AI by building intelligent systems that learn continuously from real-world execution. We transform complex enterprise workflows and business context into dynamic, interactive environments where AI agents can safely learn, adapt, and improve. By combining high-fidelity simulation environments with state-of-the-art training loops, we build specialized models that solve long-horizon, complex tasks with unmatched reliability. We’re building the most talent-dense AI team in Europe to make this happen.
You’ll develop post-training methods and own the performance of the systems that run them. The work spans learning algorithms, distributed training, and GPU performance, with responsibility for both model improvements and efficient experimentation.
What you’ll work on
• Implement and improve supervised fine-tuning, preference optimization, and RL methods such as PPO, GRPO and SDPO.
• Own the training loop from rollout generation and reward computation through policy updates, weight synchronization, and checkpointing.
• Improve throughput, GPU utilization, and memory efficiency through batching, parallelism, communication overlap, and asynchronous execution.
• Profile bottlenecks and investigate numerical instability, stale-policy effects, and training-inference discrepancies. Verify that optimizations preserve learning behavior.
• Occasionally deploy models in client environments and hill-climb alongside their teams: analyze real failures, refine rewards and training data, and iterate on task success, latency, and inference cost. Relevant technologies
• Training: PyTorch, Miles, NVIDIA Molt, verl, FSDP, and Megatron-LM
• Rollout generation: vLLM, SGLang, Ray and asynchronous execution frameworks.
• Performance/Kernels: NCCL, CUDA, Triton and or CuTE
• Profiling: PyTorch Profiler, NVIDIA Nsight Systems About you You bring strong programming skills, quantitative reasoning, and an interest in how learning algorithms interact with the systems underneath them. You can move from an experiment to a profiler trace and investigate what the evidence shows. We welcome both experienced specialists and generalists who learn exceptionally quickly. Experience with every method or framework above is not required.
What we offer
• Competitive compensation and equity.
• A fast-paced environment combining frontier research with impactful real-world applications.
• Visa sponsorship and relocation support for candidates joining us in Paris or London.
• A flexible hybrid setup, with a preference for working together in person.
• Implement and improve supervised fine-tuning, preference optimization, and RL methods such as PPO, GRPO and SDPO.
• Own the training loop from rollout generation and reward computation through policy updates, weight synchronization, and checkpointing.
• Improve throughput, GPU utilization, and memory efficiency through batching, parallelism, communication overlap, and asynchronous execution.
• Profile bottlenecks and investigate numerical instability, stale-policy effects, and training-inference discrepancies. Verify that optimizations preserve learning behavior.
• Occasionally deploy models in client environments and hill-climb alongside their teams: analyze real failures, refine rewards and training data, and iterate on task success, latency, and inference cost. Relevant technologies
• Training: PyTorch, Miles, NVIDIA Molt, verl, FSDP, and Megatron-LM
• Rollout generation: vLLM, SGLang, Ray and asynchronous execution frameworks.
• Performance/Kernels: NCCL, CUDA, Triton and or CuTE
• Profiling: PyTorch Profiler, NVIDIA Nsight Systems About you You bring strong programming skills, quantitative reasoning, and an interest in how learning algorithms interact with the systems underneath them. You can move from an experiment to a profiler trace and investigate what the evidence shows. We welcome both experienced specialists and generalists who learn exceptionally quickly. Experience with every method or framework above is not required.
What we offer
• Competitive compensation and equity.
• A fast-paced environment combining frontier research with impactful real-world applications.
• Visa sponsorship and relocation support for candidates joining us in Paris or London.
• A flexible hybrid setup, with a preference for working together in person.