Benchmark Engineer, AI Safety

Il y a 5 jours

Paris, Nouvelle-Aquitaine, France Aisafety Temps plein

White Circle is seeking a research engineer to build and maintain an internal benchmark suite for single- and multi-turn content and agentic safety. You’ll work with the team on projects studying agent behaviours in the wild and help evolve evals for new features and data.

The role involves Python development, production-grade benchmarks, and collaboration with the product team to ensure robust evaluation of White Circle’s AI safety stack.

#J-18808-Ljbffr