Benchmark Engineer, AI Safety

Il y a 1 jour

Paris, Nouvelle-Aquitaine, France Aisafety Temps plein 80 000 € - 110 000 € Contrat

White Circle is seeking a research engineer to build and maintain an internal benchmark suite for single- and multi-turn content and agentic safety. You’ll work with the team on projects studying agent behaviours in the wild and help evolve evals for new features and data.

The role involves Python development, production-grade benchmarks, and collaboration with the product team to ensure robust evaluation of White Circle’s AI safety stack.