Remote LLM Evaluation Scientist

Il y a 4 jours

Paris, Île-de-France Anyone AI Temps plein 90 000 € - 140 000 € Contrat

Anyone AI Labs is seeking a dedicated Research Scientist to advance how frontier models are evaluated. You will design benchmarks, validate methodologies, and build evaluation packages across reasoning, coding, agents, and multi-modal systems.

You will lead expert recruitment, coordinate with labs and CEOs, and push for public benchmarks and papers at venues like NeurIPS and ICLR. Fluent English required; Spanish is a plus.