Remote AI Benchmark Engineer
Enregistrez cette offre et organisez votre recherche
Créez un compte gratuit pour enregistrer des offres d'emploi, créer des alertes et revenir à cette liste depuis votre tableau de bord.
LILT is seeking experienced native-speaking software engineers to design, build, and validate multilingual benchmarks for Terminal-Bench tasks. You will create high-signal tasks and assets in your native language to measure robustness without English translation crutches.
This is a remote, freelance opportunity. You’ll work with a global team, set your own schedule, and receive competitive rates with prompt payments. No fixed hours; project-based collaboration across languages and locales.