Multimodal ML Engineer
Il y a 6 heures
Paris, Ile-de-France
Sokoni Kwetu
Temps plein
Gratuit avec email ou Google
Enregistrez cette offre et organisez votre recherche
Créez un compte gratuit pour enregistrer des offres d'emploi, créer des alertes et revenir à cette liste depuis votre tableau de bord.
Gratuit avec email ou Google
En continuant, vous acceptez nos Conditions d’utilisation & Politique de confidentialité.
We're looking for a Multimodal ML Engineer to join White Circle, an AI Safety company building the safety, reliability, and optimization layer for AI systems through natural-language policies it automatically tests, enforces, and improves at scale. Backed by $70M (Series A) from top funds and senior leaders at OpenAI, Anthropic, HuggingFace, Mistral, DeepMind, and others, White Circle processes 100M+ API calls monthly and fine-tunes and trains its own LLMs to run faster and cheaper than open or proprietary models.
You willTrain and fine-tune large-scale multimodal models (vision-language, audio, speech, video) from scratch and from pretrained checkpoints.
Design experiments, build multimodal data pipelines, and train MoE architectures.
Build alignment pipelines (SFT, DPO, GRPO), optimize for production (quantization, distillation, streaming), and deploy end-to-end.
Define evaluation metrics that actually matter for the product.
Requirements3+ years training large-scale multimodal models.
Strong PyTorch and distributed training experience (DeepSpeed, FSDP).
Deep familiarity with multimodal architectures – LLaVA, Qwen-VL, InternVL, Audio Flamingo, Whisper, HuBERT, Conformer or similar.
Hands-on RLHF/alignment across modalities (GRPO, DPO, reward modeling).
Both audio and video experience required – sequence modeling for each, plus large-scale dataset curation and production inference optimization.
Relocation to Paris or London (hybrid) required.
BonusAudio signal processing fundamentals – spectrograms, mel features, noise reduction.
MoE architecture experience.
We offer
Competitive salary + equity.
Official employment, visa and relocation help.
Find more English Speaking Jobs in France on Arbeitnow
You willTrain and fine-tune large-scale multimodal models (vision-language, audio, speech, video) from scratch and from pretrained checkpoints.
Design experiments, build multimodal data pipelines, and train MoE architectures.
Build alignment pipelines (SFT, DPO, GRPO), optimize for production (quantization, distillation, streaming), and deploy end-to-end.
Define evaluation metrics that actually matter for the product.
Requirements3+ years training large-scale multimodal models.
Strong PyTorch and distributed training experience (DeepSpeed, FSDP).
Deep familiarity with multimodal architectures – LLaVA, Qwen-VL, InternVL, Audio Flamingo, Whisper, HuBERT, Conformer or similar.
Hands-on RLHF/alignment across modalities (GRPO, DPO, reward modeling).
Both audio and video experience required – sequence modeling for each, plus large-scale dataset curation and production inference optimization.
Relocation to Paris or London (hybrid) required.
BonusAudio signal processing fundamentals – spectrograms, mel features, noise reduction.
MoE architecture experience.
We offer
Competitive salary + equity.
Official employment, visa and relocation help.
Find more English Speaking Jobs in France on Arbeitnow