FOUNDING DATA ENGINEER
Enregistrez cette offre et organisez votre recherche
Créez un compte gratuit pour enregistrer des offres d'emploi, créer des alertes et revenir à cette liste depuis votre tableau de bord.
About
Our vision
The promise of AI agents is still unmet in complex, real-world business environments. Our mission is to make it possible for agents to reliably accomplish concrete and useful tasks.
That’s why we’re building Alexandria, the world’s largest library of real-world business workflow datasets.
Over the next 6 months, we’ll acquire and anonymize 300+ full data ecosystems from real-world companies. This will allow us to become the leader of real business workflow data.
This summer we need to build our elite squad to move extremely fast and confirm our leading position on company data.
Our values
- Trust, the foundation of every relationship, internal and external. We extend it by default and value the ownership that comes with it.
- Ambition, we commit, we move fast, with tenacity and efficiency.
- Collective, team first, low ego.
- Kindness, at the heart of every interaction.
We’re building a solid, aligned, and motivated team.
Our traction speaks for itself: we’ve signed contracts with some of the largest AI labs in the world, alongside recent partnerships with several smaller labs, all while being 100% bootstrapped. This summer, we’re joining Y Combinator, which is accelerating everything.
Job Description
We’re ooak data
The promise of AI agents is still unmet in complex, real-world business environments. Our mission is to make it possible for agents to reliably accomplish concrete and useful tasks.
That’s why we’re building Alexandria, the world’s largest library of real-world business workflow datasets.
Over the next 6 months, we’ll acquire and anonymize 300+ full data ecosystems from real-world companies. This will allow us to become the leader of real business workflow data.
This summer we need to build our elite squad to move extremely fast and confirm our leading position on company data.
Our values
- Trust, the foundation of every relationship, internal and external. We extend it by default and value the ownership that comes with it.
- Ambition, we commit, we move fast, with tenacity and efficiency.
- Collective, team first, low ego.
- Kindness, at the heart of every interaction.
We’re building a solid, aligned, and motivated team.
The job offer
Paris (city center), on‑site with 1 to 2 days WFH
Start: as soon as you are available
Contract: full‑time (CDI), open to freelancing for the summer
Team
- 3 co‑founders (Pierre-Louis, CEO; Grégoire, CPO & COO; Thomas, CTO)
- 6 FTEs: 3 ML engineer and 3 ops + 5 freelances
Your missions
As a Data Engineer and a founding member of the tech team, you’ll play a central role in how we design, structure, and scale the data foundations behind our products. You’ll report directly to Thomas (CTO & co‑founder).
- Design, build, and scale our data ingestion, transformation, and structuring pipelines at scale.
- Contribute across every stage of our data platform:
- The anonymization pipeline (combining algorithmic approaches with human verification) to build digital twins of companies.
- The multi‑modal data enrichment & curation pipeline.
- The datasets feeding our reinforcement learning (RL) environments.
- Ensure data quality, lineage, versioning, and observability across the whole stack.
- Define our technical standards and take an active part in the structuring architecture decisions (tech choices, design patterns, scalability, DataOps).
- Work closely with the product, ML, and business teams to scope and ship new data projects.
- Mentor and grow the future members of the tech team, building a strong data engineering culture at Ooak Data.
Who we’re looking for
- 5-10 years of experience in data engineering or data‑intensive software development.
- Strong coding experience in Python and SQL, with a plus for Rust or TypeScript.
- Expertise building ETL/ELT pipelines and orchestration (Airflow, Dagster, dbt, or equivalent).
- Experience with data warehouses and lakes (BigQuery, Snowflake, or equivalent) and stre