Bioinformatics Data Engineer

Il y a 2 jours

Paris, Nouvelle-Aquitaine, France Rivercell Temps plein 70 000 € - 110 000 €/an

We are looking for a Bioinformatics Data Engineer to join our team and take charge of all the data used to train and evaluate our AI virtual cell models: processing, harmonization, and versioning.

In your daily work, you will write and run the bioinformatics pipelines that turn raw sequencing and imaging data into versioned, analysis-ready datasets.

You will work hand-in-hand with our ML scientists, biologists, chemists and engineers. You will be an integral part of an ambitious and exciting French startup project from its inception.

  • Build our omics pipelines, starting with single-cell, from FASTQ to analysis-ready matrices: alignment and quantification, CRISPR guide assignment, cell calling, ambient RNA and doublet handling, and per-run quality control

  • Build the imaging pipelines for live-cell brightfield and fluorescence time-lapse data: illumination correction, segmentation, tracking, and feature and embedding extraction

  • Internalize and harmonize relevant publicly available datasets with Rivercell's own data

  • Set up dataset versioning, lineage, and release management, so that every dataset behind a model, a paper, or a benchmark can be traced to its raw data and pipeline version and every experiment is reproducible

  • Build the tooling that generates and audits contamination-controlled train/test splits for our benchmarks

  • Run the data infrastructure: cloud storage and compute for tens of terabytes, cost control, and data loading fast enough to keep GPUs busy

  • Build the data side of our closed lab-in-a-loop, in which models propose experiments, the lab runs them, and the results return as retraining-ready data

  • Opportunity to work on a cutting-edge project in the field of biotech

  • Opportunity to join an ambitious team and impactful startup among the first employees

  • Competitive salary and equity package

  • Exciting opportunities for personal and professional growth within the team

You have an MSc or PhD in bioinformatics, computational biology, computer science, or a related field You have 3+ years of experience building and running bioinformatics pipelines in production, in industry, a core facility, or a large consortium You have processed single-cell RNA-seq data from raw reads (Cell Ranger, STARsolo, kallisto/bustools, alevin-fry, or equivalent) and work comfortably in the AnnData / scverse ecosystem You write strong Python and apply sound engineering practice: workflow managers, containers, tests, CI, and documentation You have experience with data versioning on large datasets, with harmonizing heterogeneous datasets, and with cloud (AWS or GCP) or HPC infrastructure You have experience with image processing at scale, ideally microscopy: segmentation, tracking, and feature or embedding extraction (Cellpose, StarDist, CellProfiler, or deep-learning models) You have experience with high-performance data management for model training: array and columnar formats (Zarr, TileDB-SOMA, Parquet, OME-Zarr), sharding and streaming from object storage, and PyTorch data loaders that keep GPUs busy An autonomous, committed, enthusiastic, fun to work with person with a "can-do", creative, and challenging mindset You have excellent communication skills and are fluent in English. Professional proficiency in French is helpful but not required.