Junior Data Engineer
Enregistrez cette offre et organisez votre recherche
Créez un compte gratuit pour enregistrer des offres d'emploi, créer des alertes et revenir à cette liste depuis votre tableau de bord.
En continuant, vous acceptez nos Conditions d’utilisation & Politique de confidentialité.
About GitGuardian
GitGuardian is a global cybersecurity scale-up. The company is based in Paris, New-York City, Boston.
Among our early investors who saw our market value proposition, are the co-founder of GitHub, Scott Chacon, along with Solomon Hykes, Docker's co-founder. American and European top-tier VC firms have also invested in GitGuardian who raised its series C early 2026.
GitGuardian leads the way in securing the credential layer, helping organizations protect the secrets that let code, machines, and AI agents access systems with broad detection, identity context, and remediation at scale. Already trusted by 600K+ developers worldwide and by 50 of the fortune 500
About your team and your mission
You will join the Data & AI Engineering team, whose mission is to build the data foundation and the internal AI-engineering capability that let GitGuardian run as an AI-native company.
The team works across three pillars. On the data side, we run the data platform, the business data models, and the ingestion chain behind every team's reporting and our product analytics, with the aim of data that is used, trusted, and increasingly near real-time. On internal AI tooling, we build the shared skills, internal MCPs, and the Analytics Agent, along with the governance and evaluation standards that make these tools safe to build on. And on apps, we maintain the internal app framework that lets anyone in the company ship an application to software engineering standards, with SSO, CI/CD, and managed hosting built in.
Data is already the backbone of the company's reporting and decision process. AI tooling and internal apps are newer ground, with real freedom to shape how all business stakeholders use data on the day-to-day.
Key challenges :
Build the data foundations for AI agents. AI agents are becoming the main way people query data. You'll help design the next generation of our warehouse and data models so that agents can understand them, trust them and answer questions correctly: clear semantics, well-documented metrics and structures built for machines as much as for humans.
Expand our data coverage to sensitive domains. Finance and HR data are joining our scope. You'll help bring them into the warehouse with the right level of rigor, access control and confidentiality.
Strengthen the existing data domains. Marketing, Sales and Product Analytics already rely on us. We want to make their data more complete, reliable and easy to use, on Snowflake today and ClickHouse soon.
Make data part of everyone's daily work. Our goal is for every team at GitGuardian to use Data & AI tools every day. That depends on trustworthy data models and clearly defined business metrics.
Keep the platform running smoothly while it grows. Day-to-day data requests and run topics matter. Handling them well lets the whole squad keep delivering on the larger projects of a 12–18 month roadmap.
Your responsibilities :
Write data models as code daily (SQL, Python with Snowpark and PySpark) to deliver new data for business insights.
Own your topics end to end, from ingestion to consumption by business users.
Work with internal stakeholders (Sales, Marketing, Finance, HR, Product) to understand their needs and turn them into well-defined business metrics.
Build and orchestrate pipelines with Dagster, and deploy them with Docker and Terraform on AWS.
Take part in the run: monitor pipelines, investigate data issues, and answer data requests from across the company.
Keep the code base healthy through code reviews, tests and documentation.
Work closely with the squad's AI Engineers so that our data powers the company's internal AI tools.
Technical environment
Coding Languages: Python (Snowpark & PySpark), SQL
Analytics databases: Snowflake, ClickHouse
Visualisation: Metabase
Orchestration: Dagster
Deployment: Docker, Terraform
Cloud Provider: AWS
VCS: GitLab
CI/CD: ArgoCD
About you
If you think you match at least 70% of these criteria, please apply
Here's what we consider essential for success in this role:
A first experience in data engineering, such as an internship, apprenticeship or up to about 2 years in a role.
Strong SQL and good Python skills.
An understanding of data modeling concepts (dimensional modeling, fact and dimension tables, business metrics).
The ability to find your way in an existing, complex code base and learn from it quickly.
Comfort talking with non-technical stakeholde