Stateful Services Team Lead

Il y a 3 jours

Paris, Île-de-France Lever, Inc. Temps plein 100 000 € - 150 000 €/an

Build the infrastructure behind the next generation of AI-powered customer communications

Diabolocom is building an AI-first communication platform where AI and human agents, business workflows and communication channels work together seamlessly.

As we enter our next phase of growth, we are looking for a Stateful services Lead to scale the infrastructure foundation behind the platform.

This is a unique opportunity to take part in leading an infrastructure organization combining carrier-grade voice systems, on-premise data centers, distributed systems and AI workloads, while building the operational maturity required to support 3x growth.

You will report directly to Sergei, our CTO and lead the growth of stateful services scope — owning operational excellence, team development and execution of infrastructure initiatives, while actively contributing to technical strategy and architecture decisions.

Your mission

Your goal is to build and lead the team responsible for the reliability, scalability and evolution of Diabolocom’s stateful services — with a strong focus on databases and data stores.

You will ensure that our databases and stateful systems remain highly available, reliable and performant, while evolving our infrastructure so that maintenance, upgrades and scaling can be performed safely and with minimal to no downtime.

Starting with a small squad, you will combine hands-on technical expertise, system design and team leadership to build the foundations of a scalable Stateful Services infrastructure and the team.

You will:

  • Own the reliability and day-to-day operations of our stateful services, including PostgreSQL, ClickHouse, Redis, RabbitMQ and Kafka/Redpanda.
  • Lead the design, evolution and scaling of our database infrastructure, with 50+ PostgreSQL instances and increasingly complex data workloads.
  • Rework and expand our database clusters to enable maintenance, upgrades and migrations without service disruption.
  • Improve high availability, replication, failover, backup/restore and disaster recovery practices across our stateful systems.
  • Work closely with development teams to understand their needs, support their use of databases and help them design reliable data architectures.
  • Collaborate with SRE, networking and infrastructure teams to ensure our stateful services operate reliably within our broad infrastructure.
  • Investigate and solve complex performance, scalability and reliability issues across databases and distributed systems.
  • Establish the processes, standards and operational practices needed to run stateful services reliably at scale.
  • Stay hands-on and close to the technology: this is a leadership role where technical depth and ownership matter more than hierarchy.

Who we are looking for

We’re looking for a technically strong and hands-on Stateful Services Lead who enjoys solving complex infrastructure problems and taking ownership of systems where reliability really matters.

You have deep experience operating databases and stateful services in production, ideally in challenging, high-availability environments. You are comfortable going beyond simply maintaining systems: you understand how they work, how they fail, how to scale them, and how to design them to remain reliable as the company grows.

You’ll be the technical leader of a new squad, combining hands-on expertise with the ability to build strong engineering practices and grow a team over time.

Must Have

  • Strong hands-on experience as a DBA, Database Engineer or similar, with solid expertise in PostgreSQL.
  • Production experience with ClickHouse or another complex database/data platform.
  • Experience operating and maintaining databases at scale, including upgrades, migrations, replication, backups, recovery and performance tuning.
  • Strong understanding of high availability and reliability for stateful systems: failover, redundancy, replication, disaster recovery and minimizing or eliminating downtime during maintenance.
  • Good understanding of distributed systems, including consistency, fault tolerance, replication and the challenges of operating stateful services at scale.
  • Experience working with on-premise infrastructure and bare-metal environments. You are comfortable operating outside of a public-cloud-only environment.
  • Strong system design skills and the ability to make pragmatic technical decisions around scalability, reliability and performance.
  • Experience collaborating closely with software development teams, understanding their requirements and helping them build reliable solutions around databases and stat