Stateful Services Engineer
Enregistrez cette offre et organisez votre recherche
Créez un compte gratuit pour enregistrer des offres d'emploi, créer des alertes et revenir à cette liste depuis votre tableau de bord.
En continuant, vous acceptez nos Conditions d’utilisation & Politique de confidentialité.
Build the infrastructure behind the next generation of AI-powered customer communications
Cette opportunité vous plaît ? Postulez vite, car un grand nombre de candidatures est attendu. Faites défiler la page pour lire la description complète du poste.
the company is building an AI-first communication platform where AI and human agents, business workflows and communication channels work together seamlessly.
As we enter our next phase of growth, we are looking for a Stateful Services Engineer to help build and operate the infrastructure foundation behind the platform.
This is a unique opportunity to work on an infrastructure environment combining carrier-grade voice systems, on-premise data centers, distributed systems and AI workloads, while helping us build the reliability and operational maturity required to support 3x growth.
You will report to the newly forming Stateful Services Lead and work closely with Engineering, SRE and Infrastructure teams to improve the reliability, scalability and evolution of our stateful systems.
Your mission
Your goal is to build, operate and improve the company’s stateful services, with a strong focus on databases and data stores.
You will help ensure that our databases and stateful systems remain highly available, reliable and performant, while evolving our infrastructure so that maintenance, upgrades and scaling can be performed safely and with minimal to no downtime.
Working closely with the Stateful Services Lead, you will take ownership of specific technical projects and contribute to the evolution of our stateful infrastructure.
You will:
- Operate and maintain our stateful services, including PostgreSQL, ClickHouse, Redis, RabbitMQ and Kafka/Redpanda.
- Contribute to the design, evolution and scaling of our database infrastructure, with 50+ PostgreSQL clusters and increasingly complex data workloads.
- Work on projects aimed at making database maintenance, upgrades and migrations safer and less disruptive to our services.
- Improve high availability, replication, failover, backup/restore and disaster recovery across our stateful systems.
- Contribute to rebuilding our PostgreSQL clusters to support automatic failover between data centers and reduce the impact of maintenance operations.
- Help improving our queueing infrastructure, including projects such as quorum-based RabbitMQ.
- Help building and maintaining stable Redis infrastructure and improve the reliability of our stateful services.
- Investigate and solve complex performance, scalability and reliability issues across databases and distributed systems.
- Work closely with development teams to understand their needs, support their use of databases and help them build reliable solutions.
- Collaborate with SRE, networking and infrastructure teams to ensure that our stateful services operate reliably within our broader infrastructure.
- Contribute to monitoring, automation and operational practices that make our infrastructure easier and safer to operate.
- Participate in our on-call and incident management rotation, helping diagnose and resolve production issues when needed.
- Stay hands‑on and close to the technology: you will be expected to understand how our systems work, how they fail and how to improve them.
Who we are looking for
We’re looking for a technically strong and hands‑on Stateful Services Engineer who enjoys solving complex infrastructure problems and taking ownership of production systems where reliability really matters.
You have experience operating databases or stateful services in production and understand that running these systems is about much more than simply maintaining them. You know how to troubleshoot them, how they behave under load, how they fail, and how to make them more reliable as they scale.
You’ll work on specific technical projects defined with the Stateful Services Lead, while having the autonomy to investigate problems, propose solutions and contribute to technical decisions.
We care more about your technical depth, curiosity and ability to take ownership than the exact number of years of experience you have.
Must Have
- Strong hands‑on experience as a DBA, Database Engineer, SRE or similar, with solid experience in PostgreSQL.
- Experience operating and maintaining production databases, including upgrades, migrations, replication, backups, recovery, monitoring and performance troubleshooting.
- Good understanding of high availability and reliability for stateful systems: failover, redundancy, replication, disaster recovery and minimising downtime during maintenance.
- Good understanding of distributed systems, including replic