Confirmed Site Reliability Engineer
Enregistrez cette offre et organisez votre recherche
Créez un compte gratuit pour enregistrer des offres d'emploi, créer des alertes et revenir à cette liste depuis votre tableau de bord.
About
Swan is Europe’s embedded banking specialist. We empower software companies to embed banking features like accounts, cards, and payments directly into their products, under their own brand.Swan processes over €2.5 billion in monthly transactions for more than 150 companies - like Pennylane, Indy, Agicap, Libeo, and Lucca.
Founded in 2019, the company has received growth capital from leading investors such as Lakestar, Accel, Creandum, Bpifrance and Eight Roads. Swan is a principal member of Mastercard and a licensed financial institution, regulated by the French banking authority (ACPR).
Our mission
Banking belongs in business software
Many software companies already serve small businesses incredibly well: helping them send invoices, run payroll, manage inventory, and more. They’re on a mission to become the central hub for managing every aspect of business life.
But when it comes to financial workflows, there’s still a gap. Too many critical tasks like managing cash flow, tracking payments, or reconciling accounts happen outside the software, across spreadsheets, email threads, banking portals.
It’s a missed opportunity. Business software shouldn’t just record financial activity — it should run it.
To learn more about us: About Swan, Our story.
Job description
Working within Swan’s Core Infrastructure team, you will help ensure the reliability, scalability, security, and performance of the platforms that support our financial services. You will take ownership of well-scoped services and operational incidents, improve observability, automate repetitive tasks, and collaborate closely with development, product, and security teams.
This is an independent engineering role for someone who has developed solid operational foundations and is ready to take greater ownership of production systems. You will contribute to incident response, infrastructure improvements, service design reviews, and the continuous improvement of our reliability practices.
Main responsibilities
On a daily basis, you will:
Act as a primary responder for well-understood production incidents and participate independently in the on-call rotation.
Assess the impact of incidents, including transaction volume affected, potential revenue impact, and implications for data integrity.
Investigate operational issues using logs, metrics, dashboards, and distributed tracing, then document clear incident updates for stakeholders.
Contribute to postmortems, update runbooks, identify recurring incident patterns, and suggest preventive measures.
Create and maintain dashboards, alerts, and basic service-level indicators for the services you support.
Tune alert thresholds to reduce noise and improve the quality of operational signals.
Participate in system design reviews, with a particular focus on reliability, operability, failure modes, and production readiness.
Implement reliability improvements such as health checks, retries with exponential backoff, circuit breakers, and appropriate monitoring.
Manage cloud resources and contribute to Infrastructure as Code using tools such as Terraform.
Write automation scripts and small internal tools in Bash, Python, or Go to reduce manual toil and improve operational efficiency.
Contribute to CI/CD pipelines and automate routine maintenance tasks such as backup verification, certificate renewal, and log management.
Participate in infrastructure code reviews and help maintain high standards for safe, repeatable changes.
Support security and compliance activities, including PCI DSS controls, ISO 27001 initiatives, security remediation, data classification, and encryption requirements.
Monitor resource utilisation, provide basic capacity forecasts, and implement practical cost optimisation measures such as rightsizing resources and removing unused infrastructure.
Collaborate with development, product, and security teams to improve the resilience and operability of services.
Provide clear handovers, maintain high-quality documentation, and communicate technical topics effectively to both technical and non-technical stakeholders.
Use approved AI tools responsibly to support tasks such as code generation, documentation, and log analysis, while validating outputs and protecting sensitive information.
Your team
Core Infrastructure is responsible for building and operating the foundations that enable Swan’s products to remain reliable as the business grows. We work closely with development and other technical teams to improve system resilience, operational efficiency, and customer experience.
We value ownership,