Duplo
Banking, Finance & Insurance
Senior Site Reliability Engineer (SRE)
Share this role
About this role
Duplo—a fast-growing Banking-as-a-Service (BaaS) and business payments platform powering B2B financial services—is recruiting a Senior Site Reliability Engineer (SRE) for its office in Lagos State (Hybrid Work Model). This senior engineering leadership role focuses on designing high-availability infrastructure, defining reliability standards, leading incident response, optimizing PostgreSQL/Redis operations, and ensuring maximum resilience for mission-critical money movement systems.
Key Responsibilities-
Infrastructure Architecture & Availability: Lead technical design for high-complexity initiatives, establishing service boundaries, rollout safety, and resilience strategies across core backend payment rails.
-
Fintech Safety & Correctness: Design infrastructure with edge-case money movement safety, auditability, and compliance in mind, going beyond simple uptime.
-
Observability & Incident Management: Own metrics, logging, tracing, and alerting (Datadog, Prometheus, Grafana, ELK); drive end-to-end incident response, production debugging, and postmortems.
-
Database & Cache Management: Manage PostgreSQL operations at scale (replication, performance tuning, recovery) and maintain Redis for distributed locking, rate-limiting, and caching.
-
Automation & IaC: Write automation scripts (TypeScript or python/Go), build CI/CD pipelines, and manage Infrastructure-as-Code (Terraform) and secrets management.
-
Capacity & SLOs: Track SLIs/SLOs, lead capacity planning, optimize AWS cloud spend, and mentor junior/mid-level engineers across the organization.
-
Experience: 7+ years of hands-on experience in Site Reliability Engineering, DevOps, or Infrastructure Engineering, with deep expertise managing Kubernetes, Docker, and AWS in production.
-
Domain Exposure: Proven track record in FinTech, payment processing, or high-stakes distributed transactional systems where correctness and zero data loss are non-negotiable.
-
Core Technologies:
-
Cloud & Containers: AWS, Kubernetes, Docker.
-
IaC & Pipelines: Terraform, CI/CD tools, secrets management.
-
Databases & Caching: PostgreSQL at scale, Redis.
-
Observability: Prometheus, Grafana, Datadog, ELK stack.
-
Scripting/Tooling: Proficiency in TypeScript or another language used for internal tools and automation.
-
-
Core Skills: Root-cause analysis, cross-functional collaboration (with product, compliance, and finance teams), SLO/SLI design, and team mentorship.
This role accepts applications on an external page.