Identiko Solutions
IT & Telecoms
Lead Streaming Engineer
Share this role
About this role
We are looking for a Lead Streaming Engineer to build the ingestion path for the analytics platform of a national-scale foundational identity programme in West Africa.
You will own how data gets out of the identity system and into the lake — change data capture, the streaming buffer, stream processing, and the batch and log paths alongside them. This is the component that determines whether the entire platform can be trusted, and it is the one that touches the production identity environment.
The constraints are the interesting part. The source system cannot absorb query load, so capture is log-based. Delivery is at-least-once, so the pipeline has to deduplicate correctly or every recovery event inflates the national enrolment figures. Events arrive out of order and the source schema changes underneath you. Personal identifiers must be pseudonymised inside the pipeline, before anything is written. And there is a historical migration of over one hundred million records to land alongside the live stream.
If you have spent time explaining to people why their pipeline is silently losing or duplicating data, this role will be familiar.
Must-Have Criteria- At least seven years in data engineering, of which three or more building production streaming or change data capture pipelines.
- Production experience of log-based change data capture — Debezium or equivalent — including schema change handling and connector operation.
- Apache Kafka in production: topic and partition design, retention, consumer groups, and the practical consequences of at-least-once delivery.
- Stream processing at scale, in Apache Spark Structured Streaming, Flink or equivalent.
- Has diagnosed and resolved a real duplication or ordering defect in production, and can describe both the diagnosis and the design change.
- Workflow orchestration — Airflow or equivalent.
- Writing to an open table format, including compaction and file sizing.
- A large historical migration, in the order of one hundred million records or more.
- MOSIP schema familiarity.
- Ingestion from systems holding personal or sensitive data under a data protection regime.
- Java and Python.
- Kubernetes.
To apply, continue with your account. We will keep this job selection for you.