The Staff Data Engineer is a senior technical contributor responsible for designing and delivering scalable, reliable and secure data solutions. The role provides technical leadership across data engineering, analytical pipelines, data ontology, data matching, data cleansing, MLOps and automation.
A key focus is transforming complex source data into trusted, reusable and dashboard-ready data products, while automating their development, deployment, testing and ongoing management.
Design and build scalable batch, streaming and analytical data pipelines.
Transform source data into trusted, reusable and analysis-ready datasets.
Create curated data layers, semantic models and reusable data products for dashboards, reporting and analytics.
Automate the dashboard data lifecycle, including ingestion, transformation, testing, deployment, monitoring, refresh and issue remediation.
Design and implement agentic workflows that coordinate data preparation, quality validation, metadata generation, dashboard updates and operational support.
Build and maintain CI/CD pipelines using GitHub Actions.
Design data models, ontologies, taxonomies and common data definitions.
Implement data matching, entity resolution, deduplication and data cleansing solutions.
Build automated data quality, validation, reconciliation and observability controls.
Develop solutions for large and complex datasets using big data and distributed processing technologies.
Support the productionisation and monitoring of machine learning and open-weight models.
Develop production-grade solutions using Python and/or R.
Establish engineering standards, mentor engineers and provide technical guidance.
Ensure data solutions meet security, privacy, resilience and governance requirements.
Strong experience in data engineering and large-scale data platforms.
Strong algorithmic optimisation capabilities
Advanced Python and/or R programming skills, with strong SQL capability.
Experience building production-grade analytical, batch or streaming pipelines.
Experience preparing reusable datasets and semantic data layers for dashboards, reporting and analytics.
Strong experience implementing CI/CD pipelines using GitHub Actions.
Experience automating data and analytical product lifecycles, including testing, deployment, monitoring and release management.
Experience designing or implementing agentic workflows, AI-driven automation or intelligent orchestration.
Experience with big data and distributed processing technologies.
Experience with data modelling, semantic modelling or data ontology.
Experience with data matching, entity resolution, deduplication and data cleansing.
Experience implementing data quality and data observability controls.
Exposure to MLOps and the productionisation of analytical or machine learning models.
Strong cloud platform, software engineering, automated testing and infrastructure-as-code experience.
Demonstrated technical leadership and mentoring capability.