Data Engineer
Data Lakehouse Engineer
Location: Basildon, Essex, UK
Work Arrangement: Fully Onsite
Contract: Contract engagement
Contract Duration: 1 year
About the Role
We are looking for an experienced Data Lakehouse Engineer to support data engineering and large-scale data processing initiatives in Basildon.
The ideal candidate will have strong hands-on experience in SQL, Python, AWS Data Lake Engineering, Apache Spark/PySpark, and modern lakehouse technologies, including Apache Iceberg.
You will be responsible for working with large-scale production data, building and supporting scalable data processing solutions, and contributing to reliable and efficient data lakehouse platforms.
AWS experience is preferred, although candidates with equivalent experience on Azure or GCP will also be considered.
Key Responsibilities
- Design, develop, and maintain scalable data lake and lakehouse engineering solutions.
- Build and optimize data processing pipelines using Python, SQL, and PySpark.
- Work with AWS data engineering services, including S3, EMR, IAM, Lambda, EKS, and MWAA, where relevant to the project.
- Implement and work with Apache Iceberg-based data lakehouse solutions.
- Process and manage large-scale production datasets while considering performance, scalability, and cloud costs.
- Develop efficient data transformation and processing workflows.
- Support data quality, reliability, and security within data engineering pipelines.
- Collaborate with technical teams and stakeholders to deliver robust data solutions.
- Contribute to workflow orchestration and automation where required.
- Follow engineering best practices for maintainability, monitoring, and deployment.
Essential Skills & Experience
Core Technical Requirements
- SQL: Strong hands-on experience with SQL for data processing, transformation, and querying.
- Python: Strong practical Python development experience in data engineering environments.
- AWS Data Lake Engineering: Experience building or supporting data lake solutions on AWS.
- Apache Spark / PySpark: Experience developing and working with Spark-based data processing pipelines.
- Apache Iceberg: Practical experience with Apache Iceberg or modern lakehouse table formats and technologies.
- Production-Scale Data: Demonstrated experience handling large-scale data volumes, including performance optimization, scalability, and cost awareness.
AWS Services
Experience with some or all of the following AWS services is desirable:
- AWS IAM
- AWS Lambda
- Amazon EKS
- Amazon S3
- Amazon EMR
- Amazon MWAA
AWS is the preferred cloud platform. Equivalent experience with Azure or GCP may be considered.
Desirable Skills
- CI/CD
- Terraform / Infrastructure as Code
- Data Quality
- Data Security
- Workflow Orchestration (e.g., Apache Airflow)