AWS Data Engineer - Pyspark
Job Title: AWS PySpark Data Engineer
Location:Northampton (Hybrid)
Job Type: Contract
Industry: Banking / Financial Services
Job Summary
We are looking for an experienced AWS PySpark Data Engineer to join a high-performing data engineering team within a leading Banking client. The ideal candidate will have strong expertise in building scalable data pipelines on AWS, developing ETL solutions using PySpark and Python, and working with cloud-native data services. Experience in data warehousing, automation, and DevOps practices is highly desirable.
The successful candidate will be responsible for designing, developing, optimizing, and maintaining enterprise-grade data platforms while ensuring high performance, security, and reliability.
Key Responsibilities
- Design, develop, and maintain scalable ETL/ELT pipelines using PySpark and Python.
- Build and manage cloud-native data solutions on AWS.
- Develop data ingestion, transformation, and orchestration workflows.
- Work extensively with AWS data services to process large-scale datasets.
- Create reusable and optimized data processing frameworks.
- Develop SQL queries for data extraction, validation, and reporting.
- Monitor and optimize ETL performance and troubleshoot production issues.
- Collaborate with Solution Architects, Data Analysts, DevOps, and Business stakeholders.
- Implement Infrastructure as Code (IaC) using CloudFormation.
- Follow CI/CD best practices and version control processes.
- Participate in code reviews and ensure coding standards are maintained.
- Ensure data quality, governance, and security across the data platform.
Must-Have Technical Skills
AWS
- Amazon S3
- AWS IAM
- AWS Lambda
- AWS Glue
- AWS Step Functions
- AWS Lake Formation
- Amazon VPC & Subnets
- Amazon Athena
- Amazon Redshift
- AWS Service Catalog
- AWS CloudFormation
- Amazon SNS
- Amazon SQS
Programming & Data Engineering
- PySpark
- Python
- SQL
- Data Pipeline Development
- ETL/ELT Development
Development Tools
- Basic Unix Commands
- Shell Scripting
- Git
Good-to-Have Skills
- GitLab
- Nexus Repository
- Redshift Performance Optimization
- Snowflake
- dbt (Data Build Tool)
Required Experience
- 6+ years of experience in Data Engineering.
- Strong hands-on expertise in PySpark, Python, and AWS.
- Experience building scalable ETL pipelines and cloud-based data platforms.
- Strong SQL development and query optimization skills.
- Experience with Infrastructure as Code using CloudFormation.
- Familiarity with Agile/Scrum methodologies.
- Banking or Financial Services experience is highly desirable.
Preferred Qualifications
- Bachelor's degree in Computer Science, Information Technology, Engineering, or a related discipline.
- AWS Certification (Associate or Professional) is a plus.
- Experience working in enterprise-scale cloud environments.
Key Skills
- AWS
- PySpark
- Python
- SQL
- S3
- Glue
- Lambda
- Athena
- Redshift
- Lake Formation
- Step Functions
- CloudFormation
- IAM
- VPC
- SNS
- SQS
- Shell Scripting
- Git
- ETL
- Data Engineering
- Snowflake (Preferred)
- dbt (Preferred)
- GitLab (Preferred)
Location:Northampton (Hybrid)
Banking domain experience is highly preferred.