Principal Platform Engineer
London, England, United Kingdom Full-time Posted 2 weeks ago
We're hiring a Principal Platform Engineer to be a senior technical anchor in our Platform/Operations function. This is a 70% hands-on engineering role: you'll design, build, and operate the infrastructure that keeps a global trading platform running, while acting as a technical mentor and design authority across the wider team.
A core part of this role is leading our adoption of AI within operations — using LLM-based tooling, agentic workflows, and automation to reduce toil, accelerate incident response, and raise the operational leverage of the whole department. We're not looking for someone to write a strategy deck about AI; we're looking for someone who will build with it.
What you'll do
Platform & cloud engineering (the 70%)
Must have
A core part of this role is leading our adoption of AI within operations — using LLM-based tooling, agentic workflows, and automation to reduce toil, accelerate incident response, and raise the operational leverage of the whole department. We're not looking for someone to write a strategy deck about AI; we're looking for someone who will build with it.
What you'll do
Platform & cloud engineering (the 70%)
- Design, build, and operate infrastructure across AWS (VPC, networking, EKS), Azure (AKS, AKV, networking), with an understand of on-prem environments
- Own core networking design and implementation across within Cloud environments — connectivity, routing, DNS, load balancing, firewalls, and private links between cloud and on-prem
- Engineer for high availability: multi-region and multi-AZ architectures, failover design, capacity planning, and disaster recovery for a platform where downtime has direct market impact
- Build and maintain infrastructure-as-code (Terraform or similar), CI/CD pipelines, and Kubernetes platforms as products consumed by engineering teams
- Define and drive SLOs, error budgets, and observability standards (metrics, logging, tracing) across the platform
- Take part in post-incident reviews and help with the resulting reliability work to completion
- Continuously reduce toil through automation — if we've done it manually twice, you're already scripting it
- Identify, prototype, and productionise AI-assisted workflows across the department: incident triage and summarisation, runbook automation, log/alert analysis, change-risk assessment, internal knowledge tooling
- Use AI-assisted engineering tools (Copilot, or equivalents) as a first-class part of your own workflow, and coach the team to do the same safely and effectively
- Establish sensible guardrails for AI use in a regulated, availability-critical environment, knowing when automation should act and when it should recommend
- Act as a mentor to engineers across the Platform and Operations teams, raising the bar on design, code, and operational practice
- Be a design authority on cross-team projects: review architectures, challenge assumptions, and ensure new services are built to be operable, observable, and resilient from day one
- Contribute to the technical roadmap for the platform function, balancing reliability investment against delivery
Must have
- A background in Operations or SRE running highly available, redundant production platforms — you understand failure domains, graceful degradation, and what "five nines" costs
- Deep hands-on experience with AWS (VPC design, networking, EKS) and Azure (AKS, Key Vault, networking) — genuinely multi-cloud, not one cloud plus a certification
- Strong Networking fundamentals: TCP/IP, routing concepts, firewalls, load balancing, hybrid connectivity (Direct Connect / ExpressRoute, VPNs)
- Production Kubernetes experience at scale, including day-2 operations (upgrades, capacity, security, multi-cluster)
- Infrastructure-as-code and automation as a default working style (Terraform, Ansible, or similar; strong scripting in Python, Go, or Bash)
- Demonstrable, practical use of AI tooling to improve engineering or operational workflows — you can show us something you've automated, accelerated, or de-toiled with it
- The credibility and communication skills to mentor senior engineers and influence design decisions without formal authority
- Understanding of different database technologies
- Experience in trading, exchanges, market data, fintech, or another latency- and availability-sensitive domain