Skip to main content

Service Manager, Site Reliability Engineering

Northern Ireland, United Kingdom Full-time Posted 55 minutes ago
At Allstate, great things happen when our people work together to protect families and their belongings from life’s uncertainties. And for more than 90 years, our innovative drive has kept us a step ahead of our customers’ evolving needs. From advocating for seat belts, air bags and graduated driving laws, to being an industry leader in pricing sophistication, telematics, and, more recently, device and identity protection.

Your role in the team

The Service Manager - Site Reliability Engineering (SRE) is responsible for ensuring the reliability, availability, observability, and operational excellence of technology services while maintaining strong alignment with business objectives. This role serves as the primary owner of service health, incident and problem management, operational governance, continuous service improvement, and stakeholder engagement.

The Service Manager partners closely with engineering, platform, infrastructure, security, and business teams to deliver stable, resilient, and high-performing services that align with business objectives and customer expectations. The role drives proactive risk management, operational maturity, automation, and service improvements while ensuring adherence to established operational processes and governance standards.

The role also provides leadership across incident, problem, change, release, and service management disciplines for a team of Service Analysts to deliver resilient, secure, and customer-focused services that meet organizational goals.

Key responsibilities:

  • Own the overall health, reliability, availability, and observability of critical business applications and technology services, ensuring alignment with established service commitments and customer expectations.
  • Establish, monitor, and report on Service Level Agreements (SLAs), Service Level Objectives (SLOs), Error Budgets, availability, performance, and operational KPIs to drive service excellence.
  • Lead Major Incident Management activities, coordinating cross-functional teams during outages, ensuring effective communication, rapid service restoration, and completion of Root Cause Analysis (RCA) for critical incidents.
  • Drive Problem Management practices by identifying recurring issues, analyzing systemic failures, implementing permanent corrective actions, and reducing operational risk through preventive measures.
  • Govern Change and Release Management processes by assessing operational risk, improving change success rates, supporting production readiness reviews, and coordinating maintenance and deployment activities.
  • Ensure effective observability across services through monitoring, alerting, logging, dashboards, and operational reporting that provide actionable insights into service health and performance.
  • Promote automation and operational efficiency by reducing manual processes, eliminating repetitive tasks, and advancing Infrastructure as Code (IaC), DevOps, self-healing, and auto-remediation capabilities.
  • Partner with engineering, platform, infrastructure, security, and business teams to proactively improve service resilience, scalability, stability, and customer experience.
  • Serve as the primary operational liaison for business stakeholders and vendors, providing regular service reviews, communicating risks, managing escalations, and ensuring alignment with business objectives.
  • Ensure adherence to ITIL-based operational processes, governance standards, compliance requirements, audit obligations, and the maintenance of operational documentation, runbooks, recovery procedures, and service support artifacts.
  • Lead continuous service improvement initiatives focused on reducing Mean Time to Recovery (MTTR), increasing service stability, improving customer satisfaction, and enhancing operational maturity.
  • Foster a culture of operational excellence, accountability, proactive risk management, customer focus, and continuous improvement across the service organization.
  • Own, lead, and facilitate operational governance ceremonies, including service reviews, incident and problem management reviews, change governance forums, operational readiness assessments, stakeholder communications, and executive service reporting.
  • Create, maintain, review, and ensure compliance with operational documentation, runbooks, standard operating procedures (SOPs), knowledge articles, audit evidence, disaster recovery procedures, and certification-related artifacts.
  • Provide leadership and oversight across incident, problem, change, release, and service management disciplines while ensuring consistent execution of operational processes and governance standards.
  • Lead, mentor, coach, and develop a team of Service Analysts, fostering technical growth, accountability, collaboration, and operational excellence.
  • Provide training for team members and users while assisting in building organizational bench strength through knowledge sharing, cross-training, and professional development initiatives.
  • Support operational risk management activities by identifying service vulnerabilities, assessing impacts, developing mitigation plans, and improving overall service resilience.
  • Collaborate with engineering and platform teams to improve service reliability, scalability, security, and production readiness while supporting modernization and transformation initiatives.
  • Drive strategic operational improvements that enhance service quality, customer experience, business alignment, and long-term operational sustainability.

Essential Skills:

  • All applicants must demonstrate they have a legal right to work in the UK for employment at Allstate. Allstate is not providing sponsorship for this vacancy.
  • A minimum of 4 years of experience supporting, or improving enterprise technology services, infrastructure environments, platform operations, Site Reliability Engineering (SRE), IT Operations, or Service Management disciplines. (Or Equivalent).
  • A minimum of 2 years leading and mentoring teams within service reliability, availability, performance, and/or operational governance within a large enterprise environment.
  • Experience leading Major Incident response activities, coordinating cross-functional teams, and driving RCA efforts and corrective actions.
  • Experience implementing operational improvements, automation initiatives, risk-reduction measures, and continuous service improvement programs.
  • Strong understanding of observability practices, including monitoring, alerting, logging, dashboards, and operational reporting.
  • Knowledge of ITIL principles and IT Service Management (ITSM) processes.
  • Experience developing and maintaining operational documentation, runbooks, support procedures, recovery documentation, and knowledge articles.

Desirable Skills:

  • Experience working within a SRE, DevOps, Cloud Operations, Platform Engineering, Enterprise Operations, or production support environment.
  • Experience supporting Identity and Access Management platforms, including IAM, ISAM, IBM Verify, SailPoint, or related identity technologies.
  • Experience managing Service Level Agreements (SLAs) and operational Key Performance Indicators (KPIs). (Or Equivalent)
  • Knowledge of networking technologies, firewalls, DNS, load balancing, and enterprise infrastructure concepts.
  • Experience leading or mentoring a team of engineers. (Or Equivalent)
  • Experience with Infrastructure as Code, automation frameworks, cloud-native operational practices, self-healing systems, or auto-remediation.
  • Familiarity with API management platforms and enterprise service-integration technologies.
  • Experience supporting messaging or event-streaming platforms such as Kafka.
  • Knowledge of middleware or integration technologies such as TIBCO or comparable enterprise platforms.
  • Experience supporting cloud platforms such as Microsoft Azure, Amazon Web Services, or Google Cloud Platform.
  • Experience using ServiceNow or a comparable ITSM platform.
  • Knowledge of compliance, risk management, audit controls, operational resilience, business continuity, and disaster recovery practices.
  • Relevant professional certifications, such as ITIL, SRE, cloud, Kubernetes, security, ServiceNow, or comparable technology certifications.
  • Experience leading operational maturity assessments, service governance programs, or production readiness reviews.

Job Posting End Date: Monday 21st September 2026 (11:59pm)

#Hybrid

Skills

Defect Resolution, ERP Applications, Functional Designs, Integration Testing, Issue Management, Reliability Management, Site Reliability Engineering, Software Reliability, SRE Observability, System Reliability, Systems Reliability

Shape the Future of Insurance with Cutting-Edge Tech and a People-First Culture

Why join us?

Allstate NI is proud to be Allstate’s European Digital Centre of Excellence, a hub for innovation and engineering excellence. We’re recent winners of Best Place to Work in IT (100+ employees) and Best Use of Cloud Services at the Belfast Telegraph IT Awards, and we’ve been recognised for our community and sustainability impact with Platinum in the Northern Ireland Environmental Benchmarking Survey.

We’re a product-driven, cloud-first organisation delivering real outcomes through modern technology, a digital product-centric talent model, and a culture rooted in engineering excellence. Our teams work in cross-functional structures, guided by an outcome-based delivery approach that accelerates speed, agility, and value.

We also invest in you. At Allstate NI, your career growth matters. You’ll have access to our Continuous Learning Hub, designed to support skills development and professional advancement through tailored learning paths, certifications, and mentoring opportunities. Whether you’re deepening technical expertise or exploring leadership roles, we provide the tools and support to help you thrive.

What do you get in return?

As well as receiving a competitive annual salary, our reward package includes:

  • Corporate bonus scheme
  • Pension scheme
  • Annual performance-related pay reviews
  • Life assurance and income protection
  • Flexible working options
  • Hybrid working
  • Private medical and dental insurance
  • Access to an employee assistance programme
  • Discounted gym membership
  • Two paid volunteering days each year
  • Cycle to work scheme

Be part of a high-performing, socially responsible organisation where your work has purpose, and your growth is supported every step of the way.

Similar sponsor-licensed roles

More roles in Northern Ireland, United Kingdom with active sponsor licences.