Site Reliability Engineer/Lead
Making a difference isn’t just our purpose, it’s what motivates our people every day. For us, work is not about ticking a box. It’s about knowing that it matters. We make a meaningful difference every day by making it easier for our customers to focus on what matters in their lives.This is a great opportunity to establish and lead Site Reliability Engineering at Macmillan Shakespeare.Here’s how you will make a difference in this role… You'll start as a hands-on technical leader, designing and building our enterprise observability platform across Azure, cloud and hybrid environments. As the capability grows, you'll transition into leading and building our SRE practice, setting engineering standards, mentoring a team and influencing technology strategy across the business.If you're excited by creating platforms, automating operations, improving reliability and using modern technologies—including the adoption of AI-assisted operations—this role offers the autonomy and support to make a lasting impact.At MMS clear expectations help our people Be the Difference. Your key responsibilities in this role will include: Build our SRE capability from the ground up.Shape the enterprise observability platform strategy and technology roadmapInfluence architecture and engineering practices across the organisationWork with modern Azure cloud technologies, automation and AI-enabled operationsEnsure Platform Robustness, Compliance & Cyber SecurityGrow into a strategic leadership role while staying close to technologyPartnering with Product Owners, Engineering Managers and Platform teams to embed reliability engineering practices throughout the software development lifecycle.Join a collaborative team that values innovation, continuous improvement and engineering excellenceTo be considered for this role you will have: Demonstrated experience designing and operating enterprise monitoring, observability and operational platforms in complex hybrid/cloud environments.Experience with Infrastructure as Code (IaC), CI/CD pipelines, platform automation, and Microsoft Azure integrations, including logging, monitoring and alerting.Experience implementing monitoring, logging and distributed tracing using modern observability practices.Experience designing automated operational workflows using APIs and event-driven architectures.Strong understanding of Site Reliability Engineering (SRE), operational automation and continuous improvement.Experience leading major incident management, root cause analysis and post-incident reviews.Ability to simplify complex technical concepts and influence engineering teams through technical leadership.Strong stakeholder engagement skills across engineering, product, operations, cyber security and executive leadership.Experience developing engineering standards, governance frameworks and operational practices.DesirableExperience establishing or leading SRE, Platform Engineering or DevOps functions.Experience with enterpri...