Site Reliability Engineer IV
Workforce Classification:Hybrid Join Our Team: Do Meaningful Work and Improve People’s Lives Our purpose, to improve customers’ lives by making healthcare work better, is far from ordinary. And so are our employees. Working at Premera means you have the opportunity to drive real change by transforming healthcare.Premera is committed to being a workplace where people feel empowered to grow, innovate, and lead with purpose. By investing in our employees and fostering a culture of collaboration and continuous development, we’re able to better serve our customers. It’s this commitment that has earned us recognition as one of the best companies to work for. Learn more about our recent awards and recognitions as a greatest workplace.Learn how Premera supports our members, customers and the communities that we serve through our Healthsource blog: https://healthsource.premera.com/.Site Reliability Engineer IVJob Description SummaryAs a Site Reliability Engineer IV, you will drive reliability and operational excellence across cloud, on-premise, and hybrid platforms. You will build scalable automation and AI-powered tooling to improve system health, reduce manual effort, and accelerate incident response.Partnering with software and platform engineering teams, you will standardize CI/CD, observability, and incident management practices, enabling resilient, self-healing systems. This role is critical to scaling engineering reliability and advancing intelligent automation across enterprise platforms.This is a hybrid role, located on our campus in Mountlake Terrace, WashingtonWhat You’ll DoBuild, run, and optimize critical services across cloud, on-premise, and hybrid environments, including managed services, custom applications, and third-party integrationsDevelop automation and AI-powered tooling to reduce manual intervention, including anomaly detection, predictive alerting, and LLM-assisted diagnostics that surface actionable insightsDesign and implement end-to-end observability, telemetry, and self-healing capabilities across platformsLead cross-team efforts to drive root cause analysis, post-incident reviews, and long-term reliability improvementsDefine and drive reliability strategy, standards, and best practices across engineering teamsStandardize workflows for change management, deployment, and incident response, replacing manual processes with tooling-driven solutionsPartner with engineering and security teams to ensure deployment pipelines and automation practices meet reliability, safety, and compliance standardsInfluence adoption of modern DevOps practices including CI/CD, infrastructure-as-code, and test-driven developmentStay current on emerging technologies in AI/ML, DevOps, and platform engineering, and apply them to improve operational efficiencyParticipate in the on-call rotation and support production systems as neededThis role does not involve day-to-day coding, it requires strong techn...