Site Reliability Engineer - Disaster Recovery & Business Continuity
About Charles River Associates For over 50 years, Charles River Associates has been a premier consulting firm that offers employees a place to learn from a diverse group of consultants, industry experts, and academics. At CRA you will be exposed to leading minds who use economic, financial, and business analysis to solve complex world problems for an impressive roster of clients, including major law firms, Fortune 100 companies, and government agencies. Through a collegial environment, formal and informal training opportunities, and a broad array of professional development resources, your experience at CRA will open doors for you throughout your career. The Information Technology (ITS) department at Charles River Associates is currently a team of more than 40 professionals dedicated to enhancing, maintaining, and developing the firm's technology infrastructure and security. The team is comprised of four functions: Service Delivery & Telecom Enterprise Application Solutions Infrastructure, Networking and Cloud Solutions Information Security Information Technology staff are based in the Boston, Chicago, London, Munich, New York, Oakland, San Francisco, College Station and Washington, DC offices. Mainly a Microsoft house, CRA is looking to maximize the performance of our on-premise systems and hybrid infrastructure, meaning experience with cloud technologies is essential for this role. Position Overview The Site Reliability Engineer (SRE) helps ensure CRA’s critical business services are reliable, scalable, and performant across on-premises and cloud environments. This role blends software engineering and operations practices to reduce manual toil through automation, improve service observability, and strengthen incident response. The SRE partners closely with infrastructure, security, application, and service delivery teams to define measurable reliability targets (SLIs/SLOs), implement resilient architectures, and drive continuous improvement through blameless post-incident learning. Key Responsibilities Hands-on System Engineering experience with core enterprise infrastructure platforms and services, including Windows Server, VMware vSphere, VMware Site Recovery Manager (SRM), SAN technologies, and the Rubrik ecosystem, with the ability to understand dependencies, recovery workflows, and failure modes across on-premises and cloud environments Service Ownership & Reliability Targets: Partner with service owners to define and maintain service level indicators (SLIs) and service level objectives (SLOs) for availability, latency, and performance; track error budgets and reliability risk. Observability: Implement and continuously improve monitoring, logging, alerting, and dashboards to provide actionable, symptom-based signals and reduce mean time to detect/respond (MTTD/MTTR). Blameless Postmortems & Continuous Improvement: Facilitate post-incident reviews, identify root causes and contributing factors, and drive remediation item...