Brillio 2

Senior SRE Lead - R01564153

Guadalajara, Jalisco, MexicoFull timeLeadPosted 13 days ago
Apply on Brillio 2 →

Sign into see who you know at Brillio 2.

Senior SRE Lead
Role Overview
We are seeking a highly experienced Senior SRE Lead to lead reliability engineering and observability initiatives for critical platforms supporting GM Financials’ ecosystem, with a primary focus on Salesforce and Microsoft Azure environments. This role will be responsible for establishing and scaling SRE practices, driving operational excellence, and ensuring high availability, performance, and resilience of business-critical applications. The ideal candidate will bring deep expertise in cloud-native architecture, observability frameworks, and enterprise-scale production support, along with strong leadership capabilities in a global delivery model.
Key Responsibilities - SRE Leadership & Strategy
Lead the SRE function for Salesforce and Azure platforms, defining the roadmap and maturity model
Establish and drive SRE best practices, including SLIs, SLOs, and error budgets
Build and mentor a high-performing SRE team across onshore and offshore locations
Collaborate with GM Financial stakeholders, product teams, and engineering leadership Platform Reliability (Salesforce & Azure)
Ensure high availability, performance, and scalability of Salesforce applications and Azure-hosted services
Lead major incident management (P1/P2), including triage, stakeholder communication, and resolution
Drive root cause analysis (RCA) and implement preventive measures
Manage production stability across integrations between Salesforce and Azure services Observability & Monitoring
Design and implement end-to-end observability across Salesforce and Azure ecosystems
Establish unified monitoring across logs, metrics, and traces
Implement and optimize tools such as Azure Monitor, Application Insights, Splunk, Datadog, or similar
Define dashboards, alerting strategies, and actionable insights for proactive issue detection Automation & DevOps
Drive automation across incident response, remediation, and operational workflows
Implement Infrastructure as Code (IaC) pr...