Senior Site Reliability Engineer
8x8 connects our customers and teams globally, empowering CX leaders with performance and insights to make smarter decisions, delight customers, and drive lasting business impact.
What You'll Do
- Production Operations & Incident Response
- Own platform reliability across global UC infrastructure, driving incident response and the overall reliability strategy for your subsystem rather than resolving issues in isolation.
- Triage and resolve the hardest issues — service restarts, hung processes, infrastructure failures — and act as the senior escalation point for the NOC and for other engineers when frontline teams hit their limit.
- Execute and improve the unglamorous but essential work: scheduled maintenance, certificate renewals, log rotation — and redesign these processes so failure is prevented systemically, not handled case by case.
- Lead blameless post-mortems that produce real follow-through, and sign off on the corrective actions that come out of them.
Cross-Team Collaboration
- Work directly with Support, Sales, Sales Engineering, NOC, Professional Services, and Engineering teams across 8x8 — this team sits at the operational center of the company.
- Translate production events into clear, business-readable communication under pressure; stakeholders across the org depend on your judgment during incidents.
- Feed operational insight back into engineering — turning recurring failures and patterns into actionable bug reports, platform improvements, and influence over the architectural roadmap.
- Work closely with technical leads to align reliability and automation work with broader engineering goals, and help focus discussion on what matters most.
- Reliability Engineering & Automation
- Identify recurring manual work and build automation to eliminate it — we treat toil as a bug, not a requirement.
- Drive design for the tooling and automation in your domain; anticipate how a change in one component impacts others and account for adjacent domains in your designs.
- Understand the limits of our existing tools — and recognize when a problem exceeds those limits and deserves the effort of building a new one.
- Take on large-scale technical debt and refactoring across the subsystem, and contribute to the team's coding methodologies and best practices.
- Participate in 2-week sprint cycles to deliver automation, tooling improvements, runbook development, and infrastructure initiatives from a structured backlog. Own the functional specifications for large features and sign off on test plans.
- Address security issues as they arise — CVEs, misconfigurations, access control gaps — treated as first-class work alongside incident response.
- Define and track SLIs, SLOs, and SLAs to drive honest, data-driven conversations about where reliability investment is needed.
- Build and maintain dashboards (Grafana, OCI Log Analytics) that give the team genuine signal; tune alerting to eliminate noise — a high-noise on-call is itself a reliability...