Calix

Site Reliability Engineer

Bangalore, IndiaFull timePosted 1 day ago
Apply on Calix →

Sign into see who you know at Calix.

The Calix platform enables Communication Service Providers (CSPs) of all sizes to transform and future-proof their businesses. Through real-time data, automation, and actionable insights delivered via Calix One — our cloud-first, AI-powered platform — CSPs can simplify operations, collapse cost, and accelerate innovation. Calix One brings together the automation of everything and the experience of one, empowering customers to deliver differentiated subscriber experiences while driving acquisition, loyalty, and revenue growth. This is the Calix mission: to enable CSPs of all sizes to simplify, innovate, and grow, strengthening both their businesses and the communities they serve. We’re at the forefront of a once in a generational change in the broadband industry. Join us as we innovate, help our customers reach their potential, and connect underserved communities with unrivaled digital experiences.The Site Reliability Engineer ensures our production services remain highly available, scalable, and efficient on Google Cloud Platform. You will bridge the gap between development and operations by diving deep into application source code and cloud infrastructure to permanently engineer away underlying issues, rather than just patching symptoms. This role focuses on owning complex alert triage, reading and debugging code to solve root causes, GitOps application deployments via ArgoCD, and leveraging Grafana observability alongside advanced AIOps platforms to drive down operational toil.Key Responsibilities:Code-Level Alert Resolution: Act as the ultimate owner of complex alerts by investigating stack traces, reading application source code, and submitting code-level fixes alongside infrastructure adjustments to permanently resolve chronic issues.Infrastructure Optimization: Diagnose and resolve deep OS and distributed system bottlenecks—including CPU throttling, memory leaks, and storage constraints—across GKE worker nodes and critical data infrastructure (e.g., Kafka, Datastream).GitOps & Deployments: Deploy, roll back, and manage the lifecycle of containerized applications using ArgoCD pipeline workflows, ensuring safe and reliable release rollouts.AIOps & Automation: Utilize AI-driven operations tools and build custom Python/Go automation (e.g., PagerDuty API integrations) to interpret correlated events, reduce alert noise, and automate manager/triage workflows.Network Troubleshooting: Diagnose complex connectivity and latency issues across all network layers, isolating problems between cloud VPCs, Kubernetes overlays, and microservices.Incident Response: Participate in on-call rotations, using Grafana dashboards, AIOps suggestions, and code-level tracing to rapidly mitigate and permanently fix production container issues.What You'll Actually Do (Example Scenario):The Alert: PagerDuty pages you for elevated consumer lag on a critical Kafka topic and latency spikes in a downstream microservice.The Investigation: You use Grafana to correlate...

Also hiring in