Senior Site Reliability Engineer - Edge Computing
By bringing together next-gen technology and the finest live data available, Genius Sports is enabling a new era of sports for fans worldwide, delivering experiences that are more immersive, interactive and personalized than ever before. Learn more at geniussports.com. The Role: Senior Site Reliability Engineer – Edge Computing We are seeking a Senior Site Reliability Engineer to be part of the Edge Computing team in the Data&AI group. As we deliver real-time player tracking, sport analytics and broadcast augmentation to more customers worldwide, we are looking to scale from hundreds of sport venues to thousands. Specifically, you can expect to: Design and code end-to-end processes enabling operational staff to autonomously prepare, install and monitor all Linux servers, networking devices and cameras installed in 300+ sport venues across the world Design and code end-to-end processes enabling developers to autonomously deploy and monitor our player tracking and augmentation applications Take ownership of long-term technical efforts and articulate design choices to technical and non-technical people Collaborate closely with teammates to solve problems, share knowledge and provide actionable feedback Participate in an on-call rotation that emphasizes eliminating repeating escalations Visit our wonderful Lausanne Jordils office 4 times per week, with flexible hours Minimum Qualifications Swiss/EU/EFTA citizen or residency permit in Switzerland 5+ years experience in SRE with Linux Strong understanding of the entire Linux server stack: OS boot and installation, systemd, networking, container deployment, logging, metrics & monitoring, out-of-band management, etc... Strong experience designing robust automation processes for a large inventory of on-premises servers Proficient with Python programming and Bash scripting Ability to communicate efficiently and articulate concepts based on the audience Preferred Qualifications Experience with remote fleet management without easy physical access Experience designing large-scale automation processes for network routers and switches Strong understanding of OSI network layers 2 and 3, ability to assess network conditions at customer sites: explain packet loss or fragmentation, cabling or NIC defects, bandwidth evaluation locally and to cloud via intercontinental transit Strong experience deploying container applications on premises with Kubernetes Experience with AWS EC2, S3, VPC, IAM Experience with Nvidia GPU driver installation and monitoring Our Stack: Languages and frameworks: Python, Rust,...