Arcticwolf

Senior Reliability Developer

Eden PrairieFull timeSeniorPosted 3 days ago
Apply on Arcticwolf →

Sign into see who you know at Arcticwolf.

At Arctic Wolf, you won’t just watch the cybersecurity industry evolve – you'll help lead the change. Our global Pack is made up of people who thrive on solving hard problems, moving fast, and building technology that protects organizations around the world. We’re proud to be recognized by Forbes, CNBC, Fortune, CRN, Gartner Peer Insights and IDC MarketScape – but what matters most is the work behind it: delivering real outcomes for customers through award winning innovation like our Aurora Platform.  If you’re looking for meaningful work, smart teammates and the chance to make a real impact in a high-growth company that’s redefining security operations, Arctic Wolf is the right place for you! Our mission is simple: End Cyber Risk. We’re looking for a Senior Reliability Developer to be part of making this happen.   LOCATION: 8939 Columbine Road, Eden Prairie, MN 55347; Telecommuting permissible from any location in the US.           About the RoleResponsible for changes and updates to public/private cloud infrastructure. Owns all monitors that support AWN business services and our customers. Implements effective automate using programming languages such as Bash, Python, and JavaScript. Ensure our full infrastructure stack is resilient with day-to-day care and feeding. Write and manage Terraform modules and configuration trees for deployments into AWS and OpenStack. Manage automated patching solutions for instances deployed across multiple regions to ensure compliance with strict security requirements. Monitor internal and external TLS certificates for expiration, renewals, and re-deployment into their respective environments. Create and improve service monitoring solutions using Prometheus, Grafana, Zabbix, and various Prometheus exporters (e.g., postgres, node-exporter). Migrate monitoring for legacy production services to a new monitoring system and recreated existing monitoring within Prometheus + Grafana. Collaborate frequently with development teams to implement new features, fixes, or troubleshoot services in lab/production environments. Troubleshoot problems with microservices running within containers, administer containers in Kubernetes clusters, manage cloud components deployed in multiple AWS regions, and maintain CI/CD pipelines. Manage the Cylance AWS Amazon cloud platform through regular administration of Kubernetes cluster upgrades, cluster module upgrades, logging configurations, and resizing of EBS volumes and instances. Troubleshoot issues between interconnected services by working with various teams and resolving issues related to API calls and service unavailability. Create, improve, and maintain detailed documentation for MOPs, service architecture, runbooks, monitoring configurations, and other general resources. Plan and execute service decommissioning in both lab and production environments upon customer or service o...