DevOps Engineer (Active Secret Clearance)
“In 36 months, agentic AI systems will be an operating reality across major institutions. We intend to be central to it.” — Dr. Jim Rebesco, Cofounder and CEO, Striveworks The government’s demand for AI is growing far faster than the systems required to support it. Fewer than 15% of federal AI programs have reached sustained production, despite billions of dollars invested. The models perform in testing, but they degrade in the real world. And when performance drops, trust goes with it. Striveworks was built to solve that problem. What you’ll build Since 2018, we have delivered the most trusted AI systems operating in real-world use cases—providing a layer of assurance underneath hundreds of deployed models that monitors performance, manages drift, and sustains systems long after they leave the lab. As a DevOps Engineer supporting operations on site at Fort Carson, CO, you are the tactical edge of our engineering team. You aren’t just maintaining a platform; you play a key role in our technical success. You’ll interface directly with customers to understand their unique constraints—restricted networks, hardware limitations, urgent operational requirements—and tailor our automation and deployment strategies to move the needle for them. You will be responsible for maintaining the end-to-end life cycle of our AI platform across a diverse architectural landscape—including remotely accessible cloud environments and on-premises hardware. You’ll thrive in this role if you enjoy the challenge of debugging complex Kubernetes clusters where connectivity is a luxury, not a given. You are someone who can navigate the unique constraints of local hardware today and write the automation that ensures seamless, reliable performance across the customer’s hybrid stack tomorrow. What it’s like here We lead with trust, treat each other with respect, and use candor consistently, kindly, and constructively. We care deeply about our work, and we find genuine satisfaction in doing it well. Above all, we take ownership—because we feel the weight of collective results personally. We are looking for people who share these values and are eager to put them into practice. What we’re looking for 3–5+ years of hands-on experience in software, DevOps, site reliability, or systems engineering Proven technical leadership experience, with the communication skills and professional presence required to manage customer relationships and lead cross-functional incident responses Expertise in deploying and diagnosing microservices within K8s Experience with comprehensive observability solutions using tools such as Prometheus, Grafana, and OpenTelemetry to ensure system reliability and performance visibility Ability to design and execute testing strategies to validate application functionality, diagnose issues, and perform root cause analysis, implementing effective long-term fixes to improve stability and performance Deep proficiency in Terraform, Ansible, or similar tools to man...
Also hiring in
- TacomaApply