Bluestaq US External

Delivery Engineer - Cloud & Automation

Colorado SpringsFull time$95,000 - $140,000 / yearPosted 21 days ago
Apply on Bluestaq US External →

Sign into see who you know at Bluestaq US External.

Delivery Engineer — Cloud & Automation  Design and Integrate Systems for Mission Success    Bluestaq is hiring a Delivery Engineer to operate and evolve our Kubernetes-based platform on AWS. You will own the deployment automation, day-to-day production health, and incident response for the services that power the Unified Data Library (UDL) — a mission-critical data platform used across government and commercial space operations. This is a hands-on platform/SRE role. You will spend your time deploying infrastructure with Terraform/Terragrunt, debugging failed rollouts in EKS, chasing latency and error spikes through Grafana and traces, and remediating production issues in Kubernetes directly.  Why This Role Matters  Bluestaq delivers integrated, secure, and reliable systems for customers operating in mission-critical and commercially demanding environments. Systems engineering is at the core of how we win, deliver, and scale. Without strong systems engineering execution, our teams can't move fast, our systems can't stay secure, and our customers can't rely on the technology that supports their operations.  This role is foundational to the team and the organization. Your work spans defined features, environments, and components within a team, with growing independence, which means your decisions, your designs, and your technical judgment have a direct and measurable effect on what Bluestaq delivers. We're building a culture of ownership and technical excellence, and we need people who bring both depth and the right mindset to do it.    What Success Looks Like in the First 90 Days Comfortable deploying and updating services across our EKS environments via Terragrunt and Argo CD Independently triaging Grafana alerts and resolving common production issues Contributing improvements to runbooks, dashboards, or automation based on at least one real incident Participating in the on-call rotation   Key Responsibilities    Deploy and maintain UDL platform infrastructure on AWS EKS using Terraform and Terragrunt Manage application rollouts via Argo CD / GitOps workflows; troubleshoot failed deployments, upgrades, and sync issues Monitor production service health using Grafana dashboards, logs, and distributed traces; drive incidents to root cause Perform live remediation in Kubernetes — pod restarts, rolling updates, scaling, node draining, debugging stuck workloads Improve observability coverage, alert quality, and runbooks based on what you learn from real incidents Automate repetitive operational toil using Python, Bash, or PowerShell Partner with product, security, and engineering teams to ship changes safely into regulated environments Mentor junior engineers and share knowledge across the platform team Participate in an on-call rotation for production support You use modern AI tools with judgment: ...