Somnia

Senior Infrastructure Engineer

RemoteRemoteFull timeSeniorPosted 12 days ago
Apply on Somnia →

Sign into see who you know at Somnia.

About the Role:

We're looking for a Senior Infrastructure Engineer to build and run Somnia's key backend
services: the L1 and node fleet, RPC and indexing layers, product backends, and
developer-facing services teams depend on. An SRE-minded role: you make reliability
measurable, rollouts safe, and infrastructure repeatable, so every team can move fast without
breaking things.

You treat infrastructure as a product: automate relentlessly, measure everything, and leverage
AI to accelerate development, operations, and incident response.

Key Responsibilities:

Define and maintain SLOs, SLIs, and error budgets, plus the observability—metrics,
logs, traces and alerts—that catches regressions before users do.
Build repeatable, self-service infrastructure through infrastructure-as-code, CI/CD and
golden paths so teams can provision, deploy and recover without reinventing the wheel.
Own rollouts end-to-end—progressive delivery, canaries, safe migrations and clean
rollbacks.
Operate the systems behind Somnia's nodes, validators, RPC and indexing, tuning for
performance and cost across regions.
Lead incident response and on-call, run blameless postmortems, and continuously
harden the platform.
Partner with product and protocol teams to design and operate production-ready
services. You'll rotate between embedding with engineering teams and building the
shared platform, tooling and operational standards that underpin the wider organisation.

Requirements:

Must Have

- Strong experience operating production infrastructure at scale (cloud and/or bare metal), with deep Linux fundamentals.

- Experience with infrastructure-as-code such as Terraform or Pulumi, alongside configuration management.

- Experience running containers and orchestration platforms (Docker, Kubernetes) in production.

- Strong programming skills, ideally in Go and/or TypeScript, for building automation and internal tooling.

- Experience with observability stacks (Prometheus, Grafana, OpenTelemetry...