Reflectionai

Member of Technical Staff - Distributed Systems Engineer

New York City, New York, United StatesFull timeStaffPosted 12 days ago
Apply on Reflectionai →

Sign into see who you know at Reflectionai.

OUR MISSION

Reflection is a research lab making intelligence open and accessible for everyone to use, customize, and build on. We build open models that let anyone control their intelligence and help shape the future of AI. Our mission: make intelligence open and accessible to all.

FOUNDATIONS

VISION:

Build and operate a company-wide foundations platform that accelerates every team by providing reliable, scalable developer infrastructure, SRE capabilities, and high-throughput data ingestion tooling enabling Reflection to move faster as we scale.

WHAT THIS TEAM DOES

Build and operate the core shared services that power our research, training, and production environments. These systems form the foundational platform that multiple teams depend on for model development, deployment, and evaluation, unifying data, compute, and workflow management across the stack while enabling rapid experimentation and reliable production systems.

- Build and operate shared services that multiple teams rely on across research and production workflows.

- Define and uphold reliability targets through SLIs, SLOs, and healthy on-call practices.

- Maintain strong operational readiness with runbooks, incident playbooks, and capacity planning.

- Ensure correctness and performance under load, addressing consistency, tail latency, and failure modes.

- Develop APIs, SDKs, and internal platforms that enable high-velocity experimentation and iteration.

- Reduce operational burden through better tooling, standardization, and platform patterns that scale across teams.

WHAT YOU'LL WORK WITH

- Container Abstractions: Containers-as-a-Service, Kubernetes abstraction layers, container orchestration, reproducible environments, multi-tenant isolation.

- Distributed Systems Architecture: Sharding, replication, coordination services, high-concurrency systems, concurrency control.

- Service Development Stack: gRPC, Protobuf, Go, Rust, C++.

- Reliability & Performance:...