Gimlet

Member of Technical Staff - Distributed Systems

San Francisco, California, United StatesFull timeStaff$150,000 - $350,000 / yearPosted 13 days ago
Apply on Gimlet →

Sign into see who you know at Gimlet.

About Us

Gimlet is building the next generation of AI infrastructure: large-scale AI datacenters and the orchestration platform that coordinates them.

The future of AI will require vastly more compute than exists today. But as AI workloads become more complex and new hardware architectures emerge, simply deploying more GPUs isn't enough. The challenge is making increasingly diverse compute work together.

Gimlet's platform intelligently partitions and routes workloads across heterogeneous hardware, enabling step-function improvements in performance and efficiency. Customers deploy through production-grade APIs without needing to think about hardware selection, placement, or optimization.

We work with foundation labs, hyperscalers, and AI-native companies to power production workloads at massive scale and help define the infrastructure layer for the future of AI.

ABOUT THE ROLE

At Gimlet, we believe every hire changes the company.

As a Series A company, talent density matters more than headcount. The engineers we hire today will shape the systems, culture, and standards that define Gimlet for years to come.

The future of AI infrastructure will not be built on a single hardware platform. It will be built on systems capable of coordinating increasingly heterogeneous compute at unprecedented scale.

This role is an opportunity to help build that future. We are not optimizing for headcount, we are optimizing for talent density.

You will design and operate the distributed systems that schedule, route, and coordinate AI workloads across thousands of nodes and diverse hardware architectures.

WHAT SUCCESS LOOKS LIKE

In the first 12-18 months, you will help:

- Build scheduling and orchestration systems that coordinate workloads across heterogeneous hardware

- Improve reliability and fault tolerance for production AI infrastructure operating at scale

- Create APIS and control planes that simplify deployment for customers running mission-critical workload...