OpenAI

Technical Program Manager, Core Network & WAN Infrastructure

San Francisco, California, United StatesFull timeManager$162,000 - $335,000 / yearPosted 15 days ago
Apply on OpenAI →

Sign into see who you know at OpenAI.

About the Team

The compute infrastructure team runs the GPU fleet and large-scale compute clusters that serve the models backing ChatGPT and the API, while also supporting training workloads for our next generation models. We operate a large, modern GPU fleet and provide a unified platform for other OpenAI teams to seamlessly run production Applied AI and Research training workloads.

We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth.

About the Role

You will join an engineer-first TPM team and own end-to-end delivery of OpenAI’s WAN Infrastructure, partnering with engineers to bring clusters online across external providers and partners.

This is a hands-on infrastructure execution role. You’ll run a broad, parallel portfolio spanning hardware, cabling, optics, cloud cross-connects, and port maps—driving execution, risk management, and crisp alignment from working teams through leadership to deliver production-ready capacity at scale.

The right person combines enough technical depth to reason through physical and logical network readiness with the program discipline and ownership to keep complex builds moving and improve how we scale.

This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance.

In this role, you will:

- Lead end-to-end delivery of OpenAI’s WAN build-out and large-scale GPU clusters across an external partner ecosystem

- Drive multi-threaded network infrastructure bring-up programs across physical and logical readiness—owning plans, dependencies, and critical paths

- Partner with engineering to turn up large-scale network capacity, reducing bring-up time for our largest GPU supercomputers by working across optics, circuits

- Identify recurring bottlenecks in WAN and network infrastructure build-out, then drive fixes t...