Tensorwave

Infrastructure Engineer – Storage Platform

Las Vegas, Nevada, United StatesFull timePosted 11 days ago
Apply on Tensorwave →

Sign into see who you know at Tensorwave.

About TensorWave

Our mission is simple: deliver seamless, secure, reliable, and resilient AI compute at scale. We've built a versatile cloud platform that eliminates infrastructure barriers, empowering builders to focus on innovation instead of fighting their stack. Because breakthrough AI should move at the speed of ideas, not infrastructure.

 

About the Role

We’re looking for a Storage Platform Infrastructure Engineer to join our team during an exciting phase of growth. In this role, you’ll be responsible for ensuring that storage systems remain stable, performant, and aligned with the demands of Kubernetes, AI/ML, and high-performance compute workloads, working closely with cross-functional partners to support business objectives while upholding our standards for excellence, collaboration, and impact.

 

What You’ll Do

Platform Operations & Ownership

- Operate and maintain distributed storage platforms, including Ceph (RBD, CephFS, RGW), High-performance NAS platforms (e.g., Weka, VAST Data)

- Manage storage lifecycle operations - cluster expansion, upgrades and migrations

- Monitor and maintain storage health, including capacity utilization, data distribution and balance, cluster state and recovery operations

Performance & Reliability

- Analyze and troubleshoot storage performance across IOPS, throughput, and latency (including tail latency)

- Identify and remediate bottlenecks across disk subsystems, network paths (including RDMA where applicable), client access patterns

- Support incident response and root cause analysis for storage-related issues

- Ensure storage platforms meet performance expectations for GPU and Kubernetes workloads

Kubernetes Storage

- Operate and support Kubernetes-integrated storage - CSI drivers, StorageClasses, PersistentVolumes / PersistentVolumeClaims

- Troubleshoot storage-related issues in Kubernetes environments, including stateful workloads, performance inconsistencies, scheduling and provisioning fai...