Sieve

Member of Technical Staff, Infrastructure

San Francisco, California, United StatesFull timeStaff$150,000 - $350,000 / yearPosted 2 months ago
Apply on Sieve →Sign in to save this role

Sign in to see who you know at Sieve.

ABOUT US

Sieve is a multi-modal lab curating the world's highest-quality training datasets — spanning video, audio, images, text, and 3D. We combine exabyte-scale data infrastructure and novel multimodal understanding techniques that push the frontier of foundation models. Video alone makes up 80% of internet traffic, and across modalities, data has become the enabling medium powering creativity, communication, gaming, AR/VR, and robotics. Sieve exists to solve the biggest bottleneck in the growth of these applications: high-quality training data.

We partner with top AI labs and did $XXM last quarter alone, as a team of ~30 people. We also raised our Series A from Tier 1 firms such as Matrix Partners https://matrix.vc/, Swift Ventures https://www.swift.vc/, Y Combinator https://www.ycombinator.com/, and AI Grant https://aigrant.com/.

WHY NOW

Sieve is one of the most capital-efficient teams in AI — roughly 30 people serving the world's leading AI labs across every major data modality. You'll join early, own problems end-to-end, and watch your work ship directly into the models defining the frontier.

ABOUT THE ROLE

As an infrastructure engineer at Sieve https://www.sievedata.com/, you’ll design and engineer systems that handle the compute, scheduling, and orchestration of complex ML + ETL pipelines that need to run quickly, reliably, and cost-effectively on large sums of video.

You’re likely a good fit if you love optimizing for system uptime, have worked with cloud technologies, optimizing hyper-fast distributed systems at the scale of thousands of GPUs, and building great internal tooling and CI/CD for rapid iteration.

REQUIREMENTS

  • 3+ years of experience building foundational data infrastructure
  • Proficient in working across diverse cloud architectures
  • Designed and maintained pipelines that process petabytes of data
  • Developed robust CI/CD pipelines tailored for ML-focused teams
  • Strong coding experience with Go and Python; Ex...

Read the full posting on Sieve →