Principal Data Engineer
About Us Are you ready to build the future of the supply chain? At Gather AI, we're not just creating software; we're pioneering a new era of warehouse intelligence. We've developed a groundbreaking, vision-powered platform that uses autonomous drones and existing equipment to capture real-time data, completely digitizing workflows that have historically been manual and error-prone. This means facilities operate smarter, safer, and more efficiently, ultimately redefining "on-time, in full" delivery. If you're looking for an opportunity to contribute to truly transformative technology and make a significant impact in a vital industry, Gather AI is the place for you. We're leading the charge in the rapidly evolving robotics industry, and we invite you to join us in reshaping the global supply chain, one intelligent warehouse at a time. About the Team You'll join the Data Platform team at its inception, helping establish the foundation from day one . Today, production and analytical workloads share a single database, and every product team defines its own metrics. This team exists to fix that: designing the warehouse, the transformation layers, and the semantic model that every product and dashboard will build on going forward, in close partnership with product, engineering, and security. About the Role Most data platform roles ask you to extend something someone else has already built. This one starts with a blank canvas. As Principal Data Engineer, you'll architect Gather AI's data foundation from the ground up: separating analytical workloads from live production traffic, building a semantic layer so metrics are defined once and stay consistent everywhere, and linking structured records to real drone imagery and video with full traceability. You'll prove the model end to end on Gather's drone product, then generalize it so every new product extends the foundation instead of rebuilding it, all while working as a Principal-level individual contributor with real influence across engineering, product, and leadership. What You'll Do Architect a greenfield, multi-layer data warehouse (raw, refined, serving) that separates analytical workloads from production OLTP traffic. Deliver a governed, self-service data-access layer for internal consumers first (Product, CSM, Deployment/Operations, and Leadership) as Phase 1, ahead of customer-facing conversational analytics. Build a semantic and metrics layer so every metric, such as "scan accuracy by site," is defined once in code and stays identical across every dashboard and product, making self-service safe from metric drift. Own the quality bar: 99%+ availability SLA with freshness guarantees, 100% traceability, zero cross-tenant leakage, 99.5%+ pipeline success, and no data loss. Design tenant isolation, per-tenant cost attribution, and schema and row-level RBAC to scale toward hundreds of tenants (300+ target), not today's fleet size. Own data-ingestion correctness at the boundary with the integr...