Staff, Reliability Engineer
Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities.Join Tenstorrent as a Staff Reliability Engineer and help define the reliability strategy behind the next generation of AI computing systems. In this highly visible technical leadership role, you'll drive reliability from architecture through production, partnering across hardware, software, and manufacturing teams to build high-performance AI platforms that set the standard for uptime, durability, and quality. If you're passionate about solving complex engineering challenges and influencing products at scale, you'll have the opportunity to shape technology powering the future of AI. This role is hybrid, based out of Toronto, Canada. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting. Who You Are You've spent 8+ years in reliability engineering, ideally in high-performance computing, AI hardware, or data center systems. You're comfortable with the statistical side of the job, HALT, HASS, ALT, MTBF, Weibull analysis, and FMEA are all familiar territory. You can work through a technical problem in a thermal lab and then explain the risks and trade-offs clearly to leadership. You're good at bringing people together, mechanical, electrical, thermal, software, and System Dev & Compliance Validation teams, especially when timelines are tight. You hold a Bachelor's or Master's in Mechanical Engineering, Electrical Engineering, Reliability Engineering, or a related field. What We Need Someone to set the reliability strategy for our next-generation AI computing systems, with MTBF modeling as a core piece. A strong problem-solver who can lead root-cause investigations and drive fixes across engineering and the supply chain. Someone who summarizes findings and feeds them back to the Systems Engineering design team, reliability as an ongoing loop, not a one-time check. A close partner to System Dev & Compliance Validation, helping set hardware up for success ahead of formal testing and certification. Willingness to travel to third-party test facilities for HALT/HASS, to preplan, oversee testing, resolve DUT issues, and assess design risk in person. What You Will Learn How to build a reliability strategy from scratch for hardware that's...