exiger

Site Reliability Engineer

Jersey City, New Jersey, United States; McLean, Virginia, United States; Richmond, Virginia, United StatesFull timePosted 19 days ago
Apply on exiger →

Sign into see who you know at exiger.

Who We Are: Exiger transforms supply chains into a strategic advantage, advancing our mission to make the world a safer and more transparent place to succeed. Our AI platform, 1Exiger, delivers instant visibility into complex supplier ecosystems, leveraging proprietary data and advanced AI to surface risk, automate compliance, and unlock efficiencies and cost savings to strengthen long-term resilience. Trusted by 550+ global customers, including Fortune 500 companies and U.S. government agencies, Exiger is a recognized, award-winning leader in supply chain AI and a FedRAMP® authorized provider to the federal government.Site Reliability Engineer Location: U.S. (Hybrid) This role requires U.S. citizenship and eligibility for a U.S. security clearance.   Role Summary: Exiger is transforming how governments and global enterprises manage supply chain, defense, and geopolitical risk. Our AI-powered platform equips the world's most important institutions with the intelligence they need to protect critical infrastructure, secure national interests, and make data-driven operational decisions. From identifying counterfeit parts in defense supply chains to anticipating geopolitical risk exposure, Exiger enables mission owners to act with clarity and confidence in complex, high-stakes environments. This is our first dedicated Site Reliability Engineering hire and a founding role. You will help stand up the SRE function at Exiger: setting the standards, tooling, and practices that keep 1Exiger reliable for our 550+ customers, including Fortune 500 companies and U.S. government agencies. You will own reliability across the full service lifecycle, from design and capacity planning through deployment, monitoring, and incident response, and build the automation that lets the platform scale without scaling headcount. Because you are first, we need someone who has practiced SRE before and can bring the playbook, not learn it on the job. You will use your expertise in coding, algorithms, complexity analysis, and large-scale distributed system design to solve the reliability challenges that are unique to operating a mission-critical AI platform in regulated and government environments. SRE's culture of intellectual curiosity, problem solving and openness is key to its success. Our organization brings together people with a wide variety of backgrounds, experiences and perspectives. We encourage them to collaborate, think big and take risks in a blame-free environment. We promote self-direction to work on meaningful projects, while we also strive to create an environment that provides the support and mentorship needed to learn and grow.   What You'll Do: Establish the SRE function: define SLIs, SLOs, and error budgets, and set reliability standards that other engineering teams adopt. Build and own observability: instrument services for availability, latency, and system health, and turn that signal into actionable insight. Drive decisions with dat...