Tricentis

Senior Director, Cloud and Site Reliability Engineering

CZ - PragueFull timeDirectorPosted 22 days ago
Apply on Tricentis →

Sign into see who you know at Tricentis.

We are looking for an experienced and strategic leader to build and scale our Cloud and Site Reliability Engineering (SRE) organization. You will define and drive the cloud infrastructure strategy and operational excellence that underpins Tricentis' SaaS platform, ensuring the highest levels of availability, reliability, and performance. You will lead a team of talented Cloud Engineers and SREs, fostering a culture of excellence, automation-first thinking, and continuous improvement.What you will do:Cloud Strategy & Infrastructure LeadershipDefine and execute the cloud infrastructure roadmap to support Tricentis' SaaS platform growth, reliability, and scalability goals across AWS, Azure, and GCP.Establish cloud architecture standards and best practices including multi-cloud, hybrid-cloud, and cloud-native strategies.Drive infrastructure cost optimization and efficiency, partnering with Finance and Engineering leadership to align cloud spending with business outcomes.Lead the adoption of modern cloud technologies and emerging capabilities (AI and Agentic) to advance platform capabilities.Collaborate with peer Engineering and Product leaders to align cloud and infrastructure initiatives with product roadmap and business goals.Site Reliability Engineering & Operational ExcellenceBuild and mature the SRE function defining SLOs, SLIs, and error budgets that reflect customer expectations and business commitments.Enhance operational effectiveness through the deployment and use of agentic capabilities to scale the team to meet enhance performance and reliability of our SaaS products.Own the incident management and on-call strategy to establish effective processes for detection, response, remediation, and post-incident review improving MTTR.Champion a culture of reliability embedding SRE principles across the broader Engineering organization to reduce toil and improve system resilience. Drive automation across infrastructure provisioning, monitoring, observability, and self-healing systems.Partner with Security to ensure cloud environments meet compliance (SOC 2, ISO 27001, ISO 42001, GDPR, FedRAMP, and others as required).Engineering Execution & DeliveryWork with Engineering teams to influence infrastructure design earlier in the agentic development process, as a first-party concern design constraint through AI skills and agents.Oversee infrastructure delivery and operational readiness for all product releases, ensuring systems are observable, scalable, and fault tolerant.Drive continuous improvement in CI/CD pipelines, deployment processes, and DevOps tooling in partnership with product engineering teams.Establish and enforce infrastructure-as-code practices (Terraform, Pulumi, or equivalent) to increase consistency and reduce operational risk.Define and track key reliability, performance, and availability of metrics, reporting regularly to senior leadership on platform health.Who you are:10+ years of experience in cloud infrastructure, D...